From 86c4b9a0d8bf29a8b333b6d37d3d48ae2557e2f2 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 24 Sep 2026 08:52:42 +0200 Subject: [PATCH] 2026-09-24: 09 decisions 24-25, part 5 shipped; the whole-copy truth table; fourth suppression; rows; STATUS; evidence Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 21 + REPORT.md | 59 +- STATUS.md | 39 +- .../architecture/00-capability-map.md | 3 +- .../architecture/07-backup-architecture.md | 12 + documentation/architecture/08-alarm-ladder.md | 5 +- .../architecture/09-update-architecture.md | 24 +- .../audits/ladder-2026-09-24/README.md | 106 ++ .../part0/10-hub-0122-rollout.txt | 10 + .../part0/20-9202-before.txt | 15 + .../part0/30-9202-deploy-0.268.0.txt | 8 + .../ladder-2026-09-24/part0/40-floor-0268.txt | 6 + .../part0/41-arrival-0268.txt | 6 + .../01-v0267-install-backup-restore.json | 100 ++ .../partA/01-v0267-install-backup-restore.log | 30 + .../audits/ladder-2026-09-24/partA/01.stdout | 30 + .../partA/02-v0268-undo-after-restore.json | 205 ++++ .../partA/02-v0268-undo-after-restore.log | 69 ++ .../audits/ladder-2026-09-24/partA/02.stdout | 69 ++ .../partA/03-why-g-was-done.txt | 31 + .../audits/ladder-2026-09-24/partA/state.json | 15 + .../partB/01-round11-reproduced.json | 75 ++ .../partB/01-round11-reproduced.log | 37 + .../audits/ladder-2026-09-24/partB/01.stdout | 37 + .../partB/02-r660-positive-observables.txt | 27 + .../partB/03-r660-control.json | 8 + .../partB/03-r660-control.log | 16 + .../partB/04-states-and-banner.txt | 5 + .../partB/05-quiet-watch.txt | 2 + .../partB/06-r660-control-banner.json | 128 ++ .../ladder-2026-09-24/partB/07-teardown.log | 30 + .../partD/00-repoint-drill.txt | 10 + .../ladder-2026-09-24/partD/00-spike.json | 110 ++ .../ladder-2026-09-24/partD/00-spike.log | 27 + .../ladder-2026-09-24/partD/00-spike.stdout | 27 + .../partD/01-spike-steps-in-clone.txt | 25 + .../partD/10-romm-two-steps.json | 295 +++++ .../partD/10-romm-two-steps.log | 70 ++ .../audits/ladder-2026-09-24/partD/10.stdout | 70 ++ .../ladder-2026-09-24/partD/11-teardown.log | 27 + .../partD/20-9202-repoint-live.txt | 28 + .../redproofs/B-r659-hold.txt | 9 + .../redproofs/B-r659-page.txt | 6 + .../redproofs/C-r653-r656.txt | 9 + .../redproofs/D-catalog-gate.txt | 5 + .../redproofs/D-catalog-writer.txt | 5 + .../ladder-2026-09-24/redproofs/D-ladder.txt | 7 + .../ladder-2026-09-24/tools/fixtures.py | 1085 +++++++++++++++++ .../audits/ladder-2026-09-24/tools/liveA1.py | 64 + .../audits/ladder-2026-09-24/tools/liveA2.py | 96 ++ .../audits/ladder-2026-09-24/tools/liveB.py | 117 ++ .../ladder-2026-09-24/tools/liveC660.py | 29 + .../ladder-2026-09-24/tools/liveC660b.py | 26 + .../audits/ladder-2026-09-24/tools/liveD.py | 71 ++ .../audits/ladder-2026-09-24/tools/repoint.py | 60 + .../audits/ladder-2026-09-24/tools/spike.py | 60 + .../ladder-2026-09-24/tools/teardownB.py | 21 + .../audits/ladder-2026-09-24/tools/walk.py | 507 ++++++++ documentation/backlog/CLOSED-ITEMS.md | 14 + documentation/backlog/OPEN-ITEMS.md | 15 +- 60 files changed, 4066 insertions(+), 57 deletions(-) create mode 100644 documentation/audits/ladder-2026-09-24/README.md create mode 100644 documentation/audits/ladder-2026-09-24/part0/10-hub-0122-rollout.txt create mode 100644 documentation/audits/ladder-2026-09-24/part0/20-9202-before.txt create mode 100644 documentation/audits/ladder-2026-09-24/part0/30-9202-deploy-0.268.0.txt create mode 100644 documentation/audits/ladder-2026-09-24/part0/40-floor-0268.txt create mode 100644 documentation/audits/ladder-2026-09-24/part0/41-arrival-0268.txt create mode 100644 documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.json create mode 100644 documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.log create mode 100644 documentation/audits/ladder-2026-09-24/partA/01.stdout create mode 100644 documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.json create mode 100644 documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.log create mode 100644 documentation/audits/ladder-2026-09-24/partA/02.stdout create mode 100644 documentation/audits/ladder-2026-09-24/partA/03-why-g-was-done.txt create mode 100644 documentation/audits/ladder-2026-09-24/partA/state.json create mode 100644 documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.json create mode 100644 documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.log create mode 100644 documentation/audits/ladder-2026-09-24/partB/01.stdout create mode 100644 documentation/audits/ladder-2026-09-24/partB/02-r660-positive-observables.txt create mode 100644 documentation/audits/ladder-2026-09-24/partB/03-r660-control.json create mode 100644 documentation/audits/ladder-2026-09-24/partB/03-r660-control.log create mode 100644 documentation/audits/ladder-2026-09-24/partB/04-states-and-banner.txt create mode 100644 documentation/audits/ladder-2026-09-24/partB/05-quiet-watch.txt create mode 100644 documentation/audits/ladder-2026-09-24/partB/06-r660-control-banner.json create mode 100644 documentation/audits/ladder-2026-09-24/partB/07-teardown.log create mode 100644 documentation/audits/ladder-2026-09-24/partD/00-repoint-drill.txt create mode 100644 documentation/audits/ladder-2026-09-24/partD/00-spike.json create mode 100644 documentation/audits/ladder-2026-09-24/partD/00-spike.log create mode 100644 documentation/audits/ladder-2026-09-24/partD/00-spike.stdout create mode 100644 documentation/audits/ladder-2026-09-24/partD/01-spike-steps-in-clone.txt create mode 100644 documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.json create mode 100644 documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.log create mode 100644 documentation/audits/ladder-2026-09-24/partD/10.stdout create mode 100644 documentation/audits/ladder-2026-09-24/partD/11-teardown.log create mode 100644 documentation/audits/ladder-2026-09-24/partD/20-9202-repoint-live.txt create mode 100644 documentation/audits/ladder-2026-09-24/redproofs/B-r659-hold.txt create mode 100644 documentation/audits/ladder-2026-09-24/redproofs/B-r659-page.txt create mode 100644 documentation/audits/ladder-2026-09-24/redproofs/C-r653-r656.txt create mode 100644 documentation/audits/ladder-2026-09-24/redproofs/D-catalog-gate.txt create mode 100644 documentation/audits/ladder-2026-09-24/redproofs/D-catalog-writer.txt create mode 100644 documentation/audits/ladder-2026-09-24/redproofs/D-ladder.txt create mode 100644 documentation/audits/ladder-2026-09-24/tools/fixtures.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/liveA1.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/liveA2.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/liveB.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/liveC660.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/liveC660b.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/liveD.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/repoint.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/spike.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/teardownB.py create mode 100644 documentation/audits/ladder-2026-09-24/tools/walk.py diff --git a/CONTEXT.md b/CONTEXT.md index 4becc7ba..fb3bfdbc 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,27 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-24 (morning) — the fleet on 0.267.0 then 0.268.0; the undo after a restore; option A; the ladder on the box + +**Operator rulings (`09` §3 decisions 24–25).** 24 — the fleet takes 0.267.0 although the chaos hour's stop rule +fired (both faults predate it). 25 — R-659 option A: a held app's page names only a copy that brings the app back +WHOLE; with none, it says so + support is informed + `app_hold_no_whole_copy` (critical, operator-only); option B +(a database-only restore under existing files) is NOT built. + +**Shipped.** Controller **v0.268.0** (`206b035`, MinAgent 0.131.0): R-658 (undo selects volumes from the compose +definition; the restore labels what it creates; the remove counts unlabelled volumes), R-659 (`backup.WholeOnTier` += the restores' own refusals; `HoldAfterFailedUpdateWhole`), R-660 (fourth suppression, `UpdateHeldStacks`), R-651, +and `09` §6.4 **part 5** (one press = one tested step; `stacks.StepKey` = catalog `ladder.step_key`). Hub +**v0.122.0** (`app_hold_no_whole_copy` allow-listed + operator-only + per-app cooldown). Catalog **`5ed599c`**: +`steps/.yml` for every intermediate step (8 backfilled from the NEWEST commit naming the step's images), +gate rule 4 with decoys, the writer keeps the superseded step, R-653/R-656 in the bench, two suites un-drifted (R-663). +Floors: 0.267.0 at 05:12Z, **0.268.0 at 06:47Z**, both with declared MinAgent 0.131.0; both demo boxes arrived +healthy within ~15 s each time. Evidence and the claims the brief got wrong: `documentation/audits/ladder-2026-09-24/README.md`. + +**Measured, and now the design:** for an app with declared drive files, neither the own-unit nor the second-drive +unit restore brings it back (R-538's guard is in the shared function) — only off-site does (`07` §6 table; R-661 is +the second-drive gap). The box's `--depth 1` clone carries a template's `steps/` folder (whole tree, one commit). + ## 2026-09-24 (night shift 2026-09-23) — tests off DooPlex's Docker, the test record, 12 published steps, the chaos hour **Operator word, `09` §3 decision 21:** tonight the catalog may move every app whose within-a-major edge is diff --git a/REPORT.md b/REPORT.md index 4c774c20..984e8e49 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,34 +1,39 @@ -# REPORT — night shift 2026-09-23: DooPlex guarded, the test record, 12 published steps, the chaos hour +# REPORT — 2026-09-24: the fleet to 0.267.0 and 0.268.0; the undo after a restore; option A; the ladder -Full record: `documentation/audits/DRILL-night-2026-09-23.md`; step log `documentation/audits/night-2026-09-23/PROGRESS.md`. -Architecture read first: `09-update-architecture.md` (§3 decisions 11–20, §6.1a, §6.4), `07` §6, `08`. +Full record (not done first, claims checked, red-proofs, live proofs, teardown): +`documentation/audits/ladder-2026-09-24/README.md`. ## Not done, or changed -- **The floor stays 0.266.0** (the brief's rule: Part D found R-658 and R-659). Recommended to raise — the two - P1s are pre-existing; STATUS carries the question. -- adventurelog not moved (R-655); gitea inconclusive (R-624); 11 across-major and 12 no-route/no-fixture apps listed. -- Chaos rounds 2–5 landed their accidents outside the action (the runner waited on the wrong file; fixed from - round 5/6). The event column was rebuilt from the controller's log after the debug ring missed two holds. -- The stop rule fired at round 11 (R-659) after all twelve rounds had run. -## What shipped -- felhom-controller **v0.267.0** `80e6ad8c4772` (CI 919): R-650, R-640, R-499, R-518; deployed to 9202 only. -- app-catalog-felhom.eu: the test record + gates `6db08a5` (CI 920), R-612/R-613 `a5a729a`, harness fixes - `1263773`, 12 steps: romm ×2, ghost, wishlist, opengist, navidrome, komga (+768M), n8n, emby, kimai-db, - immich-ML, nextcloud (CI 922–937). -- felhom.eu: evidence, `09` §3 decisions 21–23 and §6.4 parts 4/6, rows R-651..R-660, four closed, capability - map, nightly rotation, STATUS. +Nothing left undone. Changed: the operator's English sentence without „please"; the hold names one (the newest) +whole copy; Part A's second failing update was not a failing edge (R-665); R-653/R-656 unit-tested only. -## Red-proofs (each seen failing) -Controller 8 (see `felhom-controller/REPORT.md`); catalog gate 3 (`B2-gate-redproofs.txt`); R-612/R-613 -before/after on 9202 through the product. +## Baselines → ends -## Gates -controller: `go test -count=1 ./...` rc 0, `controller_gates.py` OK. catalog: `catalog_gates.py --fast` OK, -`test_gate_decoys.py` 80 OK, `test_ladder_writer.py` OK, `test_catalog_gates.py` OK. felhom.eu: `repo_gates.py` -(run at commit). `unproven.py --summary`: not walked 35 of 55 (unchanged). +| repo | start | end | +|---|---|---| +| felhom-controller | `80e6ad8c4772` v0.267.0 | `206b035` v0.268.0 | +| felhom.eu | `3e58c184f62e` hub v0.121.0 | hub v0.122.0 (`e4d45a8`, manifest `500488c`) + docs | +| app-catalog-felhom.eu | `585a7cba22cf` | `5ed599c` | +| felhom-agent | `d9864a94bf62` | untouched | -## Teardown — three layers -Machine: 9202 apps removed through the product (0 containers/volumes/copies), kept drive folders removed by name, -config identical to the saved copy, live catalog; bench 9401 destroyed; 9201 24 containers before/after. -Host: `pct list` as before; template removed. Hub: not touched (no floor change). +## Hub v0.122.0 + +`app_hold_no_whole_copy` in `allowedEventTypes` + `operatorOnlyEvents` + `perAppCooldownEvents`; test pinning both +registers, red-proofed. Rolled out by ArgoCD (Synced, image read back from the pod, startup line `felhom-hub 0.122.0 +starting`). + +## Floors + +0.267.0 (05:12:14Z) and 0.268.0 (06:47:13Z), each with declared MinAgent 0.131.0, read back from the hub; both demo +boxes healthy within 30 s. `drill-r50` stays held (agent 0.129.0). + +## Documents + +`09` §3 decisions 24–25 + §6.4 part 5 SHIPPED + §6.1a note; `07` §6 whole-copy truth table; `08` §5 fourth +suppression + the crash-loop measurement; capability map row; `CONTEXT.md`; `STATUS.md`. + +## Rows + +Closed R-658, R-659, R-660, R-651, R-653, R-656, R-40, R-663. Opened R-661, R-662, R-664, R-665, R-666, R-667. +Narrowed R-450. Register 336 → 335 rows, 683,233 → 679,393 bytes. diff --git a/STATUS.md b/STATUS.md index ac788dd1..eeb5c2f6 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,25 +1,30 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-24 (morning) — the night shift. Eleven apps moved to newer versions, each with a written test result. The chaos hour found two serious faults in the undo's way back. The new controller is NOT yet on the demo boxes: one question for you below.** +**Updated 2026-09-24 (morning session). The fleet is on controller 0.268.0. The undo works after a restore. A stopped app's page tells the truth. A box that is several updates behind now climbs one tested step per press.** -**Decisions I took on my own (you may reverse them):** -1. The memory test now counts only the app's own memory, not the kernel's file cache. The cache made two healthy apps read "100 % full" with no memory kills, and would have forced bigger memory limits for no reason. The old figure is still recorded beside the new one. -2. The catalog's push check now asks the image registry, but only for images a push moves. Without it, a tested image could change before it reaches a box. Other pushes make no network call. +**Decisions I took on my own:** none. Two small changes to your wording, both below. -**What I exercised.** Tests can no longer run real Docker commands on your own machine by accident. A restore now refuses a cut-off database copy before it touches anything (proven with the real restore button). Two page texts now tell the truth: where this box's full backup really goes, and how long the backup button stops the apps (about 8 minutes). Every catalog version move now needs a written test result, and a check refuses a move without one. I moved 11 apps (12 steps), each tested twice: on a test bench with 10 minutes of memory watching, and on the scratch machine through the real Update button. The HP demo box then updated three of its own apps through the real button: all three finished and answer. +**What I did, and it worked.** +- **The demo boxes got 0.267.0**, as you said yes to. Then, after every live test passed, they got **0.268.0**. Both boxes arrived healthy about 15 seconds after each change. +- **The undo after a restore.** A restore used to rebuild an app's storage without the tag the undo looked for, so the undo copied nothing. Now the undo finds the storage by its name. A restore also puts the tag back. Tested on the scratch machine with an app restored by the OLD version, the way customer boxes are today: the failed update was undone, and data written before and after the restore came back. +- **A stopped app's page (your option A).** When no copy on the box can bring an app back with its files, the page now says so, says support is informed, and shows no restore button. You get an urgent alarm. Tested by repeating last night's exact case, in Hungarian and English. +- **Small fixes.** A stopped app no longer raises a second, false alarm. A removed app leaves no old files behind. The test bench now marks a memory test with no real load as "not proven", and it starts every run with an empty drive folder. +- **The update ladder.** One press now does one tested step, never a jump. Tested with RomM on the scratch machine: first press moved only the app, second press moved only the database engine. The data came back after each press. The page says how many steps remain. -**What broke, and whether it healed.** -- **After a restore, the automatic undo has nothing to put back.** A restore rebuilds an app's storage without the tag the undo uses to find it. The next failed update is then "undone" with no data copy. The app I tested happened to survive this. Not fixed tonight (one controller release per night). Serious. -- **A stopped app can be pointed at a restore that refuses it.** After a failed update and a failed undo, the page names a backup to restore. For an app with files on a drive, the restore refuses that backup, and on a box with no off-site copy nothing else brings the app back. The stop rule fired here; all chaos rounds had already run. Serious. -- Adventurelog's new version was not moved: our own template checks it the wrong way, and the new version needs an internet download at every start. -- Reinstalling Nextcloud over kept files never finishes, and the box only says "unhealthy". -- Smaller: a stopped app raises a second, extra alarm; the memory test once ran with no real load (fixed and re-run). +**Your wording, changed in two small ways.** +- English: I removed the word "please". The house rule for English copy forbids it. +- Hungarian: I kept your text exactly. It uses the formal "Ön" form; the rest of the screens use the informal form. That mismatch is already on the list. -**Rows.** Ten opened, four closed. The list went from 330 to 336. Wishlist and uptime-kuma were fixed in the catalog; their rows stay open, smaller. +**What I found (new, not fixed).** +- **The second drive cannot bring back an app with files, even though it holds everything.** Its restore button refuses such apps, and the file restore brings back only files, not the database. So for Nextcloud, Immich, Paperless and Calibre, only the off-site copy counts. This needs your decision. +- **A long crash loop never raises an alarm** if the app looks "up" for a moment between crashes. Seen on the scratch machine: 385 restarts, zero alarms. +- The stopped-app sentence says "do not remove the app", but the Remove button is still there. This needs your word. +- Smaller: a ladder step has no health check of its own yet; right after a restore, an update may be judged with the restored health check; one restore option ("database only") is described in the code but offered nowhere. -**What needs you — one question.** Should the demo boxes get controller 0.267.0 now? -- **Yes, raise the floor (my recommendation).** The two serious faults are also in the version the boxes run today. 0.267.0 does not cause them, and it makes restores safer. -- **No, wait.** Nothing changes on the boxes. The fixes above stay on the scratch machine only, until you say so. -If you do nothing, the demo boxes stay on 0.266.0. +**Rows.** Seven closed, six opened. The list went from 336 to 335. -**Nothing on Peti's machine, the off-site box or your own machine (beyond normal pushes) was touched. The test bench was deleted. The scratch machine is back on the real catalogue with its standing apps.** +**What needs you — two questions.** +1. **The second drive and apps with files.** (a) Build a "files plus database" restore for the second drive (my recommendation — then the second drive is a real way back, as customers would expect). (b) Rule that the second drive is files-only for these apps. If you do nothing, the page keeps naming only the off-site copy for these apps, which is true. +2. **The Remove button on a stranded app.** (a) Hide it while support is informed (my recommendation — a removal destroys what support needs). (b) Keep it and soften the sentence. If you do nothing, both stay as they are. + +**Not done.** No automatic updates yet (plan part 7). No customer box touched by hand. Nothing on Peti's machine, the off-site box or your own machine was touched, beyond the normal hub update. The scratch machine is back on the real catalog with its standing apps. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 0bdae469..d9e05706 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -104,7 +104,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | | **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. **THE NARROWING ABOVE WAS CLOSED THE NEXT DAY, 2026-09-22 — and re-widened the row.** `audits/PROBE-FIX-2026-09-22.md`. All three wrong probes were corrected in the catalog (`app-catalog-felhom.eu@793c4fb`: tandoor `8080→80`, wger `80→8000`, zipline `/api/health→/api/healthcheck`) and **red-proofed live on 9202 through the product in both directions**: at the live pin all three read `Nem egészséges` / `Not healthy` on their own app page while docker reported every container healthy and the front door served a real page; after the real sync all three read `Fut` / `Running` with no redeploy. **tandoor's edge was then re-walked with nothing else changed and ended `done` at +41.1 s**, seed read back, where the identical edge had ended `failed` at +361.9 s with the app stopped — so the tally is now **15 proven, 2 failed, 4 inconclusive**, and all fifteen are on the live catalog. A `--fast` catalog gate (`check-probe-matches-compose.py`) now refuses a probe that does not match the same service's own compose healthcheck, with four red-proofs and ten decoys including the no-PyYAML mode CI actually runs. **WHAT THIS ROW STILL CANNOT CLAIM, and the reason is exactly R-96 rule 3:** the guard is now shown correct for **47 of 53** templates. `paperless-ngx`'s probe has **never run on any box** — no container name matches its stack name, so it is silently skipped and its badge can never go red (**R-630**); and five more cannot be judged statically, one of which (`home-assistant`) is right only because its check type cannot fail (**R-631**). An absent alarm is equally consistent with healthy and with never checked. **And the sweep's ceiling, counted: 28 of the 53 templates have never been deployed by any drill (R-632).** **THAT CEILING WAS REMOVED THE SAME NIGHT, 2026-09-22 — all 28 walked (`audits/DRILL-the-28-2026-09-22.md`), so every template in the catalog has now been attempted at least once.** 26 of 28 deployed, **6 proven**, 5 inconclusive, 14 with no within-a-major edge upstream, 1 failed honestly and 2 undeployable — one of those (`plant-it`) **by design**, refused by the product's lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a **restore from its own copy, with the seed read back again** — 21 restored, and **2 were correctly REFUSED** with the sentence `07` §6.2 predicts for a class-A app whose local copy holds no file leg. **AND THE NIGHT NARROWED THIS ROW AGAIN, in the place the probe work could not reach.** `paperless-ngx` has no container matching its stack name, so no probe is ever built for it — and `verifying` does not skip: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three containers all read `healthy`. The controller's own words: *`not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app`*, at **+313.0 s**. **R-630, raised to P1.** So "the update is guarded" is now shown correct for 47 of 53 templates, wrong for none, and **actively harmful for the one template that has no probe at all**. **Two further limits on what this row may claim, both about STATE rather than health:** a `remove` sent while a restore is still running reports success and leaves a container restarting with a live public route (**R-633**) — while the product already refuses exactly that clash for `update` and for `restore`, naming the blocking operation; and an app can be **running, healthy and serving while recorded as `deployed: false`**, in which state the product refuses to remove it at all (**R-634**). In both, a person needed a shell to clear what the product could not. **ALL THREE ARE FIXED IN CONTROLLER v0.262.0 (2026-09-22), and the first is PROVEN LIVE.** *The stopped app:* `verifying` no longer loops on a probe that resolves to nothing — it settles on container state, the way an app declaring no check is judged, and says which it did. **Measured on paperless-ngx: the identical Update that ended `failed` at +313.0 s with the app stopped now ends `done` at +53.4 s**, with no `no probe container` warning in the log because the explicit `healthcheck.container` resolved the target. *The ghost:* `RemoveStack` consults the backup side's `Busy` guard — which the product already applied to `update` and to `restore` — and then WATCHES the compose project for 25 s after `down`, removing anything that carries its label and answering `verified: true/false`, because `down` returning 0 is a request rather than a result. *The unremovable app:* the refusal now asks whether anything EXISTS (containers, a compose file, an `app.yaml`) instead of reading a flag. **WHAT THIS ROW STILL MAY NOT CLAIM:** R-634's MECHANISM — why `deployed` goes false while containers run — **is not diagnosed**; only the consequence is fixed. And a probe can be right about the port and still wrong about what a 200 means: `romm` answered 200 from nginx for six hours while its workers were OOM-killed behind it (**R-635**). **"The update is guarded" has never meant "the new version runs".** | | **A failed update is UNDONE by the box itself — the previous version back with its data from seconds before the update; the app is held only if that undo fails too** | controller **v0.263.2** (the undo), **v0.264.0** + hub **v0.120.0** (the mail) | **PROVEN-LIVE (2026-09-23)** | `audits/undo-live-2026-09-23/README.md` (+ `audits/undo-bakeoff-2026-09-23/` for the method). On scratch guest 9202, through the endpoints the UI invokes: docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite in a volume), each a real migrating edge failing a deliberately wrong probe, **undone in 30–52 s** with seeds written before the backup, after it and seconds before the press ALL read back through each app's front door, ledgers equal to before; page line in hu and en. Cut-off copy → HOLD (`untouched`), prefix in the box's language; power cut (`pct stop`) during `undoing` → resumed after boot and undone; a person's press after an undo → `done`; removal deletes kept copies. **NARROWED 2026-09-23 night (`audits/DRILL-night-2026-09-23.md` Part D):** a controller `kill -9` in `verifying` → resumed after restart → undone, seeds read back (round 9); a 1 GB memory hog and a whole-box backup during `verifying` did not disturb it (rounds 6, 8). **Two holes found:** after ANY restore the app's volumes lose their compose label and the undo copies NOTHING (R-658, P1); and a held FILE-LEG app is told to restore from a copy the restore then refuses, with no other route on a box without an off-site tier (R-659, P1). | Folder copy of NAMED volumes only (`09` §3 decision 19); bind folders never touched. **The household is told (2026-09-23, `audits/undo-fleet-2026-09-23/`):** on guest 9201 the household and the operator each received ONE mail per app per outcome — undone and held, in Hungarian (vikunja) and, after one language switch, in English (glance) — with the app named in the subject; `notification_log` rows 937–946 all `sent`. **Not built:** the automatic caller (part 7) — this protects the manual button today. Leftovers: R-647. | -| **The catalog holds only TESTED steps — an image move without a proven test record (bench + box, digests the registry still serves) is refused at push time** | catalog `6db08a5` (`scripts/check-test-record*.py`, `ladder.py`, `upgrade-test.py --write-ladder`) | **PROVEN (2026-09-23 night)** | `audits/DRILL-night-2026-09-23.md` Part B/C | 16 decoy cases both ways + 3 red-proofs; the gate judged every one of the night's 12 published steps against the live registry and refused a wrong-digest control. 21 earlier moves backfilled from their records. | The BOX does not read the ladder yet (`09` §6.4 part 5): a box two steps behind still jumps (romm on demo-hp took two steps in one press). | +| **The catalog holds only TESTED steps — an image move without a proven test record (bench + box, digests the registry still serves) is refused at push time** | catalog `6db08a5` (`scripts/check-test-record*.py`, `ladder.py`, `upgrade-test.py --write-ladder`) | **PROVEN (2026-09-23 night)** | `audits/DRILL-night-2026-09-23.md` Part B/C | 16 decoy cases both ways + 3 red-proofs; the gate judged every one of the night's 12 published steps against the live registry and refused a wrong-digest control. 21 earlier moves backfilled from their records. | **Closed 2026-09-24 by the row below** (v0.268.0 reads the ladder). | +| **A box behind climbs ONE tested step per press, each with its own definition** (`09` §3 decision 14) | controller **v0.268.0** (`stacks/ladder.go`) + catalog `5ed599c` (`steps/.yml`, gate rule 4) | **PROVEN-LIVE (2026-09-24)** | `audits/ladder-2026-09-24/partD/10-romm-two-steps.json` | romm 5.3.0/11.4 → 5.3.1/11.4 → 5.3.1/11.8 in two presses on 9202, seeded account read back after each; page count 2 → 1 → none; spike confirmed v0.267.0 jumped (`partD/00-spike.json`). 6 unit tests + red-proofs. | Nothing climbs by itself yet (part 7). A step has no `.felhom.yml` of its own (R-664). | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 20a5ba47..c93a0ae4 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -307,6 +307,18 @@ the unit, the Tier-2 mirror AND the Tier-2 record, so „Teljes visszaállítás (R-486); „Mentési adatok törlése" deletes the unit, every mirror and the app's backup preferences, and never touches off-site snapshots (R-474). A removed app whose unit was kept is listed on both local backup pages with its restore since controller v0.242.0 (R-487): **the local lists are keyed on the drives, not on what is deployed** — the rule R-237 set for the off-site list — and the restore opens the unit where it sits, a data drive included. **For a bind-data app the update's tier order is second drive → off-site → own unit (R-479, v0.241.0), because the unit does not hold the files.** +**[FACT, measured from source 2026-09-24, controller v0.267.0/v0.268.0] Which copy brings an app back WHOLE +— read from the restores' own refusals, not from what the tier stores** (`09` §3 decision 25, R-659): + +| app | own unit (Tier 1) | second drive (Tier 2) | off-site (Tier 3) | +|---|---|---|---| +| no declared drive files (class B, 45 apps) | whole — the unit restore | whole — „Teljes visszaállítás" (the unit restore) | whole — the full restore | +| declared drive files (`DeclaredDriveFileLegs`: calibre-web, immich, nextcloud, paperless-ngx) | **not** — refused (R-538) | **not** — its unit restore is refused by the same guard; the file restore only ADDS missing files (R-661) | whole — „Teljes visszaállítás (fájlok + adatbázis)" | + +`backup.WholeOnTier` is this table; `TestR659_TruthTableAgreesWithTheRestoresRefusal` pins that it cannot drift +from the refusal. An app with only OPTIONAL legs (audiobookshelf, komga, romm) is class-B here: its files are +never moved by an update or an undo, and the unit restore accepts it. + ### 6.1 The four tiers, as configured on the live fleet | Tier | Location | Captures | Cadence (LIVE) | Retention (LIVE) | Encrypted | diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index 34bdf963..06e55367 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -80,7 +80,7 @@ is allowed; every down member then reads as supervised. | `stopped`, `exited` | **yes** | not running, will not recover alone | | `degraded` | **yes** | [DESIGN, R-51] a dead supervised member is as unreachable as a single app that exited — immich-server sat Exited 18 h with the app 100 % dead and no alert | | `unhealthy` | **NO** | [DESIGN] a running container whose healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop. **R-384 did not change this** — it asks a prior question instead | -| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) | +| `restarting` | **NO**, until sustained | [DESIGN, C9-F2] `restarting` is on the normal deploy path, so folding it in would alarm fleet-wide on every update. Becomes down after **5 min** (`crashLoopAfter`) of UNINTERRUPTED `restarting` — **measured 2026-09-24: a container that reads `running` for a moment between restarts resets the clock and never gets there (gokapi, 385 restarts, „0 currently down"; R-667)** | | `starting`, `deploying` | no | mid-start | | `paused` | no | a deliberate user action | | `unknown` | no | [DESIGN] fail-OPEN — never manufacture a dead-app alert from an inconclusive read | @@ -91,13 +91,14 @@ member is dead; only the excuse is missing). Both are recorded at their sites. --- -## 5. The three suppressions, all at `classifyRunStates` +## 5. The four suppressions, all at `classifyRunStates` | Suppression | Rule | Expires? | |---|---|---| | **deliberate user stop** | `StateStopped` is not down **unless** the quiesce loop reports it failed to restart that stack | n/a — lifted by `failedRestart` | | **quiesce cycle** | a stack this backup cycle stopped is exempt | **yes**, 180 s after unquiescing | | **boot grace** | no evaluation for 90 s after controller start | **yes** | +| **update hold** (v0.268.0, R-660) | an app HELD after a failed update is stopped by the product and has its own event (`app_update_held`, and `app_hold_no_whole_copy` when no copy brings it back whole); it is not "down" | **yes** — lifted with the hold (a restore, or the operator). A RESTORE hold (R-379) is deliberately not in the set: it has no event of its own | **None of them latch.** [DESIGN, R-97b + R-88 Scenario D] A permanent suppression trades a loud false alarm for a silent real one, which is the same error as an over-eager alarm, in the opposite diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 68be7b56..a549510c 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -311,6 +311,24 @@ builds them (`audits/update-rulings-2026-09-23/`). decision 17 says the box pulls the recorded digest; recording one the registry no longer serves makes every box's pull of that step fail (Scenario E). `app-catalog-felhom.eu/scripts/check-test-record-move.py`. +### 2026-09-24 — two operator rulings + +24. **The fleet takes controller v0.267.0 although the chaos hour's stop rule fired** — both faults the + rule caught (R-658, R-659) predate that release; v0.267.0 causes neither and adds R-640's protection to + the same restore path. Floor saved with MinAgent 0.131.0 at 05:12:14Z; both demo boxes on v0.267.0, + healthy, 10 s later (`audits/ladder-2026-09-24/part0/`). + +25. **R-659, option A: a held app's page names only a copy that can bring the app back WHOLE; when this box + has none, the page says so and that support is informed, and the operator gets an urgent event.** The + database-only restore under existing files (option B) is NOT built. *As built (v0.268.0 + hub + v0.122.0):* "whole" is read from the restores' OWN refusals, not from what a tier stores — measured from + source, an app with declared drive files (`DeclaredDriveFileLegs`) is refused by both the own-unit and the + second-drive unit restore (R-538's guard sits in the shared function), so for those apps only the + off-site copy counts (the second-drive gap is R-661). The hold names the newest whole copy with what it + holds; with none, `hold.update.no_whole_copy`, no Mentések button, `app_hold_no_whole_copy` (critical, + operator-only). The operator's English was used with one word dropped („please" — the house rule, + `i18n_missing_gate.py`); the Hungarian verbatim, its formal register recorded against R-516. + **RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update (`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4); R-636's louder repeated alarm. @@ -842,7 +860,9 @@ an app pinned before v0.263.2 has no such record until its next pin. **Two more, 2026-09-23 night (`audits/DRILL-night-2026-09-23.md` Part D):** the undo finds an app's volumes by their compose label, and a restore recreates them WITHOUT it — so after any restore the undo copies nothing (R-658, P1); and a held file-leg app is pointed at a restore that refuses a database-only copy, with no route left on a box without -an off-site tier (R-659, P1). A controller kill during `verifying` resumed and undid correctly (round 9). +an off-site tier (R-659, P1). A controller kill during `verifying` resumed and undid correctly (round 9). **Both fixed in v0.268.0 (2026-09-24):** the undo selects volumes from the app's compose definition and the restore +labels what it creates (R-658); the hold names only a copy that brings the app back whole, else says support is +informed (R-659, decision 25). Proven live on 9202, `audits/ladder-2026-09-24/`. The spike, as it was run by hand before any build: @@ -1081,7 +1101,7 @@ what the part can do to a household's data if it is wrong, not how likely that i | **2** | **SHIPPED — controller v0.264.0 + hub v0.120.0, proven live on 9202 and 9201 2026-09-23** (`audits/undo-fleet-2026-09-23/`). **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held: events `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, on by default, mailed in the household's language with the app named in the subject; per-app cooldown on both legs. Leftovers: R-647. | R-606, 15 | **1** | — | none | | **3** | **SHIPPED — controller v0.264.0, proven live on 9202 2026-09-23.** **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none | | **4** | **SHIPPED — catalog `6db08a5`, night 2026-09-23** (`audits/DRILL-night-2026-09-23.md`): `update_ladder:` in `.felhom.yml`, one JSON entry per line (spiked on controller v0.266.0 and v0.267.0 first — the controller ignores the key); `check-test-record.py` (static, CI) + `check-test-record-move.py` (history + the registry for moved refs only, decision 23); the only writer `upgrade-test.py --write-ladder`; the 21 moves of 2026-09-22 backfilled from their records (21 proven). Not built: `steps/.yml` (part 5 needs it once an app has two steps); `CompareImageRefs` did NOT move to the gate — the gate asks for a proven test instead, which is decision 13's own test. **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only | -| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape | +| **5** | **SHIPPED — controller v0.268.0 (`206b035`) + catalog `5ed599c`, proven live on 9202 2026-09-24** (`audits/ladder-2026-09-24/partD/`): romm 5.3.0/11.4 → 5.3.1/11.4 (its app step, from `steps/90dd9d68258286ef.yml`) → 5.3.1/11.8 (the engine step; `mariadb-upgrade` ran) in two presses, the seeded account read back after each, the page's „Hátralévő frissítési lépések" 2 → 1 → none. Step files: `templates//steps/.yml` (sha256 of `to` as canonical JSON, 16 hex; `stacks.StepKey` = `ladder.step_key`), refused absent or wrong by `check-test-record.py` rule 4, written by `--write-ladder` when a step is superseded, 8 backfilled from the NEWEST commit naming each step's images (romm's from `f4eb94f`, not the OOM-looping `15f9ebf`). The box reads them from its `--depth 1` clone (the whole tree is there — measured). One press = one step; a missing step file refuses before anything moves; an installed version matching no entry jumps, logged by name (measured live: vikunja 2.5.0). Limitation: a step has no `.felhom.yml` of its own (R-664). **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape | | **6** | **CATALOG HALF SHIPPED — night 2026-09-23:** every ladder entry carries the digest per `to` ref (`scripts/image_digest.py`, stdlib; equals Docker's `RepoDigests` on a box), and the move gate refuses a digest the registry no longer serves. **The box half (compare, render `name:tag@sha256`) is not built.** **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low | | **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe | | **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none | diff --git a/documentation/audits/ladder-2026-09-24/README.md b/documentation/audits/ladder-2026-09-24/README.md new file mode 100644 index 00000000..34161cf0 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/README.md @@ -0,0 +1,106 @@ +# 2026-09-24 — the fleet to 0.267.0 and 0.268.0; the undo after a restore; option A; the ladder on the box + +Architecture read first and named: `architecture/09-update-architecture.md` (§3 decisions 13–25, §6.1, §6.1a, +§6.4 parts 4–5), `07-backup-architecture.md` §6, `08-alarm-ladder.md` §5. **Method: endpoint-level** — every +product act is the endpoint the UI invokes; pages fetched as HTML in both languages (`?lang=`); no browser. +Scratch guest **9202** on demo-hp (Tier 0), drill catalog for every failing edge; the live catalog carried only +the step files and gates. + +## Not done, or changed + +- **Nothing not done.** Parts 0, A, B, C, D and E all ran; the interim "one step per app per day" gate was not + needed (Part D shipped). +- **The operator's English sentence lost one word, „please"** — the house rule (`i18n_missing_gate.py`, + `EN_FORBIDDEN`) refuses it. The Hungarian is verbatim; its formal register raised the R-516 ceiling 18 → 19 with + the reason written beside it. +- **The hold names ONE copy (the newest whole one), not a list.** A list would need new sentence copy; the existing + sentence names one copy with what it holds. +- **Part A's second failing update (g) did not fail:** right after a restore the stack dir's `.felhom.yml` was the + restored one (real probe port), so that press was a good update. The proof the brief asks for is (e) — an app + restored under v0.267.0, then a failing update on v0.268.0 — and it passed. (g) still shows the v0.268.0 restore's + volumes copied with no "unlabelled" warning. Filed as R-665 (not diagnosed). +- **R-653 and R-656 are unit-tested and red-proofed, not run on a bench** — no bench was provisioned this session. +- **Two catalog test suites were already red at the start** (`test_gate_decoys.py`, `test_ladder_writer.py` — last + night's pin moves under literals). Fixed and filed as R-663 (closed). +- **One of the first R-660 control's two alarm events is not attributable** (the dropped-event line names no app); + the repeat control produced exactly one. Likely gokapi (R-667), stated as inference. +- An older wiring test (`TestR475_AdapterReadsEveryTier`) asserted the adapter hands the hold the precondition copy + — the thing R-659 changes on purpose. Its assertion was replaced, with the reason in the test. + +## Claims in the brief, checked + +| the brief's claim | what was found | +|---|---| +| compose resolves volume names as `_` for every template (some set `name:`) | **No template sets `name:`**, top-level or per volume (all 53 read). Supported anyway, and tested. | +| the syncer can carry a `steps/` folder | **It does not need to.** The syncer copies two files into the stack dir; the box reads `steps/` from its catalog CLONE, which (depth 1) holds the whole tree. Measured on 9202: `steps/spike-B.yml` present in the clone after one sync (`partD/01`). | +| „saját meghajtó" never holds files for a file-leg app (from `07` §6) | **Held, and it goes further:** read from source, the SECOND-DRIVE unit restore refuses such an app too (the R-538 guard is in the shared function), and the second drive's file restore brings back no database. So only off-site brings a file app back whole (R-661). An app with only OPTIONAL legs (romm, komga, audiobookshelf) is accepted by the unit restore. | +| one press = one step is safe for an engine step whose datadir converts | **Held for MariaDB, measured once:** romm press 2 ran `mariadb-upgrade` 11.4.13 → 11.8.9 and the account read back. The undo copy covers the pre-conversion volume if the step fails (not exercised on an engine step this session). PostgreSQL majors stay refused (part 10). | +| today's press jumps A → C | **Held** on v0.267.0: vikunja 2.4.0 → 2.6.0 in one press, 2.5.0 never ran (`partD/00-spike.json`). | + +## Part 0 / Part E — the floor + +| save (UTC) | floor / MinAgent | read back | demo-felhom | demo-hp | +|---|---|---|---|---| +| 05:12:14 | 0.267.0 / 0.131.0 | `min_controller_version=0.267.0`, `min_agent=0.131.0`; hub `SERVED … from declared` | 0.267.0 healthy at +28 s | same | +| 06:47:13 | 0.268.0 / 0.131.0 | same shape | 0.268.0 healthy at +30 s | same | + +Hub 0.122.0 rolled out first (ArgoCD, manifest `500488c`, pod image read back). **Stays below:** `drill-r50` +(DOWN, agent 0.129.0 — held by MinAgent, a Tier-0 leftover). Evidence `part0/`. + +## Red-proofs (each seen failing, then restored) + +| row | mutation | failed at | file | +|---|---|---|---| +| R-658 undo | `appVolumes` returns the label selector | `copied 0 of 2 declared volumes` | `redproofs/A-r658-undo.txt` | +| R-658 restore | create without labels | `created vikunja_files WITHOUT the project label` | `A-r658-restore.txt` | +| R-658 remove report | `appVolumeSet` = labelled only | `volumes_removed = []` | `C-r651-and-A-remove-report.txt` | +| R-651 | skip deleting the applied files | both files survive; undo would probe the REMOVED install's file | same | +| R-659 hold | `WholeOnTier` always true | round-11 case names „saját meghajtó" | `B-r659-hold.txt` | +| R-659 page | gate dropped from `stacks.html` | Mentések button beside a no-whole-copy hold | `B-r659-page.txt` | +| R-659 hub | dropped from `operatorOnlyEvents` | "must be operator-only" | `B-hub-operator-only.txt` | +| R-660 | `updateHeldSet` returns nil / union line removed | "held app reported DOWN" / wiring | `C-r660.txt` | +| ladder (box) | `nextLadderStep` always the template | press 1 pinned C; C brought up after B failed; missing step file → `done` | `D-ladder.txt` | +| ladder (catalog gate) | rule 4 removed | 3 decoys pass wrongly | `D-catalog-gate.txt` | +| ladder (writer) | STEP block off | "has no definition" / file absent | `D-catalog-writer.txt` | +| R-653 / R-656 | `load_verdict` always reached / no removal | ghost's all-`err` watch reads reached; last run's config survives | `C-r653-r656.txt` | + +Green: controller `go build/vet/test` rc 0 (31 packages); `controller_gates.py` all OK; hub `go test` OK + +`repo_gates.py` OK; catalog `catalog_gates.py` OK, decoys 84 OK, writer 6 OK, bench 5 OK. + +## Live proofs on 9202 + +- **A (R-658)** — v0.267.0: vikunja installed, a pull-failing update took its own backup, restore from „helyi" → + both volumes' labels `null` (`partA/01`). v0.268.0: failing update (2.5.0 → 2.6.0, probe 8999) → log + *„carry no compose label … copied by name"*, *„will hold 2 named volume(s)"*, undone in 7 s, seeds A (before) and + B (after the restore) read back. v0.268.0 restore → labels `project`/`volume`/`version`, no compose warning + (`partA/02`). Remove → `volumes_removed` names both. +- **B (R-659)** — nextcloud installed at ladder step 1, seeded; failing update with the undo copy's marker removed + during `verifying` → HELD, `hold_no_whole_copy: true`; the sentence in hu and en exactly as ruled (ASCII + fragments, positive and negative control); no Mentések button on the app page or the list, both languages; log + *„NO copy on this box brings it back whole (seen: tier 2 …, tier 1 …; drive files declared: true)"* and + *„DROPPED event app_hold_no_whole_copy (severity critical)"* (R-620's line). Steps line „… lépések: 1" + (`partB/01`). +- **C (R-660, R-651)** — dead-app heartbeat 2 s after the hold: *„0 currently down"*; no `app_start_failed` for + the held app in ~12 min; a throwaway stopped out of band raised one (`partB/03`, `06`). Remove: `applied-compose.yml` + and `applied-meta/` present before, gone after; every undo copy gone (`partB/07`). +- **D (the ladder)** — romm installed at 5.3.0/11.4, seeded; the page read „Hátralévő frissítési lépések: 2" / + "Update steps remaining: 2". Press 1 → *„ladder — step 2 of 3 … from 90dd9d68258286ef.yml"*, pin 5.3.1/11.4, + applied == that step file, MariaDB „upgrade not required", account read back, count 1. Press 2 → *„the last step + (3 of 3)"*, pin 5.3.1/11.8, `mariadb-upgrade` 11.4.13 → 11.8.9, account read back, count gone (`partD/10`). A + second seed after press 1 was refused by RomM (403 for a non-admin) — the readback rests on seed A. + +## Teardown — three layers + +- **Machine (9202):** vikunja, nextcloud, romm, actualbudget ×2 removed THROUGH THE PRODUCT; 0 volumes, 0 undo + copies; drive folders nextcloud/romm removed by name (R-442 keeps them on 9202); 11 test images removed by name + (`/var/lib/docker` 27 % → 18 %), no prune. `controller.yaml` restored from `controller.yaml.pre-ladder0924`, git + URL read back as the live catalog, cache at live `5ed599c`. **9202 stays on v0.268.0.** Standing apps as found + (gokapi was crash-looping before the session — R-644/R-667). +- **Host (demo-hp):** no guest created; nothing else touched. +- **Hub:** no customer or appliance record created. Floor 0.268.0 is the intended end state. +- **Drill repo:** reset to live `main` `5ed599c`, `has_actions: False`, private — read back. + +## Rows + +Closed: R-658, R-659, R-660, R-651, R-653, R-656, R-40 (superseded), R-663 (filed and fixed). Opened: R-661, R-662, +R-664, R-665, R-666, R-667. Narrowed: R-450 (part 5 shipped). Register **336 → 335 rows, 683,233 → 679,393 bytes**. diff --git a/documentation/audits/ladder-2026-09-24/part0/10-hub-0122-rollout.txt b/documentation/audits/ladder-2026-09-24/part0/10-hub-0122-rollout.txt new file mode 100644 index 00000000..fab0d5e8 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/part0/10-hub-0122-rollout.txt @@ -0,0 +1,10 @@ +2026/09/24 07:29:33 [INFO] host-report from demo-hp-bb76ea (1 guests, 4 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14808 bytes) +2026/09/24 07:29:33 [INFO] DR-recipe host-half stored for customer demo-hp (host demo-hp-bb76ea, v1) +2026/09/24 07:31:48 [INFO] Offsite pool-box refreshed: 0.3% full (3.0 GB of 1.00 TB), Σ shared quota 250 GB, oversub 0.24x +2026/09/24 07:32:46 [INFO] PBS-DR box refreshed: 16.1% full (15.7 GB of 97.9 GB) +2026/09/24 07:33:27 [INFO] wgsync: pushed 4 peers to 167.233.158.164:22 +2026/09/24 07:34:43 [INFO] felhom-hub 0.122.0 starting +2026/09/24 07:34:43 [INFO] Default controller-version floor: 0.120.0 +2026/09/24 07:34:45 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000 +2026/09/24 07:34:45 [INFO] Registry version checker started (every 6h) +2026/09/24 07:34:45 [DEBUG] Registry version check: latest = 0.267.0 diff --git a/documentation/audits/ladder-2026-09-24/part0/20-9202-before.txt b/documentation/audits/ladder-2026-09-24/part0/20-9202-before.txt new file mode 100644 index 00000000..d01f5a4f --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/part0/20-9202-before.txt @@ -0,0 +1,15 @@ +privatebin privatebin/pdo:2.0.6 Up 5 hours (healthy) +paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 5 hours (healthy) +paperless-postgres postgres:16-alpine Up 5 hours (healthy) +paperless-redis redis:7-alpine Up 5 hours (healthy) +gokapi f0rc3/gokapi:v1.9.6 Restarting (1) 45 seconds ago +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.267.0 Up 7 hours (healthy) +filebrowser gtstef/filebrowser:1.3.3-stable Up 9 hours (healthy) +traefik traefik:v3.6.7 Up 9 hours +/dev/loop1 69G 18G 48G 27% /var/lib/docker +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: +actualbudget adventurelog audiobookshelf bentopdf bookstack calcom calibre-web claper code-server crafty-controller docmost emby filebrowser ghost gitea glance gokapi grafana gramps-web home-assistant homebox homepage immich jellyfin kimai komga mealie n8n navidrome nextcloud onlyoffice opengist outline paperless-ngx papra plant-it plex privatebin radarr rallly recipe-importer romm seerr sonarr sparkyfitness tandoor termix traefik uptime-kuma vaultwarden vikunja wanderer wger wishlist zipline diff --git a/documentation/audits/ladder-2026-09-24/part0/30-9202-deploy-0.268.0.txt b/documentation/audits/ladder-2026-09-24/part0/30-9202-deploy-0.268.0.txt new file mode 100644 index 00000000..94079b2b --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/part0/30-9202-deploy-0.268.0.txt @@ -0,0 +1,8 @@ +gitea.dooplex.hu/admin/felhom-controller:0.268.0 +gitea.dooplex.hu/admin/felhom-controller:0.268.0 Up 6 seconds (healthy) +2026/09/24 06:16:41 pin.go:290: [INFO] [stacks] pin adoption: 0 pinned, 4 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers) +2026/09/24 06:16:41 undo.go:656: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box +2026/09/24 06:16:41 main.go:527: [INFO] Metrics collector started (60s interval) +2026/09/24 06:16:41 scheduler.go:67: [DEBUG] [scheduler] scheduler started: periodic=8 daily=8 +2026/09/24 06:16:46 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only + diff --git a/documentation/audits/ladder-2026-09-24/part0/40-floor-0268.txt b/documentation/audits/ladder-2026-09-24/part0/40-floor-0268.txt new file mode 100644 index 00000000..abaa95c7 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/part0/40-floor-0268.txt @@ -0,0 +1,6 @@ +gitea.dooplex.hu/admin/felhom-hub:0.122.0 +2026-09-24T06:47:13Z +HTTP/1.1 303 See Other +Location: /configuration?flash=floor_set +name="min_agent" value="0.131.0" +name="min_controller_version" value="0.268.0" diff --git a/documentation/audits/ladder-2026-09-24/part0/41-arrival-0268.txt b/documentation/audits/ladder-2026-09-24/part0/41-arrival-0268.txt new file mode 100644 index 00000000..9a3559e2 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/part0/41-arrival-0268.txt @@ -0,0 +1,6 @@ +06:47:26 felhom-pve=gitea.dooplex.hu/admin/felhom-controller:0.268.0|Up 3 seconds (health: starting) demo-hp=gitea.dooplex.hu/admin/felhom-controller:0.268.0|Up 3 seconds (health: starting) +06:47:43 felhom-pve=gitea.dooplex.hu/admin/felhom-controller:0.268.0|Up 20 seconds (healthy) demo-hp=gitea.dooplex.hu/admin/felhom-controller:0.268.0|Up 20 seconds (healthy) +both-arrived +2026/09/24 08:47:14 [INFO] Global controller-version floor set to "0.268.0" (declared MinAgent "0.131.0") +2026/09/24 08:47:16 [INFO] managed floor SERVED for demo-felhom: floor 0.268.0, agent requirement "0.131.0" from declared (golden 0.258.0) +2026/09/24 08:47:17 [INFO] managed floor SERVED for demo-hp: floor 0.268.0, agent requirement "0.131.0" from declared (golden 0.258.0) diff --git a/documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.json b/documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.json new file mode 100644 index 00000000..017f100c --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.json @@ -0,0 +1,100 @@ +{ + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.267.0", + "drill_A": "58c7c938bc7c", + "deployed": true, + "seed_A": true, + "labels_after_install": "vikunja_vikunja_data {\"com.docker.compose.config-hash\":\"fe3fc6629444c7799d20fe7f96a201fcf52bd22fe1dcea326288a2a8963b48e0\",\"com.docker.compose.project\":\"vikunja\",\"com.docker.compose.version\":\"5.5.0\",\"com.docker.compose.volume\":\"vikunja_data\"}\nvikunja_vikunja_db {\"com.docker.compose.config-hash\":\"9c995441cdf40d38f9a48997f7376d7d20f88e50557ea8bca04237eeb3b2c9d1\",\"com.docker.compose.project\":\"vikunja\",\"com.docker.compose.version\":\"5.5.0\",\"com.docker.compose.volume\":\"vikunja_db\"}\n", + "drill_E": "ef42b4f34455", + "badge_wait": 22.8, + "press_backup": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "backing-up", + "label": "Biztonsági mentés készül a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 2.1, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 4.1, + "phase": "failed", + "label": "A frissítés nem sikerült", + "updating": false, + "error": "Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább.", + "hold": null + } + ], + "duration_s": 4.1, + "final_phase": "failed", + "update_error": "Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább.", + "hold_reason": null, + "state": "running" + }, + "snapshots": [ + { + "time": "2026-09-24T06:11:05Z", + "short_id": "helyi", + "tier": 1, + "drive_label": "Belső SSD (rendszer)" + } + ], + "drill_back": "087b3edb7475", + "restore": { + "ok": true, + "snapshot_id": "helyi", + "snapshots": [ + { + "time": "2026-09-24T06:11:05Z", + "short_id": "helyi", + "tier": 1, + "drive_label": "Belső SSD (rendszer)" + } + ], + "http": "HTTP/2 302", + "location": [ + "location: /backups/restore?flash=flash.restore.started" + ], + "seconds": 6.1, + "state_after": "running", + "hold_after": null, + "observables_after": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.5.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.5.0" + }, + "catalog_images": { + "vikunja": "vikunja/vikunja:2.5.99-notag" + }, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.5.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.5.0 running=true restarts=0" + ] + } + }, + "labels_after_restore_v0267": "vikunja_vikunja_data null\nvikunja_vikunja_db null\n", + "A_after_restore": true, + "seed_B_after_restore": true +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.log b/documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.log new file mode 100644 index 00000000..41ca37ea --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/01-v0267-install-backup-restore.log @@ -0,0 +1,30 @@ +08:10:26 drill 58c7c938bc7c: vikunja 2.5.0 (the install version) push rc=0 +08:10:30 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:10:35 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.5.0'} +08:10:36 vikunja: register http=200 +08:10:36 vikunja: create project http=201 +08:10:39 labels after install: +vikunja_vikunja_data {"com.docker.compose.config-hash":"fe3fc6629444c7799d20fe7f96a201fcf52bd22fe1dcea326288a2a8963b48e0","com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_data"} +vikunja_vikunja_db {"com.docker.compose.config-hash":"9c995441cdf40d38f9a48997f7376d7d20f88e50557ea8bca04237eeb3b2c9d1","com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_db"} + +08:10:40 drill ef42b4f34455: vikunja 2.5.99-notag (a tag that does not exist: backing-up runs, the pull fails, nothing moves) push rc=0 +08:11:03 [sync] the badge needed 22.8s and 4 sync+rescan rounds to catch up to vikunja/vikunja:2.5.99-notag — R-607's window, measured +08:11:03 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:11:03 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +08:11:05 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:11:06 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:11:07 + 4.1s phase=failed label=A frissítés nem sikerült err=Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább. hold=None +08:11:07 snapshots offered: [{"time": "2026-09-24T06:11:05Z", "short_id": "helyi", "tier": 1, "drive_label": "Bels\u0151 SSD (rendszer)"}] +08:11:08 drill 087b3edb7475: vikunja 2.5.0 (back to the install version) push rc=0 +08:11:13 [R] restoring vikunja from snapshot 'helyi' (of 1 offered) +08:11:13 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started'] +08:11:13 + 0.0s restore (True, None, None) +08:11:19 + 6.1s restore (False, None, None) +08:11:19 [R] after restore: state=running hold=None phase=failed +08:11:25 labels after the v0.267.0 restore: +vikunja_vikunja_data null +vikunja_vikunja_db null + +08:11:26 vikunja: readback of the seeded project http=200 ok=True +08:11:26 vikunja: register http=200 +08:11:26 vikunja: create project http=201 diff --git a/documentation/audits/ladder-2026-09-24/partA/01.stdout b/documentation/audits/ladder-2026-09-24/partA/01.stdout new file mode 100644 index 00000000..41ca37ea --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/01.stdout @@ -0,0 +1,30 @@ +08:10:26 drill 58c7c938bc7c: vikunja 2.5.0 (the install version) push rc=0 +08:10:30 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:10:35 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.5.0'} +08:10:36 vikunja: register http=200 +08:10:36 vikunja: create project http=201 +08:10:39 labels after install: +vikunja_vikunja_data {"com.docker.compose.config-hash":"fe3fc6629444c7799d20fe7f96a201fcf52bd22fe1dcea326288a2a8963b48e0","com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_data"} +vikunja_vikunja_db {"com.docker.compose.config-hash":"9c995441cdf40d38f9a48997f7376d7d20f88e50557ea8bca04237eeb3b2c9d1","com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_db"} + +08:10:40 drill ef42b4f34455: vikunja 2.5.99-notag (a tag that does not exist: backing-up runs, the pull fails, nothing moves) push rc=0 +08:11:03 [sync] the badge needed 22.8s and 4 sync+rescan rounds to catch up to vikunja/vikunja:2.5.99-notag — R-607's window, measured +08:11:03 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:11:03 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +08:11:05 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:11:06 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:11:07 + 4.1s phase=failed label=A frissítés nem sikerült err=Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább. hold=None +08:11:07 snapshots offered: [{"time": "2026-09-24T06:11:05Z", "short_id": "helyi", "tier": 1, "drive_label": "Bels\u0151 SSD (rendszer)"}] +08:11:08 drill 087b3edb7475: vikunja 2.5.0 (back to the install version) push rc=0 +08:11:13 [R] restoring vikunja from snapshot 'helyi' (of 1 offered) +08:11:13 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started'] +08:11:13 + 0.0s restore (True, None, None) +08:11:19 + 6.1s restore (False, None, None) +08:11:19 [R] after restore: state=running hold=None phase=failed +08:11:25 labels after the v0.267.0 restore: +vikunja_vikunja_data null +vikunja_vikunja_db null + +08:11:26 vikunja: readback of the seeded project http=200 ok=True +08:11:26 vikunja: register http=200 +08:11:26 vikunja: create project http=201 diff --git a/documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.json b/documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.json new file mode 100644 index 00000000..076a0ab6 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.json @@ -0,0 +1,205 @@ +{ + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.268.0", + "probe_port_real": 3456, + "labels_before": "vikunja_vikunja_data null\nvikunja_vikunja_db null\n", + "drill_e": "17a5e25f99bc", + "badge_e": 4.5, + "press_e": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 1.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "copying", + "label": "Az adatok másolása a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 4.1, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 94.3, + "phase": "undoing", + "label": "Visszaállítás az előző változatra…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 102.5, + "phase": "undone", + "label": "Visszaállítva az előző változatra", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 102.5, + "final_phase": "undone", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "log_e": "2026/09/24 06:17:27 update.go:722: [INFO] [stacks] update vikunja: ladder — the installed version vikunja=vikunja/vikunja:2.5.0 matches no update_ladder entry (1 entries) — an app older than the ladder has no record to climb; the catalog's current definition\n2026/09/24 06:17:27 undo.go:387: [WARN] [stacks] update vikunja: volume(s) [vikunja_vikunja_data vikunja_vikunja_db] carry no compose label (recreated by a restore before v0.268.0 — R-658) — copied by name\n2026/09/24 06:17:28 undo.go:273: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB\n2026/09/24 06:17:30 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260924T061729Z in 431ms\n2026/09/24 06:17:30 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260924T061729Z in 428ms\n2026/09/24 06:19:01 undo.go:501: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy))\n2026/09/24 06:19:09 undo.go:549: [INFO] [stacks] update vikunja: UNDONE in 7s — the previous version is running on the data from before the update (the app's health check passed)\n", + "read_e": { + "A": true, + "B": true + }, + "obs_e": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.5.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.5.0" + }, + "catalog_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.5.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.5.0 running=true restarts=0" + ] + }, + "restore_f": { + "ok": true, + "snapshot_id": "helyi", + "snapshots": [ + { + "time": "2026-09-24T06:16:42Z", + "short_id": "helyi", + "tier": 1, + "drive_label": "Belső SSD (rendszer)" + } + ], + "http": "HTTP/2 302", + "location": [ + "location: /backups/restore?flash=flash.restore.started" + ], + "seconds": 97.2, + "state_after": "unhealthy", + "hold_after": null, + "observables_after": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.5.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.5.0" + }, + "catalog_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.5.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.5.0 running=true restarts=0" + ] + } + }, + "labels_after_restore_v0268": "vikunja_vikunja_data {\"com.docker.compose.project\":\"vikunja\",\"com.docker.compose.version\":\"5.5.0\",\"com.docker.compose.volume\":\"vikunja_data\"}\nvikunja_vikunja_db {\"com.docker.compose.project\":\"vikunja\",\"com.docker.compose.version\":\"5.5.0\",\"com.docker.compose.volume\":\"vikunja_db\"}\n", + "log_f": "2026/09/24 06:19:17 restore.go:152: [INFO] [backup] Restoring Docker volume vikunja_vikunja_data for vikunja\n2026/09/24 06:19:18 restore.go:152: [INFO] [backup] Restoring Docker volume vikunja_vikunja_db for vikunja\n", + "compose_warn_f": "", + "read_f": { + "A": true + }, + "press_g": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "backing-up", + "label": "Biztonsági mentés készül a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 2.1, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 4.1, + "phase": "copying", + "label": "Az adatok másolása a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 5.2, + "phase": "starting", + "label": "Indítás az új verzióval…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 6.2, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 11.3, + "phase": "done", + "label": "Frissítve", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 11.3, + "final_phase": "done", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "log_g": "2026/09/24 06:21:07 update.go:722: [INFO] [stacks] update vikunja: ladder — the installed version vikunja=vikunja/vikunja:2.5.0 matches no update_ladder entry (1 entries) — an app older than the ladder has no record to climb; the catalog's current definition\n2026/09/24 06:21:09 tier2.go:425: [INFO] [backup] Tier 2 copied vikunja → /mnt/felhom-drives/scratch_hdd/backups/secondary/vikunja (2.6 MB, 0 leg(s), 0s)\n2026/09/24 06:21:10 undo.go:273: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB\n2026/09/24 06:21:12 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260924T062112Z in 481ms\n2026/09/24 06:21:13 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260924T062112Z in 462ms\n", + "read_g": { + "A": true, + "C": true + }, + "drill_h": "418e8ed2aeb2", + "remove_h": "200", + "leftovers_h": "total 24\ndrwxr-xr-x 3 root root 4096 Sep 24 06:21 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4419 Sep 24 06:19 .felhom.yml\n-rw-r--r-- 1 root root 1280 Sep 24 06:21 docker-compose.yml\ndrwxr-xr-x 12 root root 4096 Sep 24 06:19 hold-logs\nno vikunja volumes\n" +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.log b/documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.log new file mode 100644 index 00000000..d5302058 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/02-v0268-undo-after-restore.log @@ -0,0 +1,69 @@ +08:17:21 labels before (unlabelled after the v0.267.0 restore): +vikunja_vikunja_data null +vikunja_vikunja_db null + +08:17:22 drill 17a5e25f99bc: vikunja 2.6.0 + probe 8999 (the failing edge) push rc=0 +08:17:27 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:17:27 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:17:28 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:17:30 + 3.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:17:31 + 4.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:19:01 + 94.3s phase=undoing label=Visszaállítás az előző változatra… err=None hold=None +08:19:10 + 102.5s phase=undone label=Visszaállítva az előző változatra err=None hold=None +08:19:13 controller log (e): +2026/09/24 06:17:27 update.go:722: [INFO] [stacks] update vikunja: ladder — the installed version vikunja=vikunja/vikunja:2.5.0 matches no update_ladder entry (1 entries) — an app older than the ladder has no record to climb; the catalog's current definition +2026/09/24 06:17:27 undo.go:387: [WARN] [stacks] update vikunja: volume(s) [vikunja_vikunja_data vikunja_vikunja_db] carry no compose label (recreated by a restore before v0.268.0 — R-658) — copied by name +2026/09/24 06:17:28 undo.go:273: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB +2026/09/24 06:17:30 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260924T061729Z in 431ms +2026/09/24 06:17:30 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260924T061729Z in 428ms +2026/09/24 06:19:01 undo.go:501: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/24 06:19:09 undo.go:549: [INFO] [stacks] update vikunja: UNDONE in 7s — the previous version is running on the data from before the update (the app's health check passed) + +08:19:13 vikunja: readback of the seeded project http=200 ok=True +08:19:14 vikunja: readback of the seeded project http=200 ok=True +08:19:14 READBACK after the undo (e): {'A': True, 'B': True} +08:19:17 [R] restoring vikunja from snapshot 'helyi' (of 1 offered) +08:19:17 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started'] +08:19:17 + 0.0s restore (True, None, None) +08:20:54 + 97.2s restore (False, None, None) +08:20:54 [R] after restore: state=unhealthy hold=None phase=undone +08:21:00 labels after the v0.268.0 restore: +vikunja_vikunja_data {"com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_data"} +vikunja_vikunja_db {"com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_db"} + +08:21:06 compose warnings after the restore (want none): (none) +08:21:07 vikunja: readback of the seeded project http=200 ok=True +08:21:07 READBACK after the v0.268.0 restore (f): {'A': True} +08:21:07 vikunja: register http=200 +08:21:07 vikunja: create project http=201 +08:21:07 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:21:08 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +08:21:10 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:21:11 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:21:12 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:21:13 + 5.2s phase=starting label=Indítás az új verzióval… err=None hold=None +08:21:14 + 6.2s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:21:19 + 11.3s phase=done label=Frissítve err=None hold=None +08:21:22 controller log (g): +2026/09/24 06:21:07 update.go:722: [INFO] [stacks] update vikunja: ladder — the installed version vikunja=vikunja/vikunja:2.5.0 matches no update_ladder entry (1 entries) — an app older than the ladder has no record to climb; the catalog's current definition +2026/09/24 06:21:09 tier2.go:425: [INFO] [backup] Tier 2 copied vikunja → /mnt/felhom-drives/scratch_hdd/backups/secondary/vikunja (2.6 MB, 0 leg(s), 0s) +2026/09/24 06:21:10 undo.go:273: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB +2026/09/24 06:21:12 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260924T062112Z in 481ms +2026/09/24 06:21:13 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260924T062112Z in 462ms + +08:21:23 vikunja: readback of the seeded project http=200 ok=True +08:21:23 vikunja: readback of the seeded project http=200 ok=True +08:21:23 READBACK after the second undo (g): {'A': True, 'C': True} +08:21:24 drill 418e8ed2aeb2: vikunja back to 2.5.0 + its real probe push rc=0 +08:21:24 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +08:21:56 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +08:22:04 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +08:22:07 after remove: +total 24 +drwxr-xr-x 3 root root 4096 Sep 24 06:21 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 4419 Sep 24 06:19 .felhom.yml +-rw-r--r-- 1 root root 1280 Sep 24 06:21 docker-compose.yml +drwxr-xr-x 12 root root 4096 Sep 24 06:19 hold-logs +no vikunja volumes + diff --git a/documentation/audits/ladder-2026-09-24/partA/02.stdout b/documentation/audits/ladder-2026-09-24/partA/02.stdout new file mode 100644 index 00000000..d5302058 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/02.stdout @@ -0,0 +1,69 @@ +08:17:21 labels before (unlabelled after the v0.267.0 restore): +vikunja_vikunja_data null +vikunja_vikunja_db null + +08:17:22 drill 17a5e25f99bc: vikunja 2.6.0 + probe 8999 (the failing edge) push rc=0 +08:17:27 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:17:27 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:17:28 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:17:30 + 3.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:17:31 + 4.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:19:01 + 94.3s phase=undoing label=Visszaállítás az előző változatra… err=None hold=None +08:19:10 + 102.5s phase=undone label=Visszaállítva az előző változatra err=None hold=None +08:19:13 controller log (e): +2026/09/24 06:17:27 update.go:722: [INFO] [stacks] update vikunja: ladder — the installed version vikunja=vikunja/vikunja:2.5.0 matches no update_ladder entry (1 entries) — an app older than the ladder has no record to climb; the catalog's current definition +2026/09/24 06:17:27 undo.go:387: [WARN] [stacks] update vikunja: volume(s) [vikunja_vikunja_data vikunja_vikunja_db] carry no compose label (recreated by a restore before v0.268.0 — R-658) — copied by name +2026/09/24 06:17:28 undo.go:273: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB +2026/09/24 06:17:30 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260924T061729Z in 431ms +2026/09/24 06:17:30 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260924T061729Z in 428ms +2026/09/24 06:19:01 undo.go:501: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/24 06:19:09 undo.go:549: [INFO] [stacks] update vikunja: UNDONE in 7s — the previous version is running on the data from before the update (the app's health check passed) + +08:19:13 vikunja: readback of the seeded project http=200 ok=True +08:19:14 vikunja: readback of the seeded project http=200 ok=True +08:19:14 READBACK after the undo (e): {'A': True, 'B': True} +08:19:17 [R] restoring vikunja from snapshot 'helyi' (of 1 offered) +08:19:17 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started'] +08:19:17 + 0.0s restore (True, None, None) +08:20:54 + 97.2s restore (False, None, None) +08:20:54 [R] after restore: state=unhealthy hold=None phase=undone +08:21:00 labels after the v0.268.0 restore: +vikunja_vikunja_data {"com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_data"} +vikunja_vikunja_db {"com.docker.compose.project":"vikunja","com.docker.compose.version":"5.5.0","com.docker.compose.volume":"vikunja_db"} + +08:21:06 compose warnings after the restore (want none): (none) +08:21:07 vikunja: readback of the seeded project http=200 ok=True +08:21:07 READBACK after the v0.268.0 restore (f): {'A': True} +08:21:07 vikunja: register http=200 +08:21:07 vikunja: create project http=201 +08:21:07 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:21:08 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +08:21:10 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:21:11 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:21:12 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:21:13 + 5.2s phase=starting label=Indítás az új verzióval… err=None hold=None +08:21:14 + 6.2s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:21:19 + 11.3s phase=done label=Frissítve err=None hold=None +08:21:22 controller log (g): +2026/09/24 06:21:07 update.go:722: [INFO] [stacks] update vikunja: ladder — the installed version vikunja=vikunja/vikunja:2.5.0 matches no update_ladder entry (1 entries) — an app older than the ladder has no record to climb; the catalog's current definition +2026/09/24 06:21:09 tier2.go:425: [INFO] [backup] Tier 2 copied vikunja → /mnt/felhom-drives/scratch_hdd/backups/secondary/vikunja (2.6 MB, 0 leg(s), 0s) +2026/09/24 06:21:10 undo.go:273: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB +2026/09/24 06:21:12 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260924T062112Z in 481ms +2026/09/24 06:21:13 undo.go:423: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260924T062112Z in 462ms + +08:21:23 vikunja: readback of the seeded project http=200 ok=True +08:21:23 vikunja: readback of the seeded project http=200 ok=True +08:21:23 READBACK after the second undo (g): {'A': True, 'C': True} +08:21:24 drill 418e8ed2aeb2: vikunja back to 2.5.0 + its real probe push rc=0 +08:21:24 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +08:21:56 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +08:22:04 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +08:22:07 after remove: +total 24 +drwxr-xr-x 3 root root 4096 Sep 24 06:21 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 4419 Sep 24 06:19 .felhom.yml +-rw-r--r-- 1 root root 1280 Sep 24 06:21 docker-compose.yml +drwxr-xr-x 12 root root 4096 Sep 24 06:19 hold-logs +no vikunja volumes + diff --git a/documentation/audits/ladder-2026-09-24/partA/03-why-g-was-done.txt b/documentation/audits/ladder-2026-09-24/partA/03-why-g-was-done.txt new file mode 100644 index 00000000..f7e55866 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/03-why-g-was-done.txt @@ -0,0 +1,31 @@ +2026/09/24 06:19:17 handlers.go:1701: [WARN] [web] Restore requested (async): stack=vikunja, snapshot=helyi from 172.18.0.9:57016 +2026/09/24 06:19:17 restore_unit.go:379: [INFO] [backup] Restoring vikunja from recovery unit /mnt/sys_drive/felhom-data/backups/primary/vikunja: images=1, secrets recovered=1/1, data_keys=0 +2026/09/24 06:19:17 manager.go:1265: [INFO] [stacks] Stopping stack: vikunja +2026/09/24 06:19:17 restore.go:152: [INFO] [backup] Restoring Docker volume vikunja_vikunja_data for vikunja +2026/09/24 06:19:18 restore.go:152: [INFO] [backup] Restoring Docker volume vikunja_vikunja_db for vikunja +2026/09/24 06:19:18 restore.go:195: [INFO] [backup] Restored 2 Docker volume(s) for vikunja +2026/09/24 06:19:18 deploy.go:645: [INFO] [stacks] Redeploying vikunja from recovery unit with 3 env vars +2026/09/24 06:19:18 pin.go:93: [INFO] [stacks] pin vikunja: vikunja=vikunja/vikunja:2.5.0 +2026/09/24 06:19:21 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:19:31 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:19:41 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:19:51 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:20:01 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:20:21 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:20:41 healthprobe.go:190: [WARN] Health probe vikunja: API GET :8999/api/v1/info → Get "http://vikunja:8999/api/v1/info": dial tcp 172.18.0.5:8999: connect: connection refused +2026/09/24 06:20:52 restore_unit.go:458: [WARN] [backup] vikunja restored but health check failed: stack vikunja did not reach running state within 1m30s after restore +2026/09/24 06:20:52 restore_unit.go:464: [INFO] [backup] Restore-from-unit completed: vikunja — 2 volume(s) of 2 listed, 0 database(s) of 0 listed +2026/09/24 06:20:52 handlers.go:1713: [INFO] [web] Restore completed (async): stack=vikunja in 1m35.21599499s (volumes 2/2, dbs 0/0) +2026/09/24 06:21:07 update.go:733: [INFO] [stacks] update vikunja: no usable copy on any tier — younger than 24h0m0s and not older than this install's deploy (2026-09-24T06:19:18Z) (found: Tier 2 (second drive) at 2026-09-24T06:11:05Z (10m0s old); Tier 1 (own recovery unit) at 2026-09-24T06:16:42Z (4m0s old)) — backing up first +2026/09/24 06:21:07 update_guard.go:360: [INFO] [backup] update pre-backup for vikunja: starting (DB dump → volume dump → unit capture → Tier 2) +2026/09/24 06:21:08 backup.go:918: [INFO] [backup] Stopping vikunja for safe volume dump +2026/09/24 06:21:08 manager.go:1265: [INFO] [stacks] Stopping stack: vikunja +2026/09/24 06:21:09 update_guard.go:416: [INFO] [backup] update pre-backup for vikunja: recovery unit captured (0 database dump(s)) +2026/09/24 06:21:09 update.go:763: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) [] +2026/09/24 06:21:10 update.go:1213: [INFO] [stacks] update vikunja: phase pinning +2026/09/24 06:21:10 pin.go:370: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0) +2026/09/24 06:21:13 healthprobe.go:190: [WARN] Health probe vikunja: API GET :3456/api/v1/info → Get "http://vikunja:3456/api/v1/info": dial tcp 172.18.0.5:3456: connect: connection refused +2026/09/24 06:21:18 update.go:884: [INFO] [stacks] update vikunja: healthy after 5s (the app's health check passed) +2026/09/24 06:21:18 update.go:892: [INFO] [stacks] update vikunja: DONE in 11s +2026/09/24 06:21:24 manager.go:1265: [INFO] [stacks] Stopping stack: vikunja + diff --git a/documentation/audits/ladder-2026-09-24/partA/state.json b/documentation/audits/ladder-2026-09-24/partA/state.json new file mode 100644 index 00000000..c7dad9e2 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partA/state.json @@ -0,0 +1,15 @@ +{ + "A": { + "user": "drill9ba9d9", + "pw": "", + "title": "drill-d60cb3a1c9", + "pid": 2 + }, + "B": { + "user": "drilld26feb", + "pw": "", + "title": "drill-4821c514ac", + "pid": 4 + }, + "sub": "vikunja-la" +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.json b/documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.json new file mode 100644 index 00000000..522f6ddd --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.json @@ -0,0 +1,75 @@ +{ + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.268.0", + "drill_1": "6a41e807228f", + "deployed": true, + "seed_A": true, + "drill_2": "a6d49ca9eabc", + "badge_wait": 4.4, + "steps_left_line_hu": [ + [ + "1", + "Hátralévő frissítési lépések: 1" + ] + ], + "steps_left_line_en": [ + [ + "1", + "Update steps remaining: 1" + ] + ], + "phases": [ + [ + 0.0, + "backing-up" + ], + [ + 18.9, + "safety-dump" + ], + [ + 21.0, + "pulling" + ], + [ + 22.6, + "copying" + ], + [ + 36.8, + "starting" + ], + [ + 43.1, + "verifying" + ], + [ + 133.4, + "undoing" + ], + [ + 136.0, + "failed" + ] + ], + "cutoff": "copy: nextcloud_nextcloud_db_data.pre-update-20260924T062548Z\ndata\n", + "api": { + "state": "stopped", + "update_phase": "failed", + "hold_reason": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást.", + "hold_no_whole_copy": true, + "update_error": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást." + }, + "app_page_hu": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást.", + "app_page_hu_restore_button": false, + "stacks_page_hu_restore_button": false, + "app_page_en": "The update of nextcloud did not work, and neither did the automatic undo. This box has no copy that can bring the app back together with its files. Felhom support has been told — until then, do not restart or remove the app.", + "app_page_en_restore_button": false, + "stacks_page_en_restore_button": false, + "ascii_checks": { + "hu has 'gyfelszolgalat' (positive)": true, + "hu has 'Visszaallithato a Mentesek' (must be absent)": false, + "en has 'Felhom support has been told' (positive)": true + }, + "log": "2026/09/24 06:25:42 notifier.go:1259: [DEBUG] notifier disabled: dropped event crossdrive_completed (severity info)\n2026/09/24 06:26:42 notifier.go:1259: [DEBUG] notifier disabled: dropped event health_change (severity warning)\n2026/09/24 06:27:37 undo.go:501: [WARN] [stacks] update nextcloud: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy))\n2026/09/24 06:27:40 update.go:937: [ERROR] [stacks] update nextcloud: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{nextcloud_nextcloud_db_data nextcloud_nextcloud_db_data.pre-update-20260924T062548Z} {nextcloud_nextcloud_html nextcloud_nextcloud_html.pre-update-20260924T062548Z} {nextcloud_nextcloud_redis_data nextcloud_nextcloud_redis_data.pre-update-20260924T062548Z}]\n2026/09/24 06:27:40 update_guard.go:705: [ERROR] [backup] nextcloud is HELD STOPPED after a failed update and NO copy on this box brings it back whole (seen: [tier 2 at 2026-09-24T06:25:42Z tier 1 at 2026-09-24T06:25:43Z]; drive files declared: true; undo: \"untouched\") — support must act (R-659)\n2026/09/24 06:27:40 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only\n2026/09/24 06:27:40 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_hold_no_whole_copy (severity critical) — further app_hold_no_whole_copy events are logged at DEBUG only\n", + "debug_ring_events": [] +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.log b/documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.log new file mode 100644 index 00000000..f64cc7b1 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/01-round11-reproduced.log @@ -0,0 +1,37 @@ +08:23:32 drill 6a41e807228f: nextcloud template = ladder step 1's own definition (34.0.1 / mariadb 12.3) push rc=0 +08:23:40 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud'] +08:23:40 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH'] +08:23:41 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:25:06 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.1-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'} +08:25:15 nextcloud: occ user:add :: The account "drill10a42a" was created successfully Display name set to "drill10a42a" +08:25:18 nextcloud: seeded user drill10a42a +08:25:19 drill a6d49ca9eabc: nextcloud back to the head (34.0.4) + probe port 8999 push rc=0 +08:25:24 steps-left line hu=[('1', 'Hátralévő frissítési lépések: 1')] en=[('1', 'Update steps remaining: 1')] +08:25:24 Update -> 202 +08:25:24 + 0.0s phase=backing-up +08:25:43 + 18.9s phase=safety-dump +08:25:45 + 21.0s phase=pulling +08:25:47 + 22.6s phase=copying +08:26:01 + 36.8s phase=starting +08:26:07 + 43.1s phase=verifying +08:26:13 >>> finished-marker removed from one undo copy: +copy: nextcloud_nextcloud_db_data.pre-update-20260924T062548Z +data + +08:27:37 + 133.4s phase=undoing +08:27:40 + 136.0s phase=failed +08:27:40 API: {"state": "stopped", "update_phase": "failed", "hold_reason": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást.", "hold_no_whole_copy": true, "update_error": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást."} +08:27:40 [hu] app page held block: 'A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást.' restore-button app=False list=False +08:27:40 [en] app page held block: 'The update of nextcloud did not work, and neither did the automatic undo. This box has no copy that can bring the app back together with its files. Felhom support has been told — until then, do not restart or remove the app.' restore-button app=False list=False +08:27:40 checks: {"hu has 'gyfelszolgalat' (positive)": true, "hu has 'Visszaallithato a Mentesek' (must be absent)": false, "en has 'Felhom support has been told' (positive)": true} +08:30:13 controller log: +2026/09/24 06:25:42 notifier.go:1259: [DEBUG] notifier disabled: dropped event crossdrive_completed (severity info) +2026/09/24 06:26:42 notifier.go:1259: [DEBUG] notifier disabled: dropped event health_change (severity warning) +2026/09/24 06:27:37 undo.go:501: [WARN] [stacks] update nextcloud: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/24 06:27:40 update.go:937: [ERROR] [stacks] update nextcloud: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{nextcloud_nextcloud_db_data nextcloud_nextcloud_db_data.pre-update-20260924T062548Z} {nextcloud_nextcloud_html nextcloud_nextcloud_html.pre-update-20260924T062548Z} {nextcloud_nextcloud_redis_data nextcloud_nextcloud_redis_data.pre-update-20260924T062548Z}] +2026/09/24 06:27:40 update_guard.go:705: [ERROR] [backup] nextcloud is HELD STOPPED after a failed update and NO copy on this box brings it back whole (seen: [tier 2 at 2026-09-24T06:25:42Z tier 1 at 2026-09-24T06:25:43Z]; drive files declared: true; undo: "untouched") — support must act (R-659) +2026/09/24 06:27:40 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only +2026/09/24 06:27:40 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_hold_no_whole_copy (severity critical) — further app_hold_no_whole_copy events are logged at DEBUG only + +08:30:14 debug ring (events since the press): + diff --git a/documentation/audits/ladder-2026-09-24/partB/01.stdout b/documentation/audits/ladder-2026-09-24/partB/01.stdout new file mode 100644 index 00000000..f64cc7b1 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/01.stdout @@ -0,0 +1,37 @@ +08:23:32 drill 6a41e807228f: nextcloud template = ladder step 1's own definition (34.0.1 / mariadb 12.3) push rc=0 +08:23:40 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud'] +08:23:40 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH'] +08:23:41 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:25:06 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.1-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'} +08:25:15 nextcloud: occ user:add :: The account "drill10a42a" was created successfully Display name set to "drill10a42a" +08:25:18 nextcloud: seeded user drill10a42a +08:25:19 drill a6d49ca9eabc: nextcloud back to the head (34.0.4) + probe port 8999 push rc=0 +08:25:24 steps-left line hu=[('1', 'Hátralévő frissítési lépések: 1')] en=[('1', 'Update steps remaining: 1')] +08:25:24 Update -> 202 +08:25:24 + 0.0s phase=backing-up +08:25:43 + 18.9s phase=safety-dump +08:25:45 + 21.0s phase=pulling +08:25:47 + 22.6s phase=copying +08:26:01 + 36.8s phase=starting +08:26:07 + 43.1s phase=verifying +08:26:13 >>> finished-marker removed from one undo copy: +copy: nextcloud_nextcloud_db_data.pre-update-20260924T062548Z +data + +08:27:37 + 133.4s phase=undoing +08:27:40 + 136.0s phase=failed +08:27:40 API: {"state": "stopped", "update_phase": "failed", "hold_reason": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást.", "hold_no_whole_copy": true, "update_error": "A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást."} +08:27:40 [hu] app page held block: 'A(z) nextcloud frissítése nem sikerült, és az automatikus visszaállítás sem. Ezen a dobozon nincs olyan másolat, amely az alkalmazást a fájljaival együtt vissza tudná hozni. A Felhom ügyfélszolgálatát értesítettük — kérjük, addig ne indítsa újra és ne törölje az alkalmazást.' restore-button app=False list=False +08:27:40 [en] app page held block: 'The update of nextcloud did not work, and neither did the automatic undo. This box has no copy that can bring the app back together with its files. Felhom support has been told — until then, do not restart or remove the app.' restore-button app=False list=False +08:27:40 checks: {"hu has 'gyfelszolgalat' (positive)": true, "hu has 'Visszaallithato a Mentesek' (must be absent)": false, "en has 'Felhom support has been told' (positive)": true} +08:30:13 controller log: +2026/09/24 06:25:42 notifier.go:1259: [DEBUG] notifier disabled: dropped event crossdrive_completed (severity info) +2026/09/24 06:26:42 notifier.go:1259: [DEBUG] notifier disabled: dropped event health_change (severity warning) +2026/09/24 06:27:37 undo.go:501: [WARN] [stacks] update nextcloud: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/24 06:27:40 update.go:937: [ERROR] [stacks] update nextcloud: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{nextcloud_nextcloud_db_data nextcloud_nextcloud_db_data.pre-update-20260924T062548Z} {nextcloud_nextcloud_html nextcloud_nextcloud_html.pre-update-20260924T062548Z} {nextcloud_nextcloud_redis_data nextcloud_nextcloud_redis_data.pre-update-20260924T062548Z}] +2026/09/24 06:27:40 update_guard.go:705: [ERROR] [backup] nextcloud is HELD STOPPED after a failed update and NO copy on this box brings it back whole (seen: [tier 2 at 2026-09-24T06:25:42Z tier 1 at 2026-09-24T06:25:43Z]; drive files declared: true; undo: "untouched") — support must act (R-659) +2026/09/24 06:27:40 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only +2026/09/24 06:27:40 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_hold_no_whole_copy (severity critical) — further app_hold_no_whole_copy events are logged at DEBUG only + +08:30:14 debug ring (events since the press): + diff --git a/documentation/audits/ladder-2026-09-24/partB/02-r660-positive-observables.txt b/documentation/audits/ladder-2026-09-24/partB/02-r660-positive-observables.txt new file mode 100644 index 00000000..7353e8f7 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/02-r660-positive-observables.txt @@ -0,0 +1,27 @@ +2026/09/24 06:23:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:24:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:24:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:25:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:25:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:26:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:26:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:27:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:27:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:27:42 main.go:1951: [INFO] [deadapp] check alive: 20 scans since boot, 4 deployed app(s) evaluated, 0 currently down +2026/09/24 06:28:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:28:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:29:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:29:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:30:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting + +200 ['entries', 'total'] +200 {'timestamp': '2026-09-24T06:27:41Z', 'level': 'DEBUG', 'message': '[stacks] RunHealthProbes: collected 0 targets (2 skipped not due, 0 skipped no container)', 'source': 'healthprobe.go:104'} +{"timestamp": "2026-09-24T06:27:42Z", "level": "INFO", "message": "[deadapp] check alive: 20 scans since boot, 4 deployed app(s) evaluated, 0 currently down", "source": "main.go:1951"} +{"timestamp": "2026-09-24T06:28:11Z", "level": "DEBUG", "message": "[scheduler] job deadapp-check: execution starting", "source": "scheduler.go:67"} +{"timestamp": "2026-09-24T06:28:41Z", "level": "DEBUG", "message": "[scheduler] job deadapp-check: execution starting", "source": "scheduler.go:67"} +{"timestamp": "2026-09-24T06:29:11Z", "level": "DEBUG", "message": "[scheduler] job deadapp-check: execution starting", "source": "scheduler.go:67"} +{"timestamp": "2026-09-24T06:29:41Z", "level": "DEBUG", "message": "[scheduler] job deadapp-check: execution starting", "source": "scheduler.go:67"} +{"timestamp": "2026-09-24T06:30:11Z", "level": "DEBUG", "message": "[scheduler] job deadapp-check: execution starting", "source": "scheduler.go:67"} +2026/09/24 06:27:42 main.go:1951: [INFO] [deadapp] check alive: 20 scans since boot, 4 deployed app(s) evaluated, 0 currently down +gokapi Restarting (1) 28 seconds ago + diff --git a/documentation/audits/ladder-2026-09-24/partB/03-r660-control.json b/documentation/audits/ladder-2026-09-24/partB/03-r660-control.json new file mode 100644 index 00000000..d1584d2c --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/03-r660-control.json @@ -0,0 +1,8 @@ +{ + "events": [ + "2026-09-24T06:31:41Z notifier disabled: dropped event app_start_failed (severity warning)", + "2026-09-24T06:31:41Z notifier disabled: dropped event app_start_failed (severity warning)" + ], + "log": "2026/09/24 06:31:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting\n2026/09/24 06:31:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting\n2026/09/24 06:31:41 notifier.go:1259: [DEBUG] notifier disabled: dropped event app_start_failed (severity warning)\n2026/09/24 06:31:41 notifier.go:1259: [DEBUG] notifier disabled: dropped event app_start_failed (severity warning)\n", + "nextcloud_state": "stopped" +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partB/03-r660-control.log b/documentation/audits/ladder-2026-09-24/partB/03-r660-control.log new file mode 100644 index 00000000..c4654494 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/03-r660-control.log @@ -0,0 +1,16 @@ +08:31:02 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:31:12 [1] deployed, controller state=running, pinned={'actualbudget': 'actualbudget/actual-server:26.9.0'} +08:31:16 stopping the throwaway's container out of band (not through the product): actualbudget +08:31:46 events/heartbeats since the control started: + 2026-09-24T06:31:41Z notifier disabled: dropped event app_start_failed (severity warning) + 2026-09-24T06:31:41Z notifier disabled: dropped event app_start_failed (severity warning) +08:31:49 controller log lines: +2026/09/24 06:31:11 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:31:41 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/24 06:31:41 notifier.go:1259: [DEBUG] notifier disabled: dropped event app_start_failed (severity warning) +2026/09/24 06:31:41 notifier.go:1259: [DEBUG] notifier disabled: dropped event app_start_failed (severity warning) + +08:31:49 nextcloud meanwhile: state=stopped held=True +08:31:49 [X] stop -> 200 {'ok': True, 'message': 'Stack actualbudget stop completed'} +08:32:21 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'actualbudget', 'volumes_removed': ['actualbudget_actualbudget_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd +08:32:30 [X] after remove: deployed=False leftovers='/opt/docker/stacks/actualbudget' diff --git a/documentation/audits/ladder-2026-09-24/partB/04-states-and-banner.txt b/documentation/audits/ladder-2026-09-24/partB/04-states-and-banner.txt new file mode 100644 index 00000000..2ebc2943 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/04-states-and-banner.txt @@ -0,0 +1,5 @@ +gokapi restarting running +nextcloud stopped held running +paperless-ngx running running +privatebin running stopped +[] diff --git a/documentation/audits/ladder-2026-09-24/partB/05-quiet-watch.txt b/documentation/audits/ladder-2026-09-24/partB/05-quiet-watch.txt new file mode 100644 index 00000000..13037c45 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/05-quiet-watch.txt @@ -0,0 +1,2 @@ +2026-09-24T06:34:42Z [stacks] ScanStacks: found stack "gokapi" deployed=true composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026-09-24T06:34:42Z [stacks] ScanStacks: found stack "nextcloud" deployed=true composePath=/opt/docker/stacks/nextcloud/docker-compose.yml diff --git a/documentation/audits/ladder-2026-09-24/partB/06-r660-control-banner.json b/documentation/audits/ladder-2026-09-24/partB/06-r660-control-banner.json new file mode 100644 index 00000000..0c8f9c18 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/06-r660-control-banner.json @@ -0,0 +1,128 @@ +[ + { + "t": 10, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 15, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 20, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 25, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 30, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 35, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 40, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 45, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 50, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 55, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 60, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 65, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 70, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 75, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 80, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 85, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 90, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + }, + { + "t": 95, + "banner": [], + "events": [ + "2026-09-24T06:37:11Z notifier disabled: dropped event app_start_failed (severity warning)" + ] + } +] \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partB/07-teardown.log b/documentation/audits/ladder-2026-09-24/partB/07-teardown.log new file mode 100644 index 00000000..7202852d --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partB/07-teardown.log @@ -0,0 +1,30 @@ +08:39:36 before remove: +nextcloud_nextcloud_db_data +nextcloud_nextcloud_db_data.pre-update-20260924T062548Z +nextcloud_nextcloud_html +nextcloud_nextcloud_html.pre-update-20260924T062548Z +nextcloud_nextcloud_redis_data +nextcloud_nextcloud_redis_data.pre-update-20260924T062548Z +app.yaml +applied-compose.yml +applied-meta +docker-compose.yml +hold-logs + +08:39:36 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'} +08:39:41 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó +08:39:41 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +08:40:10 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], +08:40:18 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud' +08:40:21 after remove: +no nextcloud volumes (incl. undo copies) +total 32 +drwxr-xr-x 3 root root 4096 Sep 24 06:40 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 9182 Sep 24 06:25 .felhom.yml +-rw-r--r-- 1 root root 4593 Sep 24 06:25 docker-compose.yml +drwxr-xr-x 4 root root 4096 Sep 24 06:27 hold-logs +ls: cannot access '/mnt/felhom-drives/scratch_hdd/appdata/nextcloud': No such file or directory +/mnt/felhom-drives/scratch_hdd/userdata/nextcloud + +08:40:24 drive folders removed by name (R-442 keeps them on 9202): ls: cannot access '/mnt/felhom-drives/scratch_hdd/userdata/nextcloud': No such file or directory diff --git a/documentation/audits/ladder-2026-09-24/partD/00-repoint-drill.txt b/documentation/audits/ladder-2026-09-24/partD/00-repoint-drill.txt new file mode 100644 index 00000000..f2c89786 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/00-repoint-drill.txt @@ -0,0 +1,10 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + diff --git a/documentation/audits/ladder-2026-09-24/partD/00-spike.json b/documentation/audits/ladder-2026-09-24/partD/00-spike.json new file mode 100644 index 00000000..6b082bf7 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/00-spike.json @@ -0,0 +1,110 @@ +{ + "deployed_A": true, + "A": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.4.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.4.0" + }, + "catalog_images": null, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.4.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.4.0 running=true restarts=0" + ] + }, + "badge_wait_s": 10.5, + "cache_steps": "ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache/templates/vikunja/': No such file or directory\nls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache/templates/vikunja/steps/': No such file or directory\napp.yaml\napplied-compose.yml\napplied-meta\ndocker-compose.yml\nhold-logs\n", + "badges": { + "hu": [ + { + "title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", + "text": "Frissítés elérhető — 2 napja" + } + ], + "en": [ + { + "title": "A newer version of this app is available. Select the Update button to start it.", + "text": "Update available — 2 days ago" + } + ] + }, + "press": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "backing-up", + "label": "Biztonsági mentés készül a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 2.1, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 4.1, + "phase": "copying", + "label": "Az adatok másolása a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 12.3, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 32.8, + "phase": "done", + "label": "Frissítve", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 32.9, + "final_phase": "done", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "after": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "catalog_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.6.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.6.0 running=true restarts=0" + ] + } +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partD/00-spike.log b/documentation/audits/ladder-2026-09-24/partD/00-spike.log new file mode 100644 index 00000000..6ab7fc84 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/00-spike.log @@ -0,0 +1,27 @@ +07:44:29 drill commit 15a4107ef0c0: SPIKE vikunja A=2.4.0 (push rc=0) +07:45:48 [sync] the badge NEVER caught up to vikunja/vikunja:2.4.0 in 78.9s — catalog_images = None +07:45:48 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +07:45:58 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.4.0'} +07:46:03 drill commit 4d557385d146: SPIKE vikunja B=2.5.0 + steps/spike-B.yml (push rc=0) +07:46:03 drill commit 1cb96ca718dd: SPIKE vikunja C=2.6.0 (push rc=0) +07:46:14 [sync] the badge needed 10.5s and 2 sync+rescan rounds to catch up to vikunja/vikunja:2.6.0 — R-607's window, measured +07:46:18 cache: +ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache/templates/vikunja/': No such file or directory +ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache/templates/vikunja/steps/': No such file or directory +app.yaml +applied-compose.yml +applied-meta +docker-compose.yml +hold-logs + +07:46:18 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +07:46:18 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +07:46:20 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +07:46:21 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +07:46:22 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +07:46:30 + 12.3s phase=verifying label=Működés ellenőrzése… err=None hold=None +07:46:51 + 32.8s phase=done label=Frissítve err=None hold=None +07:46:54 after: {"pinned_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.6.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.6.0 running=true restarts=0"]} +07:46:55 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +07:47:26 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +07:47:35 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' diff --git a/documentation/audits/ladder-2026-09-24/partD/00-spike.stdout b/documentation/audits/ladder-2026-09-24/partD/00-spike.stdout new file mode 100644 index 00000000..6ab7fc84 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/00-spike.stdout @@ -0,0 +1,27 @@ +07:44:29 drill commit 15a4107ef0c0: SPIKE vikunja A=2.4.0 (push rc=0) +07:45:48 [sync] the badge NEVER caught up to vikunja/vikunja:2.4.0 in 78.9s — catalog_images = None +07:45:48 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +07:45:58 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.4.0'} +07:46:03 drill commit 4d557385d146: SPIKE vikunja B=2.5.0 + steps/spike-B.yml (push rc=0) +07:46:03 drill commit 1cb96ca718dd: SPIKE vikunja C=2.6.0 (push rc=0) +07:46:14 [sync] the badge needed 10.5s and 2 sync+rescan rounds to catch up to vikunja/vikunja:2.6.0 — R-607's window, measured +07:46:18 cache: +ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache/templates/vikunja/': No such file or directory +ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache/templates/vikunja/steps/': No such file or directory +app.yaml +applied-compose.yml +applied-meta +docker-compose.yml +hold-logs + +07:46:18 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +07:46:18 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +07:46:20 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +07:46:21 + 3.1s phase=pulling label=Új verzió letöltése… err=None hold=None +07:46:22 + 4.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +07:46:30 + 12.3s phase=verifying label=Működés ellenőrzése… err=None hold=None +07:46:51 + 32.8s phase=done label=Frissítve err=None hold=None +07:46:54 after: {"pinned_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.6.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.6.0 running=true restarts=0"]} +07:46:55 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +07:47:26 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +07:47:35 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' diff --git a/documentation/audits/ladder-2026-09-24/partD/01-spike-steps-in-clone.txt b/documentation/audits/ladder-2026-09-24/partD/01-spike-steps-in-clone.txt new file mode 100644 index 00000000..6d1b5438 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/01-spike-steps-in-clone.txt @@ -0,0 +1,25 @@ +1cb96ca718dd +1 +true +/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/vikunja/: +total 24 +drwxr-xr-x 3 root root 4096 Sep 24 05:46 . +drwxr-xr-x 55 root root 4096 Sep 24 05:44 .. +-rw-r--r-- 1 root root 4419 Sep 24 05:44 .felhom.yml +-rw-r--r-- 1 root root 1280 Sep 24 05:46 docker-compose.yml +drwxr-xr-x 2 root root 4096 Sep 24 05:46 steps + +/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/vikunja/steps/: +total 12 +drwxr-xr-x 2 root root 4096 Sep 24 05:46 . +drwxr-xr-x 3 root root 4096 Sep 24 05:46 .. +-rw-r--r-- 1 root root 1280 Sep 24 05:46 spike-B.yml +total 32 +drwxr-xr-x 4 root root 4096 Sep 24 05:47 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 4419 Sep 23 22:15 .felhom.yml +-rw-r--r-- 1 root root 1280 Sep 24 05:46 applied-compose.yml +drwxr-xr-x 2 root root 4096 Sep 24 05:46 applied-meta +-rw-r--r-- 1 root root 1280 Sep 24 05:46 docker-compose.yml +drwxr-xr-x 11 root root 4096 Sep 23 21:05 hold-logs + diff --git a/documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.json b/documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.json new file mode 100644 index 00000000..c1873bcb --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.json @@ -0,0 +1,295 @@ +{ + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.268.0", + "drill_install": "95d83b7ab698", + "deployed": true, + "seed_A": true, + "drill_head": "0e4e12aed1f2", + "badge_wait": 4.4, + "before": { + "obs": { + "pinned_images": { + "romm": "rommapp/romm:5.3.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "installed_images": { + "romm": "rommapp/romm:5.3.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "catalog_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.8", + "romm-redis": "redis:7-alpine" + }, + "live_compose_image_lines": [ + "image: rommapp/romm:5.3.0", + "image: mariadb:11.4", + "image: redis:7-alpine" + ], + "docker_inspect": [ + "romm rommapp/romm:5.3.0 running=true restarts=0", + "romm-redis redis:7-alpine running=true restarts=0", + "romm-db mariadb:11.4 running=true restarts=0" + ] + }, + "steps_line": { + "hu": [ + [ + "2", + "Hátralévő frissítési lépések: 2" + ] + ], + "en": [ + [ + "2", + "Update steps remaining: 2" + ] + ] + }, + "badges": { + "hu": [ + { + "title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", + "text": "Frissítés elérhető — 1 napja" + } + ], + "en": [ + { + "title": "A newer version of this app is available. Select the Update button to start it.", + "text": "Update available — 1 day ago" + } + ] + } + }, + "press1": { + "press": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 2.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "copying", + "label": "Az adatok másolása a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 10.3, + "phase": "starting", + "label": "Indítás az új verzióval…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 21.6, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 57.4, + "phase": "done", + "label": "Frissítve", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 57.5, + "final_phase": "done", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "obs": { + "pinned_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "installed_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "catalog_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.8", + "romm-redis": "redis:7-alpine" + }, + "live_compose_image_lines": [ + "image: rommapp/romm:5.3.1", + "image: mariadb:11.4", + "image: redis:7-alpine" + ], + "docker_inspect": [ + "romm rommapp/romm:5.3.1 running=true restarts=0", + "romm-redis redis:7-alpine running=true restarts=0", + "romm-db mariadb:11.4 running=true restarts=0" + ] + }, + "readback_A": true, + "log": "2026/09/24 06:42:31 update.go:722: [INFO] [stacks] update romm: ladder — step 2 of 3: romm=rommapp/romm:5.3.0, romm-db=mariadb:11.4, romm-redis=redis:7-alpine → romm=rommapp/romm:5.3.1, romm-db=mariadb:11.4, romm-redis=redis:7-alpine, from 90dd9d68258286ef.yml\n2026/09/24 06:42:33 pin.go:370: [INFO] [stacks] update romm: pin advanced to the catalog's current definition (romm=rommapp/romm:5.3.1, romm-db=mariadb:11.4, romm-redis=redis:7-alpine)\n2026/09/24 06:43:28 update.go:884: [INFO] [stacks] update romm: healthy after 35s (the app's health check passed)\n2026/09/24 06:43:28 update.go:892: [INFO] [stacks] update romm: DONE in 57s\n", + "db_log": "2026-09-24 08:42:42+02:00 [Note] [Entrypoint]: MariaDB upgrade not required\nVersion: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution\n", + "files": "26: image: rommapp/romm:5.3.1\n86: memory: 768M\n102: image: mariadb:11.4\n123: memory: 384M\n132: image: redis:7-alpine\n145: memory: 128M\n---\napplied == steps/90dd9d68258286ef.yml\napplied != template\n", + "steps_line": { + "hu": [ + [ + "1", + "Hátralévő frissítési lépések: 1" + ] + ], + "en": [ + [ + "1", + "Update steps remaining: 1" + ] + ] + }, + "badges": { + "hu": [ + { + "title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", + "text": "Frissítés elérhető — 1 napja" + } + ], + "en": [ + { + "title": "A newer version of this app is available. Select the Update button to start it.", + "text": "Update available — 1 day ago" + } + ] + } + }, + "press2": { + "press": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "safety-dump", + "label": "Adatbázis pillanatkép…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 2.1, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, + "phase": "copying", + "label": "Az adatok másolása a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 15.4, + "phase": "starting", + "label": "Indítás az új verzióval…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 31.8, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 67.7, + "phase": "done", + "label": "Frissítve", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 67.7, + "final_phase": "done", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "obs": { + "pinned_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.8", + "romm-redis": "redis:7-alpine" + }, + "installed_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.8", + "romm-redis": "redis:7-alpine" + }, + "catalog_images": { + "romm": "rommapp/romm:5.3.1", + "romm-db": "mariadb:11.8", + "romm-redis": "redis:7-alpine" + }, + "live_compose_image_lines": [ + "image: rommapp/romm:5.3.1", + "image: mariadb:11.8", + "image: redis:7-alpine" + ], + "docker_inspect": [ + "romm-db mariadb:11.8 running=true restarts=0", + "romm rommapp/romm:5.3.1 running=true restarts=0", + "romm-redis redis:7-alpine running=true restarts=0" + ] + }, + "readback_A": true, + "log": "2026/09/24 06:43:46 update.go:722: [INFO] [stacks] update romm: ladder — the last step (3 of 3) — the catalog's current definition\n2026/09/24 06:43:48 pin.go:370: [INFO] [stacks] update romm: pin advanced to the catalog's current definition (romm=rommapp/romm:5.3.1, romm-db=mariadb:11.8, romm-redis=redis:7-alpine)\n2026/09/24 06:44:53 update.go:884: [INFO] [stacks] update romm: healthy after 35s (the app's health check passed)\n2026/09/24 06:44:53 update.go:892: [INFO] [stacks] update romm: DONE in 1m7s\n", + "db_log": "Version: '11.8.9-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution\n2026-09-24 08:44:05+02:00 [Note] [Entrypoint]: Starting mariadb-upgrade\nThe --upgrade-system-tables option was used, user tables won't be touched.\nMajor version upgrade detected from 11.4.13-MariaDB to 11.8.9-MariaDB. Check required!\n2026-09-24 08:44:12+02:00 [Note] [Entrypoint]: Finished mariadb-upgrade\nVersion: '11.8.9-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution\n", + "files": "26: image: rommapp/romm:5.3.1\n86: memory: 768M\n102: image: mariadb:11.8\n123: memory: 384M\n132: image: redis:7-alpine\n145: memory: 128M\n---\napplied != steps/90dd…\napplied == template\n", + "steps_line": { + "hu": [], + "en": [] + }, + "badges": { + "hu": [ + { + "title": "Ez az alkalmazás a legfrissebb elérhető változatot futtatja.", + "text": "Naprakész" + } + ], + "en": [ + { + "title": "This app is running the newest version available.", + "text": "Up to date" + } + ] + } + } +} \ No newline at end of file diff --git a/documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.log b/documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.log new file mode 100644 index 00000000..39e328a0 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/10-romm-two-steps.log @@ -0,0 +1,70 @@ +08:41:05 drill 95d83b7ab698: romm template = step 1's own definition (5.3.0 / mariadb 11.4) for the install push rc=0 +08:41:12 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm'] +08:41:12 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH'] +08:41:12 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:42:13 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} +08:42:22 romm: POST /api/users http=201 +08:42:23 drill 0e4e12aed1f2: romm back to the head (5.3.1 / mariadb 11.8) push rc=0 +08:42:31 before: steps line {"hu": [["2", "Hátralévő frissítési lépések: 2"]], "en": [["2", "Update steps remaining: 2"]]} badges {"hu": [{"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", "text": "Frissítés elérhető — 1 napja"}], "en": [{"title": "A newer version of this app is available. Select the Update button to start it.", "text": "Update available — 1 day +08:42:31 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:42:31 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:42:33 + 2.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:42:34 + 3.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:42:41 + 10.3s phase=starting label=Indítás az új verzióval… err=None hold=None +08:42:53 + 21.6s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:43:28 + 57.4s phase=done label=Frissítve err=None hold=None +08:43:37 romm: login as the seeded user http=200 ok=True +08:43:43 romm: POST /api/users http=403 +08:43:43 romm: refused {"detail":"Forbidden"} +08:43:46 PRESS 1: phase=done pinned={'romm': 'rommapp/romm:5.3.1', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} readback A=True +2026/09/24 06:42:31 update.go:722: [INFO] [stacks] update romm: ladder — step 2 of 3: romm=rommapp/romm:5.3.0, romm-db=mariadb:11.4, romm-redis=redis:7-alpine → romm=rommapp/romm:5.3.1, romm-db=mariadb:11.4, romm-redis=redis:7-alpine, from 90dd9d68258286ef.yml +2026/09/24 06:42:33 pin.go:370: [INFO] [stacks] update romm: pin advanced to the catalog's current definition (romm=rommapp/romm:5.3.1, romm-db=mariadb:11.4, romm-redis=redis:7-alpine) +2026/09/24 06:43:28 update.go:884: [INFO] [stacks] update romm: healthy after 35s (the app's health check passed) +2026/09/24 06:43:28 update.go:892: [INFO] [stacks] update romm: DONE in 57s + +DB: 2026-09-24 08:42:42+02:00 [Note] [Entrypoint]: MariaDB upgrade not required +Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution + +files: 26: image: rommapp/romm:5.3.1 +86: memory: 768M +102: image: mariadb:11.4 +123: memory: 384M +132: image: redis:7-alpine +145: memory: 128M +--- +applied == steps/90dd9d68258286ef.yml +applied != template + +steps line: {'hu': [('1', 'Hátralévő frissítési lépések: 1')], 'en': [('1', 'Update steps remaining: 1')]} +08:43:46 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:43:46 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:43:48 + 2.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:43:50 + 3.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:44:02 + 15.4s phase=starting label=Indítás az új verzióval… err=None hold=None +08:44:18 + 31.8s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:44:54 + 67.7s phase=done label=Frissítve err=None hold=None +08:45:02 romm: login as the seeded user http=200 ok=True +08:45:11 PRESS 2: phase=done pinned={'romm': 'rommapp/romm:5.3.1', 'romm-db': 'mariadb:11.8', 'romm-redis': 'redis:7-alpine'} readback A=True +2026/09/24 06:43:46 update.go:722: [INFO] [stacks] update romm: ladder — the last step (3 of 3) — the catalog's current definition +2026/09/24 06:43:48 pin.go:370: [INFO] [stacks] update romm: pin advanced to the catalog's current definition (romm=rommapp/romm:5.3.1, romm-db=mariadb:11.8, romm-redis=redis:7-alpine) +2026/09/24 06:44:53 update.go:884: [INFO] [stacks] update romm: healthy after 35s (the app's health check passed) +2026/09/24 06:44:53 update.go:892: [INFO] [stacks] update romm: DONE in 1m7s + +DB: Version: '11.8.9-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution +2026-09-24 08:44:05+02:00 [Note] [Entrypoint]: Starting mariadb-upgrade +The --upgrade-system-tables option was used, user tables won't be touched. +Major version upgrade detected from 11.4.13-MariaDB to 11.8.9-MariaDB. Check required! +2026-09-24 08:44:12+02:00 [Note] [Entrypoint]: Finished mariadb-upgrade +Version: '11.8.9-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution + +files: 26: image: rommapp/romm:5.3.1 +86: memory: 768M +102: image: mariadb:11.8 +123: memory: 384M +132: image: redis:7-alpine +145: memory: 128M +--- +applied != steps/90dd… +applied == template + +steps line: {'hu': [], 'en': []} diff --git a/documentation/audits/ladder-2026-09-24/partD/10.stdout b/documentation/audits/ladder-2026-09-24/partD/10.stdout new file mode 100644 index 00000000..39e328a0 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/10.stdout @@ -0,0 +1,70 @@ +08:41:05 drill 95d83b7ab698: romm template = step 1's own definition (5.3.0 / mariadb 11.4) for the install push rc=0 +08:41:12 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm'] +08:41:12 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH'] +08:41:12 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +08:42:13 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} +08:42:22 romm: POST /api/users http=201 +08:42:23 drill 0e4e12aed1f2: romm back to the head (5.3.1 / mariadb 11.8) push rc=0 +08:42:31 before: steps line {"hu": [["2", "Hátralévő frissítési lépések: 2"]], "en": [["2", "Update steps remaining: 2"]]} badges {"hu": [{"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", "text": "Frissítés elérhető — 1 napja"}], "en": [{"title": "A newer version of this app is available. Select the Update button to start it.", "text": "Update available — 1 day +08:42:31 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:42:31 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:42:33 + 2.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:42:34 + 3.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:42:41 + 10.3s phase=starting label=Indítás az új verzióval… err=None hold=None +08:42:53 + 21.6s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:43:28 + 57.4s phase=done label=Frissítve err=None hold=None +08:43:37 romm: login as the seeded user http=200 ok=True +08:43:43 romm: POST /api/users http=403 +08:43:43 romm: refused {"detail":"Forbidden"} +08:43:46 PRESS 1: phase=done pinned={'romm': 'rommapp/romm:5.3.1', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} readback A=True +2026/09/24 06:42:31 update.go:722: [INFO] [stacks] update romm: ladder — step 2 of 3: romm=rommapp/romm:5.3.0, romm-db=mariadb:11.4, romm-redis=redis:7-alpine → romm=rommapp/romm:5.3.1, romm-db=mariadb:11.4, romm-redis=redis:7-alpine, from 90dd9d68258286ef.yml +2026/09/24 06:42:33 pin.go:370: [INFO] [stacks] update romm: pin advanced to the catalog's current definition (romm=rommapp/romm:5.3.1, romm-db=mariadb:11.4, romm-redis=redis:7-alpine) +2026/09/24 06:43:28 update.go:884: [INFO] [stacks] update romm: healthy after 35s (the app's health check passed) +2026/09/24 06:43:28 update.go:892: [INFO] [stacks] update romm: DONE in 57s + +DB: 2026-09-24 08:42:42+02:00 [Note] [Entrypoint]: MariaDB upgrade not required +Version: '11.4.13-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution + +files: 26: image: rommapp/romm:5.3.1 +86: memory: 768M +102: image: mariadb:11.4 +123: memory: 384M +132: image: redis:7-alpine +145: memory: 128M +--- +applied == steps/90dd9d68258286ef.yml +applied != template + +steps line: {'hu': [('1', 'Hátralévő frissítési lépések: 1')], 'en': [('1', 'Update steps remaining: 1')]} +08:43:46 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +08:43:46 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +08:43:48 + 2.1s phase=pulling label=Új verzió letöltése… err=None hold=None +08:43:50 + 3.1s phase=copying label=Az adatok másolása a frissítés előtt… err=None hold=None +08:44:02 + 15.4s phase=starting label=Indítás az új verzióval… err=None hold=None +08:44:18 + 31.8s phase=verifying label=Működés ellenőrzése… err=None hold=None +08:44:54 + 67.7s phase=done label=Frissítve err=None hold=None +08:45:02 romm: login as the seeded user http=200 ok=True +08:45:11 PRESS 2: phase=done pinned={'romm': 'rommapp/romm:5.3.1', 'romm-db': 'mariadb:11.8', 'romm-redis': 'redis:7-alpine'} readback A=True +2026/09/24 06:43:46 update.go:722: [INFO] [stacks] update romm: ladder — the last step (3 of 3) — the catalog's current definition +2026/09/24 06:43:48 pin.go:370: [INFO] [stacks] update romm: pin advanced to the catalog's current definition (romm=rommapp/romm:5.3.1, romm-db=mariadb:11.8, romm-redis=redis:7-alpine) +2026/09/24 06:44:53 update.go:884: [INFO] [stacks] update romm: healthy after 35s (the app's health check passed) +2026/09/24 06:44:53 update.go:892: [INFO] [stacks] update romm: DONE in 1m7s + +DB: Version: '11.8.9-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 0 mariadb.org binary distribution +2026-09-24 08:44:05+02:00 [Note] [Entrypoint]: Starting mariadb-upgrade +The --upgrade-system-tables option was used, user tables won't be touched. +Major version upgrade detected from 11.4.13-MariaDB to 11.8.9-MariaDB. Check required! +2026-09-24 08:44:12+02:00 [Note] [Entrypoint]: Finished mariadb-upgrade +Version: '11.8.9-MariaDB-ubu2404' socket: '/run/mysqld/mysqld.sock' port: 3306 mariadb.org binary distribution + +files: 26: image: rommapp/romm:5.3.1 +86: memory: 768M +102: image: mariadb:11.8 +123: memory: 384M +132: image: redis:7-alpine +145: memory: 128M +--- +applied != steps/90dd… +applied == template + +steps line: {'hu': [], 'en': []} diff --git a/documentation/audits/ladder-2026-09-24/partD/11-teardown.log b/documentation/audits/ladder-2026-09-24/partD/11-teardown.log new file mode 100644 index 00000000..307553e0 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/11-teardown.log @@ -0,0 +1,27 @@ +08:45:32 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'} +08:45:37 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +08:45:37 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +08:46:04 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat +08:46:12 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm' +08:45:32 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'} +08:45:37 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +08:45:37 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +08:46:04 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat +08:46:12 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm' +no-romm-volumes +total 36 +drwxr-xr-x 3 root root 4096 Sep 24 06:46 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 12798 Sep 23 22:15 .felhom.yml +-rw-r--r-- 1 root root 5771 Sep 24 06:43 docker-compose.yml +drwxr-xr-x 11 root root 4096 Sep 23 10:13 hold-logs +ls: cannot access '/mnt/felhom-drives/scratch_hdd/userdata/romm': No such file or directory +felhom-controller Up 29 minutes (healthy) +privatebin Up 6 hours (healthy) +paperless-webserver Up 6 hours (healthy) +paperless-postgres Up 6 hours (healthy) +paperless-redis Up 6 hours (healthy) +gokapi Restarting (1) 58 seconds ago +filebrowser Up 10 hours (healthy) +traefik Up 10 hours + diff --git a/documentation/audits/ladder-2026-09-24/partD/20-9202-repoint-live.txt b/documentation/audits/ladder-2026-09-24/partD/20-9202-repoint-live.txt new file mode 100644 index 00000000..b6476f56 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/partD/20-9202-repoint-live.txt @@ -0,0 +1,28 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +/dev/loop1 69G 18G 48G 27% /var/lib/docker +0 +removed vikunja/vikunja:2.4.0 +removed vikunja/vikunja:2.5.0 +removed vikunja/vikunja:2.6.0 +removed nextcloud:34.0.1-apache +removed nextcloud:34.0.4-apache +removed mariadb:12.3 +removed rommapp/romm:5.3.0 +removed rommapp/romm:5.3.1 +removed mariadb:11.4 +removed mariadb:11.8 +removed actualbudget/actual-server:26.9.0 +/dev/loop1 69G 12G 54G 18% /var/lib/docker +ref: refs/heads/main +5ed599cd5f50 + +drill=5ed599cd5f50 live=5ed599cd5f50 +drill has_actions: False private: True diff --git a/documentation/audits/ladder-2026-09-24/redproofs/B-r659-hold.txt b/documentation/audits/ladder-2026-09-24/redproofs/B-r659-hold.txt new file mode 100644 index 00000000..4eb591ab --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/redproofs/B-r659-hold.txt @@ -0,0 +1,9 @@ +--- FAIL: TestR659_TheHoldNamesOnlyAWholeCopy (0.01s) + --- FAIL: TestR659_TheHoldNamesOnlyAWholeCopy/files_/_unit_only_(round_11) (0.00s) + r659_whole_copy_test.go:95: no whole copy exists, yet the hold names "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) nextcloud frissítése 2026-09-23 23:42-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 21:42 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." (none=false) + --- FAIL: TestR659_TheHoldNamesOnlyAWholeCopy/files_/_second_drive_only (0.00s) + r659_whole_copy_test.go:95: no whole copy exists, yet the hold names "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) nextcloud frissítése 2026-09-23 23:42-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 03:42 — ez a másolat a beállításokat, az adatbázist és a fájlokat tartalmazza." (none=false) + --- FAIL: TestR659_TheHoldNamesOnlyAWholeCopy/files_/_unit_+_second_drive (0.00s) + r659_whole_copy_test.go:95: no whole copy exists, yet the hold names "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) nextcloud frissítése 2026-09-23 23:42-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 21:42 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." (none=false) + --- FAIL: TestR659_TheHoldNamesOnlyAWholeCopy/files_/_off-site_+_unit (0.00s) +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.019s diff --git a/documentation/audits/ladder-2026-09-24/redproofs/B-r659-page.txt b/documentation/audits/ladder-2026-09-24/redproofs/B-r659-page.txt new file mode 100644 index 00000000..176bcef6 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/redproofs/B-r659-page.txt @@ -0,0 +1,6 @@ +--- FAIL: TestR659_HeldPage_NoWholeCopyOffersNoRestoreButton (0.24s) + r659_held_page_test.go:33: stacks: a hold naming no whole copy still offers the Mentések button — the restore there would refuse +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.252s +FAIL +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.283s diff --git a/documentation/audits/ladder-2026-09-24/redproofs/C-r653-r656.txt b/documentation/audits/ladder-2026-09-24/redproofs/C-r653-r656.txt new file mode 100644 index 00000000..48c4eaa3 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/redproofs/C-r653-r656.txt @@ -0,0 +1,9 @@ +FAIL +FAIL: test_the_apps_own_folders_are_cleared_and_said (__main__.ClearScratch.test_the_apps_own_folders_are_cleared_and_said) +AssertionError: True is not false : the last run's files are still there +FAIL: test_every_request_errored_is_inconclusive (__main__.LoadVerdict.test_every_request_errored_is_inconclusive) +AssertionError: 'reached' != 'inconclusive' +FAIL: test_under_half_is_inconclusive_and_none_is_inconclusive (__main__.LoadVerdict.test_under_half_is_inconclusive_and_none_is_inconclusive) +AssertionError: 'reached' != 'inconclusive' +FAILED (failures=3) +OK diff --git a/documentation/audits/ladder-2026-09-24/redproofs/D-catalog-gate.txt b/documentation/audits/ladder-2026-09-24/redproofs/D-catalog-gate.txt new file mode 100644 index 00000000..362dcfb0 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/redproofs/D-catalog-gate.txt @@ -0,0 +1,5 @@ +-- probe-matches-compose: the decoys (the label moves, the fact does not) + XX FACT: a two-step ladder with NO steps/ file for the first step rc=0 (expected 1) + XX DECOY: the steps/ file has the right NAME and names the head's image rc=0 (expected 1) + XX DECOY: the step's definition sits beside the template under another name rc=0 (expected 1) +test-record gate — 53 template(s) read, 25 carry a ladder, 0 convicted diff --git a/documentation/audits/ladder-2026-09-24/redproofs/D-catalog-writer.txt b/documentation/audits/ladder-2026-09-24/redproofs/D-catalog-writer.txt new file mode 100644 index 00000000..449b026f --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/redproofs/D-catalog-writer.txt @@ -0,0 +1,5 @@ +FAIL +FAIL: test_a_second_step_keeps_the_first_steps_definition (__main__.WriterTest.test_a_second_step_keeps_the_first_steps_definition) +AssertionError: False is not true : WROTE navidrome: {'navidrome': 'deluan/navidrome:0.64.1-next'} -> {'navidrome': 'deluan/navidrome:0.64.1-after'} peak 41.0% marks {'files_may_change': False, 'needs_person': None, 'memory_tight': False} +FAILED (failures=1) +OK diff --git a/documentation/audits/ladder-2026-09-24/redproofs/D-ladder.txt b/documentation/audits/ladder-2026-09-24/redproofs/D-ladder.txt new file mode 100644 index 00000000..d8a6ad62 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/redproofs/D-ladder.txt @@ -0,0 +1,7 @@ +--- FAIL: TestLadder_TwoPressesTwoSteps (0.01s) + ladder_test.go:83: press 1 pinned nextcloud:34.0.1-apache, want the tested step B nextcloud:33.0.0-apache — one press must be one step +--- FAIL: TestLadder_FailedStepStopsTheLadder (0.01s) + ladder_test.go:152: C was brought up after B failed: ups=[nextcloud:34.0.1-apache nextcloud:31.0.14-apache] +--- FAIL: TestLadder_MissingStepFileRefusesBeforeAnythingMoves (0.01s) + ladder_test.go:169: phase="done" key="", want failed/pin_failed +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.051s diff --git a/documentation/audits/ladder-2026-09-24/tools/fixtures.py b/documentation/audits/ladder-2026-09-24/tools/fixtures.py new file mode 100644 index 00000000..28cd756c --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/fixtures.py @@ -0,0 +1,1085 @@ +#!/usr/bin/env python3 +"""Box-side seed/verify fixtures for walk.py, guest 9202. + +THE ONE RULE (R-156), carried verbatim from `app-catalog-felhom.eu/scripts/upgrade_fixtures.py`: +*nothing is ever seeded into a volume by hand.* Every seed here goes in through the app's OWN +interface — its HTTP API through the household's real front door (traefik, `Host: .`), +or its own CLI running inside its own container. A raw SQL INSERT or a planted file is never used. + +If an app has no non-browser route, its fixture returns None and the edge is recorded +`inconclusive — no non-browser seed route`, WITH WHAT WAS TRIED. That is a result, not a gap. + +Each fixture: + seed(w, sub, say) -> an opaque token, or None + verify(w, sub, tok, say) -> True / False +verify() must ask the APP, never the filesystem: a migration is supposed to rewrite files. +Where a fixture can prove itself (a negative control that must read as absent) it does so on EVERY +call, so a readback that has broken into always saying "found" fails instead of passing everything. +""" +import base64, json, re, secrets, time + + +def _gx(w, container, *cmd, timeout=240): + """Run a command inside the app's OWN container on 9202 (its own CLI, not our SQL).""" + import shlex + line = " ".join(shlex.quote(c) for c in cmd) + return w.guest(f"docker exec {container} {line} 2>&1", timeout=timeout) + + +# ============================================================================================= +class PrivateBin: + """PrivateBin's own JSON API. A paste is a POST and reading it back is a GET — an + application-level round trip. File-backed, no database: this single seed IS the file half.""" + sub = "paste" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200",)): + return None + marker = "upg-" + secrets.token_hex(8) + ct = base64.b64encode(marker.encode()).decode() + body = json.dumps({ + "v": 2, + "adata": [[base64.b64encode(secrets.token_bytes(16)).decode(), + base64.b64encode(secrets.token_bytes(8)).decode(), + 100000, 256, 128, "aes", "gcm", "none"], "plaintext", 0, 0], + "ct": ct, "meta": {"expire": "never"}}) + rc, code, out = w.app_curl(sub, "/", "-H", "X-Requested-With: JSONHttpRequest", + "-H", "Content-Type: application/json", + data=body, method="POST") + try: + j = json.loads(out) + except Exception: + say(f" privatebin: POST returned non-JSON (http {code}): {out[:200]}") + return None + if j.get("status") != 0 or not j.get("id"): + say(f" privatebin: POST refused: {out[:250]}") + return None + say(f" privatebin: seeded paste id={j['id']}") + return {"id": j["id"], "marker": ct} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200",), tries=36): + return False + # negative control, every call: a paste id that cannot exist must NOT read back + rc, code, out = w.app_curl(sub, "/?pasteid=" + secrets.token_hex(8), + "-H", "X-Requested-With: JSONHttpRequest") + if t["marker"] in out: + say(" privatebin: READBACK UNUSABLE — a paste id that cannot exist returned the marker") + return False + rc, code, out = w.app_curl(sub, "/?pasteid=" + t["id"], + "-H", "X-Requested-With: JSONHttpRequest") + got = code == "200" and t["marker"] in out + say(f" privatebin: readback http={code} marker_present={got}") + return got + + +# ============================================================================================= +class Docmost: + """Docmost's own REST API: create the first workspace+user, then prove the account survives by + asking the app to AUTHENTICATE it. Login is version-stable across the API churn.""" + sub = "docs" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302", "404")): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + body = json.dumps({"workspaceName": "drill", "name": "drill", "email": email, "password": pw}) + rc, code, out = w.app_curl(sub, "/api/auth/setup", "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" docmost: /api/auth/setup http={code} rc={rc}") + if code not in ("200", "201"): + say(f" docmost: setup refused: {out[:250]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "404"), tries=36): + return False + # negative control: a password that was never set must NOT authenticate + bad = json.dumps({"email": t["email"], "password": "definitely-" + secrets.token_hex(8)}) + rc, code, _ = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" docmost: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"email": t["email"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" docmost: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" docmost: login body {out[:200]}") + return ok + + +# ============================================================================================= +class BookStack: + """BookStack mints no API token without a browser, so BOTH halves go through `php artisan` — + BookStack's OWN CLI, inside its own container, against its own User model. + + The exit code carries no information here (`bookstack:reset-mfa` exits 1 for a user it FOUND + and for one it did not), so the discriminator is the OUTPUT: the positive sentence required and + the not-found sentence required absent. The negative control runs on every verify. + + LIMITATION (R-460): this seeds the DATABASE half only. The FILE half needs the API token the + app cannot mint headlessly — so a bookstack edge is at best HALF-proven here. + """ + sub = "wiki" + + def _artisan(self, w, *args): + for path in ("/app/www/artisan", "/var/www/html/artisan"): + out = _gx(w, "bookstack", "php", path, *args) + if "Could not open input file" not in out: + return " ".join(out.split()) + return " ".join(out.split()) + + def _lookup(self, w, email): + out = self._artisan(w, "bookstack:reset-mfa", f"--email={email}") + found = f"Email: {email}" in out + missing = "could not be found" in out + if found == missing: + return None, out + return found, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + out = self._artisan(w, "bookstack:create-admin", f"--email={email}", + f"--name=drill-{secrets.token_hex(3)}", f"--password={pw}") + say(f" bookstack: artisan create-admin :: {out[:140]}") + if "successfully created" not in out: + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" bookstack: the app never served /login") + return False + absent, _ = self._lookup(w, f"nobody-{secrets.token_hex(6)}@gate.invalid") + if absent is not False: + say(f" bookstack: READBACK UNUSABLE — an email that cannot exist did not read absent ({absent})") + return False + found, out = self._lookup(w, t["email"]) + say(f" bookstack: readback of the seeded account found={found} :: {out[:140]}") + return found is True + + +# ============================================================================================= +class Gitea: + """Gitea's own admin CLI creates the first user; its own REST API (basic auth) then creates a + repository and reads it back. Both are the app's own interfaces.""" + sub = "git" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = _gx(w, "gitea", "su", "git", "-c", + f"gitea admin user create --username {user} --password {pw} " + f"--email {user}@gate.invalid --admin --must-change-password=false") + say(f" gitea: admin user create :: {' '.join(out.split())[:140]}") + if "has been successfully created" not in out and "successfully created" not in out: + return None + repo = "drillrepo" + secrets.token_hex(3) + rc, code, body = w.app_curl(sub, "/api/v1/user/repos", "-u", f"{user}:{pw}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": repo, "private": True}), method="POST") + say(f" gitea: create repo http={code}") + if code not in ("201", "200"): + say(f" gitea: repo refused {body[:200]}") + return None + return {"user": user, "pw": pw, "repo": repo} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + rc, code, _ = w.app_curl(sub, f"/api/v1/repos/{t['user']}/nope{secrets.token_hex(4)}", + "-u", f"{t['user']}:{t['pw']}") + if code == "200": + say(" gitea: READBACK UNUSABLE — a repo that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/repos/{t['user']}/{t['repo']}", + "-u", f"{t['user']}:{t['pw']}") + ok = code == "200" and t["repo"] in body + say(f" gitea: readback of the seeded repo http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Navidrome: + """Navidrome's own REST API: create the first admin through /auth/createAdmin, then prove the + account survives by logging in through the same door.""" + sub = "music" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, out = w.app_curl(sub, "/auth/createAdmin", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), method="POST") + say(f" navidrome: createAdmin http={code}") + if code not in ("200", "201"): + say(f" navidrome: refused {out[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" navidrome: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"username": t["user"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" navidrome: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vaultwarden: + """Vaultwarden's own account API: register an account, then prove it survives by asking the app + to issue a token for it (its own login endpoint, the household's own route).""" + sub = "vault" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/alive", want=("200",)): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + # Vaultwarden stores an already-hashed master key; the value is opaque to the server. + key = base64.b64encode(secrets.token_bytes(32)).decode() + body = json.dumps({"email": email, "name": "drill", "masterPasswordHash": key, + "key": "0." + base64.b64encode(secrets.token_bytes(48)).decode(), + "kdf": 0, "kdfIterations": 600000}) + rc, code, out = w.app_curl(sub, "/api/accounts/register", + "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" vaultwarden: register http={code}") + if code not in ("200", "204"): + say(f" vaultwarden: refused {out[:250]}") + return None + return {"email": email, "key": key} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/alive", want=("200",), tries=36): + return False + def login(pwhash): + return w.app_curl(sub, "/identity/connect/token", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=("grant_type=password&scope=api%20offline_access" + f"&client_id=web&deviceType=9&deviceIdentifier=drill" + f"&deviceName=drill&username={t['email']}&password={pwhash}"), + method="POST") + rc, code, _ = login(base64.b64encode(secrets.token_bytes(32)).decode()) + if code == "200": + say(" vaultwarden: READBACK UNUSABLE — a wrong master key authenticated") + return False + rc, code, out = login(t["key"].replace("+", "%2B").replace("=", "%3D").replace("/", "%2F")) + ok = code == "200" and "access_token" in out + say(f" vaultwarden: token for the seeded account http={code} ok={ok}") + if not ok: + say(f" vaultwarden: body {out[:200]}") + return ok + + +# ============================================================================================= +class Django: + """A Django app's OWN management CLI, inside its own container, against its own User model. + + Same category as BookStack's `php artisan`: the app's own code and its own ORM, never a raw SQL + INSERT and never a planted file (R-156). `createsuperuser --noinput` is Django's own documented + non-interactive route, and the readback asks the SAME ORM whether the account exists. + + THE FIXTURE PROVES ITSELF ON EVERY CALL: each verify() also asks for a username that cannot + exist and requires the answer False. A readback that has broken into always saying True + therefore fails instead of passing everything. + + LIMITATION, recorded rather than papered over: this seeds the DATABASE half only. An app whose + data is also FILES (adventurelog's images) has a file half this fixture does not touch. + """ + + def __init__(self, container, sub, ready_path="/", ready=("200", "302", "301", "404"), + python="python", workdir=None): + # `python` and `workdir` are per-app because the image decides them: adventurelog's + # interpreter is on PATH, tandoor ships a VENV and the bare `python` cannot import Django + # at all ("Couldn't import Django. Are you sure it's installed…"). Measured, not guessed. + self.container = container + self.sub = sub + self.ready_path = ready_path + self.ready = ready + self.python = python + self.workdir = workdir + + def _wd(self): + return f"-w {self.workdir} " if self.workdir else "" + + def _manage(self, w, code): + # -c is passed to `manage.py shell`; the app's own shell, its own ORM. + return w.guest( + f"docker exec {self._wd()}{self.container} {self.python} manage.py shell " + f"-c {json.dumps(code)} 2>&1", timeout=300) + + def _exists(self, w, username): + # ONE LINE, semicolon-separated. A `\n` inside a double-quoted shell argument reaches + # python as a literal backslash-n and is a SyntaxError — which is exactly how the first + # adventurelog run read as `inconclusive`. The fixture refused to guess, which is right, + # but the instrument was the thing that was broken. + out = self._manage(w, ( + "from django.contrib.auth import get_user_model; " + f"print('DRILL_ANSWER=' + str(get_user_model().objects.filter(username={username!r}).exists()))" + )) + m = re.search(r"DRILL_ANSWER=(True|False)", out) + return (m.group(1) == "True") if m else None, " ".join(out.split())[-300:] + + def seed(self, w, sub, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -e DJANGO_SUPERUSER_PASSWORD={pw} {self._wd()}{self.container} " + f"{self.python} manage.py createsuperuser --noinput " + f"--username {user} --email {user}@gate.invalid 2>&1", timeout=300) + say(f" {self.container}: createsuperuser :: {' '.join(out.split())[:160]}") + got, detail = self._exists(w, user) + if got is not True: + say(f" {self.container}: the account did not appear in the app's own ORM :: {detail[:200]}") + return None + say(f" {self.container}: seeded superuser {user}") + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + say(f" {self.container}: the app never served {self.ready_path}") + return False + absent, detail = self._exists(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" {self.container}: READBACK UNUSABLE — a username that cannot exist did not " + f"read as absent ({absent}) :: {detail[:200]}") + return False + found, detail = self._exists(w, t["user"]) + say(f" {self.container}: readback of the seeded account found={found}") + if found is not True: + say(f" {self.container}: :: {detail[:250]}") + return found is True + + +# ============================================================================================= +class Nextcloud: + """Nextcloud's OWN admin CLI, `occ`, inside its own container: its own code, its own user + backend. Not a SQL INSERT and not a planted file (R-156). + + `occ user:info` is the readback, and it PROVES ITSELF on every call: a uid that cannot exist + must answer "user not found". A readback that has broken into always succeeding therefore + fails instead of passing everything. + + This is the app chosen for the MariaDB engine-major edge (`09` §3 decision 5, R-469 lifted): + the app image does NOT move, only the `mariadb:` sidecar, so the edge carries exactly one + migration and a failure is readable. + """ + sub = "cloud" + + def _occ(self, w, *args, timeout=420): + import shlex + line = " ".join(shlex.quote(a) for a in args) + return w.guest(f"docker exec -u www-data nextcloud php occ {line} 2>&1", timeout=timeout) + + def _info(self, w, uid): + out = self._occ(w, "user:info", uid) + flat = " ".join(out.split()) + if "user not found" in flat.lower() or "could not be found" in flat.lower(): + return False, flat + if f"user_id: {uid}" in flat or f"- user_id: {uid}" in flat or f"user_id: {uid}" in out: + return True, flat + return None, flat + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + return None + # /status.php answers 200 while the image's own first-run install is still going — measured + # 2026-09-23 night on a loaded bench: `occ` then says "Nextcloud is not installed". Ask occ. + for i in range(60): + st = self._occ(w, "status", timeout=120) + if "installed: true" in st: + break + time.sleep(5) + else: + say(f" nextcloud: occ status never said installed: {' '.join(st.split())[:160]}") + return None + uid = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -u www-data -e OC_PASS={pw} nextcloud php occ user:add " + f"--password-from-env --display-name={uid} {uid} 2>&1", timeout=420) + say(f" nextcloud: occ user:add :: {' '.join(out.split())[:160]}") + got, flat = self._info(w, uid) + if got is not True: + say(f" nextcloud: the account did not appear via occ user:info :: {flat[:220]}") + return None + say(f" nextcloud: seeded user {uid}") + return {"uid": uid, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + say(" nextcloud: the app never served /status.php") + return False + absent, flat = self._info(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" nextcloud: READBACK UNUSABLE — a uid that cannot exist did not read absent " + f"({absent}) :: {flat[:200]}") + return False + found, flat = self._info(w, t["uid"]) + say(f" nextcloud: readback of the seeded user found={found}") + if found is not True: + say(f" nextcloud: :: {flat[:250]}") + return found is True + + +# ============================================================================================= +class Grafana: + """Grafana's own HTTP API as the admin the DEPLOY created. The password is the one the + controller showed the household — read from the app's own `app.yaml`, not invented — and the + data (a folder) goes in and comes back through the app's own REST API.""" + sub = "grafana" + + def _auth(self, w, name="grafana"): + # app.yaml stores this ENCRYPTED (`ENC:…`), so it cannot be read back off the box — which + # is correct, and is why the harness uses the value IT generated for the deploy. + pw = (w.GENERATED.get(name) or {}).get("GF_SECURITY_ADMIN_PASSWORD") or "admin" + return f"admin:{pw}" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return None + au = self._auth(w) + title = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/folders", "-u", au, + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="POST") + say(f" grafana: create folder http={code}") + if code not in ("200", "201"): + say(f" grafana: refused {body[:220]}") + return None + try: + uid = json.loads(body)["uid"] + except Exception: + say(f" grafana: no uid in {body[:200]}") + return None + return {"uid": uid, "title": title} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return False + au = self._auth(w) + rc, code, _ = w.app_curl(sub, "/api/folders/nope" + secrets.token_hex(5), "-u", au) + if code == "200": + say(" grafana: READBACK UNUSABLE — a folder uid that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/folders/{t['uid']}", "-u", au) + ok = code == "200" and t["title"] in body + say(f" grafana: readback of the seeded folder http={code} ok={ok}") + return ok + + +# ============================================================================================= +class AudiobookShelf: + """audiobookshelf's own /init endpoint creates the first root account; its own /login proves + the account survived. Both are the app's own API.""" + sub = "audiobooks" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/init", "-H", "Content-Type: application/json", + data=json.dumps({"newRoot": {"username": user, "password": pw}}), + method="POST") + say(f" audiobookshelf: /init http={code}") + if code not in ("200", "204"): + say(f" audiobookshelf: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code == "200": + say(" audiobookshelf: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": t["pw"]}), + method="POST") + ok = code == "200" and t["user"] in body + say(f" audiobookshelf: login as the seeded root http={code} ok={ok}") + return ok + + +# ============================================================================================= +class ActualBudget: + """Actual's own bootstrap API sets the server password; its own login proves it survived.""" + sub = "budget" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/account/bootstrap", + "-H", "Content-Type: application/json", + data=json.dumps({"password": pw}), method="POST") + say(f" actualbudget: /account/bootstrap http={code} :: {body[:140]}") + if code not in ("200", "201") or '"status":"ok"' not in body: + return None + return {"pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def login(p): + return w.app_curl(sub, "/account/login", "-H", "Content-Type: application/json", + data=json.dumps({"loginMethod": "password", "password": p}), + method="POST") + rc, code, body = login("wrong-" + secrets.token_hex(6)) + if '"status":"ok"' in body: + say(" actualbudget: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = '"status":"ok"' in body + say(f" actualbudget: login with the seeded password http={code} ok={ok}") + if not ok: + say(f" actualbudget: body {body[:200]}") + return ok + + +# ============================================================================================= +class Mealie: + """Mealie ships a documented first-run admin. We log in as it through the app's own OAuth-style + token endpoint, create a recipe through the app's own API, and read the recipe back.""" + sub = "recipes" + + def _token(self, w, sub, pw="MyPassword"): + rc, code, body = w.app_curl( + sub, "/api/auth/token", "-H", "Content-Type: application/x-www-form-urlencoded", + data=f"username=changeme%40example.com&password={pw}", method="POST") + if code != "200": + return None, f"http={code} {body[:200]}" + try: + return json.loads(body)["access_token"], "" + except Exception: + return None, body[:200] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return None + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate as the first-run admin :: {why}") + return None + name = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/recipes", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": name}), method="POST") + say(f" mealie: create recipe http={code}") + if code not in ("200", "201"): + say(f" mealie: refused {body[:220]}") + return None + slug = body.strip().strip('"') + return {"slug": slug, "name": name} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return False + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate after the update :: {why}") + return False + rc, code, _ = w.app_curl(sub, "/api/recipes/nope" + secrets.token_hex(5), + "-H", f"Authorization: Bearer {tok}") + if code == "200": + say(" mealie: READBACK UNUSABLE — a slug that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/recipes/{t['slug']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["name"] in body + say(f" mealie: readback of the seeded recipe http={code} ok={ok}") + return ok + + +# ============================================================================================= +class N8n: + """n8n's own owner-setup API creates the first account; its own login proves it survived.""" + sub = "auto" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill" + secrets.token_hex(8) + "1" + rc, code, body = w.app_curl(sub, "/rest/owner/setup", "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "firstName": "drill", + "lastName": "drill", "password": pw}), + method="POST") + say(f" n8n: /rest/owner/setup http={code}") + if code not in ("200", "201"): + say(f" n8n: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return False + def login(p): + return w.app_curl(sub, "/rest/login", "-H", "Content-Type: application/json", + data=json.dumps({"emailOrLdapLoginId": t["email"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" n8n: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" and t["email"] in body + say(f" n8n: login as the seeded owner http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Zipline: + """Zipline's own setup/login API. Zipline 4 creates the first user through its own endpoint.""" + sub = "img" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/healthcheck", want=("200",), tries=90): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=30): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + for path in ("/api/auth/register", "/api/auth/setup"): + rc, code, body = w.app_curl(sub, path, "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), + method="POST") + say(f" zipline: {path} http={code} :: {body[:160]}") + if code in ("200", "201"): + return {"user": user, "pw": pw} + say(" zipline: neither register nor setup accepted a first user") + return None + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=60): + return False + def login(p): + return w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" zipline: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" + say(f" zipline: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vikunja: + """Vikunja's own REST API: register a user, log in, create a project, read the project back. + Four calls, all the app's own front door.""" + sub = "tasks" + + def _token(self, w, sub, t, pw=None): + rc, code, body = w.app_curl(sub, "/api/v1/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], + "password": pw or t["pw"]}), method="POST") + if code != "200": + return None, f"http={code} {body[:160]}" + try: + return json.loads(body)["token"], "" + except Exception: + return None, body[:160] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/v1/register", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw, + "email": f"{user}@gate.invalid"}), + method="POST") + say(f" vikunja: register http={code}") + if code not in ("200", "201"): + say(f" vikunja: refused {body[:220]}") + return None + t = {"user": user, "pw": pw} + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: could not log in after registering :: {why}") + return None + title = "drill-" + secrets.token_hex(5) + # Vikunja CREATES with PUT, not POST — a POST answers `405 Method Not Allowed`, which + # reads like a broken fixture and is really the wrong verb. Measured 2026-09-21. + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="PUT") + say(f" vikunja: create project http={code}") + if code not in ("200", "201"): + say(f" vikunja: project refused {body[:220]}") + return None + t["title"] = title + t["pid"] = json.loads(body).get("id") + return t + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return False + bad, why = self._token(w, sub, t, pw="wrong-" + secrets.token_hex(6)) + if bad: + say(" vikunja: READBACK UNUSABLE — a wrong password authenticated") + return False + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: the seeded account no longer authenticates :: {why}") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{t['pid']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["title"] in body + say(f" vikunja: readback of the seeded project http={code} ok={ok}") + return ok + + +# ============================================================================================= +class OpenGist: + """Opengist's own sign-up and sign-in FORMS. + + Two things had to be measured. Its sign-up is CSRF-protected: a bare POST answers 500 with an + HTML page, which reads like a broken app and is really a missing token — fetch the form, keep + its cookie, send its `_csrf` back. And its REST API refuses the account's own password + (`401 {"message":"Bad crendentials"}`) because it wants a token the app will not mint without a + browser. So the SEEDED DATA is the account itself and the READBACK is a real sign-in, which is + the same shape the docmost and navidrome fixtures use. + + LIMITATION, recorded rather than papered over: this is the DATABASE half. A gist's CONTENT is + not seeded, because that needs the API token above. + """ + sub = "gist" + + def _form(self, w, sub, path, jar, fields): + rc, code, html = w.app_curl(sub, path, "-b", jar, "-c", jar) + m = re.search(r'name="_csrf"[^>]*value="([^"]+)"', html or "") + if not m: + return None, f"no _csrf on {path} (http={code})" + body = "&".join([f"_csrf={m.group(1)}"] + [f"{k}={v}" for k, v in fields.items()]) + rc, code, out = w.app_curl(sub, path, "-b", jar, "-c", jar, + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + return code, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, out = self._form(w, sub, "/register", jar, {"username": user, "password": pw}) + say(f" opengist: /register (with its own _csrf) http={code}") + if code not in ("200", "302", "303"): + say(f" opengist: refused {str(out)[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + # Wait for the LOGIN FORM, not for the root page. Measured 2026-09-21: immediately after a + # successful update the root answers while /login does not yet carry its `_csrf`, so the + # sign-in silently fails and the app looks like it lost the account. It had not. + # 1.15 moved every page under `/-/` (`/-/login`, `/-/all`; `/login` answers 404) — measured + # 2026-09-23 night. Ask the app which shape it serves instead of assuming one. + pre, home_path = "", "/" + for _ in range(72): + if w.app_curl(sub, "/-/login")[1] == "200": + pre, home_path = "/-", "/-/all" + break + if w.app_curl(sub, "/login")[1] == "200": + break + time.sleep(5) + else: + say(" opengist: neither /login nor /-/login came back after the update") + return False + say(f" opengist: sign-in form at {pre}/login") + for _ in range(24): + rc, code, html = w.app_curl(sub, pre + "/login") + if code == "200" and '_csrf' in (html or ""): + break + time.sleep(5) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, pre + "/login", jar, + {"username": t["user"], "password": "wrong-" + secrets.token_hex(5)}) + rc, c2, home = w.app_curl(sub, home_path, "-b", jar) + if t["user"] in (home or ""): + say(" opengist: READBACK UNUSABLE — a wrong password signed in") + return False + jar2 = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, pre + "/login", jar2, {"username": t["user"], "password": t["pw"]}) + rc, c2, home = w.app_curl(sub, home_path, "-b", jar2) + signed_in = t["user"] in (home or "") + # The ACCOUNT's own public page is the readback that does not depend on a cookie: 1.15 marks its + # session cookie Secure, so a plain-HTTP bench cannot send it back (measured 2026-09-23 night). + # A user that was never created must 404 on the same call, or the readback proves nothing. + rc, pc, prof = w.app_curl(sub, "/" + t["user"]) + rc, nc, _ = w.app_curl(sub, "/nobody" + secrets.token_hex(4)) + profile = pc == "200" and t["user"] in (prof or "") + if nc == "200": + say(" opengist: READBACK UNUSABLE — a never-created user's page answered 200") + return None + say(f" opengist: account page /{t['user']} http={pc} found={profile} (never-created user {nc}); " + f"sign-in http={code} name_on_page={signed_in}") + return profile + + +# ============================================================================================= +class Papra: + """Papra's own e-mail sign-up and sign-in endpoints.""" + sub = "papra" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/auth/sign-up/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "password": pw, + "name": "drill"}), method="POST") + say(f" papra: sign-up http={code}") + if code not in ("200", "201"): + say(f" papra: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def signin(p): + return w.app_curl(sub, "/api/auth/sign-in/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": t["email"], "password": p}), method="POST") + rc, code, _ = signin("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" papra: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = signin(t["pw"]) + ok = code == "200" + say(f" papra: sign-in as the seeded account http={code} ok={ok}") + return ok + + +# ============================================================================================= +class HomeAssistant: + """Home Assistant's own onboarding API creates the owner account and hands back a code the + same API exchanges for a token. Both are the app's own documented non-browser route.""" + sub = "ha" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/onboarding/users", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "name": "drill", "username": user, + "password": pw, "language": "en"}), + method="POST") + say(f" home-assistant: /api/onboarding/users http={code}") + if code not in ("200", "201"): + say(f" home-assistant: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def _login(self, w, sub, user, pw): + """The app's own login flow: start it, then answer it. A 200 with a step_id of + `mfa`/`init` means the credentials were REFUSED; only `create_entry` is a pass.""" + rc, code, body = w.app_curl(sub, "/auth/login_flow", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "handler": ["homeassistant", None], + "redirect_uri": f"https://{sub}.felhom.invalid/", + "type": "authorize"}), method="POST") + if code not in ("200", "201"): + return None, f"flow start http={code} {body[:160]}" + try: + fid = json.loads(body)["flow_id"] + except Exception: + return None, body[:160] + rc, code, body = w.app_curl(sub, f"/auth/login_flow/{fid}", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "username": user, "password": pw}), + method="POST") + try: + j = json.loads(body) + except Exception: + return None, body[:160] + return (j.get("result") if j.get("type") == "create_entry" else None), body[:200] + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return False + bad, why = self._login(w, sub, t["user"], "wrong-" + secrets.token_hex(6)) + if bad: + say(" home-assistant: READBACK UNUSABLE — a wrong password authenticated") + return False + good, why = self._login(w, sub, t["user"], t["pw"]) + ok = bool(good) + say(f" home-assistant: login as the seeded owner ok={ok}") + if not ok: + say(f" home-assistant: {why}") + return ok + + +# ============================================================================================= +class Romm: + """RomM's own user API, driven the way RomM's own front end drives it. + + Three things had to be measured rather than guessed, and each one answered a 403 or a 422 that + looked like a different fault: RomM sets a **`romm_csrftoken` cookie** on any GET and requires + it back in an **`x-csrftoken` header** (a bare POST is `403 CSRF token verification failed`, + which reads like an auth problem); the fields go in the **JSON body**, not the query string (a + query-string POST is `422 Field required` for every field it was just given); and `email` is + required alongside username, password and role. + + On a fresh install with no admin the first `POST /api/users` is accepted unauthenticated; + afterwards it is not — which is what makes the readback (`POST /api/login` as that user) a real + authentication rather than a repeat of the seed. + + LIMITATION: this is the DATABASE half. RomM's other half is the ROM library on the drive, which + this does not populate. + """ + sub = "arcade" + + def _csrf(self, w, sub): + jar = f"/tmp/romm-{secrets.token_hex(4)}.jar" + w.app_curl(sub, "/api/heartbeat", "-c", jar) + out = w.sh(["bash", "-lc", f"grep -i csrf {jar} | awk '{{print $7}}'"]).stdout or "" + return jar, out.strip() + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + jar, tok = self._csrf(w, sub) + if not tok: + say(" romm: no romm_csrftoken cookie was set on /api/heartbeat") + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl( + sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": pw, "role": "admin"}), method="POST") + say(f" romm: POST /api/users http={code}") + if code not in ("200", "201"): + say(f" romm: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + return False + jar, tok = self._csrf(w, sub) + rc, code, _ = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:wrong-{secrets.token_hex(5)}", method="POST") + if code == "200": + say(" romm: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:{t['pw']}", method="POST") + ok = code == "200" + say(f" romm: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" romm: body {body[:200]}") + return ok + + + +# ============================================================================================= +class Wishlist: + """Wishlist's own SvelteKit FORM actions (added night 2026-09-23, R-612's app). Sign-up at + /signup, then prove the account survived by signing in at /login — and by a wrong password being + REFUSED on the same call, so a readback that always says "ok" fails instead of passing. + SvelteKit refuses a cross-site form post: the Origin must be the app's own https origin.""" + sub = "wishlist" + + def _post(self, w, sub, path, body): + origin = "https://" + getattr(w, "host", lambda s: f"{s}.{w.DOMAIN}")(sub) # the Host the request carries + rc, code, out = w.app_curl(sub, path, "-H", f"Origin: {origin}", "-H", "x-sveltekit-action: true", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + try: + return code, json.loads(out) + except Exception: + return code, {"type": "unparsed", "raw": (out or "")[:200]} + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/signup", want=("200",), tries=72): + return None + u = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(8) + code, j = self._post(w, sub, "/signup", + f"name=Drill&username={u}&email={u}%40example.invalid&password={pw}&tokenId=") + say(f" wishlist: /signup http={code} type={j.get('type')}") + if j.get("type") not in ("success", "redirect"): + self.tried = f"POST /signup -> {code} {str(j)[:150]}" + return None + return {"u": u, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" wishlist: /login never came back") + return False + c1, bad = self._post(w, sub, "/login", f"username={t['u']}&password=wrong-{secrets.token_hex(5)}") + if bad.get("type") != "failure": + say(f" wishlist: READBACK UNUSABLE — a wrong password was not refused ({bad.get('type')})") + return None + c2, good = self._post(w, sub, "/login", f"username={t['u']}&password={t['pw']}") + ok = good.get("type") in ("success", "redirect") + say(f" wishlist: sign-in as the seeded user type={good.get('type')} ok={ok} (wrong password refused)") + return ok + +FIXTURES = { + "home-assistant": HomeAssistant(), + "romm": Romm(), + "vikunja": Vikunja(), + "opengist": OpenGist(), + "papra": Papra(), + "mealie": Mealie(), + "n8n": N8n(), + "zipline": Zipline(), + "grafana": Grafana(), + "audiobookshelf": AudiobookShelf(), + "actualbudget": ActualBudget(), + "nextcloud": Nextcloud(), + "adventurelog": Django("adventurelog", "travel", "/admin/login/"), + "tandoor": Django("tandoor", "recipes", "/accounts/login/", + python="/opt/recipes/venv/bin/python", workdir="/opt/recipes"), + "privatebin": PrivateBin(), + "docmost": Docmost(), + "bookstack": BookStack(), + "gitea": Gitea(), + "navidrome": Navidrome(), + "vaultwarden": Vaultwarden(), + "wishlist": Wishlist(), +} diff --git a/documentation/audits/ladder-2026-09-24/tools/liveA1.py b/documentation/audits/ladder-2026-09-24/tools/liveA1.py new file mode 100644 index 00000000..c737b2ef --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/liveA1.py @@ -0,0 +1,64 @@ +#!/usr/bin/env python3 +"""Part A, first half — on controller v0.267.0 (the field's version): vikunja installed, seeded (A), +backed up by the product (the update's own `backing-up` leg, then a pull that fails: nothing moves), +RESTORED from its own unit — which recreates its volumes WITHOUT compose labels (R-658) — and seeded +again (B, after the restore). Evidence only; every act is a product endpoint.""" +import json, re, sys, time +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +ST = "../partA/state.json" +F = fx.Vikunja() +SUB = "vikunja-la" +out = {"controller": w.guest("cat /etc/felhom-controller-image").strip()} +w.login() + + +def drill_set(ref, msg): + import fcntl + with open(f"{w.SC}/drill.lock", "w") as lk: + fcntl.flock(lk, fcntl.LOCK_EX) + w.sh(["git", "-C", w.DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120) + p = f"{w.DRILL}/templates/vikunja/docker-compose.yml" + s = open(p).read() + s = re.sub(r"image: vikunja/vikunja:\S+", f"image: vikunja/vikunja:{ref}", s) + open(p, "w").write(s) + w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"LIVE-A vikunja {ref}: {msg}"]) + r = w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + h = w.sh(["git", "-C", w.DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + w.say(f"drill {h}: vikunja {ref} ({msg}) push rc={r.returncode}") + return h + + +def labels(): + return w.guest("for v in $(docker volume ls -q | grep '^vikunja_'); do echo \"$v $(docker volume inspect $v --format '{{json .Labels}}')\"; done") + + +out["drill_A"] = drill_set("2.5.0", "the install version") +w.sync_rescan() +out["deployed"] = w.deploy("vikunja", SUB) +w.wait_app(SUB, "/", tries=60, delay=3) +A = F.seed(w, SUB, w.say) +out["seed_A"] = bool(A) +out["labels_after_install"] = labels() +w.say("labels after install:\n" + out["labels_after_install"]) + +out["drill_E"] = drill_set("2.5.99-notag", "a tag that does not exist: backing-up runs, the pull fails, nothing moves") +out["badge_wait"] = w.sync_rescan("vikunja", "vikunja/vikunja:2.5.99-notag") +out["press_backup"] = w.press_update("vikunja") +out["snapshots"] = w.snapshots("vikunja") +w.say("snapshots offered: " + json.dumps(out["snapshots"])[:400]) +out["drill_back"] = drill_set("2.5.0", "back to the install version") +w.sync_rescan() + +out["restore"] = w.restore("vikunja") +out["labels_after_restore_v0267"] = labels() +w.say("labels after the v0.267.0 restore:\n" + out["labels_after_restore_v0267"]) +w.wait_app(SUB, "/", tries=60, delay=3) +out["A_after_restore"] = bool(F.verify(w, SUB, A, w.say)) +B = F.seed(w, SUB, w.say) +out["seed_B_after_restore"] = bool(B) +json.dump({"A": A, "B": B, "sub": SUB}, open(ST, "w")) +json.dump(out, open("../partA/01-v0267-install-backup-restore.json", "w"), indent=2, ensure_ascii=False, default=str) +open("../partA/01-v0267-install-backup-restore.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/liveA2.py b/documentation/audits/ladder-2026-09-24/tools/liveA2.py new file mode 100644 index 00000000..d7ec75df --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/liveA2.py @@ -0,0 +1,96 @@ +#!/usr/bin/env python3 +"""Part A, second half — on controller v0.268.0. vikunja's volumes are UNLABELLED (restored under v0.267.0, +liveA1). (e) a failing update (real move 2.5.0 -> 2.6.0 + a probe port the app does not answer): the undo +must copy BOTH volumes and seeds A and B must read back. (f) a restore under v0.268.0: the recreated +volumes carry compose's labels. (g) seed C, the same failing update again: undone, both copied, A B C read +back. (h) remove: volumes_removed names both, no applied-compose.yml / applied-meta/ left (R-651).""" +import json, re, sys, time +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +S = json.load(open("../partA/state.json")) +F, SUB = fx.Vikunja(), S["sub"] +out = {"controller": w.guest("cat /etc/felhom-controller-image").strip()} +w.login() + + +def drill(edit, msg): + import fcntl + with open(f"{w.SC}/drill.lock", "w") as lk: + fcntl.flock(lk, fcntl.LOCK_EX) + w.sh(["git", "-C", w.DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120) + edit() + w.sh(["git", "-C", w.DRILL, "commit", "-qam", "LIVE-A " + msg]) + r = w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + h = w.sh(["git", "-C", w.DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + w.say(f"drill {h}: {msg} push rc={r.returncode}") + return h + + +def set_vik(ref, port): + p = f"{w.DRILL}/templates/vikunja/docker-compose.yml" + s = re.sub(r"image: vikunja/vikunja:\S+", f"image: vikunja/vikunja:{ref}", open(p).read()) + open(p, "w").write(s) + fy = f"{w.DRILL}/templates/vikunja/.felhom.yml" + f = open(fy).read() + f2 = re.sub(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )\d+", lambda m: m.group(1) + str(port), f, count=1) + assert f2 != f or str(port) in f, "probe port not found" + open(fy, "w").write(f2) + + +def labels(): + return w.guest("for v in $(docker volume ls -q | grep '^vikunja_'); do echo \"$v $(docker volume inspect $v --format '{{json .Labels}}')\"; done") + + +def ctl_log(since): + return w.guest(f"docker logs --since {since} felhom-controller 2>&1 | grep -E 'vikunja' | grep -E 'undo copy will hold|carry no compose label|ladder|copied|UNDONE|UNDO|RestoreStack|Restoring Docker volume|was not created|volume .* already exists|created WITHOUT' | tail -30") + + +def readall(tag, seeds): + w.wait_app(SUB, "/", tries=60, delay=3) + r = {k: bool(F.verify(w, SUB, v, w.say)) for k, v in seeds.items()} + w.say(f"READBACK {tag}: {r}") + return r + + +PORT_REAL = int(re.search(r"port: (\d+)", open(f"{w.DRILL}/templates/vikunja/.felhom.yml").read().split("healthcheck:")[1]).group(1)) +out["probe_port_real"] = PORT_REAL +out["labels_before"] = labels() +w.say("labels before (unlabelled after the v0.267.0 restore):\n" + out["labels_before"]) + +# (e) +t0 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +out["drill_e"] = drill(lambda: set_vik("2.6.0", 8999), "vikunja 2.6.0 + probe 8999 (the failing edge)") +out["badge_e"] = w.sync_rescan("vikunja", "vikunja/vikunja:2.6.0") +out["press_e"] = w.press_update("vikunja") +out["log_e"] = ctl_log(t0) +w.say("controller log (e):\n" + out["log_e"]) +out["read_e"] = readall("after the undo (e)", {"A": S["A"], "B": S["B"]}) +out["obs_e"] = w.observables("vikunja") + +# (f) +t1 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +out["restore_f"] = w.restore("vikunja") +out["labels_after_restore_v0268"] = labels() +w.say("labels after the v0.268.0 restore:\n" + out["labels_after_restore_v0268"]) +out["log_f"] = ctl_log(t1) +out["compose_warn_f"] = w.guest(f"docker logs --since {t1} felhom-controller 2>&1 | grep -iE 'not created by Docker Compose|already exists|Recreate' | tail -5") +w.say("compose warnings after the restore (want none): " + (out["compose_warn_f"].strip() or "(none)")) +out["read_f"] = readall("after the v0.268.0 restore (f)", {"A": S["A"]}) + +# (g) +C = F.seed(w, SUB, w.say) +t2 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +out["press_g"] = w.press_update("vikunja") +out["log_g"] = ctl_log(t2) +w.say("controller log (g):\n" + out["log_g"]) +out["read_g"] = readall("after the second undo (g)", {"A": S["A"], "C": C}) + +# (h) +out["drill_h"] = drill(lambda: set_vik("2.5.0", PORT_REAL), "vikunja back to 2.5.0 + its real probe") +out["remove_h"] = w.remove("vikunja") +out["leftovers_h"] = w.guest("ls -la /opt/docker/stacks/vikunja/; docker volume ls -q | grep -E '^vikunja' || echo 'no vikunja volumes'") +w.say("after remove:\n" + out["leftovers_h"]) +json.dump(out, open("../partA/02-v0268-undo-after-restore.json", "w"), indent=2, ensure_ascii=False, default=str) +open("../partA/02-v0268-undo-after-restore.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/liveB.py b/documentation/audits/ladder-2026-09-24/tools/liveB.py new file mode 100644 index 00000000..2000e8d6 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/liveB.py @@ -0,0 +1,117 @@ +#!/usr/bin/env python3 +"""Part B — R-659 on controller v0.268.0: chaos round 11 reproduced. nextcloud (declared drive files), +a failing update (34.0.1 -> 34.0.4 + a probe port it does not answer), the undo copy CUT OFF during +`verifying` (its finished-marker removed — `09` §6.1a's measured case) -> HOLD. No off-site tier on 9202. +Expected: the no-whole-copy sentence in hu and en, no Mentések button beside it, the operator event +(DROPPED at WARN on this hub-less box — R-620), and NO app_start_failed after the hold (R-660).""" +import html as H, json, re, sys, time +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +F, SUB, APP = fx.Nextcloud(), "cloud-lb", "nextcloud" +CAT = f"{w.DRILL}/templates/{APP}" +out = {"controller": w.guest("cat /etc/felhom-controller-image").strip()} +w.login() + + +def drill(edit, msg): + import fcntl + with open(f"{w.SC}/drill.lock", "w") as lk: + fcntl.flock(lk, fcntl.LOCK_EX) + w.sh(["git", "-C", w.DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120) + edit() + w.sh(["git", "-C", w.DRILL, "commit", "-qam", "LIVE-B " + msg]) + r = w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + h = w.sh(["git", "-C", w.DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + w.say(f"drill {h}: {msg} push rc={r.returncode}") + return h + + +HEAD = open(f"{CAT}/docker-compose.yml").read() +STEP1 = open(f"{CAT}/steps/eef2e4afe1218021.yml").read() +FY = open(f"{CAT}/.felhom.yml").read() + + +def set_compose(body): + return lambda: open(f"{CAT}/docker-compose.yml", "w").write(body) + + +def set_probe(port): + def _e(): + f = open(f"{CAT}/.felhom.yml").read() + open(f"{CAT}/.felhom.yml", "w").write(re.sub(r"(healthcheck:\n\s+checks:\n\s+- type: api\n\s+port: )\d+", lambda m: m.group(1) + str(port), f, count=1)) + return _e + + +out["drill_1"] = drill(set_compose(STEP1), "nextcloud template = ladder step 1's own definition (34.0.1 / mariadb 12.3)") +w.sync_rescan() +out["deployed"] = w.deploy(APP, SUB) +A = F.seed(w, SUB, w.say) +out["seed_A"] = bool(A) +out["drill_2"] = drill(lambda: (set_compose(HEAD)(), set_probe(8999)()), "nextcloud back to the head (34.0.4) + probe port 8999") +out["badge_wait"] = w.sync_rescan(APP, "nextcloud:34.0.4-apache", tries=60) +h = H.unescape(w.page(f"/apps/{APP}")) +out["steps_left_line_hu"] = re.findall(r'data-ladder-steps="(\d+)">([^<]*)<', h) +out["steps_left_line_en"] = re.findall(r'data-ladder-steps="(\d+)">([^<]*)<', H.unescape(w.page(f"/apps/{APP}?lang=en"))) +w.say(f"steps-left line hu={out['steps_left_line_hu']} en={out['steps_left_line_en']}") + +cut = {} + + +def on_phase(ph): + if ph == "verifying" and not cut: + cut["out"] = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={APP} | head -1); echo "copy: $c" +docker run --rm -v $c:/c alpine sh -c 'rm -f /c/felhom-undo-complete; ls /c | head -5'""") + w.say(" >>> finished-marker removed from one undo copy:\n" + cut["out"]) + + +t0 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +code, d = w.ctl("POST", f"/api/stacks/{APP}/update") +w.say(f"Update -> {code}") +phases, seen, tp = [], None, time.time() +while time.time() - tp < 1500: + st = w.stack(APP) + ph = st.get("update_phase") + if ph != seen and ph: + seen = ph + phases.append((round(time.time() - tp, 1), ph)) + w.say(f" +{phases[-1][0]:6.1f}s phase={ph}") + on_phase(ph) + if not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - tp > 3: + break + time.sleep(0.5) +held_at = time.time() +out["phases"], out["cutoff"] = phases, cut.get("out") +st = w.stack(APP) +out["api"] = {k: st.get(k) for k in ("state", "update_phase", "hold_reason", "hold_no_whole_copy", "update_error")} +w.say("API: " + json.dumps(out["api"], ensure_ascii=False)) +for lang in ("hu", "en"): + page = H.unescape(w.page(f"/apps/{APP}?lang={lang}")) + m = re.search(r'data-held="true">(.*?)', page, re.S) + out[f"app_page_{lang}"] = m.group(1).strip() if m else None + out[f"app_page_{lang}_restore_button"] = bool(m and 'href="/backups/apps"' in m.group(1)) + lst = H.unescape(w.page(f"/stacks?lang={lang}")) + m2 = re.search(r'data-held="true">(.*?)', lst, re.S) + out[f"stacks_page_{lang}_restore_button"] = bool(m2 and 'href="/backups/apps"' in m2.group(1)) + w.say(f"[{lang}] app page held block: {out[f'app_page_{lang}']!r} restore-button app={out[f'app_page_{lang}_restore_button']} list={out[f'stacks_page_{lang}_restore_button']}") +# ASCII-fragment checks, with controls (the ui-hungarian rule) +hu = out["app_page_hu"] or "" +out["ascii_checks"] = {"hu has 'gyfelszolgalat' (positive)": "gyfélszolgálat" in hu, + "hu has 'Visszaallithato a Mentesek' (must be absent)": "Visszaállítható a Mentések" in hu, + "en has 'Felhom support has been told' (positive)": "Felhom support has been told" in (out["app_page_en"] or "")} +w.say("checks: " + json.dumps(out["ascii_checks"])) + +# R-660: two and a half minutes of the dead-app scan after the hold +while time.time() - held_at < 150: + time.sleep(10) +out["log"] = w.guest(f"docker logs --since {t0} felhom-controller 2>&1 | grep -E 'nextcloud|DROPPED|dropped event' | grep -E 'NO copy|DROPPED|dropped event|HELD|UNDO|app_start_failed|no_whole' | tail -30") +w.say("controller log:\n" + out["log"]) +c2, dl = w.ctl("GET", "/api/debug/logs?level=DEBUG&lines=4000") +ents = ((dl.get("data") or {}).get("entries") or []) if isinstance(dl, dict) else [] +out["debug_ring_events"] = [e.get("timestamp", "") + " " + e.get("message", "") for e in ents + if ("app_start_failed" in e.get("message", "") or "app_hold_no_whole_copy" in e.get("message", "") + or "app_update_held" in e.get("message", "")) and e.get("timestamp", "") >= t0] +w.say("debug ring (events since the press):\n " + "\n ".join(out["debug_ring_events"])) +json.dump(out, open("../partB/01-round11-reproduced.json", "w"), indent=2, ensure_ascii=False, default=str) +open("../partB/01-round11-reproduced.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/liveC660.py b/documentation/audits/ladder-2026-09-24/tools/liveC660.py new file mode 100644 index 00000000..158e2460 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/liveC660.py @@ -0,0 +1,29 @@ +#!/usr/bin/env python3 +"""R-660 positive control: a throwaway app (actualbudget) stopped OUT OF BAND must raise app_start_failed +while the held nextcloud (held since 06:27:40Z) must not. Then the throwaway is removed through the product.""" +import json, sys, time +sys.path.insert(0, ".") +import walk as w +w.login() +t0 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +ok = w.deploy("actualbudget", "budget-c660") +w.wait_app("budget-c660", "/", tries=40, delay=3) +w.say("stopping the throwaway's container out of band (not through the product): " + w.guest("docker stop actualbudget 2>&1").strip()) +found = None +for i in range(24): + time.sleep(10) + c, d = w.ctl("GET", "/api/debug/logs?level=DEBUG&lines=6000") + ents = ((d.get("data") or {}).get("entries") or []) if isinstance(d, dict) else [] + evs = [e["timestamp"] + " " + e["message"] for e in ents if e.get("timestamp", "") >= t0 and + ("app_start_failed" in e.get("message", "") or "currently down" in e.get("message", ""))] + if any("app_start_failed" in x for x in evs): + found = evs + break +w.say("events/heartbeats since the control started:\n " + "\n ".join(found or evs)) +names = w.guest(f"docker logs --since {t0} felhom-controller 2>&1 | grep -iE 'app_start_failed|not running|DeadApp|nem fut' | tail -8") +w.say("controller log lines:\n" + names) +st = w.stack("nextcloud") +w.say(f"nextcloud meanwhile: state={st.get('state')} held={bool(st.get('hold_reason'))}") +w.remove("actualbudget") +json.dump({"events": found or evs, "log": names, "nextcloud_state": st.get("state")}, open("../partB/03-r660-control.json", "w"), indent=2) +open("../partB/03-r660-control.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/liveC660b.py b/documentation/audits/ladder-2026-09-24/tools/liveC660b.py new file mode 100644 index 00000000..4f2f78cc --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/liveC660b.py @@ -0,0 +1,26 @@ +import sys, time, re, html as H, json +sys.path.insert(0, ".") +import walk as w +w.login() +w.deploy("actualbudget", "budget-c660b") +w.wait_app("budget-c660b", "/", tries=40, delay=3) +t0 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +w.say("out-of-band stop: " + w.guest("docker stop actualbudget 2>&1").strip()) +seen = [] +for i in range(20): + time.sleep(5) + h = H.unescape(w.page("/")) + txt = re.sub(r"\s+", " ", re.sub("<[^>]+>", " ", h)) + for name in ("actualbudget", "Actual", "nextcloud", "Nextcloud", "gokapi", "Gokapi"): + pass + m = re.findall(r"(nem fut[^.]{0,200})", txt) + c, d = w.ctl("GET", "/api/debug/logs?level=DEBUG&lines=6000") + ents = ((d.get("data") or {}).get("entries") or []) + ev = [e["timestamp"] + " " + e["message"] for e in ents if e.get("timestamp", "") >= t0 and "app_start_failed" in e["message"]] + if m or ev: + seen.append({"t": i * 5, "banner": m[:3], "events": ev}) + w.say(f"+{i*5}s banner={m[:3]} events={len(ev)}") + if ev and m: + break +json.dump(seen, open("../partB/06-r660-control-banner.json", "w"), indent=2, ensure_ascii=False) +w.remove("actualbudget") diff --git a/documentation/audits/ladder-2026-09-24/tools/liveD.py b/documentation/audits/ladder-2026-09-24/tools/liveD.py new file mode 100644 index 00000000..ce5cf39a --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/liveD.py @@ -0,0 +1,71 @@ +#!/usr/bin/env python3 +"""Part D live — the ladder on controller v0.268.0. romm installed at 5.3.0 / mariadb 11.4 (the ladder's +step-1 definition, steps/7f9c6a74b5891f8e.yml, served as the drill template for the install only); the +drill template then back to the head (5.3.1 / 11.8). Two presses: press 1 must apply ONLY the app step +(5.3.1 / 11.4, from steps/90dd9d68258286ef.yml), press 2 the engine step (11.4 -> 11.8, the template). +Seeded data read back after each press; the page's steps-left line read in both languages each time.""" +import html as H, json, re, sys, time +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +F, SUB, APP = fx.Romm(), "arcade-ld", "romm" +CAT = f"{w.DRILL}/templates/{APP}" +out = {"controller": w.guest("cat /etc/felhom-controller-image").strip()} +w.login() + + +def drill(body, msg): + import fcntl + with open(f"{w.SC}/drill.lock", "w") as lk: + fcntl.flock(lk, fcntl.LOCK_EX) + w.sh(["git", "-C", w.DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120) + open(f"{CAT}/docker-compose.yml", "w").write(body) + w.sh(["git", "-C", w.DRILL, "commit", "-qam", "LIVE-D " + msg]) + r = w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + h = w.sh(["git", "-C", w.DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + w.say(f"drill {h}: {msg} push rc={r.returncode}") + return h + + +def steps_line(): + r = {} + for lang, q in (("hu", ""), ("en", "?lang=en")): + r[lang] = re.findall(r'data-ladder-steps="(\d+)">([^<]*)<', H.unescape(w.page(f"/apps/{APP}{q}"))) + return r + + +def box_files(): + return w.guest(f"grep -nE 'image:|memory:' /opt/docker/stacks/{APP}/docker-compose.yml; echo ---; cmp -s /opt/docker/stacks/{APP}/applied-compose.yml /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/{APP}/steps/90dd9d68258286ef.yml && echo 'applied == steps/90dd9d68258286ef.yml' || echo 'applied != steps/90dd…'; cmp -s /opt/docker/stacks/{APP}/applied-compose.yml /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/{APP}/docker-compose.yml && echo 'applied == template' || echo 'applied != template'") + + +HEAD = open(f"{CAT}/docker-compose.yml").read() +STEP1 = open(f"{CAT}/steps/7f9c6a74b5891f8e.yml").read() +out["drill_install"] = drill(STEP1, "romm template = step 1's own definition (5.3.0 / mariadb 11.4) for the install") +w.sync_rescan() +out["deployed"] = w.deploy(APP, SUB) +A = F.seed(w, SUB, w.say) +out["seed_A"] = bool(A) +out["drill_head"] = drill(HEAD, "romm back to the head (5.3.1 / mariadb 11.8)") +out["badge_wait"] = w.sync_rescan(APP, "mariadb:11.8", tries=60) +out["before"] = {"obs": w.observables(APP), "steps_line": steps_line(), "badges": w.badges(APP)} +w.say("before: steps line " + json.dumps(out["before"]["steps_line"], ensure_ascii=False) + " badges " + json.dumps(out["before"]["badges"], ensure_ascii=False)[:300]) + +for n in (1, 2): + t0 = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + p = w.press_update(APP) + obs = w.observables(APP) + w.wait_app(SUB, "/api/heartbeat", want=("200",), tries=60, delay=5) + rb = bool(F.verify(w, SUB, A, w.say)) + log = w.guest(f"docker logs --since {t0} felhom-controller 2>&1 | grep -E 'update romm' | grep -E 'ladder|pin advanced|DONE|UNDO|healthy after' | tail -8") + dblog = w.guest(f"docker logs --since {t0} romm-db 2>&1 | grep -iE 'upgrade|mariadb-upgrade|Version:' | tail -6") + B = F.seed(w, SUB, w.say) if n == 1 else None + out[f"press{n}"] = {"press": p, "obs": obs, "readback_A": rb, "log": log, "db_log": dblog, + "files": box_files(), "steps_line": steps_line(), "badges": w.badges(APP)} + if B: + out["seed_B"] = B + w.say(f"PRESS {n}: phase={p.get('final_phase')} pinned={obs.get('pinned_images')} readback A={rb}\n{log}\nDB: {dblog}\nfiles: {out[f'press{n}']['files']}\nsteps line: {out[f'press{n}']['steps_line']}") +if out.get("seed_B"): + out["readback_B_after_press2"] = bool(F.verify(w, SUB, out["seed_B"], w.say)) +json.dump(out, open("../partD/10-romm-two-steps.json", "w"), indent=2, ensure_ascii=False, default=str) +open("../partD/10-romm-two-steps.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/repoint.py b/documentation/audits/ladder-2026-09-24/tools/repoint.py new file mode 100644 index 00000000..c216a42f --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/repoint.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config. + +`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is +`controller.yaml.pre-ladder0924` (NOT the older `.pre-28`, which a restore must never pick up). +""" +import re, sys, io +sys.path.insert(0, '.') +import walk as w + +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" + + +def creds(): + for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"): + m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l) + if m: + return m.group(1), m.group(2) + raise SystemExit("no admin credential") + + +def to_drill(): + u, t = creds() + print(w.guest(f""" +set -e +test -f {VOL}/controller.yaml.pre-ladder0924 || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-ladder0924 +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml" +s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +if not re.search(r'^update:', s, re.M): + s += "update:\\n health_timeout: 90s\\n" +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -A2 '^update:' {VOL}/controller.yaml +""")) + + +def restore(): + print(w.guest(f""" +set -e +cp -p {VOL}/controller.yaml.pre-ladder0924 {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -c '^update:' {VOL}/controller.yaml || true +""")) + + +if __name__ == "__main__": + to_drill() if sys.argv[1] == "drill" else restore() diff --git a/documentation/audits/ladder-2026-09-24/tools/spike.py b/documentation/audits/ladder-2026-09-24/tools/spike.py new file mode 100644 index 00000000..cabbac95 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/spike.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Part D spike (30 min) on guest 9202, controller v0.267.0, drill catalog. + +1. an app two steps behind (vikunja A=2.4.0 -> B=2.5.0 -> C=2.6.0): does ONE press jump A -> C today? +2. does a `templates//steps/` folder reach the box? (it is read from the catalog clone) +Evidence only; every act is a product endpoint.""" +import json, os, sys, time +sys.path.insert(0, ".") +import walk as w + +DRILL = w.DRILL +CACHE = "/var/lib/docker/volumes/felhom-controller-data/_data/catalog-cache" + + +def commit(msg, edit): + edit() + w.sh(["git", "-C", DRILL, "add", "-A"]) + w.sh(["git", "-C", DRILL, "commit", "-q", "-m", msg]) + r = w.sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = w.sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + w.say(f"drill commit {h}: {msg} (push rc={r.returncode})") + return h + + +def set_img(ref): + p = f"{DRILL}/templates/vikunja/docker-compose.yml" + s = open(p).read() + import re + s = re.sub(r"image: vikunja/vikunja:\S+", f"image: vikunja/vikunja:{ref}", s) + open(p, "w").write(s) + + +out = {} +commit("SPIKE vikunja A=2.4.0", lambda: set_img("2.4.0")) +w.login() +w.sync_rescan("vikunja", "vikunja/vikunja:2.4.0") +ok = w.deploy("vikunja", "vikunja-sp") +out["deployed_A"] = ok +out["A"] = w.observables("vikunja") + + +def step_b(): + set_img("2.5.0") + os.makedirs(f"{DRILL}/templates/vikunja/steps", exist_ok=True) + open(f"{DRILL}/templates/vikunja/steps/spike-B.yml", "w").write( + open(f"{DRILL}/templates/vikunja/docker-compose.yml").read()) + + +commit("SPIKE vikunja B=2.5.0 + steps/spike-B.yml", step_b) +commit("SPIKE vikunja C=2.6.0", lambda: set_img("2.6.0")) +out["badge_wait_s"] = w.sync_rescan("vikunja", "vikunja/vikunja:2.6.0") +out["cache_steps"] = w.guest(f"git -C {CACHE} rev-parse --short=12 HEAD; git -C {CACHE} rev-list --count HEAD; ls -la {CACHE}/templates/vikunja/ {CACHE}/templates/vikunja/steps/ 2>&1; ls /opt/docker/stacks/vikunja/") +w.say("cache:\n" + out["cache_steps"]) +out["badges"] = w.badges("vikunja") +out["press"] = w.press_update("vikunja") +out["after"] = w.observables("vikunja") +w.say("after: " + json.dumps(out["after"])) +w.remove("vikunja") +json.dump(out, open("../partD/00-spike.json", "w"), indent=2, ensure_ascii=False) +open("../partD/00-spike.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/teardownB.py b/documentation/audits/ladder-2026-09-24/tools/teardownB.py new file mode 100644 index 00000000..462f9364 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/teardownB.py @@ -0,0 +1,21 @@ +import sys, json, re +sys.path.insert(0, ".") +import walk as w +w.login() +import fcntl +CAT = f"{w.DRILL}/templates/nextcloud" +with open(f"{w.SC}/drill.lock", "w") as lk: + fcntl.flock(lk, fcntl.LOCK_EX) + w.sh(["git", "-C", w.DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120) + f = open(f"{CAT}/.felhom.yml").read() + open(f"{CAT}/.felhom.yml", "w").write(re.sub(r"(healthcheck:\n\s+checks:\n\s+- type: api\n\s+port: )8999", r"\g<1>80", f, count=1)) + w.sh(["git", "-C", w.DRILL, "commit", "-qam", "LIVE-B teardown: nextcloud real probe back"]) + w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) +before = w.guest("docker volume ls -q | grep -E '^nextcloud' ; ls /opt/docker/stacks/nextcloud/") +w.say("before remove:\n" + before) +code = w.remove("nextcloud") +after = w.guest("docker volume ls -q | grep -E '^nextcloud' || echo 'no nextcloud volumes (incl. undo copies)'; ls -la /opt/docker/stacks/nextcloud/; ls -d /mnt/felhom-drives/scratch_hdd/userdata/nextcloud /mnt/felhom-drives/scratch_hdd/appdata/nextcloud 2>&1") +w.say("after remove:\n" + after) +rm = w.guest("rm -rf /mnt/felhom-drives/scratch_hdd/userdata/nextcloud /mnt/felhom-drives/scratch_hdd/appdata/nextcloud && ls -d /mnt/felhom-drives/scratch_hdd/userdata/nextcloud 2>&1") +w.say("drive folders removed by name (R-442 keeps them on 9202): " + rm.strip()) +open("../partB/07-teardown.log", "w").write("\n".join(w.LOG) + "\n") diff --git a/documentation/audits/ladder-2026-09-24/tools/walk.py b/documentation/audits/ladder-2026-09-24/tools/walk.py new file mode 100644 index 00000000..9d23c023 --- /dev/null +++ b/documentation/audits/ladder-2026-09-24/tools/walk.py @@ -0,0 +1,507 @@ +#!/usr/bin/env python3 +"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints. + +EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses: + POST /api/stacks//deploy · POST /api/sync · POST /api/stacks/rescan + POST /api/stacks//update · POST /api/stacks//remove +and reads GET /api/stacks/. No controller code exists for it. + +The walk, per `09` §6.4 and the update-night brief §4: + 1 deploy from the DRILL catalog at the LIVE pin + 2 seed through the app's OWN front door (R-156: never a volume, never SQL) + 3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing + 4 „Mentés most" + 5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages + 6 press the guarded Update, record every phase with timestamps + 7 read the seed back through the front door + 8 the four version observables side by side + 9 write the verdict record in `09`'s JSON shape + +`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`. +""" +import argparse, json, os, re, subprocess, sys, time +from datetime import datetime, timezone + +SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/25947a06-40b7-43dd-8334-2519e02142da/scratchpad" +EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/ladder-2026-09-24" +DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest. +GUEST = os.environ.get("GUEST", "9202") +BASE = {"9202": "https://192.168.0.114", "9201": "https://192.168.0.138"}[GUEST] +DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu") +HOSTHDR = f"Host: felhom.{DOMAIN}" +HP = "demo-hp" + +LOG = [] + + +def say(*a): + line = " ".join(str(x) for x in a) + ts = datetime.now().strftime("%H:%M:%S") + print(f"{ts} {line}", flush=True) + LOG.append(f"{ts} {line}") + + +def sh(args, timeout=300, inp=None): + try: + return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp) + except (subprocess.TimeoutExpired, OSError) as e: + return subprocess.CompletedProcess(args, 124, "", f"{e}") + + +def guest(script, timeout=600): + """Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting).""" + # ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w.sh swapped scripts under + # two concurrent walks (memory: guest-helper-shares-one-tmp-file). + import secrets as _s + t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh" + r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP, + f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; " + f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"], + timeout=timeout, inp=script) + return r.stdout or "" + + +def login(): + pw = open(f"{SC}/.ctlpw").read().strip() + sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR, + "-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"]) + h = open(f"{SC}/hdr{os.getpid()}.txt").read() + m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I) + if not m: + sys.exit("login failed: no session cookie") + open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0)) + r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"]) + c = re.search(r'/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN. + + Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because + a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password + (grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only + reason this was safe to discover by running it (live-probes rule). + + A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on + the scratch drive first — the same act the drive browser performs for a household. + """ + code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields") + fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or [] + values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub} + made = [] + for f in fields: + ev, ty = f.get("env_var"), f.get("type") + if ev in values: + continue + # `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses + # when the caller sends none, deliberately ("the user needs to know their password"), + # while `.felhom.yml` declares `required: false` and the API serves that verbatim. A + # caller that trusts the contract gets a 400. Measured tonight on grafana; filed. + if not f.get("required") and ty != "password": + continue # the controller generates the optional secrets itself + if ty == "path": + p = f"{DRIVE}/{name}" + values[ev] = p + made.append(p) + elif ty in ("secret", "password"): + import secrets as _s + values[ev] = "Drill-" + _s.token_hex(12) + GENERATED.setdefault(name, {})[ev] = values[ev] + elif f.get("default"): + values[ev] = f["default"] + else: + values[ev] = f"drill-{name}" + if made: + guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made)) + say(f" [1] made the drive paths this app requires: {made}") + extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")] + if extra: + say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}") + return values + + +def deploy(name, sub, extra_values=None): + st = stack(name) + if st.get("deployed"): + say(f" [1] {name} already deployed — reusing") + return True + values = deploy_values(name, sub) + if extra_values: + values.update(extra_values) + code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values}) + say(f" [1] deploy -> {code} {str(d)[:120]}") + if code != "202": + return False + # WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the + # container `healthy` while the controller's own state read `unhealthy` — a gate on `running` + # alone therefore times out on an app that is up. The state is RECORDED rather than required; + # the real gate is the fixture's own `wait_app`, which asks whether the APP answers. + seen = None + for _ in range(90): + time.sleep(5) + st = stack(name) + seen = st.get("state") + # `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21: + # tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm + # read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the + # deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark + # (`runComposeDeploy` writes it), so that is what to wait for. + pins = (st.get("app_config") or {}).get("pinned_images") + if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"): + say(f" [1] deployed, controller state={seen}, " + f"pinned={(st.get('app_config') or {}).get('pinned_images')}") + if seen != "running": + say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, " + f"not treated as a failure; the fixture's front-door wait is the real gate") + return True + say(f" [1] never became deployed (last controller state={seen!r})") + return False + + +def backup_now(name): + """R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever. + + `POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and + restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product + has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup + exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one + app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its + phase list. A seed written "after the backup" is therefore written before the update's own backup + — the undo's last-second copy is still the one that must bring it back.""" + say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)") + return None + +def drill_bump(app, frm, to, service_hint=None): + """Serialised across concurrent walks: one git working tree, one lock.""" + import fcntl + with open(f"{SC}/drill.lock", "w") as lk: + fcntl.flock(lk, fcntl.LOCK_EX) + sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120) + return _drill_bump(app, frm, to, service_hint) + + +def _drill_bump(app, frm, to, service_hint=None): + """Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates). + + `frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in + two images (adventurelog's backend and frontend) moves both in one edge, while its engine + sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move + every service that carries the app's own version and no others. + """ + comp = f"{DRILL}/templates/{app}/docker-compose.yml" + fy = f"{DRILL}/templates/{app}/.felhom.yml" + s = open(comp).read() + froms = [x.strip() for x in frm.split(",") if x.strip()] + tos = [x.strip() for x in to.split(",") if x.strip()] + if len(froms) != len(tos): + say(f" [5] from/to lists differ in length: {froms} vs {tos}") + return None + for f1, t1 in zip(froms, tos): + if f"image: {f1}" not in s: + say(f" [5] FROM ref not found in compose: {f1}") + return None + s = s.replace(f"image: {f1}", f"image: {t1}") + open(comp, "w").write(s) + f = open(fy).read() + today = datetime.now().strftime("%Y-%m-%d") + f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M) + open(fy, "w").write(f) + sh(["git", "-C", DRILL, "add", "-A"]) + sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"]) + r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})") + return h + + +def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5): + """Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP. + + R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and + `catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not + enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the + Update that followed moved nothing and still reported "Frissitve". So when the caller knows + which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the + NUMBER R-607 asks for and has never had. + """ + t0 = time.time() + ctl("POST", "/api/sync") + time.sleep(2) + ctl("POST", "/api/stacks/rescan") + time.sleep(2) + if not expect_app or not expect_ref: + return None + for i in range(tries): + cat = stack(expect_app).get("catalog_images") or {} + if expect_ref in cat.values(): + waited = round(time.time() - t0, 1) + if i: + say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up " + f"to {expect_ref} — R-607's window, measured") + return waited + time.sleep(delay) + ctl("POST", "/api/sync") + time.sleep(1) + ctl("POST", "/api/stacks/rescan") + say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — " + f"catalog_images = {stack(expect_app).get('catalog_images')}") + return None + + +def badges(name): + out = {} + for lang, suffix in (("hu", ""), ("en", "?lang=en")): + h = page(f"/apps/{name}{suffix}") + m = re.findall(r']*title="([^"]*)"[^>]*>([^<]*)<', h) + out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3] + return out + + +def press_update(name, poll=1.0, cap_s=1800): + code, d = ctl("POST", f"/api/stacks/{name}/update") + say(f" [6] Update -> {code} {str(d)[:220]}") + if code not in ("202", "200"): + return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0} + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < cap_s: + st = stack(name) + ph = st.get("update_phase") + if ph != seen: + seen = ph + rec = {"t": round(time.time() - t0, 1), "phase": ph, + "label": st.get("update_phase_label"), "updating": st.get("updating"), + "error": st.get("update_error"), "hold": st.get("hold_reason")} + phases.append(rec) + say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} " + f"err={rec['error']} hold={rec['hold']}") + if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3: + break + time.sleep(poll) + st = stack(name) + return {"accepted": True, "http": code, "phases": phases, + "duration_s": round(time.time() - t0, 1), + "final_phase": st.get("update_phase"), "update_error": st.get("update_error"), + "hold_reason": st.get("hold_reason"), "state": st.get("state")} + + +def observables(name): + st = stack(name) + ac = st.get("app_config") or {} + live = guest(f""" +grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//' +echo '---inspect---' +for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do + echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}' +done +""") + a, _, b = live.partition("---inspect---") + return { + "pinned_images": ac.get("pinned_images"), + "installed_images": {k: (v.get("ref") if isinstance(v, dict) else v) + for k, v in (ac.get("installed_images") or {}).items()}, + "catalog_images": st.get("catalog_images"), + "live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()], + "docker_inspect": [x for x in b.strip().splitlines() if x.strip()], + } + + +def app_logs(name, lines=400): + """The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is + one string with escaped newlines — a scan over the envelope sees a single enormous line and + finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96 + rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines.""" + code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}") + if isinstance(d, dict): + data = d.get("data") + if isinstance(data, dict) and isinstance(data.get("logs"), str): + return data["logs"] + if isinstance(d.get("_raw"), str): + return d["_raw"] + return str(d) + + +def write_verdict(rec, appdir): + os.makedirs(appdir, exist_ok=True) + p = os.path.join(appdir, "verdict.json") + json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False) + say(f" [9] verdict {rec['verdict']} -> {p}") + + +def remove(name): + """Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint + refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up.""" + c1, d1 = ctl("POST", f"/api/stacks/{name}/stop") + say(f" [X] stop -> {c1} {str(d1)[:100]}") + for _ in range(24): + time.sleep(5) + if stack(name).get("state") != "running": + break + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": True, "remove_backups": True}) + say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}") + if code == "409": + # R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive + # path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202 + # `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits + # this. The household's other choice — remove the app, KEEP the data — is accepted, and the + # harness takes it, then tidies its own directory by name at teardown. + say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)" + " — removing the app and KEEPING the drive data instead") + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": False, "remove_backups": True}) + say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}") + time.sleep(5) + st = stack(name) + left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; " + f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'") + say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}") + return code + + +def app_env(name, key): + """Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the + app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller + shows them the same value. Data still goes in through the app's own front door.""" + out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1") + if ":" in out: + return out.split(":", 1)[1].strip().strip('"').strip("'") + return "" + + +def snapshots(name): + """The restorable copies the backups page offers for this app.""" + code, d = ctl("GET", f"/api/backup/snapshots?stack={name}") + data = d.get("data") if isinstance(d, dict) else None + if isinstance(data, dict): + for k in ("snapshots", "items", "restore_points"): + if isinstance(data.get(k), list): + return data[k] + return data if isinstance(data, list) else [] + + +def restore(name, snapshot_id=None, wait_s=1200): + """The household's own way out: the „Visszaállítás a mentésből" button on the backups page. + + A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`, + `snapshot_id` — because that is the button the sentence tells them to press. + """ + snaps = snapshots(name) + if snapshot_id is None: + if not snaps: + say(f" [R] no restorable copy offered for {name}") + return {"ok": False, "why": "no snapshot offered", "snapshots": snaps} + first = snaps[0] + snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id") + say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)") + sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip() + csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip() + r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}", + "-X", "POST", + "--data-urlencode", f"_csrf={csrf}", + "--data-urlencode", f"stack_name={name}", + "--data-urlencode", f"snapshot_id={snapshot_id}", + f"{BASE}/backup/restore"], timeout=180) + head = (r.stdout or "").split("\n")[0].strip() + loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")] + say(f" [R] POST /backup/restore -> {head} {loc[:1]}") + t0 = time.time() + last = None + while time.time() - t0 < wait_s: + code, d = ctl("GET", "/api/backup/restore-status") + dd = d.get("data") or {} + cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message")) + if cur != last: + say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}") + last = cur + if not dd.get("running", False) and time.time() - t0 > 5: + break + time.sleep(2) + st = stack(name) + say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} " + f"phase={st.get('update_phase')}") + return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps, + "http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1), + "state_after": st.get("state"), "hold_after": st.get("hold_reason"), + "observables_after": observables(name)} diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 3b0b240f..3359feca 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -362,3 +362,17 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-640** | **A truncated PostgreSQL copy loaded with rc 0 into an empty database (P2, narrowed to the restore paths).** Closed in **v0.267.0**: `appbackup.CheckDumpComplete` (the engine's end marker); the unit restore and the off-site restore refuse before the first mutation, every replay checks again before any load. **Rule:** a dump is judged by its END — the header and a `CREATE TABLE` say nothing about whether it finished. | **CLOSED 2026-09-23 — three red-proofs + PROVEN-LIVE on 9202 (the household's restore button refused a half-length docmost copy; containers untouched; the whole copy then restored)** | `audits/night-2026-09-23/A2-*`, `E1-r640-live.*` | | **R-499** | **Every driveless app was told its data was „already in the full system backup (PBS)" (P2).** Closed in **v0.267.0**: the Tier-2 page's sentence has four branches from the box's own whole-system backup target (own drive / same disk / drive gone / cannot ask); „(PBS)" and „nincs külön teendő" only where true. | **CLOSED 2026-09-23 — two red-proofs; live on 9202 (the `unknown` branch, hu + en, matching `/api/storage/backup-target`)** | `audits/night-2026-09-23/A4-*`, `A7-*` | | **R-626** | **A removed app came back (P2).** Measured on v0.266.0, NOT reproduced: navidrome removed through the product, 390 s of `docker events` (the remove's destroy seen, no create), a controller restart at +150 s, a guest reboot after → no container, no volume. The two known creators (the restore/remove race R-633, the backup/deploy race R-634) are fixed. **Rule kept from the row:** a check that runs once, immediately, cannot see a thing created just after it — watch a window. Leftover found and filed: R-651. | **CLOSED 2026-09-23 — by measurement (a positive observable: the destroy events)** | `audits/night-2026-09-23/A3-*` | + +## 2026-09-24 — the undo after a restore, the held app's page, the ladder (controller v0.268.0, hub v0.122.0, catalog `5ed599c`) + +| Row | What | Closed | Full text | +|---|---|---|---| +| **R-658** | **After a restore, the undo copied NOTHING (P1).** v0.268.0 (`206b035`): the undo selects volumes from the rendered compose file (`DeclaredVolumeNames`: `name:` else `_`, each checked to exist); the label is a logged cross-check. The unit restore creates volumes WITH compose's project/volume/version labels (never a guessed config-hash); the remove counts unlabelled declared volumes. Live on 9202: restored under v0.267.0 → labels `null`; on v0.268.0 a failing update copied both by name and seeds A and B (B written after the restore) read back; a v0.268.0 restore → labels present, no compose warning. **Rule:** never select an app's volumes by the compose label. | v0.268.0, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-659** | **A held app's page named a way back the restore refused (P1).** Operator ruling 2026-09-24 (`09` §3 decision 25, option A). v0.268.0 + hub v0.122.0: the hold names the newest WHOLE copy (`WholeOnTier` asks the refusal's own predicate); with none, `hold.update.no_whole_copy` in the household's language, no Mentések button, and `app_hold_no_whole_copy` (critical, operator-only). Live on 9202: round 11 reproduced (nextcloud, cut-off undo copy) — both languages, no button on either page, the event's R-620 line. **Rule:** a sentence may name a copy as a way back only when the restore for that copy would accept it. Left open beside it: R-661, R-666. | v0.268.0 / hub v0.122.0, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-660** | **A held app also raised `app_start_failed` (P3).** v0.268.0: a fourth suppression set at `classifyRunStates` — the update-held apps (`backup.UpdateHeldStacks`; a RESTORE hold is not in it). Live: after the hold no `app_start_failed` in ~12 min of scans, while a throwaway stopped out of band raised exactly one. | v0.268.0, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-651** | **A removed app left `applied-compose.yml` and `applied-meta/` (P3).** v0.268.0: remove deletes both (the catalog mirror and `hold-logs/` stay — the latter is evidence). Live: present before, gone after, on nextcloud and vikunja. | v0.268.0, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-653** | **The memory watch wrote `proven` over a watch whose load never reached the app (P3).** Catalog `5ed599c`: `load_verdict` — `reached` only when at least half the requests got an HTTP answer, else the edge is `inconclusive`. Unit-tested and red-proofed; no bench run this session. | catalog, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-656** | **The bench re-used an app's scratch drive folder (P3).** Catalog `5ed599c`: `clear_scratch_folders` removes the app's own `${HDD_PATH}`/`${USERDATA_PATH}`/`${IMPORT_PATH}` bind folders before FROM and says so; never a bare root, never outside. Unit-tested and red-proofed; no bench run this session. | catalog, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-40** | **The update path could not express a multi-hop upgrade (P2).** Superseded by `09` §3 decisions 13–14 and shipped as §6.4 part 5 (v0.268.0): the catalog records every tested step with its own definition, and one press climbs one. **Rule kept:** a hop exists for a box only when the catalog holds it as a TESTED step — a >1-major catalog move without its intermediate steps is still a jump for a box below it. | v0.268.0, 2026-09-24 | `git show 500488cad673:documentation/backlog/OPEN-ITEMS.md` | +| **R-663** | **Two catalog test suites had been red since the night of 2026-09-23 (P3).** Filed and fixed the same session: `test_gate_decoys.py` (kimai-db 11.6 → 11.8) and `test_ladder_writer.py` (navidrome 0.64.0 → 0.64.1) typed the live pins as literals; they now READ them. 84 decoy cases OK. **Rule:** a fixture that names a live pin reads it; the catalog moves under it every night. | catalog `5ed599c`, 2026-09-24 | this entry | + diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b636ec06..c55bd4b2 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -633,7 +633,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-31** | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-32** | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | | **R-35** | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | -| **R-40** | **[P2-HIGH] The update path cannot express a MULTI-HOP major upgrade.** A template pin is a single value; the customer's update button pulls whatever the catalog now says. For apps whose upstream forbids version skipping this produces a broken upgrade. Nextcloud is explicit: *"You cannot skip major releases. Please re-run the upgrade until you have reached the highest available release."* Campaign 7 moved its template **31 → 34** (a fresh deploy validates fine — 302, 3/3 healthy), so an existing 31 customer pressing update would attempt a jump Nextcloud refuses. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Origin: CAMPAIGN 7 (`audits/CAMPAIGN-7-catalog-sweep-2026-07-19.md` §7 F7). Not nextcloud-only — any app with sequential-major rules (gitea, tandoor, outline…) has the same shape. Directions: a per-app `upgrade_path:`/`max_hop:` in `.felhom.yml` that the update button walks in stages; or refuse-and-explain when the installed major is >1 behind; or pin an intermediate "stepping-stone" tag. **Until this exists, a >1-major catalog bump is safe for NEW deploys and unsafe for the update button** — which is exactly the asymmetry the campaign's MAJOR flag was meant to record but cannot enforce | CC | | **R-49** | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-76** | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | | **R-78** | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | @@ -667,7 +666,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | | **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. **-- RULED 2026-09-23 (`09` §3 decision 17):** YES — the catalog records the image digest of every pin at push time; the box compares against it and, where the catalog carries one, pulls **that exact image**, which makes a floating tag reproducible, not only the badge honest. *Pull-by-digest while the definition names a tag is a claim to verify in the build, not a ruling on mechanism.* **-- 2026-09-23:** pull-by-digest MEASURED on 9202 — Docker and Compose both pull and run `redis:7-alpine@sha256:858f…` and refuse a digest that does not exist. **Build trap, read from source:** `splitImageRef` returns "unorderable" for any ref containing `@` (`stacks/updateorder.go:134`), so a digest-carrying pin must have its digest split off before ordering or every such app reads Unknown. `09` §6.4 part 6. **— NIGHT 2026-09-23:** the CATALOG half of the close shipped: every ladder entry records the digest the registry served for each `to` ref (`scripts/image_digest.py`), and the move gate refuses a digest the registry no longer serves. The box does not compare it yet (`09` §6.4 part 6, box half). | **READY TO BUILD — owner: CC; `09` §6.4** | -| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. **-- SPIKED 2026-09-23 (`audits/update-rulings-2026-09-23/`):** the undo works by hand on three real migrating edges and needs eight product additions (R-637..R-642); **the ladder is measured absent** — one press on a box two steps behind jumped vikunja 2.3.0 → 2.5.0 in 9.5 s and 2.4.0 never ran, and the box cannot see intermediate steps at all because its catalog clone is `--depth 1` (`sync.go:283`/`:300`, one commit visible on both demo guests). The ladder's recommended format is an `update_ladder:` list in `.felhom.yml` with each intermediate step's own definition, NOT the git history (romm's image-moving commit is the definition that OOM-looped). The chain's update leg has ≤15 min as ruled (R-643). Build order `09` §6.4. **— NIGHT 2026-09-23: `09` §6.4 part 4 SHIPPED** (catalog `6db08a5`): the test record `update_ladder:` + two gates + the only writer + the 21-move backfill; 12 more steps published with records. Parts 5 (the box climbs), 6-box-half and 7 remain; the romm press on demo-hp showed today's jump live — 5.3.0 → 5.3.1 AND mariadb 11.4 → 11.8 in one press (both tested steps; `done`). | **READY TO BUILD — owner: CC; `09` §6.4 part by part, each part returns to the operator for go/no-go** | +| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. **-- SPIKED 2026-09-23 (`audits/update-rulings-2026-09-23/`):** the undo works by hand on three real migrating edges and needs eight product additions (R-637..R-642); **the ladder is measured absent** — one press on a box two steps behind jumped vikunja 2.3.0 → 2.5.0 in 9.5 s and 2.4.0 never ran, and the box cannot see intermediate steps at all because its catalog clone is `--depth 1` (`sync.go:283`/`:300`, one commit visible on both demo guests). The ladder's recommended format is an `update_ladder:` list in `.felhom.yml` with each intermediate step's own definition, NOT the git history (romm's image-moving commit is the definition that OOM-looped). The chain's update leg has ≤15 min as ruled (R-643). Build order `09` §6.4. **— NIGHT 2026-09-23: `09` §6.4 part 4 SHIPPED** (catalog `6db08a5`): the test record `update_ladder:` + two gates + the only writer + the 21-move backfill; 12 more steps published with records. Parts 5 (the box climbs), 6-box-half and 7 remain; the romm press on demo-hp showed today's jump live — 5.3.0 → 5.3.1 AND mariadb 11.4 → 11.8 in one press (both tested steps; `done`). **— 2026-09-24: `09` §6.4 PART 5 SHIPPED** (controller v0.268.0 `206b035`, catalog `5ed599c`): one press = one tested step, each step's own definition at `templates//steps/.yml`, proven live on 9202 (romm 5.3.0/11.4 → 5.3.1/11.4 → 5.3.1/11.8 in two presses, `audits/ladder-2026-09-24/partD/`). Left in this row: part 6's box half, part 7 (the automatic leg), part 10 (PostgreSQL majors). | **READY TO BUILD — owner: CC; `09` §6.4 part by part, each part returns to the operator for go/no-go** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. **-- RULED 2026-09-23 (`09` §3 decision 18):** the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. **Built later, when the fleet grows** — Q7's recommendation, confirmed. | **RULED — build deferred until the fleet grows; owner: CC** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | @@ -803,16 +802,16 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** | | **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** | | **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | -| **R-651** | **[P3-LOW] A removed app leaves `applied-compose.yml` and `applied-meta/` behind in its stack directory.** MEASURED 2026-09-23 night on 9202 (v0.266.0, R-626's run): after `POST /api/stacks/navidrome/remove` answered 200 with `verified: true`, `/opt/docker/stacks/navidrome/` still held `applied-compose.yml` and `applied-meta/` written by the deploy (the directory itself is the catalog mirror and belongs there). Seen again after every remove of the night (`after remove: leftovers='/opt/docker/stacks/'`). **Consequence not measured:** whether a later reinstall of the same app reads a stale `applied-meta/` (the undo's OLD `.felhom.yml`, v0.263.2) before its own deploy overwrites it. **Needs:** the remove deletes both; a test that removes, reinstalls, fails an update and asserts the undo uses the reinstall's file. Evidence: `audits/night-2026-09-23/A3-r626-verdict.json`. | **READY — P3; owner: CC (controller)** | | **R-652** | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** | -| **R-653** | **[P3-LOW] The memory watch's load did not reach box-fixture apps: every request errored.** MEASURED 2026-09-23 night: the load generator followed each app's redirect to `https://.gate.invalid`, which does not resolve, so ghost's and nextcloud's first ten-minute watches recorded **11 797 and 9 427 requests, every one `err`** — memory measured, but idle. Fixed the same night in `upgrade-test.py` (no redirects; the app's own Host header): ghost's re-run served **10 538 × 200**. **Why a row though fixed:** the harness wrote `proven` over a watch with zero effective load and nothing flagged it. **Needs:** the watch reports INCONCLUSIVE when fewer than half its requests reached the app. Evidence: `apps/ghost/bench-noload/`, `apps/ghost/bench/`. | **READY — P3; owner: CC (catalog harness)** | | **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** | | **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** | -| **R-656** | **[P3-LOW] The bench keeps an app's drive folder between runs, so a re-run of the same app starts on the last run's files.** MEASURED 2026-09-23 night: nextcloud's second bench run never installed (`occ status: installed: false` for 5 min) because `/srv/felhom-gate/hdd/appdata/nextcloud` still held the first run's `config/` and data — `compose down -v` removes named volumes, not bind-mounted host folders. Cleared by hand and re-run. Other re-runs tonight (romm's engine step after its app step, komga, immich) reused their drive folders too; their seeds live in databases in named volumes, so their verdicts stand, but the class is unguarded. **Needs:** the bench clears the app's scratch drive folder before FROM (and says so), or runs each edge under a fresh folder. | **READY — P3; owner: CC (catalog harness)** | | **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** | -| **R-658** | **[P1-HIGH] After a restore, an app's volumes lose their compose labels — so the next update's automatic undo copies NOTHING.** FOUND 2026-09-23 night by the chaos hour on 9202 (v0.267.0). The unit restore recreates each named volume with a bare `docker volume create ` (`internal/backup/restore.go:154`), which carries no `com.docker.compose.project` label; the undo selects the volumes it copies by exactly that label (`internal/stacks/undo.go:156`, `ProjectVolumes`). **Measured, with a control:** the three apps restored in the chaos hour (docmost, navidrome, vikunja) have **0** labeled volumes of 3 / 2 / 2; the three never restored (romm, nextcloud, adventurelog) have all of theirs labeled. **Consequence, seen live in round 9:** vikunja's broken update logged `the undo copy will hold 0 named volume(s), 0.0 MiB`, failed health, and was „UNDONE" — the OLD binary put back on the NEW version's migrated data, with no copy restored. vikunja happens to start on migrated data (the morning spike), so it read back; **an app whose old version refuses migrated data would HOLD, and one that starts but mis-reads it would run on a database the old code does not understand, reported as a successful undo.** The same label filter is used by the remove path (`delete.go:994`, `projectVolumes`) — **measured at teardown:** the remove of the three restored apps DID delete their volumes (compose's `down --volumes` finds them by name) but answered `volumes_removed: []` — the report is wrong, not the act. **Needs (not fixed tonight — one controller release per session):** the restore creates volumes WITH the project's labels (or `compose create` does it); the undo and the remove resolve an app's volumes from its compose file's declared names, not only the label; a test that restores, then fails an update, and asserts the undo copied every declared volume. Evidence: `audits/night-2026-09-23/chaos/round-09.*`, `chaos/09-volumes-unlabeled-after-restore.txt`. | **READY — P1; owner: CC (controller)** | -| **R-659** | **[P1-HIGH] A held app's page names a way back that the restore then REFUSES — and no screen offers another.** MEASURED 2026-09-24 00:00 on 9202 (v0.267.0), chaos round 11: nextcloud's update failed, its undo copy was cut off (the drill removes the finished-marker, `09` §6.1a's measured case), and the product HELD it — honestly, in both languages: *„… Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 22:55 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."* The household presses exactly that restore (`POST /backup/restore`, snapshot `helyi`, the ONE copy offered) and is refused: *„Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"."* This box has no remote copy. The app stays stopped and held; Update and Start refuse a held app. **So the hold sentence and the restore disagree about the same copy, and on a box without an off-site tier nothing on any page brings the app back.** Both refusals are individually right (the restore will not pour a database over files it cannot also restore; the hold will not start a half-migrated app); together they strand a file-leg app. navidrome, held the same way in round 8, came back through the same button because its copy was accepted. **The brief's stop rule fired here** (a state neither the product nor its screens can recover): Part D had finished all twelve rounds; nothing further was injected. **Needs:** the hold must name only a route the restore will accept for that app (or say plainly that none exists on this box and the operator is needed); and a decision on what a household with no off-site tier may do for a file-leg app whose undo failed. Evidence: `audits/night-2026-09-23/chaos/round-11.*`, `chaos/round-11-recovery.*`. | **READY — P1; owner: operator (the promise) / CC (the build)** | -| **R-660** | **[P3-LOW] A held app also raises `app_start_failed`, seconds after its own `app_update_held`.** MEASURED 2026-09-23 night on 9202 (chaos rounds 8 and 11, from the controller's full log): `app_update_held (error)` at 20:56:04 and 21:42:39, then `app_start_failed (warning)` at 20:56:15 and 21:42:52 — each at the next status refresh after the hold stopped the app. The `DROPPED event` line names no app, so the attribution is by timing, not by name; nothing else changed state at those moments. The hold is a deliberate stop BY THE PRODUCT, a fourth way to stop an app that `classifyRunStates`' three suppressions (`08` §5) do not know — the same class as the nightly volume-dump false alarm (fixed v0.224.0: "a third way to stop an app needs a third suppression set"). The operator gets two alarms for one fact. **Needs:** a held app is not "down" to the app-down ladder; a test that holds an app and asserts no `app_start_failed`. Evidence: `audits/night-2026-09-23/E8-events-*`. | **READY — P3; owner: CC (controller)** | +| **R-661** | **[P2-MEDIUM] The second drive holds a file app WHOLE, and no single action brings the app back from it.** MEASURED FROM SOURCE 2026-09-24 (v0.267.0/v0.268.0): the Tier-2 mirror carries the unit AND the drive files, but „Teljes visszaállítás" there is `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`, whose R-538 guard refuses any app with declared drive files; the file restore (`RestoreTier2Files`) only ADDS missing files and replays no database. So for nextcloud/immich/paperless-ngx/calibre-web the second drive is not a way back after a failed update — which is why R-659's hold (v0.268.0) names only the off-site copy for those apps (`backup.WholeOnTier`). **Needs:** a decision on a Tier-2 "files + database" restore (the off-site full restore's shape), or a ruling that the second drive is files-only for these apps. Evidence: `audits/ladder-2026-09-24/README.md` (the truth table); `backup/update_guard.go` R-659 block. | **WAITING-ON-OPERATOR — P2; owner: operator (the promise) / CC (the build)** | +| **R-662** | **[P3-LOW] The unit restore's documented second step („csak az adatbázist és a beállításokat", `accept_missing_files=1`) is dead.** MEASURED FROM SOURCE 2026-09-24: no template sends the field (`grep accept_missing_files internal/web/templates` → 0), and when a request does, the web handler skips its own refusal but calls `RestoreFromRecoveryUnit`, which passes `UnitRestoreOptions{}` — so the backup layer refuses anyway. Nothing ever calls `AcceptMissingFiles: true` outside a test. **Needs:** either wire it end to end (with its own confirm, R-48) or remove the flag and the comment that promises it. Evidence: `web/handlers.go` backupRestoreHandler; `backup/restore_unit.go` L226. | **READY — P3; owner: CC (controller)** | +| **R-664** | **[P3-LOW] A ladder step has no `.felhom.yml` of its own.** v0.268.0 pins a step's own COMPOSE file (`steps/.yml`); its health probe, memory check and the recorded `applied-meta/` still come from the catalog's CURRENT `.felhom.yml`. Harmless while a step's probe and limits equal the head's (true for all 8 step files today); wrong the day a step's probe differs. Also: the log line for a step pin says „pin advanced to the catalog's current definition" (it names the step file one line earlier). **Needs:** `steps/.felhom.yml` (or the probe in the ladder entry) and the log line naming the source. Evidence: `audits/ladder-2026-09-24/partD/10-romm-two-steps.log`. | **READY — P3; owner: CC (controller + catalog)** | +| **R-665** | **[P3-LOW] An update pressed soon after a restore was judged with the RESTORED `.felhom.yml`.** OBSERVED 2026-09-24 on 9202 (v0.268.0), not diagnosed: vikunja restored from its unit at 06:19; the drill catalog's `.felhom.yml` named probe port 8999 (a deliberately failing edge); the periodic probe used 8999 until 06:20:41, and the update pressed at 06:21:07 probed 3456 (the unit's own file) and ended `done`. The restore apparently writes the unit's `.felhom.yml` into the stack dir, and the next catalog sync has not yet put the catalog's back. Consequence bounded (the OLD probe, ≤ 15 min), but an update's verdict should not depend on how long ago a restore ran. **Needs:** a measurement of what the restore writes and when the sync overwrites it. Evidence: `audits/ladder-2026-09-24/partA/03-why-g-was-done.txt`. | **READY — P3; owner: CC (controller)** | +| **R-666** | **[P3-LOW] A held app whose box has no whole copy is told „ne törölje az alkalmazást" — and its card still offers Eltávolítás.** v0.268.0 (R-659) hides the Mentések button beside the no-whole-copy sentence; the Remove button on the stacks card stays (the household may remove its own app — a promise, not CC's to take away). Measured live on 9202: the page and the sentence disagree. **Needs:** an operator word — hide Remove while support is informed, or soften the sentence. Evidence: `audits/ladder-2026-09-24/partB/01-round11-reproduced.json`. | **WAITING-ON-OPERATOR — P3; owner: operator** | +| **R-667** | **[P2-MEDIUM] A crash-looping app whose container reads `running` for a moment between restarts never reaches the crash-loop alarm.** MEASURED 2026-09-24 on 9202 (v0.268.0): `gokapi` (R-644) at **385 restarts**, `restarting_since` = 06:46:30Z at 06:47 — the clock resets whenever a scan catches the container up — while the dead-app heartbeat said *„4 deployed app(s) evaluated, 0 currently down"*. `Stack.CrashLooping` needs `crashLoopAfter` (5 min) of CONTINUOUS `restarting`. Likely also the unattributed second `app_start_failed` at 06:31:41Z (a scan that caught gokapi `exited`): the dropped-event line names no app, so this is inference. **Needs:** crash-loop judged on the container's RestartCount growth over a window, not on an uninterrupted state; a test with a flapping fixture. Evidence: `audits/ladder-2026-09-24/partB/04-states-and-banner.txt`, `02-r660-positive-observables.txt`. | **READY — P2; owner: CC (controller)** |