diff --git a/CONTEXT.md b/CONTEXT.md index a6dbd571..4febc269 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,28 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-23 (late evening) — clean-up: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0); floor 0.265.0 + +**R-634 mechanism (reproduced on demand on 9202):** the in-memory `Deployed` flag is true from the moment a +deploy is ACCEPTED (`stacks/deploy.go:395`), and `stackAdapter.ListDeployedStacks` filtered on it alone, so +every backup leg saw a DEPLOYING app; the volume leg's `DumpAppVolumesSafe` ran `compose down` and a second +`compose up -d` in the middle of the deploy's own `up`; both failed and `runComposeDeploy`'s failure branch +wrote `Deployed=false` over whatever containers the race left. **Fix:** deploying apps leave the list; the +volume leg re-asks (`deployingReporter`, optional interface — no fake churn) right before the stop and SKIPs; +`StopStack`/`StartStack` return `ErrStackDeploying` for every caller. **Open:** R-649 (operator question: a +deploy that fails on its own and leaves containers). +**R-625:** held → badge `badge.update.held` (`tag-error`), no Update button (the list already hid it). +**R-636:** `ScanOOMKilled` reads `memory.events` `oom_kill` / `memory.max` / `memory.peak` by `docker exec +cat` for FLAGGED containers only (the controller's own cgroupns cannot see others); the notifier keeps 30 min +of samples per `container|startedAt` and sends `app_oom_storm` (error) once at ≥20; hub: allow-list + +`operatorOnlyEvents` + `perAppCooldownEvents`. `08` §6.2 has the rung. +**R-647:** web `readerFuncs(lang, prev)` apply to EVERY language (holdText / updateErrorText / held badge +title) — same bytes on a Hungarian box; held error key `stacks.UpdateErrorKeyHeld`; API uses +`Manager.UpdateErrorFor` / `HoldReasonFor` (guards' optional `HoldForLang`). +**R-648:** there is NO per-app backup endpoint; drills press nothing and rely on the update's `backing-up`. +**R-650:** a test that reaches real docker acts on DooPlex — measured once (my draft), unguarded. +Evidence: `audits/cleanup-2026-09-23/README.md`. + ## 2026-09-23 (evening) — the household is told: controller v0.264.0 + hub v0.120.0 (`09` §6.4 parts 2–3); floor 0.264.0 **Events.** `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, sent by the diff --git a/REPORT.md b/REPORT.md index f7b187e1..b8aeb90c 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,11 +1,13 @@ -# REPORT — the undo reaches the fleet, and the household is TOLD (controller v0.264.0, hub v0.120.0) +# REPORT — clean-up evening: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0) -2026-09-23 (evening). Repos touched: **felhom-controller** (v0.264.0, `bc27894`), **felhom.eu** (hub -v0.120.0 `d3b2848`, manifest `7caa24f`, docs, register, evidence), **admin/app-catalog-drill** (drill -commits, reset to live `main` at the end). **felhom-agent, app-catalog-felhom.eu: untouched.** -Architecture read and named: `documentation/architecture/09-update-architecture.md` (§3 decision 15, -§6.4). Evidence: `documentation/audits/undo-fleet-2026-09-23/README.md`. **Method: endpoint-level** -(the endpoints the UI invokes; mails read from the catch-all mailbox; no browser). +2026-09-23 (late evening). Baselines verified live: controller `0a3026180ae8` (v0.264.0), agent +`d9864a94bf62`, felhom.eu `21f17ed32bdc` (hub v0.120.0), catalog `cfcfe5278428` — as the brief said. +Repos touched: **felhom-controller** (v0.265.0 `0054d4b`), **felhom.eu** (hub v0.121.0, manifest, docs, +register, evidence), **admin/app-catalog-drill** (drill commits, reset to live `main`). +**felhom-agent, app-catalog-felhom.eu: untouched.** Architecture read and named: +`09-update-architecture.md` §6.1/§6.1a/§6.4, `08-alarm-ladder.md` §4/§6.2, `02-controller-module-map.md` +(stacks ↔ backup seam). Evidence: `documentation/audits/cleanup-2026-09-23/README.md`. **Method: +endpoint-level** (the endpoints the UI invokes; no browser). --- @@ -13,104 +15,90 @@ Architecture read and named: `documentation/architecture/09-update-architecture. | item | state | |---|---| -| Part 0 — floor 0.263.2 | **done**, read back (`00-floor-0.263.2.txt`) | -| Part 1 — R-646 startup pass | **done**; live: 3 recorded on 9202, 10 on 9201, 0 skipped | -| Part 2 — two events, hub v0.120.0 first | **done**; hub deployed 11:49Z, before any box ran v0.264.0 | -| Part 3 — R-606 | **done, with leftovers → R-647** (a reader in the other language than the box sees a HELD sentence in the box's language; the English mail's raw `Note:` line keeps the Hungarian `copy_holds` phrase) | -| Part 4 — R-620 | **done** | -| Part 5 — live proof | **done**; all proofs passed → **floor raised to 0.264.0** | -| "standing apps untouched" | **NOT fully met.** The harness's „Mentés most" is a whole-box backup: on 9201 it stopped and restarted 9 of the 10 standing apps for their volume dumps, twice (a few seconds each). All healthy afterwards, files byte-identical. **R-648** filed. | -| the brief's Hungarian undone line | **changed:** it ends „semmi nem veszett el, nincs teendőd." — the brief's „teendője nincs" is the formal „ön" voice; the product speaks „te" everywhere and the voice gate ratchets the formal form (18). | -| the customer cooldown | **changed shape:** a NEW customer-side register (`perAppCustomerCooldownEvents`) instead of putting the customer leg of R-389's register to work — that would also have changed `app_start_failed`'s household grain, which nobody asked for (pinned by a test). | -| a box-side seed + two page toggles | **added, not in the brief:** the box pushes its own `EnabledEvents` to the hub on every notifications-page save, so the hub migration alone would be undone by the next save. | +| Part 1 — R-634 | **cause found in ~25 min (cap 2 h), fixed, proven live.** One design question left → **R-649** (operator) | +| Part 2 — R-625 | **done.** The Update button was ALREADY hidden for a held app on the list (since v0.238.0); the defect was the badge only. Now pinned by a test with a positive control. The badge applies to **both** hold kinds (a failed restore holds the app stopped too; the way back is also a restore) — the brief said `update_failed` only | +| Part 3 — R-636 | **done, and the brief's rule changed shape:** "the same key re-fires ≥ 20 times" cannot work — `OOMKilled` is sticky, so the key re-fires every 30 s after ONE kill (a hiccup would storm after 10 min). The controller now counts real kills (kernel `oom_kill`); the 20-in-30-min threshold is applied to THAT | +| Part 3 — hub live `curl` | **not done:** posting a synthetic event with a real customer's key would put a fabricated alarm in the operator's record; the hub side is proven by unit tests | +| Part 4 — R-647 | **done** | +| Part 5 — R-648 | **done — no per-app endpoint exists**, so the drill presses nothing | +| Part 6 — floor | **raised to 0.265.0** — every live proof passed | +| a mistake of mine | my first draft of the R-634 backup test drove the real dump and **created an empty volume `outline_outline_data` on DooPlex**; my inspection of it pulled `alpine:3.20`. Both verified new and unused, removed by name. Tests rewritten on seams. **R-650** filed (the class is unguarded) | **Claims in the brief, checked:** -1. *"No app-update event exists today"* — **true for apps.** The only near-misses: `controller_updated` is - the controller's OWN self-update, and `update_available` exists only as a debug-page test option — never - emitted, and not in the hub's allow-list (it would 400). -2. *"`enabled_events` is the household's only mail gate"* — **wrong.** Four more gates stand in the way: - the `operatorOnlyEvents` register, a blocked customer, a missing e-mail address, and the cooldown — - keyed `customer:type`, which would have merged two apps' mails into one. And the box OVERWRITES - the hub's list on each page save (above). -3. *"9201 can follow the drill catalog without disturbing its standing apps"* — **true for the catalog, - false for the harness.** Measured: the 10 standing apps' compose + `.felhom.yml` were byte-identical - before, during and after (the drill differed from live only in the throwaway apps). The disturbance - came from the harness's whole-box backup press (R-648), not from the catalog. +1. *"sparkyfitness reproduces alone at today's pin"* — **wrong.** Deployed alone in 47.3 s on v0.264.0 + (`10-*`). The 2026-09-22 "alone" walk pressed the whole-box backup itself — the same race. +2. *"a whole-box backup touches a DEPLOYING stack at all (never read)"* — **true, and it is THE cause.** +3. *"20 kills in 30 minutes separates a storm from a hiccup"* — **true only if kills are counted, and they + were not.** RomM's real rate: 4,530 kills in 6 h ≈ **375 per 30 min**; a hiccup is 1–3. 20 kept, applied + to the kernel's counter. Live: 21 kills 2.5 min after start → one storm. +4. *"a per-app backup endpoint exists"* — **wrong.** `router.go` has `/backup/run` and `/backup/tier2` + only; the per-app backup (`RunAppBackupNow`) is reachable only through the guarded update. -## 2. Floor raises, read back from the hub +## 2. R-634 — the mechanism, at `file:line` -| raise | read back | fleet | +1. `stacks/deploy.go:395` — `DeployStack` sets the in-memory `Deployed=true` at ACCEPT (no-stale-button UX). +2. `cmd/controller/main.go:2556` — `ListDeployedStacks` filtered on that flag only → a deploying app is in + every backup leg's list. +3. `backup/backup.go:693` → `DumpAppVolumesSafe` — `StopStack` (`compose down`) in the middle of the deploy, + tar of half-made volumes, `StartStack` (a SECOND `compose up -d`, `backup.go:918`). +4. Both `up` calls fail together; `stacks/deploy.go:421-435` writes `Deployed=false` without asking whether + containers exist. Which `up` wins decides the end: running under „not deployed" (2026-09-22) or + `Created` (tonight, 14:39:48). **No deploy time limit exists** — shape (i) ruled out. + +**Fix:** deploying apps leave the list; the volume leg re-asks right before the stop (SKIP, never FAIL); +`StopStack`/`StartStack` refuse a deploying stack (`ErrStackDeploying`) — the backstop for quiesce, +restore, export, storage and the Stop button, which all had the same exposure. +**Live (`32-*`, `33-*`):** outline's image removed so the deploy must pull, backup at +5 s: the backup ran +15:23:07–15:23:31, stopped gokapi, paperless-ngx and privatebin, **never outline**; `deployed +successfully (took 48.9s)`. The first fix run is kept and labelled NOT A RACE (images cached, 6.3 s). + +## 3. Red-proofs (each seen failing, mutation restored, `=== RUN` checked) + +| # | mutation | test → failure | |---|---|---| -| 0.263.2 / MinAgent 0.131.0 | `min_controller_version=0.263.2`, `min_agent=0.131.0` | `00-floor-0.263.2.txt` | -| **0.264.0 / MinAgent 0.131.0** (14:35:08) | `min_controller_version=0.264.0`, `min_agent=0.131.0`; hub log `managed floor SERVED for demo-felhom: floor 0.264.0 … from declared` | demo-felhom 0.263.2 → **0.264.0 in 20 s**; demo-hp 9201 already 0.264.0 (`50-floor-0.264.0.txt`) | +| 1 | backup skip removed | `TestR634_VolumeLegNeverStopsADeployingApp` — `dumped=[outline]` | +| 2 | StopStack guard removed | `TestR634_StopAndStartRefuseADeployingStack` — `exit code -1` instead of `ErrStackDeploying` | +| 3 | held badge removed (v0.264.0's shape) | `TestR625_AHeldAppSaysStoppedRestoreNeeded` — badge missing, „Frissítés elérhető" present | +| 4 | reader funcs off | `TestR647_AReaderInTheOtherLanguageReadsTheHoldInTheirs` — English hold for the Hungarian reader | +| 5 | Hungarian phrase wired back | `TestR647_HeldEventCarriesTheCopyHoldsKey` | +| 6 | health status passed as severity | `TestR647_DisabledHealthChangeNamesTheSeverity` — `severity warn` | +| 7 | escalation removed | `TestR636_TwentyKillsInThirtyMinutesIsOneStorm` — `kills=20 … storms=0, want 1` | +| 8 | counter never read | `TestR636_ScanReadsTheCounterOfFlaggedContainersOnly` — `Kills:-1` | +| 9 | hub: not operator-only | `TestAppOOMStormIsAllowlistedAndOperatorOnly` — "must be operator-only" | +| 10 | hub: not allow-listed | same test — "must be in allowedEventTypes" | +| 11 | hub: no per-app cooldown | `TestR389_TheAllowListHasExactlyOneMember` | -## 3. The hub migration — rows changed (one-time, add-only) +Also pinned without a mutation run: 19 kills → no storm, 200 → one, one hiccup over 6 h → none, 20 kills +over 3 h → none, an unreadable counter → never. Gates: controller `go test ./...` rc=0 + +`controller_gates.py` rc=0; hub `go test ./...` rc=0 + `repo_gates.py` rc=0. Parity: `app_info_deployed` +and `stacks_full` regenerated — measured diff **exactly one line each, the new held badge**. -| customer | before | after | -|---|---|---| -| `demo-felhom` | 10 types | the same 10 + `app_update_undone`, `app_update_held` | -| `demo-hp` | `null` (nothing on) | `["app_update_undone","app_update_held"]` | -| `_resend-rotation-test` | `[]` | `["app_update_undone","app_update_held"]` | +## 4. Live proofs in both languages (9202) -Guard row `seed_app_update_events_v1 = 2026-09-23T11:49:03Z`. E-mail column never selected. +- **R-625 + R-647 (1)** (`44-*`), every box/reader pair: hu reader „Megállítva — visszaállítás szükséges", + title „A frissítés nem sikerült, és az automatikus visszaállítás sem."; en reader "Stopped — restore + needed", title "The update did not succeed, and the automatic undo did not either."; with the box in + ENGLISH the Hungarian reader reads Hungarian (page and API). No vikunja Update button; other apps keep + theirs (control). `POST …/update` → **409 `held`**. +- **R-636** (`45-*`, `46-*`): romm at 320M — kills 8, 13, **21 → `OOM STORM — 21 kills in 30 min (limit + 320M, peak 320M)`** + `DROPPED event app_oom_storm (severity error)`; one storm line at 49 kills. +- **R-648** (`43-*`): vikunja's update ran its own `backing-up` phase. -## 4. The mails (9201, customer demo-hp) and their `notification_log` rows +## 5. Rows -| row | time (UTC) | type | leg | subject as received | -|---|---|---|---|---| -| 937 / 938 | 12:15 | undone | op / household | op `⚠️ demo-hp: app_update_undone` · hh **`Figyelmeztetés: vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut`** | -| 939 / 940 | 12:16 | held | op / household | op `🔴 demo-hp: app_update_held` · hh **`Hiba: vikunja: az alkalmazás leállítva, visszaállítás szükséges`** | -| 942 / 943 | 12:29 | undone | op / household | hh **`Warning: glance: the update did not work; the app runs on its previous version`** | -| 944 / 946 | 12:31 | held | op / household | hh **`Error: glance: the app is stopped and needs a restore`** | +**Closed (5):** R-634, R-625, R-636, R-647, R-648. **Opened (2):** R-649 (operator question — a deploy that +fails on its own), R-650 (tests that reach real docker act on DooPlex). **Open rows 334 → 331.** +`09` §6.4 parts 8 and 9 marked SHIPPED; `08` §6.2 has the storm rung and the two per-app grains. -All 8 `sent`. The box language was switched to English once, between the two rounds (the hub had two -reports in English before the second round). Household bodies carry the box's own sentence (the undone -line / the hold sentence) in the household's language; the operator's carry the Hungarian wire text. -Verbatim bodies: `41-mails-as-received.md`. **Per-app cooldown, live:** row 943 went out 14 minutes after -row 938, inside the 6-hour customer cooldown. +## 6. Floor -## 5. The pages in both languages (9202, then 9201) +0.265.0 / MinAgent 0.131.0, read back from the hub (`min_controller_version=0.265.0`, +`min_agent=0.131.0`; `managed floor SERVED` for demo-felhom and demo-hp); both on 0.265.0 in 20 s. -- undone: hu *„A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan - visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en *"The update of vikunja - at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically — - nothing was lost."* -- held, box hu: `?lang=hu` → the Hungarian hold sentence; `?lang=en` → *"The update did not succeed, and the - automatic undo did not either. … own drive, 2026-09-23 13:58 — this copy holds the settings, the database - and the data volumes."* — the whole sentence English now (before v0.264.0 only the prefix was). -- `GET /api/stacks/vikunja?lang=en`: label *"The update did not succeed"*, error and hold in English. -- **Leftover (R-647):** box en + `?lang=hu` showed the English hold. +## 7. Teardown — three layers -## 6. Red-proofs (each seen failing, then restored) - -**Controller (9):** backfill stores nothing → `TestR646_…` *"the current app must have its record"*; -`dropped()` not called → `TestR620_…` *"want exactly TWO WARN lines, got 0"*; undone event not emitted → -*"exactly ONE app_update_undone event, got []"*; held event not emitted → *"…got []"*; page -`updateErrorText` not localised → `TestR606_UpdateSentencesFollowTheReader`; hold sentence ignores the -language → `undo_hold_test.go:62`; `UpdateErrorIn` returns the Hungarian → *"English = „A frissítés nem -indult el…""*; box seed not run → *"got [backup_failed offbox_enlarge_blocked]"*; sink not wired in main → -*"SetUpdateEventSink is never called"* (first mutation did not compile; retried with a compiling one). -**Hub (5):** seed call does nothing → *"c1: enabled_events = [backup_failed], want […]"*; customer per-app -suffix removed → *"want 3 household mails …, got 2"*; allow-list entry removed → *"must be in -allowedEventTypes"*; app-named headline removed → subject *"%s: a frissítés…"*; seed default removed → -*"defaultSeedEvents must include app_update_undone"* (first attempt matched no test; retried with a new -test that names both types). Every `-run` was checked for `=== RUN`. - -Gates: controller `go test ./...` rc=0, `controller_gates.py` rc=0; hub `go build/vet/test` rc=0, -`repo_gates.py` rc=0. The R-389 fence test was widened on purpose, with its reason; the settings-page -parity fixture was regenerated (measured diff: exactly the two new toggles). - -## 7. Rows - -**Closed (3, moved to `CLOSED-ITEMS.md`):** R-606, R-620, R-646. **Opened (2):** R-647 (the three -leftovers), R-648 (the whole-box backup press). **Open rows 335 → 334.** `09` §6.4 parts 2 and 3 marked -SHIPPED; the capability map's undo row now carries the mail; STATUS rewritten. - -## 8. Teardown — three layers - -- **machine:** 9202 and 9201 — throwaway apps removed through the product (0 volumes, 0 undo copies), - test images removed, `controller.yaml` identical to the saved copy, catalog on live `cfcfe52`, 9201's - language back to `hu`, standing apps deployed and healthy. Both on 0.264.0 (= the floor). -- **host:** nothing provisioned on demo-hp. -- **hub:** v0.120.0 (planned); floor 0.264.0. Scratchpad password files deleted; hub DB copies deleted. -- **drill repo:** reset to live `main`, read back from the remote. +- **machine (9202):** all four test apps removed through the product; romm's drive folder by name (R-442 + fence, as always); 0 volumes, 0 undo copies; test images removed by name where unused; config identical to + the saved copy; live catalog; language `hu`; recorders stopped. 9201 not touched. +- **host:** nothing provisioned. **hub:** v0.121.0 + floor (planned). **DooPlex:** the volume and image + above, removed. **drill repo:** reset to live `main`. Password files deleted from the scratchpad. diff --git a/STATUS.md b/STATUS.md index 7a5930a1..4428a33c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,23 +1,25 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-23 (evening) — the undo is on every demo machine now, and the household gets a mail when an update fails. The floor is raised.** +**Updated 2026-09-23 (late evening) — clean-up evening. The "runs but not installed" fault has its cause found and fixed. The fleet has the new version.** -**Decisions I took on my own: none.** One small deviation from your brief: the Hungarian mail line ends "nincs teendőd", not "teendője nincs", because the product speaks to the household as "te" everywhere. +**Decisions I took on my own: none.** One question for you is below. -**What the household gets now.** When an update fails and the machine puts the app back, the household gets one mail: "docmost: the update did not work; the app runs on its previous version". When the undo fails too, they get one mail that says the app is stopped and needs a restore, and which backup to use. The mail is in the household's own language. The operator gets the same two events. Two apps on the same night give two mails, not one. +**The top fault: an app could run while the box said "not installed".** Found and fixed. The cause: the "back up now" button backs up every app. It also caught an app that was still installing. It stopped that app in the middle of the install and started it again. The install then failed and wrote "not installed", while the app's containers sometimes kept running. I made this happen on purpose on the scratch machine, twice. With the fix, the backup leaves an installing app alone. I proved that on the same machine: the backup stopped three other apps and did not touch the one installing, and the install finished normally. Sparkyfitness did not fail on its own this time. It installed in 47 seconds. -**Proven on real machines, through the same buttons the page uses:** -- On the demo-hp machine: 4 mails to the household and 4 to the operator, first in Hungarian, then in English after one language switch. All arrived. The hub log shows all 8 as sent. -- On the scratch machine: the update messages on both app pages show in Hungarian and English. A machine with no hub now writes one warning line per kind of lost message. -- At start, the new version stores the health-check file of each up-to-date app, so its first undo asks the right question. It did this for 3 apps on one machine and 10 on the other. -- The floor is 0.264.0. The N100 demo machine updated itself within 20 seconds. +**A stopped app now says so.** When an update failed and the app is stopped, the badge says "Stopped — restore needed", in Hungarian or English. There is no Update button. Proven on the scratch machine in both languages. -**What was not clean.** -- My test "back up now" button backs up the whole machine. On demo-hp it stopped and restarted 9 of the 10 standing apps for a few seconds, twice. All came back healthy, with unchanged files. So "standing apps untouched" is not fully true. I filed a row to fix the test method. -- Three small leftovers, filed as one row: a reader who forces the other language sees a stopped app's sentence in the machine's language; the raw detail line of an English mail has one Hungarian phrase; and two of my log lines use unclear words. +**A long memory problem is now louder.** The box now counts real memory kills. If one app is killed 20 times in 30 minutes, you get one louder alarm ("error"), once. Households do not get it. Proven with RomM on the scratch machine: the alarm came at 21 kills and stayed one alarm at 49. -**Rows.** Three closed, two new. The list went from 335 to 334. +**The three small mail and log leftovers are fixed.** A reader who picks the other language now reads a stopped app's message in their own language. -**What needs you: nothing.** If you do nothing, the fleet keeps the new version. The leftovers are small and wait for a free evening. +**The test method is fixed.** My tests no longer press "back up now". An app's update makes its own backup instead. -**Nothing on your own machine, Peti's machine or the off-site box was touched, except the hub update. The test apps are gone from both demo machines. Both are back on the real catalogue.** +**Rows.** Five closed, two new. The list went from 334 to 331. The floor is 0.265.0, and both demo machines updated themselves within 20 seconds. + +**What was not clean.** While I wrote one test, it created an empty storage volume on your own machine. I saw it at once, checked it was new and unused, and removed it. Nothing else was touched. I filed a row so tests cannot do this again. + +**What needs you — one question.** When an install fails for its own reasons, some of its containers can stay running while the box says "not installed". The household can remove them, so nothing is stuck. What should happen? +- **Remove what the failed install started (my pick):** "failed" then always means nothing runs. The household presses Install again. +- **Leave it as it is:** nothing changes. The page can say "not installed" over running containers until the household removes them. + +**Nothing on Peti's machine or the off-site box was touched. On your own machine, only the hub was updated, plus the test mistake above. The demo-hp machine was not touched by hand. The scratch machine is back to its three standing apps and the real catalogue.** diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index 828c78d3..34bdf963 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -211,6 +211,22 @@ minutes later (90 m instead of 60). **One value, everywhere:** both staleness ch hardcoded 30 m / 1 h until then; pinned by `TestControllerStatus_FollowsConfiguredThreshold`). A running hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`). +**An out-of-memory storm gets its own, louder rung (R-636, controller v0.265.0 / hub v0.121.0, 2026-09-23).** +`app_oom` stays exactly as it was: `warning`, operator-only, ONCE per container run (R-514) — the once +is what stops a crash loop from mailing thousands of times (R-629). Beside it, `app_oom_storm`: + +| event | severity | who | minted by | why that audience | +|---|---|---|---|---| +| `app_oom_storm` | **error** | **operator only** | controller v0.265.0, when the kernel's `oom_kill` counter of the SAME container run rises by **≥ 20 within 30 min**; once per run | raw container names and memory figures; the household's side is the dashboard tag | + +**Why a counter and not the flag:** Docker's `OOMKilled` is sticky — true for the whole run after ONE +kill — so "the key re-fired N times" measures only how long ago the first kill was. The kernel's +`memory.events` `oom_kill` counts kills. **Why 20 in 30 minutes:** RomM's measured rate on 2026-09-22 +was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202 (2026-09-23): RomM at 320M +reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:** +the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose +`OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms. + **Two event types added 2026-09-17, with who receives them:** | event | severity | who | minted by | why that audience | @@ -221,6 +237,8 @@ hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`). | Family | Grain | Key carries | Why | |---|---|---|---| | app down (`app_start_failed`) | **per APP** | `…:` | no digest exists for it — see below | +| update outcome (`app_update_undone`, `app_update_held`) | **per APP**, both legs | `…:` | hub v0.120.0 — one mail per app per outcome; the household leg has its own register (`perAppCustomerCooldownEvents`) | +| OOM storm (`app_oom_storm`) | **per APP**, operator | `…:` | hub v0.121.0 — no digest; the controller already sends it at most once per container run | | backup run (`backup_run_failures`) | per RUN | `…:` | a digest already lists every failing app; one per run | | tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:` | the tiers fail independently and mean different things | | everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index a65b71b6..f17cddad 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -1049,8 +1049,8 @@ what the part can do to a household's data if it is wrong, not how likely that i | **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape | | **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low | | **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe | -| **8** | **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none | -| **9** | **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none | +| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none | +| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none | | **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside | | **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows | diff --git a/documentation/audits/cleanup-2026-09-23/10-1a-sparkyfitness-alone.txt b/documentation/audits/cleanup-2026-09-23/10-1a-sparkyfitness-alone.txt new file mode 100644 index 00000000..7e76d527 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/10-1a-sparkyfitness-alone.txt @@ -0,0 +1,8 @@ +16:35:59 deploy sparkyfitness -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +16:35:59 + 0.0s sparkyfitness: deployed=True deploying=True state=not_deployed err='None' +16:36:08 + 9.1s sparkyfitness: deployed=True deploying=True state=deploying err='None' +16:36:29 + 30.3s sparkyfitness: deployed=True deploying=True state=degraded err='None' +16:36:48 + 48.4s sparkyfitness: deployed=True deploying=False state=starting err='None' +16:37:00 + 60.5s sparkyfitness: deployed=True deploying=False state=running err='None' +16:38:00 final: {"deployed": true, "deploying": false, "state": "running", "deploy_error": null, "containers": [{"name": "sparkyfitness", "image": "codewithcj/sparkyfitness:v0.17.3", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "sparkyfitness-server", "image": "codewithcj/sparkyfitness_server:v0.17.3", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "sparkyfitness-db", "image": "postgres:15-alpine", "state": "running", "status": "Up About a minute (healthy)"}]} +rc=0 diff --git a/documentation/audits/cleanup-2026-09-23/11-1a-guest-logs.txt b/documentation/audits/cleanup-2026-09-23/11-1a-guest-logs.txt new file mode 100644 index 00000000..00584ae2 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/11-1a-guest-logs.txt @@ -0,0 +1,155 @@ +== controller log (sparkyfitness lines + deploy/stop/start) +2026/09/23 14:35:59 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness/deploy-fields +2026/09/23 14:35:59 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness/deploy-fields (path=/stacks/sparkyfitness/deploy-fields) +2026/09/23 14:35:59 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/sparkyfitness/deploy +2026/09/23 14:35:59 router.go:82: [DEBUG] [api] POST /api/stacks/sparkyfitness/deploy (path=/stacks/sparkyfitness/deploy) +2026/09/23 14:35:59 router.go:438: [INFO] [api] Deploy requested for stack: sparkyfitness +2026/09/23 14:35:59 router.go:82: [DEBUG] [api] deployStack: name=sparkyfitness contentLength=70 +2026/09/23 14:35:59 deploy.go:253: [DEBUG] Deploy sparkyfitness: received 2 user values +2026/09/23 14:35:59 deploy.go:258: [DEBUG] SUBDOMAIN = "sparkyfitness" +2026/09/23 14:35:59 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 4 encrypted, 4 sensitive fields +2026/09/23 14:35:59 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness +2026/09/23 14:35:59 deploy.go:383: [INFO] [stacks] Deploying stack sparkyfitness with 6 env vars: [APP_DB_PASSWORD, API_ENCRYPTION_KEY, BETTER_AUTH_SECRET, DOMAIN, SUBDOMAIN, DB_PASSWORD] +2026/09/23 14:35:59 manager.go:1533: [INFO] [stacks] Deploying stack sparkyfitness — checking 3 images... +2026/09/23 14:35:59 manager.go:1537: [DEBUG] codewithcj/sparkyfitness_server:v0.17.3 — not found locally, will pull +2026/09/23 14:35:59 manager.go:1537: [DEBUG] codewithcj/sparkyfitness:v0.17.3 — not found locally, will pull +2026/09/23 14:35:59 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=sparkyfitness, waiting 5s for state refresh +2026/09/23 14:35:59 manager.go:1414: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/sparkyfitness) +2026/09/23 14:35:59 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:35:59 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:02 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:02 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:04 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=sparkyfitness integrationsFound=0 +2026/09/23 14:36:05 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:05 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:08 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:08 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:11 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:11 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:14 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:14 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:17 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/09/23 14:36:17 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:17 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:20 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:20 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:23 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:23 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:26 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:26 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:27 manager.go:815: [DEBUG] [stacks] restart-policy of down member "sparkyfitness" = "unless-stopped" +2026/09/23 14:36:29 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:29 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:32 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:32 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:35 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:35 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:37 healthprobe.go:99: [DEBUG] [stacks] RunHealthProbes: sparkyfitness has a health check but no container to probe — candidates: [sparkyfitness(stopped) sparkyfitness-server(stopped) sparkyfitness-db(starting)] +2026/09/23 14:36:38 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:38 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:41 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:41 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:45 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:45 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:46 deploy.go:451: [INFO] [stacks] Stack sparkyfitness deployed successfully (took 47.3s) +2026/09/23 14:36:46 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 4 encrypted, 4 sensitive fields +2026/09/23 14:36:46 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness +2026/09/23 14:36:47 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 0 encrypted, 4 sensitive fields +2026/09/23 14:36:47 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness +2026/09/23 14:36:47 installed.go:408: [INFO] [stacks] installed-images sparkyfitness: recorded 3 service(s) (sparkyfitness-db=postgres:15-alpine (sha256:f7d23353e1b1…), sparkyfitness-frontend=codewithcj/sparkyfitness:v0.17.3 (sha256:46d90e46bd87…), sparkyfitness-server=codewithcj/sparkyfitness_server:v0.17.3 (sha256:6aa7d9832324…)) +2026/09/23 14:36:47 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 0 encrypted, 4 sensitive fields +2026/09/23 14:36:47 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness +2026/09/23 14:36:47 pin.go:93: [INFO] [stacks] pin sparkyfitness: sparkyfitness-db=postgres:15-alpine, sparkyfitness-frontend=codewithcj/sparkyfitness:v0.17.3, sparkyfitness-server=codewithcj/sparkyfitness_server:v0.17.3 +2026/09/23 14:36:48 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:48 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:50 manager.go:1414: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/sparkyfitness) +2026/09/23 14:36:50 manager.go:1494: [INFO] [stacks] Stack sparkyfitness post-start status: +2026/09/23 14:36:50 manager.go:1497: [INFO] [stacks] sparkyfitness codewithcj/sparkyfitness:v0.17.3 running Up 3 seconds (health: starting) +2026/09/23 14:36:50 manager.go:1497: [INFO] [stacks] sparkyfitness-db postgres:15-alpine running Up 24 seconds (healthy) +2026/09/23 14:36:50 manager.go:1497: [INFO] [stacks] sparkyfitness-server codewithcj/sparkyfitness_server:v0.17.3 running Up 19 seconds (healthy) +2026/09/23 14:36:51 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:51 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:54 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:54 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:36:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:36:57 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:37:00 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:37:00 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:37:00 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:37:00 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +2026/09/23 14:37:03 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness +2026/09/23 14:37:03 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness) +== docker events (sparkyfitness) +1790174166 image pull codewithcj/sparkyfitness +1790174183 image pull codewithcj/sparkyfitness_server +1790174183 network create sparkyfitness_sparkyfitness-internal +1790174185 container create sparkyfitness-db +1790174185 container create sparkyfitness-server +1790174185 container create sparkyfitness +1790174185 network connect sparkyfitness_sparkyfitness-internal +1790174185 container start sparkyfitness-db +1790174190 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174190 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174190 container exec_die sparkyfitness-db +1790174190 container health_status: healthy sparkyfitness-db +1790174191 network connect sparkyfitness_sparkyfitness-internal +1790174191 container start sparkyfitness-server +1790174196 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174196 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174196 container exec_die sparkyfitness-server +1790174200 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174200 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174200 container exec_die sparkyfitness-db +1790174201 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174201 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174201 container exec_die sparkyfitness-server +1790174206 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174206 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174206 container exec_die sparkyfitness-server +1790174206 container health_status: healthy sparkyfitness-server +1790174206 network connect sparkyfitness_sparkyfitness-internal +1790174206 container start sparkyfitness +1790174210 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174210 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174210 container exec_die sparkyfitness-db +1790174211 container exec_create: wget --spider -q http://127.0.0.1:80/ sparkyfitness +1790174211 container exec_start: wget --spider -q http://127.0.0.1:80/ sparkyfitness +1790174211 container exec_die sparkyfitness +1790174211 container health_status: healthy sparkyfitness +1790174220 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174220 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174220 container exec_die sparkyfitness-db +1790174230 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174230 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174230 container exec_die sparkyfitness-db +1790174236 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174236 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174236 container exec_die sparkyfitness-server +1790174240 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174240 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174240 container exec_die sparkyfitness-db +1790174241 container exec_create: wget --spider -q http://127.0.0.1:80/ sparkyfitness +1790174241 container exec_start: wget --spider -q http://127.0.0.1:80/ sparkyfitness +1790174242 container exec_die sparkyfitness +1790174250 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174250 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174250 container exec_die sparkyfitness-db +1790174260 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174260 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174260 container exec_die sparkyfitness-db +1790174266 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174266 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server +1790174266 container exec_die sparkyfitness-server +1790174270 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174270 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174271 container exec_die sparkyfitness-db +1790174272 container exec_create: wget --spider -q http://127.0.0.1:80/ sparkyfitness +1790174272 container exec_start: wget --spider -q http://127.0.0.1:80/ sparkyfitness +1790174272 container exec_die sparkyfitness +1790174281 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174281 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174281 container exec_die sparkyfitness-db +1790174291 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174291 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db +1790174291 container exec_die sparkyfitness-db +grep: write error: Broken pipe diff --git a/documentation/audits/cleanup-2026-09-23/12-1b-outline-with-backup.txt b/documentation/audits/cleanup-2026-09-23/12-1b-outline-with-backup.txt new file mode 100644 index 00000000..2638ef7a --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/12-1b-outline-with-backup.txt @@ -0,0 +1,7 @@ +16:39:08 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +16:39:08 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None' +16:39:17 + 9.1s outline: deployed=True deploying=True state=deploying err='None' +16:39:51 + 42.4s outline: deployed=False deploying=False state=deploying err='exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021' +16:39:57 + 48.4s outline: deployed=False deploying=False state=stopped err='exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021' +16:40:57 final: {"deployed": false, "deploying": false, "state": "stopped", "deploy_error": "exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021cb Pulling fs layer 0B\n f25fb26220a8 Pulling fs layer 0B\n 4d027270bfe6 Pulling fs layer 0B\n d256567eb8f6 Pulling fs layer 0B\n bc77c6bc12ba Pulling fs layer 0B\n 167140e5ddd9 Pulling fs layer 0B\n b6b3029da283 Pulling fs layer 0B\n 5d5585bcfc4f Pulling fs layer 0B\n 5726d50585da Pulling fs layer 0B\n 4a58d711ba36 +rc=0 diff --git a/documentation/audits/cleanup-2026-09-23/12-1b-press.txt b/documentation/audits/cleanup-2026-09-23/12-1b-press.txt new file mode 100644 index 00000000..2cf00aa0 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/12-1b-press.txt @@ -0,0 +1 @@ +16:39:28 WHOLE-BOX BACKUP pressed 20 s after the deploy was accepted -> 200 {'ok': True, 'message': 'Mentés elindítva'} diff --git a/documentation/audits/cleanup-2026-09-23/13-1b-guest-logs.txt b/documentation/audits/cleanup-2026-09-23/13-1b-guest-logs.txt new file mode 100644 index 00000000..53af7fa0 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/13-1b-guest-logs.txt @@ -0,0 +1,132 @@ +== controller log: outline + backup legs +2026/09/23 14:39:08 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline/deploy-fields +2026/09/23 14:39:08 router.go:82: [DEBUG] [api] GET /api/stacks/outline/deploy-fields (path=/stacks/outline/deploy-fields) +2026/09/23 14:39:08 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/outline/deploy +2026/09/23 14:39:08 router.go:82: [DEBUG] [api] POST /api/stacks/outline/deploy (path=/stacks/outline/deploy) +2026/09/23 14:39:08 router.go:438: [INFO] [api] Deploy requested for stack: outline +2026/09/23 14:39:08 router.go:82: [DEBUG] [api] deployStack: name=outline contentLength=64 +2026/09/23 14:39:08 deploy.go:253: [DEBUG] Deploy outline: received 2 user values +2026/09/23 14:39:08 deploy.go:258: [DEBUG] SUBDOMAIN = "outline" +2026/09/23 14:39:08 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields +2026/09/23 14:39:08 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 14:39:08 deploy.go:383: [INFO] [stacks] Deploying stack outline with 5 env vars: [DOMAIN, SUBDOMAIN, SECRET_KEY, UTILS_SECRET, DB_PASSWORD] +2026/09/23 14:39:08 manager.go:1533: [INFO] [stacks] Deploying stack outline — checking 3 images... +2026/09/23 14:39:08 manager.go:1537: [DEBUG] outlinewiki/outline:1.9.1 — not found locally, will pull +2026/09/23 14:39:08 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=outline, waiting 5s for state refresh +2026/09/23 14:39:08 manager.go:1414: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline) +2026/09/23 14:39:08 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:08 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:11 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:11 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:13 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=outline integrationsFound=0 +2026/09/23 14:39:14 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:14 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:17 backup.go:465: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [gokapi, outline, privatebin] +2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml +2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml +2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml +2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml +2026/09/23 14:39:17 recovery_unit.go:224: [INFO] [backup] Recovery unit captured for outline → /mnt/sys_drive/felhom-data/backups/primary/outline (images=3, secrets-referenced=3, data_keys=0, portable-carried=3/3, withheld=0) +2026/09/23 14:39:17 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:17 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:20 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:20 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:23 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:23 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:26 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:26 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:28 backup.go:520: [INFO] [backup] Starting database dump run +2026/09/23 14:39:29 backup.go:908: [INFO] [backup] Stopping gokapi for safe volume dump +2026/09/23 14:39:29 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:29 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:30 backup.go:809: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_config → 3.0 KB +2026/09/23 14:39:31 backup.go:809: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_data → 99.5 KB +2026/09/23 14:39:31 backup.go:918: [INFO] [backup] Restarting gokapi after volume dump +2026/09/23 14:39:32 backup.go:908: [INFO] [backup] Stopping outline for safe volume dump +2026/09/23 14:39:32 manager.go:1234: [DEBUG] [stacks] StopStack outline: current state=deploying deployed=true containers=0 +2026/09/23 14:39:32 manager.go:1237: [INFO] [stacks] Stopping stack: outline +2026/09/23 14:39:32 manager.go:1414: [DEBUG] Running: docker compose down (in /opt/docker/stacks/outline) +2026/09/23 14:39:32 manager.go:1246: [INFO] [stacks] Stack outline stopped successfully (took 0.1s) +2026/09/23 14:39:32 backup.go:784: [DEBUG] [backup] Dumping volume outline_outline_data for outline +2026/09/23 14:39:32 backup.go:809: [INFO] [backup] Volume dump: outline/outline_outline_data → 1.5 KB +2026/09/23 14:39:32 backup.go:784: [DEBUG] [backup] Dumping volume outline_outline_postgres_data for outline +2026/09/23 14:39:32 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:32 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:33 backup.go:809: [INFO] [backup] Volume dump: outline/outline_outline_postgres_data → 1.5 KB +2026/09/23 14:39:33 backup.go:784: [DEBUG] [backup] Dumping volume outline_outline_redis_data for outline +2026/09/23 14:39:33 backup.go:809: [INFO] [backup] Volume dump: outline/outline_outline_redis_data → 1.5 KB +2026/09/23 14:39:33 backup.go:918: [INFO] [backup] Restarting outline after volume dump +2026/09/23 14:39:33 manager.go:1152: [DEBUG] [stacks] StartStack outline: current state=deploying deployed=true +2026/09/23 14:39:33 manager.go:1155: [INFO] [stacks] Starting stack: outline +2026/09/23 14:39:33 manager.go:1162: [DEBUG] [stacks] StartStack outline: prepared 11 env vars for compose +2026/09/23 14:39:33 manager.go:1414: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline) +2026/09/23 14:39:36 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:36 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:39 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:39 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:42 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:42 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:45 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:45 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:48 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:48 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:48 manager.go:1423: [ERROR] [stacks] Command failed: docker compose up -d (in /opt/docker/stacks/outline) — exit code 1 (took 40.0s) +2026/09/23 14:39:48 manager.go:1429: [ERROR] [stacks] stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:39:48 deploy.go:424: [ERROR] [stacks] Stack outline deploy failed after 40.0s: exit code 1 +stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:39:48 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields +2026/09/23 14:39:48 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 14:39:48 manager.go:1423: [ERROR] [stacks] Command failed: docker compose up -d (in /opt/docker/stacks/outline) — exit code 1 (took 15.3s) +2026/09/23 14:39:48 manager.go:1429: [ERROR] [stacks] stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:39:48 manager.go:1166: [ERROR] [stacks] Stack outline start failed after 15.3s: exit code 1 +stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:39:48 backup.go:921: [ERROR] [backup] Failed to restart outline after volume dump: starting stack outline: exit code 1 +stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:39:48 backup.go:740: [ERROR] [backup] Volume dump failed for outline: volume dump OK but restart failed for outline: starting stack outline: exit code 1 +stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:39:48 backup.go:908: [INFO] [backup] Stopping paperless-ngx for safe volume dump +2026/09/23 14:39:51 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:51 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:54 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:54 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:56 backup.go:809: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 543.5 KB +2026/09/23 14:39:56 backup.go:809: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 67.8 MB +2026/09/23 14:39:57 backup.go:809: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 490.0 KB +2026/09/23 14:39:57 backup.go:918: [INFO] [backup] Restarting paperless-ngx after volume dump +2026/09/23 14:39:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:57 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:39:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:39:57 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:00 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:00 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:03 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:03 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:06 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:06 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:08 backup.go:908: [INFO] [backup] Stopping privatebin for safe volume dump +2026/09/23 14:40:09 backup.go:809: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB +2026/09/23 14:40:09 backup.go:918: [INFO] [backup] Restarting privatebin after volume dump +2026/09/23 14:40:09 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:09 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:09 backup.go:638: [WARN] [backup] some backup steps failed (FAIL outline volumes: volume dump OK but restart failed for outline: starting stack outline: exit code 1 +stderr: Image outlinewiki/outline:1.9.1 Pulling +2026/09/23 14:40:12 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:12 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:15 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:15 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:17 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/23 14:40:18 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:18 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +2026/09/23 14:40:21 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline +2026/09/23 14:40:21 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline) +== docker events: outline +1790174388 image pull outlinewiki/outline +1790174388 image pull outlinewiki/outline +1790174388 network create outline_outline-internal +1790174388 container create outline-redis +1790174388 container create outline-postgres +== now +outline-postgres Created +outline-redis Created +deployed: false +desired_state: running diff --git a/documentation/audits/cleanup-2026-09-23/20-hub-0.121.0-deploy.txt b/documentation/audits/cleanup-2026-09-23/20-hub-0.121.0-deploy.txt new file mode 100644 index 00000000..22cab8a6 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/20-hub-0.121.0-deploy.txt @@ -0,0 +1,3 @@ +# hub v0.121.0 deploy 2026-09-23T15:18:56Z +gitea.dooplex.hu/admin/felhom-hub:0.121.0 +sync=Synced health=Healthy rev=3add9fa678b89973462a4b6ae3a3c35d2e642b9d diff --git a/documentation/audits/cleanup-2026-09-23/30-fix-attempt1-NOT-A-RACE-images-cached.txt b/documentation/audits/cleanup-2026-09-23/30-fix-attempt1-NOT-A-RACE-images-cached.txt new file mode 100644 index 00000000..3295362f --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/30-fix-attempt1-NOT-A-RACE-images-cached.txt @@ -0,0 +1,10 @@ +17:20:26 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +17:20:26 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None' +17:20:29 + 3.0s outline: deployed=True deploying=True state=stopped err='None' +17:20:35 + 9.1s outline: deployed=True deploying=False state=starting err='None' +17:20:56 + 30.3s outline: deployed=True deploying=False state=stopped err='None' +17:20:56 + 30.3s outline: deployed=True deploying=False state=degraded err='None' +17:21:03 + 36.4s outline: deployed=True deploying=False state=starting err='None' +17:21:18 + 51.5s outline: deployed=True deploying=False state=running err='None' +17:21:57 final: {"deployed": true, "deploying": false, "state": "running", "deploy_error": null, "containers": [{"name": "outline", "image": "outlinewiki/outline:1.9.1", "state": "running", "status": "Up 54 seconds (healthy)"}, {"name": "outline-postgres", "image": "postgres:16-alpine", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "outline-redis", "image": "redis:7-alpine", "state": "running", "status": "Up About a minute (healthy)"}]} +rc=0 diff --git a/documentation/audits/cleanup-2026-09-23/30-fix-attempt1-press.txt b/documentation/audits/cleanup-2026-09-23/30-fix-attempt1-press.txt new file mode 100644 index 00000000..4d8d7947 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/30-fix-attempt1-press.txt @@ -0,0 +1 @@ +17:20:46 WHOLE-BOX BACKUP pressed 20 s after the deploy was accepted -> 200 {'ok': True, 'message': 'Mentés elindítva'} diff --git a/documentation/audits/cleanup-2026-09-23/31-fix-guest-logs.txt b/documentation/audits/cleanup-2026-09-23/31-fix-guest-logs.txt new file mode 100644 index 00000000..84d0e06e --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/31-fix-guest-logs.txt @@ -0,0 +1,67 @@ +2026/09/23 15:19:56 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/23 15:19:56 backup.go:1187: [INFO] [backup] Backup status cache refreshed +2026/09/23 15:19:57 sync.go:416: [DEBUG] [sync] outline/docker-compose.yml: hash match, skipped +2026/09/23 15:19:57 sync.go:416: [DEBUG] [sync] outline/.felhom.yml: hash match, skipped +2026/09/23 15:20:26 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/outline/deploy +2026/09/23 15:20:26 router.go:81: [DEBUG] [api] POST /api/stacks/outline/deploy (path=/stacks/outline/deploy) +2026/09/23 15:20:26 router.go:431: [INFO] [api] Deploy requested for stack: outline +2026/09/23 15:20:26 router.go:81: [DEBUG] [api] deployStack: name=outline contentLength=64 +2026/09/23 15:20:26 deploy.go:253: [DEBUG] Deploy outline: received 2 user values +2026/09/23 15:20:26 deploy.go:258: [DEBUG] SUBDOMAIN = "outline" +2026/09/23 15:20:26 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields +2026/09/23 15:20:26 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:20:26 deploy.go:383: [INFO] [stacks] Deploying stack outline with 5 env vars: [DOMAIN, SUBDOMAIN, SECRET_KEY, UTILS_SECRET, DB_PASSWORD] +2026/09/23 15:20:26 manager.go:1546: [INFO] [stacks] Deploying stack outline — checking 3 images... +2026/09/23 15:20:26 manager.go:1552: [DEBUG] outlinewiki/outline:1.9.1 — found locally +2026/09/23 15:20:26 manager.go:1427: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline) +2026/09/23 15:20:26 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=outline, waiting 5s for state refresh +2026/09/23 15:20:31 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=outline integrationsFound=0 +2026/09/23 15:20:32 deploy.go:451: [INFO] [stacks] Stack outline deployed successfully (took 6.3s) +2026/09/23 15:20:32 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields +2026/09/23 15:20:32 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:20:33 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields +2026/09/23 15:20:33 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:20:33 installed.go:408: [INFO] [stacks] installed-images outline: recorded 3 service(s) (outline=outlinewiki/outline:1.9.1 (sha256:9fe2cbdcecce…), outline-postgres=postgres:16-alpine (sha256:721873c34ceb…), outline-redis=redis:7-alpine (sha256:858f009f9709…)) +2026/09/23 15:20:33 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields +2026/09/23 15:20:33 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:20:33 pin.go:93: [INFO] [stacks] pin outline: outline=outlinewiki/outline:1.9.1, outline-postgres=postgres:16-alpine, outline-redis=redis:7-alpine +2026/09/23 15:20:36 manager.go:1427: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/outline) +2026/09/23 15:20:36 manager.go:1507: [INFO] [stacks] Stack outline post-start status: +2026/09/23 15:20:36 manager.go:1510: [INFO] [stacks] outline outlinewiki/outline:1.9.1 running Up 3 seconds (health: starting) +2026/09/23 15:20:36 manager.go:1510: [INFO] [stacks] outline-postgres postgres:16-alpine running Up 9 seconds (healthy) +2026/09/23 15:20:36 manager.go:1510: [INFO] [stacks] outline-redis redis:7-alpine running Up 9 seconds (healthy) +2026/09/23 15:20:46 dbdump.go:147: [DEBUG] DiscoverDatabases: docker ps output: a0affd38f0c9 outline outline outlinewiki/outline:1.9.1 +10d80fc75f8c outline-postgres outline postgres:16-alpine +fdd9eaa9a3bd outline-redis outline redis:7-alpine +2026/09/23 15:20:46 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container outline (image=outlinewiki/outline:1.9.1, not a database) +2026/09/23 15:20:46 dbdump.go:176: [DEBUG] DiscoverDatabases: found postgres container: outline-postgres (id=10d80fc75f8c) +2026/09/23 15:20:46 dbdump.go:196: [DEBUG] DiscoverDatabases: outline-postgres → stack=outline, dbUser=outline, dbName=outline +2026/09/23 15:20:46 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container outline-redis (image=redis:7-alpine, not a database) +2026/09/23 15:20:46 backup.go:548: [INFO] [backup] Discovered 2 database(s): outline-postgres(postgres), paperless-postgres(postgres) +2026/09/23 15:20:46 dbdump.go:264: [DEBUG] DumpOne: starting dump for container=outline-postgres, stack=outline, dbType=postgres, dest=/mnt/sys_drive/felhom-data/backups/primary/outline/db-dumps/outline-postgres.sql +2026/09/23 15:20:46 dbdump.go:301: [DEBUG] DumpOne: pg_dump command: docker exec 10d80fc75f8c pg_dump -U outline -d outline --clean --if-exists --no-owner --no-privileges +2026/09/23 15:20:46 [WARN] [backup] ValidateDump: /mnt/sys_drive/felhom-data/backups/primary/outline/db-dumps/outline-postgres.sql is structurally valid (32 tables) but its accounts table has NO rows — the dump may predate the customer's data +2026/09/23 15:20:46 dbdump.go:403: [DEBUG] DumpOne: completed outline-postgres → outline-postgres.sql (size=98.6 KB, valid=true, tables=32, duration=254ms) +2026/09/23 15:20:46 dbdump.go:408: [INFO] [backup] DB dump: outline-postgres → outline-postgres.sql (98.6 KB, 254ms, 32 tables) +2026/09/23 15:20:48 backup.go:918: [INFO] [backup] Stopping outline for safe volume dump +2026/09/23 15:20:48 manager.go:1247: [DEBUG] [stacks] StopStack outline: current state=starting deployed=true containers=3 +2026/09/23 15:20:48 manager.go:1250: [INFO] [stacks] Stopping stack: outline +2026/09/23 15:20:48 manager.go:1427: [DEBUG] Running: docker compose down (in /opt/docker/stacks/outline) +2026/09/23 15:20:55 manager.go:1259: [INFO] [stacks] Stack outline stopped successfully (took 6.2s) +2026/09/23 15:20:55 backup.go:794: [DEBUG] [backup] Dumping volume outline_outline_redis_data for outline +2026/09/23 15:20:55 backup.go:819: [INFO] [backup] Volume dump: outline/outline_outline_redis_data → 2.5 KB +2026/09/23 15:20:55 backup.go:794: [DEBUG] [backup] Dumping volume outline_outline_data for outline +2026/09/23 15:20:55 backup.go:819: [INFO] [backup] Volume dump: outline/outline_outline_data → 1.5 KB +2026/09/23 15:20:55 backup.go:794: [DEBUG] [backup] Dumping volume outline_outline_postgres_data for outline +2026/09/23 15:20:56 backup.go:819: [INFO] [backup] Volume dump: outline/outline_outline_postgres_data → 49.3 MB +2026/09/23 15:20:56 backup.go:928: [INFO] [backup] Restarting outline after volume dump +2026/09/23 15:20:56 manager.go:1158: [DEBUG] [stacks] StartStack outline: current state=stopped deployed=true +2026/09/23 15:20:56 manager.go:1161: [INFO] [stacks] Starting stack: outline +2026/09/23 15:20:56 manager.go:1168: [DEBUG] [stacks] StartStack outline: prepared 11 env vars for compose +== now +outline Up 57 seconds (healthy) +outline-postgres Up About a minute (healthy) +outline-redis Up About a minute (healthy) +deployed: true +desired_state: running +grep: write error: Broken pipe diff --git a/documentation/audits/cleanup-2026-09-23/32-fix-outline-with-backup.txt b/documentation/audits/cleanup-2026-09-23/32-fix-outline-with-backup.txt new file mode 100644 index 00000000..7ab715ec --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/32-fix-outline-with-backup.txt @@ -0,0 +1,8 @@ +17:23:02 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +17:23:02 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None' +17:23:08 + 6.1s outline: deployed=True deploying=True state=deploying err='None' +17:23:47 + 45.4s outline: deployed=True deploying=True state=degraded err='None' +17:23:54 + 51.4s outline: deployed=True deploying=False state=starting err='None' +17:24:27 + 84.8s outline: deployed=True deploying=False state=running err='None' +17:25:27 final: {"deployed": true, "deploying": false, "state": "running", "deploy_error": null, "containers": [{"name": "outline", "image": "outlinewiki/outline:1.9.1", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "outline-postgres", "image": "postgres:16-alpine", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "outline-redis", "image": "redis:7-alpine", "state": "running", "status": "Up About a minute (healthy)"}]} +rc=0 diff --git a/documentation/audits/cleanup-2026-09-23/32-fix-press.txt b/documentation/audits/cleanup-2026-09-23/32-fix-press.txt new file mode 100644 index 00000000..e9009e05 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/32-fix-press.txt @@ -0,0 +1 @@ +17:23:07 WHOLE-BOX BACKUP pressed 5 s after the deploy was accepted -> 200 {'ok': True, 'message': 'Mentés elindítva'} diff --git a/documentation/audits/cleanup-2026-09-23/33-fix-guest-logs.txt b/documentation/audits/cleanup-2026-09-23/33-fix-guest-logs.txt new file mode 100644 index 00000000..d013a882 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/33-fix-guest-logs.txt @@ -0,0 +1,46 @@ +10d80fc75f8c outline-postgres outline postgres:16-alpine +fdd9eaa9a3bd outline-redis outline redis:7-alpine +2026/09/23 15:23:02 router.go:81: [DEBUG] [api] POST /api/stacks/outline/deploy (path=/stacks/outline/deploy) +2026/09/23 15:23:02 router.go:431: [INFO] [api] Deploy requested for stack: outline +2026/09/23 15:23:02 router.go:81: [DEBUG] [api] deployStack: name=outline contentLength=64 +2026/09/23 15:23:02 deploy.go:253: [DEBUG] Deploy outline: received 2 user values +2026/09/23 15:23:02 deploy.go:258: [DEBUG] SUBDOMAIN = "outline" +2026/09/23 15:23:02 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields +2026/09/23 15:23:02 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:23:02 deploy.go:383: [INFO] [stacks] Deploying stack outline with 5 env vars: [UTILS_SECRET, DB_PASSWORD, DOMAIN, SUBDOMAIN, SECRET_KEY] +2026/09/23 15:23:02 manager.go:1546: [INFO] [stacks] Deploying stack outline — checking 3 images... +2026/09/23 15:23:02 manager.go:1550: [DEBUG] outlinewiki/outline:1.9.1 — not found locally, will pull +2026/09/23 15:23:02 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=outline, waiting 5s for state refresh +2026/09/23 15:23:02 manager.go:1427: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline) +2026/09/23 15:23:07 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=outline integrationsFound=0 +2026/09/23 15:23:07 backup.go:918: [INFO] [backup] Stopping gokapi for safe volume dump +2026/09/23 15:23:08 backup.go:819: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_config → 3.0 KB +2026/09/23 15:23:09 backup.go:819: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_data → 99.5 KB +2026/09/23 15:23:10 backup.go:918: [INFO] [backup] Stopping paperless-ngx for safe volume dump +2026/09/23 15:23:17 backup.go:819: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 580.0 KB +2026/09/23 15:23:18 backup.go:819: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 67.8 MB +2026/09/23 15:23:18 backup.go:819: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 512.5 KB +2026/09/23 15:23:30 backup.go:918: [INFO] [backup] Stopping privatebin for safe volume dump +2026/09/23 15:23:31 backup.go:819: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB +2026/09/23 15:23:46 manager.go:815: [DEBUG] [stacks] restart-policy of down member "outline" = "unless-stopped" +2026/09/23 15:23:51 deploy.go:451: [INFO] [stacks] Stack outline deployed successfully (took 48.9s) +2026/09/23 15:23:51 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields +2026/09/23 15:23:51 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:23:51 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields +2026/09/23 15:23:51 installed.go:408: [INFO] [stacks] installed-images outline: recorded 3 service(s) (outline=outlinewiki/outline:1.9.1 (sha256:9fe2cbdcecce…), outline-postgres=postgres:16-alpine (sha256:721873c34ceb…), outline-redis=redis:7-alpine (sha256:858f009f9709…)) +2026/09/23 15:23:51 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:23:51 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields +2026/09/23 15:23:51 pin.go:93: [INFO] [stacks] pin outline: outline=outlinewiki/outline:1.9.1, outline-postgres=postgres:16-alpine, outline-redis=redis:7-alpine +2026/09/23 15:23:51 [INFO] [stacks] SaveAppConfig: saved config for outline +2026/09/23 15:23:54 manager.go:1427: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/outline) +2026/09/23 15:23:54 manager.go:1507: [INFO] [stacks] Stack outline post-start status: +2026/09/23 15:23:54 manager.go:1510: [INFO] [stacks] outline outlinewiki/outline:1.9.1 running Up 3 seconds (health: starting) +2026/09/23 15:23:54 manager.go:1510: [INFO] [stacks] outline-postgres postgres:16-alpine running Up 9 seconds (healthy) +2026/09/23 15:23:54 manager.go:1510: [INFO] [stacks] outline-redis redis:7-alpine running Up 9 seconds (healthy) +2026/09/23 15:23:56 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=true composePath=/opt/docker/stacks/outline/docker-compose.yml +== now +outline Up About a minute (healthy) +outline-postgres Up About a minute (healthy) +outline-redis Up About a minute (healthy) +deployed: true +desired_state: running diff --git a/documentation/audits/cleanup-2026-09-23/40-9202-repoint.txt b/documentation/audits/cleanup-2026-09-23/40-9202-repoint.txt new file mode 100644 index 00000000..40494c7b --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/40-9202-repoint.txt @@ -0,0 +1,14 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + +17:27:51 catalog cache: d4eeff6 DRILL vikunja: vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.3.0 + +17:27:53 romm drill memory: 86: memory: 320M # DRILL (R-636 live proof): too small on purpose — workers are OOM-killed in a loop + diff --git a/documentation/audits/cleanup-2026-09-23/41-9202-romm-deploy.txt b/documentation/audits/cleanup-2026-09-23/41-9202-romm-deploy.txt new file mode 100644 index 00000000..45cec547 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/41-9202-romm-deploy.txt @@ -0,0 +1,5 @@ +17:28:14 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm'] +17:28:14 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH'] +17:28:14 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +17:29:34 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} +17:29:34 romm deployed: True diff --git a/documentation/audits/cleanup-2026-09-23/42-9202-vikunja-prep.txt b/documentation/audits/cleanup-2026-09-23/42-9202-vikunja-prep.txt new file mode 100644 index 00000000..7b01166d --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/42-9202-vikunja-prep.txt @@ -0,0 +1,13 @@ +17:29:52 === prep vikunja: deploy at the drill FROM pin +17:29:52 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +17:29:57 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'} +17:29:57 deployed: True +17:30:03 vikunja: register http=200 +17:30:03 vikunja: create project http=201 +17:30:03 vikunja: readback of the seeded project http=200 ok=True +17:30:03 C1 A: True +17:30:03 [4] backup press SKIPPED for vikunja (R-648: whole-box only; the update's backing-up phase backs up vikunja alone) +17:30:03 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill1a7c22","created":"2026-09 +17:30:04 vikunja B: project readback=True attachment content readback=True +17:30:04 B reads back: True +17:30:07 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db'] diff --git a/documentation/audits/cleanup-2026-09-23/43-9202-vikunja-held.txt b/documentation/audits/cleanup-2026-09-23/43-9202-vikunja-held.txt new file mode 100644 index 00000000..5cf6ce34 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/43-9202-vikunja-held.txt @@ -0,0 +1,29 @@ +17:30:19 [5] drill commit 21b56e4891e3: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0) +17:30:24 badge caught up after 4.5 s +17:30:24 === vikunja: a cut-off undo copy +17:30:24 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +17:30:24 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… +17:30:26 + 2.1s phase=safety-dump label=Adatbázis pillanatkép… +17:30:27 + 3.2s phase=pulling label=Új verzió letöltése… +17:30:32 + 7.9s phase=copying label=Az adatok másolása a frissítés előtt… +17:30:33 + 9.5s phase=starting label=Indítás az új verzióval… +17:30:34 + 10.0s phase=verifying label=Működés ellenőrzése… +17:30:37 >>> finished-marker removed from one copy: +copy: vikunja_vikunja_data.pre-update-20260923T153032Z +total 12 +drwxr-xr-x 3 root root 4096 Sep 23 15:30 . +drwxr-xr-x 1 root root 4096 Sep 23 15:30 .. +drwxr-xr-x 2 root root 4096 Sep 23 15:30 data +-rw-r--r-- 1 root root 0 Sep 23 15:30 felhom-undo-complete +data + +17:32:05 + 100.9s phase=undoing label=Visszaállítás az előző változatra… +17:32:05 + 101.5s phase=failed label=A frissítés nem sikerült +17:32:05 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' +17:32:09 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 17:32 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: second drive, 2026-09-23 17:30 — this copy holds the settings, the database and the data volumes."}} +17:32:09 box language -> en: http 302 +17:32:09 PAGE (box language en): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 17:32 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: second drive, 2026-09-23 17:30 — this copy holds the settings, the database and the data volumes."}} +17:32:09 box language -> hu: http 302 +17:32:11 copies kept: vikunja_vikunja_data.pre-update-20260923T153032Z +vikunja_vikunja_db.pre-update-20260923T153032Z + diff --git a/documentation/audits/cleanup-2026-09-23/44-9202-held-badge-and-button.txt b/documentation/audits/cleanup-2026-09-23/44-9202-held-badge-and-button.txt new file mode 100644 index 00000000..0bcdbf27 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/44-9202-held-badge-and-button.txt @@ -0,0 +1,16 @@ +17:32:35 box language -> hu: http 302 +17:32:36 box=hu reader=hu: badge=('Megállítva — visszaállítás szükséges', 'A frissítés nem sikerült, és az automatikus visszaállítás sem.') +17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True +17:32:36 API: phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új v' +17:32:36 box=hu reader=en: badge=('Stopped — restore needed', 'The update did not succeed, and the automatic undo did not either.') +17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True +17:32:36 API: phase=failed err='The update did not succeed, and the automatic undo did not either. The data is a' +17:32:36 box language -> en: http 302 +17:32:36 box=en reader=hu: badge=('Megállítva — visszaállítás szükséges', 'A frissítés nem sikerült, és az automatikus visszaállítás sem.') +17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True +17:32:36 API: phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új v' +17:32:36 box=en reader=en: badge=('Stopped — restore needed', 'The update did not succeed, and the automatic undo did not either.') +17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True +17:32:36 API: phase=failed err='The update did not succeed, and the automatic undo did not either. The data is a' +17:32:36 the button's endpoint still refuses: 409 {'ok': False, 'data': {'reason': 'held'}, 'error': 'The update did not succeed, and the automatic undo did not either. The data is as the new version left it. T +17:32:36 box language -> hu: http 302 diff --git a/documentation/audits/cleanup-2026-09-23/45-9202-oom-storm.txt b/documentation/audits/cleanup-2026-09-23/45-9202-oom-storm.txt new file mode 100644 index 00000000..700668ef --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/45-9202-oom-storm.txt @@ -0,0 +1,13 @@ +15:32:53 +== the scan's own lines (first 4) +2026/09/23 15:30:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 8, limit 320M, peak 320M) +2026/09/23 15:30:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 13, limit 320M, peak 320M) +2026/09/23 15:31:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 21, limit 320M, peak 320M) +2026/09/23 15:31:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 28, limit 320M, peak 320M) +== storm + dropped lines +2026/09/23 15:30:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom (severity warning) — further app_oom events are logged at DEBUG only +2026/09/23 15:31:10 notifier.go:738: [ERROR] [notify] romm: container romm OOM STORM — 21 kills in 30 min (limit 320M, peak 320M) +2026/09/23 15:31:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom_storm (severity error) — further app_oom_storm events are logged at DEBUG only +== kernel counter now +oom_kill 45 +started=2026-09-23T15:28:39.049958056Z restarts=0 oom=true diff --git a/documentation/audits/cleanup-2026-09-23/46-9202-logs-before-teardown.txt b/documentation/audits/cleanup-2026-09-23/46-9202-logs-before-teardown.txt new file mode 100644 index 00000000..6d5e1bdf --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/46-9202-logs-before-teardown.txt @@ -0,0 +1,104 @@ +== 1b (v0.264.0) docker events for outline +1790174388 image pull outlinewiki/outline +1790174388 image pull outlinewiki/outline +1790174388 network create outline_outline-internal +1790174388 container create outline-redis +1790174388 container create outline-postgres +1790174712 container destroy outline-postgres +1790174712 container destroy outline-redis +1790174712 network destroy outline_outline-internal +== fix run (v0.265.0) docker events for outline +1790176826 network create outline_outline-internal +1790176826 container create outline-redis +1790176826 container create outline-postgres +1790176826 container create outline +1790176827 network connect outline_outline-internal +1790176827 network connect outline_outline-internal +1790176827 container start outline-redis +1790176827 container start outline-postgres +1790176832 container exec_create: redis-cli ping outline-redis +1790176832 container exec_start: redis-cli ping outline-redis +1790176832 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176832 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176832 container exec_die outline-redis +1790176832 container health_status: healthy outline-redis +1790176832 container exec_die outline-postgres +1790176832 container health_status: healthy outline-postgres +1790176832 network connect outline_outline-internal +1790176832 container start outline +1790176837 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176837 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176837 container exec_die outline +1790176842 container exec_create: redis-cli ping outline-redis +1790176842 container exec_start: redis-cli ping outline-redis +1790176842 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176842 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176842 container exec_die outline-redis +1790176842 container exec_die outline-postgres +1790176842 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176842 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176843 container exec_die outline +1790176846 container exec_create: pg_dump -U outline -d outline --clean --if-exists --no-owner --no-privileges outline-postgres +1790176846 container exec_start: pg_dump -U outline -d outline --clean --if-exists --no-owner --no-privileges outline-postgres +1790176846 container exec_die outline-postgres +1790176848 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176848 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176848 container exec_die outline +1790176848 container kill outline +1790176852 container exec_create: redis-cli ping outline-redis +1790176852 container exec_start: redis-cli ping outline-redis +1790176852 container exec_die outline-redis +1790176852 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176852 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176852 container exec_die outline-postgres +1790176853 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176853 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +1790176853 container exec_die outline +1790176854 network disconnect outline_outline-internal +1790176854 container stop outline +1790176854 container die outline +1790176854 container destroy outline +1790176854 container kill outline-redis +1790176854 container kill outline-postgres +1790176854 network disconnect outline_outline-internal +1790176854 container stop outline-redis +1790176854 container die outline-redis +1790176854 container destroy outline-redis +1790176854 network disconnect outline_outline-internal +1790176854 container stop outline-postgres +1790176854 container die outline-postgres +1790176854 container destroy outline-postgres +1790176855 network destroy outline_outline-internal +1790176856 network create outline_outline-internal +1790176856 container create outline-redis +1790176856 container create outline-postgres +1790176856 container create outline +1790176856 network connect outline_outline-internal +1790176856 network connect outline_outline-internal +1790176856 container start outline-postgres +1790176856 container start outline-redis +1790176861 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176861 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres +1790176861 container exec_create: redis-cli ping outline-redis +1790176861 container exec_start: redis-cli ping outline-redis +1790176861 container exec_die outline-postgres +1790176861 container health_status: healthy outline-postgres +1790176861 container exec_die outline-redis +1790176861 container health_status: healthy outline-redis +1790176862 network connect outline_outline-internal +1790176862 container start outline +1790176867 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline +== storm log: every OOM/STORM line +2026/09/23 15:30:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 8, limit 320M, peak 320M) +2026/09/23 15:30:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom (severity warning) — further app_oom events are logged at DEBUG only +2026/09/23 15:30:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 13, limit 320M, peak 320M) +2026/09/23 15:31:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 21, limit 320M, peak 320M) +2026/09/23 15:31:10 notifier.go:738: [ERROR] [notify] romm: container romm OOM STORM — 21 kills in 30 min (limit 320M, peak 320M) +2026/09/23 15:31:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom_storm (severity error) — further app_oom_storm events are logged at DEBUG only +2026/09/23 15:31:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 28, limit 320M, peak 320M) +2026/09/23 15:32:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 34, limit 320M, peak 320M) +2026/09/23 15:32:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 40, limit 320M, peak 320M) +2026/09/23 15:33:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 49, limit 320M, peak 320M) +== kernel counter at teardown +oom_kill 49 +recorders left: 0 diff --git a/documentation/audits/cleanup-2026-09-23/49-9202-teardown.txt b/documentation/audits/cleanup-2026-09-23/49-9202-teardown.txt new file mode 100644 index 00000000..a871e0d1 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/49-9202-teardown.txt @@ -0,0 +1,52 @@ +17:33:26 === remove vikunja through the product +17:33:26 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +17:33:58 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +17:34:06 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +17:34:06 === remove romm through the product +17:34:17 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'} +17:34:22 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +17:34:22 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +17:34:49 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat +17:34:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm' +17:34:58 undo copies left: 0 +17:35:00 volumes left: 0 +17:35:02 romm drive folder: removed-by-name +17:35:10 rmi vikunja/vikunja:2.3.0 +rmi vikunja/vikunja:2.6.0 +rmi rommapp/romm:5.3.0 +rmi mariadb:11.4 +rmi outlinewiki/outline:1.9.1 +KEPT postgres:16-alpine (in use by 1) +rmi codewithcj/sparkyfitness:v0.17.3 +rmi codewithcj/sparkyfitness_server:v0.17.3 +rmi postgres:15-alpine +rmi gitea.dooplex.hu/admin/felhom-controller:0.264.0 +r634-scratch-removed + +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +17:36:00 catalog cache: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462) +https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + +17:36:02 controller.yaml vs saved: IDENTICAL + +17:36:04 language: hu + +17:36:04 deployed: ['gokapi', 'paperless-ngx', 'privatebin'] +17:36:06 felhom-controller Up 54 seconds (healthy) +filebrowser Up 5 hours (healthy) +gokapi Restarting (1) 9 seconds ago +paperless-postgres Up 12 minutes (healthy) +paperless-redis Up 12 minutes (healthy) +paperless-webserver Up 12 minutes (healthy) +privatebin Up 12 minutes (healthy) +traefik Up 5 hours +gitea.dooplex.hu/admin/felhom-controller:0.265.0 + diff --git a/documentation/audits/cleanup-2026-09-23/50-floor-0.265.0.txt b/documentation/audits/cleanup-2026-09-23/50-floor-0.265.0.txt new file mode 100644 index 00000000..c67b56ec --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/50-floor-0.265.0.txt @@ -0,0 +1,16 @@ +=== FLOOR -> 0.265.0, MinAgent 0.131.0 declared (2026-09-23 17:36:28) +POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set +--- read back from the hub: +"min_controller_version" value="0.265.0" +"min_agent" value="0.131.0" +"min_agent" value="0.131.0" +--- watching the two demo boxes (up to 3 min) + +10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up 3 hours (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up 3 hours (healthy) + +20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.265.0 Up 7 seconds (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.265.0 Up 10 seconds (healthy) +--- hub log: +2026/09/23 17:36:31 [INFO] managed floor SERVED for demo-felhom: floor 0.265.0, agent requirement "0.131.0" from declared (golden 0.258.0) +2026/09/23 17:36:31 [INFO] managed floor SERVED for demo-hp: floor 0.265.0, agent requirement "0.131.0" from declared (golden 0.258.0) +--- hosts page: +demo-felhom-8363b5 +demo-hp-bb76ea +drill-r50-0a4f9a diff --git a/documentation/audits/cleanup-2026-09-23/README.md b/documentation/audits/cleanup-2026-09-23/README.md new file mode 100644 index 00000000..44bf496b --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/README.md @@ -0,0 +1,50 @@ +# Clean-up evening — R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0), 2026-09-23 + +**Method: endpoint-level** on scratch guest 9202 (the endpoints the UI invokes; pages as HTML with +`?lang=hu|en`); `docker events` + the controller log recorded inside the guest from before the first +call (`setsid docker events/logs -f` to a file — never piped through a buffering command). No browser. +Harness copied from `undo-fleet-2026-09-23/`; **`walk.py backup_now` presses nothing now (R-648)**. + +## Part 1 — R-634, diagnosed before any code + +| # | what | result | evidence | +|---|---|---|---| +| 1a | `sparkyfitness` alone, live pin, v0.264.0 | **did NOT reproduce**: deployed in 47.3 s, running | `10-*`, `11-*` | +| 1b | `outline` + whole-box backup at +20 s, v0.264.0 | **reproduced**: `StopStack outline: state=deploying deployed=true containers=0` (14:39:32) → `Restarting outline after volume dump` (14:39:33) → both `compose up` fail at 14:39:48 → `deployed=false`, containers `Created` | `12-*`, `13-*` | +| fix, try 1 | same, v0.265.0 | **not a race** — images cached, deploy done in 6.3 s before the backup came; kept as a record, not as proof | `30-fix-attempt1-*` | +| fix, try 2 | outline image removed by name, backup at +5 s, v0.265.0 | **the backup ran 15:23:07–15:23:31 across the deploy, stopped gokapi, paperless-ngx, privatebin and never outline**; deploy `successfully (took 48.9s)`, `deployed=true`, running | `32-*`, `33-*` | + +Mechanism at `file:line`: `stacks/deploy.go:395` (in-memory `Deployed=true` at accept) → +`cmd/controller/main.go:2556` `ListDeployedStacks` (flag only) → `backup/backup.go:693` `runVolumeDumps` → +`DumpAppVolumesSafe` `StopStack` + `StartStack` (`backup.go:908/918`) → `stacks/deploy.go:421-435` failure +branch writes `Deployed=false`. There is no deploy time limit (the brief's shape (i) is ruled out). + +## Parts 2–4 (live, 9202, drill catalog: romm 320M/4 workers, vikunja at 2.3.0) + +| proof | result | evidence | +|---|---|---| +| R-625 badge, 4 box/reader pairs | hu reader: „Megállítva — visszaállítás szükséges", title „A frissítés nem sikerült, és az automatikus visszaállítás sem."; en reader: "Stopped — restore needed", title "The update did not succeed, and the automatic undo did not either."; no „Frissítés elérhető"/"Update available"; **no vikunja Update button** on the list while other apps keep theirs (control); API `POST …/update` → **409 `held`** | `44-*` | +| R-647 (1) | box **en**, `?lang=hu`: the hold in Hungarian on the app page, and `update_error` Hungarian in the API | `43-*`, `44-*` | +| R-648 | vikunja's update ran its own `backing-up` phase; no whole-box press | `42-*`, `43-*` | +| R-636 | romm: kills 8 → 13 → **21 → `OOM STORM — 21 kills in 30 min (limit 320M, peak 320M)`** + `DROPPED event app_oom_storm (severity error)` (R-620 witness, 9202 has no hub); still ONE storm at 49 kills | `45-*`, `46-*` | + +Hub v0.121.0 deployed first (`20-*`). The hub side of R-636 is proven by unit tests only: posting a +synthetic event with a real customer's key would put a fabricated alarm into the operator's record. + +## Floor + +0.265.0 / MinAgent 0.131.0, read back; demo-hp and demo-felhom both on 0.265.0 within 20 s (`50-*`). + +## Teardown — three layers + +- **machine (9202):** outline, sparkyfitness, vikunja, romm removed through the product (romm's + with-data remove refused 409 on the drive path, R-442 as in every drill — its drive folder removed by + name); 0 volumes, 0 undo copies; test images removed by name only where no container used them + (`postgres:16-alpine` kept — in use); `controller.yaml` identical to `.pre-cleanup`; catalog on live + `cfcfe52`; language `hu`; recorders stopped and `/root/r634` removed; standing apps as at the start. +- **host:** nothing provisioned on demo-hp. **9201 not touched** (no backup press; it reached 0.265.0 by + the floor). +- **hub:** v0.121.0 (planned); floor 0.265.0. +- **DooPlex:** one empty volume `outline_outline_data` and one `alpine:3.20` image were created by MY first + test draft / my inspection of it; both verified unused and removed by name (R-650). +- **drill repo:** reset to live `main`, read back from the remote. diff --git a/documentation/audits/cleanup-2026-09-23/bakeoff.py b/documentation/audits/cleanup-2026-09-23/bakeoff.py new file mode 100644 index 00000000..7bd759f9 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/bakeoff.py @@ -0,0 +1,266 @@ +#!/usr/bin/env python3 +"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on +the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202, +drill catalog. + + F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure. + D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and + recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads. + +EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync, +rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is +lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own +Start supplies the app's secrets (they are never decrypted here). + +Usage: python3 bakeoff.py stages: prep fcopy break update undoF undoD cutoff state +""" +import json, os, sys, time +sys.path.insert(0, ".") +import walk as w +from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \ + romm_verify_b, vik_seed_b, vik_verify_b + +EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999), + "romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999), + "vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)} +UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost", + "vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja", + "romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"} +SUFFIX = ".pre-undo" + + +def seed_b(app, sub, A): + return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A) + + +def verify_b(app, sub, A, B): + r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B) + return all(r.values()) if isinstance(r, dict) else r + + +def db_state(app): + """The table count and the migration ledger, asked of the engine (or the SQLite file, read-only).""" + if app == "docmost": + return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; " + "docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip() + if app == "romm": + q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '" + base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45") + alembic = q("select version_num from alembic_version") + return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip() + return w.guest("""python3 - <<'PY' +import sqlite3 +c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True) +t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0] +m=c.execute("select count(*), max(id) from migration").fetchone() +print(f"tables={t} ledger={m[0]} newest={m[1]}") +PY""").strip() + + +def volumes(app): + return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split() + if not v.endswith(SUFFIX)] + + +def containers(app): + return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split() + + +def stage_prep(app): + s = {"app": app, "sub": SUB[app]} + w.say(f"=== prep {app}: deploy at the drill FROM pin") + w.say("deployed:", w.deploy(app, s["sub"])) + s["seedA"] = FX[app].seed(w, s["sub"], w.say) + w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say)) + w.backup_now(app) + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60) + s["seedB"] = seed_b(app, s["sub"], s["seedA"]) + w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"])) + s["volumes"] = volumes(app) + w.say("named volumes:", s["volumes"]) + save(app, s) + + +def stage_fcopy(app): + """F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and + separately a tar, for the rate), start. Measures the EXTRA downtime and the disk.""" + s = load(app) + s["state_before_update"] = db_state(app) + w.say("db state before the update:", s["state_before_update"]) + vols = s["volumes"] + w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml")) + script = f""" +set -u +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}') +t0=$(date +%s.%N) +docker stop $cs >/dev/null +t1=$(date +%s.%N) +for v in {' '.join(vols)}; do + docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1 + docker volume create "$v{SUFFIX}" >/dev/null + a=$(date +%s.%N) + docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v" + b=$(date +%s.%N) + bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1) + files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l') + cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1) + cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l') + mkdir -p /root/bk; c=$(date +%s.%N) + docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N) + tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar + python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))" +done +t2=$(date +%s.%N) +docker start $cs >/dev/null +t3=$(date +%s.%N) +python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))" +df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}' +""" + out = w.guest(script, timeout=1800) + w.say(out) + t0 = time.time() + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + w.say(f"front door answering again {round(time.time()-t0,1)}s after the start") + s["fcopy"] = out + s["fcopy_at"] = ts() + save(app, s) + + +def stage_break(app): + s = load(app) + frm, to, p0, p1 = EDGE[app] + s["break_commit"] = break_edge(app, frm, to, p0, p1) + w.say("badge caught up after", w.sync_rescan(app, to), "s") + save(app, s) + + +def stage_update(app): + s = load(app) + r = w.press_update(app, poll=1.0) + w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False)) + s.setdefault("updates", []).append(r) + save(app, s) + + +LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold' +docker restart felhom-controller >/dev/null; sleep 15""" + +# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre--compose.yml, taken +# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s +# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new +# version again. That is R-639 seen live, and why the product undo keeps its own copies. +PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml +grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }} +cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml +sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4""" + + +def lift_and_pinback(app): + frm, to, _, _ = EDGE[app] + t = time.time() + w.say(w.guest(LIFT.format(app=app))) + w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to))) + w.login() + return round(time.time() - t, 1) + + +def stage_undoF(app): + s = load(app) + w.say(f"=== undo by F: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + vols = s["volumes"] + script = f""" +t0=$(date +%s.%N) +cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null +for v in {' '.join(vols)}; do + docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v" +done +python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))" +""" + w.say(w.guest(script, timeout=1800)) + t0 = time.time() + c, d = w.ctl("POST", f"/api/stacks/{app}/start") + w.say("product start ->", c) + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_undoD(app): + """D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy.""" + s = load(app) + w.say(f"=== undo by D: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + if app == "vikunja": + w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it " + "is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.") + return + c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse + time.sleep(8) + if app == "docmost": + load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1""" + marker = "-- PostgreSQL database dump complete" + else: + load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1""" + marker = "-- Dump completed" + script = f""" +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B" +docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null +t0=$(date +%s.%N) +if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi +{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out +python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))" +docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null +""" + w.say(w.guest(script, timeout=900)) + t0 = time.time() + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_cutoff(app): + """The cut-off copy, both methods, DETECTED before anything is loaded or swapped. + F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker + the build would write only after cp exits 0. D: the safety dump cut in half — the marker check.""" + s = load(app) + big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0)) + script = f""" +v={big} +docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null +# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured +# on docmost 2026-09-23, rc 137 with a complete copy and a marker). +docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null +sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1 +echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)" +echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap" +docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1) +if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql + echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load" + rm -f /root/cut.sql +else echo "D: no safety dump exists for this app (R-641)"; fi +""" + w.say(f"=== cut-off copy: {app} (largest volume {big})") + out = w.guest(script, timeout=600) + w.say(out) + s["cutoff"] = out + save(app, s) + + +if __name__ == "__main__": + w.login() + {"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update, + "undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff, + "state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/cleanup-2026-09-23/fixtures.py b/documentation/audits/cleanup-2026-09-23/fixtures.py new file mode 100644 index 00000000..3a681e96 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/fixtures.py @@ -0,0 +1,1008 @@ +#!/usr/bin/env python3 +"""Box-side seed/verify fixtures for walk.py, guest 9202. + +THE ONE RULE (R-156), carried verbatim from `app-catalog-felhom.eu/scripts/upgrade_fixtures.py`: +*nothing is ever seeded into a volume by hand.* Every seed here goes in through the app's OWN +interface — its HTTP API through the household's real front door (traefik, `Host: .`), +or its own CLI running inside its own container. A raw SQL INSERT or a planted file is never used. + +If an app has no non-browser route, its fixture returns None and the edge is recorded +`inconclusive — no non-browser seed route`, WITH WHAT WAS TRIED. That is a result, not a gap. + +Each fixture: + seed(w, sub, say) -> an opaque token, or None + verify(w, sub, tok, say) -> True / False +verify() must ask the APP, never the filesystem: a migration is supposed to rewrite files. +Where a fixture can prove itself (a negative control that must read as absent) it does so on EVERY +call, so a readback that has broken into always saying "found" fails instead of passing everything. +""" +import base64, json, re, secrets, time + + +def _gx(w, container, *cmd, timeout=240): + """Run a command inside the app's OWN container on 9202 (its own CLI, not our SQL).""" + import shlex + line = " ".join(shlex.quote(c) for c in cmd) + return w.guest(f"docker exec {container} {line} 2>&1", timeout=timeout) + + +# ============================================================================================= +class PrivateBin: + """PrivateBin's own JSON API. A paste is a POST and reading it back is a GET — an + application-level round trip. File-backed, no database: this single seed IS the file half.""" + sub = "paste" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200",)): + return None + marker = "upg-" + secrets.token_hex(8) + ct = base64.b64encode(marker.encode()).decode() + body = json.dumps({ + "v": 2, + "adata": [[base64.b64encode(secrets.token_bytes(16)).decode(), + base64.b64encode(secrets.token_bytes(8)).decode(), + 100000, 256, 128, "aes", "gcm", "none"], "plaintext", 0, 0], + "ct": ct, "meta": {"expire": "never"}}) + rc, code, out = w.app_curl(sub, "/", "-H", "X-Requested-With: JSONHttpRequest", + "-H", "Content-Type: application/json", + data=body, method="POST") + try: + j = json.loads(out) + except Exception: + say(f" privatebin: POST returned non-JSON (http {code}): {out[:200]}") + return None + if j.get("status") != 0 or not j.get("id"): + say(f" privatebin: POST refused: {out[:250]}") + return None + say(f" privatebin: seeded paste id={j['id']}") + return {"id": j["id"], "marker": ct} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200",), tries=36): + return False + # negative control, every call: a paste id that cannot exist must NOT read back + rc, code, out = w.app_curl(sub, "/?pasteid=" + secrets.token_hex(8), + "-H", "X-Requested-With: JSONHttpRequest") + if t["marker"] in out: + say(" privatebin: READBACK UNUSABLE — a paste id that cannot exist returned the marker") + return False + rc, code, out = w.app_curl(sub, "/?pasteid=" + t["id"], + "-H", "X-Requested-With: JSONHttpRequest") + got = code == "200" and t["marker"] in out + say(f" privatebin: readback http={code} marker_present={got}") + return got + + +# ============================================================================================= +class Docmost: + """Docmost's own REST API: create the first workspace+user, then prove the account survives by + asking the app to AUTHENTICATE it. Login is version-stable across the API churn.""" + sub = "docs" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302", "404")): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + body = json.dumps({"workspaceName": "drill", "name": "drill", "email": email, "password": pw}) + rc, code, out = w.app_curl(sub, "/api/auth/setup", "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" docmost: /api/auth/setup http={code} rc={rc}") + if code not in ("200", "201"): + say(f" docmost: setup refused: {out[:250]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "404"), tries=36): + return False + # negative control: a password that was never set must NOT authenticate + bad = json.dumps({"email": t["email"], "password": "definitely-" + secrets.token_hex(8)}) + rc, code, _ = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" docmost: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"email": t["email"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" docmost: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" docmost: login body {out[:200]}") + return ok + + +# ============================================================================================= +class BookStack: + """BookStack mints no API token without a browser, so BOTH halves go through `php artisan` — + BookStack's OWN CLI, inside its own container, against its own User model. + + The exit code carries no information here (`bookstack:reset-mfa` exits 1 for a user it FOUND + and for one it did not), so the discriminator is the OUTPUT: the positive sentence required and + the not-found sentence required absent. The negative control runs on every verify. + + LIMITATION (R-460): this seeds the DATABASE half only. The FILE half needs the API token the + app cannot mint headlessly — so a bookstack edge is at best HALF-proven here. + """ + sub = "wiki" + + def _artisan(self, w, *args): + for path in ("/app/www/artisan", "/var/www/html/artisan"): + out = _gx(w, "bookstack", "php", path, *args) + if "Could not open input file" not in out: + return " ".join(out.split()) + return " ".join(out.split()) + + def _lookup(self, w, email): + out = self._artisan(w, "bookstack:reset-mfa", f"--email={email}") + found = f"Email: {email}" in out + missing = "could not be found" in out + if found == missing: + return None, out + return found, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + out = self._artisan(w, "bookstack:create-admin", f"--email={email}", + f"--name=drill-{secrets.token_hex(3)}", f"--password={pw}") + say(f" bookstack: artisan create-admin :: {out[:140]}") + if "successfully created" not in out: + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" bookstack: the app never served /login") + return False + absent, _ = self._lookup(w, f"nobody-{secrets.token_hex(6)}@gate.invalid") + if absent is not False: + say(f" bookstack: READBACK UNUSABLE — an email that cannot exist did not read absent ({absent})") + return False + found, out = self._lookup(w, t["email"]) + say(f" bookstack: readback of the seeded account found={found} :: {out[:140]}") + return found is True + + +# ============================================================================================= +class Gitea: + """Gitea's own admin CLI creates the first user; its own REST API (basic auth) then creates a + repository and reads it back. Both are the app's own interfaces.""" + sub = "git" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = _gx(w, "gitea", "su", "git", "-c", + f"gitea admin user create --username {user} --password {pw} " + f"--email {user}@gate.invalid --admin --must-change-password=false") + say(f" gitea: admin user create :: {' '.join(out.split())[:140]}") + if "has been successfully created" not in out and "successfully created" not in out: + return None + repo = "drillrepo" + secrets.token_hex(3) + rc, code, body = w.app_curl(sub, "/api/v1/user/repos", "-u", f"{user}:{pw}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": repo, "private": True}), method="POST") + say(f" gitea: create repo http={code}") + if code not in ("201", "200"): + say(f" gitea: repo refused {body[:200]}") + return None + return {"user": user, "pw": pw, "repo": repo} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + rc, code, _ = w.app_curl(sub, f"/api/v1/repos/{t['user']}/nope{secrets.token_hex(4)}", + "-u", f"{t['user']}:{t['pw']}") + if code == "200": + say(" gitea: READBACK UNUSABLE — a repo that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/repos/{t['user']}/{t['repo']}", + "-u", f"{t['user']}:{t['pw']}") + ok = code == "200" and t["repo"] in body + say(f" gitea: readback of the seeded repo http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Navidrome: + """Navidrome's own REST API: create the first admin through /auth/createAdmin, then prove the + account survives by logging in through the same door.""" + sub = "music" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, out = w.app_curl(sub, "/auth/createAdmin", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), method="POST") + say(f" navidrome: createAdmin http={code}") + if code not in ("200", "201"): + say(f" navidrome: refused {out[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" navidrome: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"username": t["user"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" navidrome: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vaultwarden: + """Vaultwarden's own account API: register an account, then prove it survives by asking the app + to issue a token for it (its own login endpoint, the household's own route).""" + sub = "vault" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/alive", want=("200",)): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + # Vaultwarden stores an already-hashed master key; the value is opaque to the server. + key = base64.b64encode(secrets.token_bytes(32)).decode() + body = json.dumps({"email": email, "name": "drill", "masterPasswordHash": key, + "key": "0." + base64.b64encode(secrets.token_bytes(48)).decode(), + "kdf": 0, "kdfIterations": 600000}) + rc, code, out = w.app_curl(sub, "/api/accounts/register", + "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" vaultwarden: register http={code}") + if code not in ("200", "204"): + say(f" vaultwarden: refused {out[:250]}") + return None + return {"email": email, "key": key} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/alive", want=("200",), tries=36): + return False + def login(pwhash): + return w.app_curl(sub, "/identity/connect/token", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=("grant_type=password&scope=api%20offline_access" + f"&client_id=web&deviceType=9&deviceIdentifier=drill" + f"&deviceName=drill&username={t['email']}&password={pwhash}"), + method="POST") + rc, code, _ = login(base64.b64encode(secrets.token_bytes(32)).decode()) + if code == "200": + say(" vaultwarden: READBACK UNUSABLE — a wrong master key authenticated") + return False + rc, code, out = login(t["key"].replace("+", "%2B").replace("=", "%3D").replace("/", "%2F")) + ok = code == "200" and "access_token" in out + say(f" vaultwarden: token for the seeded account http={code} ok={ok}") + if not ok: + say(f" vaultwarden: body {out[:200]}") + return ok + + +# ============================================================================================= +class Django: + """A Django app's OWN management CLI, inside its own container, against its own User model. + + Same category as BookStack's `php artisan`: the app's own code and its own ORM, never a raw SQL + INSERT and never a planted file (R-156). `createsuperuser --noinput` is Django's own documented + non-interactive route, and the readback asks the SAME ORM whether the account exists. + + THE FIXTURE PROVES ITSELF ON EVERY CALL: each verify() also asks for a username that cannot + exist and requires the answer False. A readback that has broken into always saying True + therefore fails instead of passing everything. + + LIMITATION, recorded rather than papered over: this seeds the DATABASE half only. An app whose + data is also FILES (adventurelog's images) has a file half this fixture does not touch. + """ + + def __init__(self, container, sub, ready_path="/", ready=("200", "302", "301", "404"), + python="python", workdir=None): + # `python` and `workdir` are per-app because the image decides them: adventurelog's + # interpreter is on PATH, tandoor ships a VENV and the bare `python` cannot import Django + # at all ("Couldn't import Django. Are you sure it's installed…"). Measured, not guessed. + self.container = container + self.sub = sub + self.ready_path = ready_path + self.ready = ready + self.python = python + self.workdir = workdir + + def _wd(self): + return f"-w {self.workdir} " if self.workdir else "" + + def _manage(self, w, code): + # -c is passed to `manage.py shell`; the app's own shell, its own ORM. + return w.guest( + f"docker exec {self._wd()}{self.container} {self.python} manage.py shell " + f"-c {json.dumps(code)} 2>&1", timeout=300) + + def _exists(self, w, username): + # ONE LINE, semicolon-separated. A `\n` inside a double-quoted shell argument reaches + # python as a literal backslash-n and is a SyntaxError — which is exactly how the first + # adventurelog run read as `inconclusive`. The fixture refused to guess, which is right, + # but the instrument was the thing that was broken. + out = self._manage(w, ( + "from django.contrib.auth import get_user_model; " + f"print('DRILL_ANSWER=' + str(get_user_model().objects.filter(username={username!r}).exists()))" + )) + m = re.search(r"DRILL_ANSWER=(True|False)", out) + return (m.group(1) == "True") if m else None, " ".join(out.split())[-300:] + + def seed(self, w, sub, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -e DJANGO_SUPERUSER_PASSWORD={pw} {self._wd()}{self.container} " + f"{self.python} manage.py createsuperuser --noinput " + f"--username {user} --email {user}@gate.invalid 2>&1", timeout=300) + say(f" {self.container}: createsuperuser :: {' '.join(out.split())[:160]}") + got, detail = self._exists(w, user) + if got is not True: + say(f" {self.container}: the account did not appear in the app's own ORM :: {detail[:200]}") + return None + say(f" {self.container}: seeded superuser {user}") + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + say(f" {self.container}: the app never served {self.ready_path}") + return False + absent, detail = self._exists(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" {self.container}: READBACK UNUSABLE — a username that cannot exist did not " + f"read as absent ({absent}) :: {detail[:200]}") + return False + found, detail = self._exists(w, t["user"]) + say(f" {self.container}: readback of the seeded account found={found}") + if found is not True: + say(f" {self.container}: :: {detail[:250]}") + return found is True + + +# ============================================================================================= +class Nextcloud: + """Nextcloud's OWN admin CLI, `occ`, inside its own container: its own code, its own user + backend. Not a SQL INSERT and not a planted file (R-156). + + `occ user:info` is the readback, and it PROVES ITSELF on every call: a uid that cannot exist + must answer "user not found". A readback that has broken into always succeeding therefore + fails instead of passing everything. + + This is the app chosen for the MariaDB engine-major edge (`09` §3 decision 5, R-469 lifted): + the app image does NOT move, only the `mariadb:` sidecar, so the edge carries exactly one + migration and a failure is readable. + """ + sub = "cloud" + + def _occ(self, w, *args, timeout=420): + import shlex + line = " ".join(shlex.quote(a) for a in args) + return w.guest(f"docker exec -u www-data nextcloud php occ {line} 2>&1", timeout=timeout) + + def _info(self, w, uid): + out = self._occ(w, "user:info", uid) + flat = " ".join(out.split()) + if "user not found" in flat.lower() or "could not be found" in flat.lower(): + return False, flat + if f"user_id: {uid}" in flat or f"- user_id: {uid}" in flat or f"user_id: {uid}" in out: + return True, flat + return None, flat + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + return None + uid = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -u www-data -e OC_PASS={pw} nextcloud php occ user:add " + f"--password-from-env --display-name={uid} {uid} 2>&1", timeout=420) + say(f" nextcloud: occ user:add :: {' '.join(out.split())[:160]}") + got, flat = self._info(w, uid) + if got is not True: + say(f" nextcloud: the account did not appear via occ user:info :: {flat[:220]}") + return None + say(f" nextcloud: seeded user {uid}") + return {"uid": uid, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + say(" nextcloud: the app never served /status.php") + return False + absent, flat = self._info(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" nextcloud: READBACK UNUSABLE — a uid that cannot exist did not read absent " + f"({absent}) :: {flat[:200]}") + return False + found, flat = self._info(w, t["uid"]) + say(f" nextcloud: readback of the seeded user found={found}") + if found is not True: + say(f" nextcloud: :: {flat[:250]}") + return found is True + + +# ============================================================================================= +class Grafana: + """Grafana's own HTTP API as the admin the DEPLOY created. The password is the one the + controller showed the household — read from the app's own `app.yaml`, not invented — and the + data (a folder) goes in and comes back through the app's own REST API.""" + sub = "grafana" + + def _auth(self, w, name="grafana"): + # app.yaml stores this ENCRYPTED (`ENC:…`), so it cannot be read back off the box — which + # is correct, and is why the harness uses the value IT generated for the deploy. + pw = (w.GENERATED.get(name) or {}).get("GF_SECURITY_ADMIN_PASSWORD") or "admin" + return f"admin:{pw}" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return None + au = self._auth(w) + title = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/folders", "-u", au, + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="POST") + say(f" grafana: create folder http={code}") + if code not in ("200", "201"): + say(f" grafana: refused {body[:220]}") + return None + try: + uid = json.loads(body)["uid"] + except Exception: + say(f" grafana: no uid in {body[:200]}") + return None + return {"uid": uid, "title": title} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return False + au = self._auth(w) + rc, code, _ = w.app_curl(sub, "/api/folders/nope" + secrets.token_hex(5), "-u", au) + if code == "200": + say(" grafana: READBACK UNUSABLE — a folder uid that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/folders/{t['uid']}", "-u", au) + ok = code == "200" and t["title"] in body + say(f" grafana: readback of the seeded folder http={code} ok={ok}") + return ok + + +# ============================================================================================= +class AudiobookShelf: + """audiobookshelf's own /init endpoint creates the first root account; its own /login proves + the account survived. Both are the app's own API.""" + sub = "audiobooks" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/init", "-H", "Content-Type: application/json", + data=json.dumps({"newRoot": {"username": user, "password": pw}}), + method="POST") + say(f" audiobookshelf: /init http={code}") + if code not in ("200", "204"): + say(f" audiobookshelf: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code == "200": + say(" audiobookshelf: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": t["pw"]}), + method="POST") + ok = code == "200" and t["user"] in body + say(f" audiobookshelf: login as the seeded root http={code} ok={ok}") + return ok + + +# ============================================================================================= +class ActualBudget: + """Actual's own bootstrap API sets the server password; its own login proves it survived.""" + sub = "budget" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/account/bootstrap", + "-H", "Content-Type: application/json", + data=json.dumps({"password": pw}), method="POST") + say(f" actualbudget: /account/bootstrap http={code} :: {body[:140]}") + if code not in ("200", "201") or '"status":"ok"' not in body: + return None + return {"pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def login(p): + return w.app_curl(sub, "/account/login", "-H", "Content-Type: application/json", + data=json.dumps({"loginMethod": "password", "password": p}), + method="POST") + rc, code, body = login("wrong-" + secrets.token_hex(6)) + if '"status":"ok"' in body: + say(" actualbudget: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = '"status":"ok"' in body + say(f" actualbudget: login with the seeded password http={code} ok={ok}") + if not ok: + say(f" actualbudget: body {body[:200]}") + return ok + + +# ============================================================================================= +class Mealie: + """Mealie ships a documented first-run admin. We log in as it through the app's own OAuth-style + token endpoint, create a recipe through the app's own API, and read the recipe back.""" + sub = "recipes" + + def _token(self, w, sub, pw="MyPassword"): + rc, code, body = w.app_curl( + sub, "/api/auth/token", "-H", "Content-Type: application/x-www-form-urlencoded", + data=f"username=changeme%40example.com&password={pw}", method="POST") + if code != "200": + return None, f"http={code} {body[:200]}" + try: + return json.loads(body)["access_token"], "" + except Exception: + return None, body[:200] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return None + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate as the first-run admin :: {why}") + return None + name = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/recipes", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": name}), method="POST") + say(f" mealie: create recipe http={code}") + if code not in ("200", "201"): + say(f" mealie: refused {body[:220]}") + return None + slug = body.strip().strip('"') + return {"slug": slug, "name": name} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return False + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate after the update :: {why}") + return False + rc, code, _ = w.app_curl(sub, "/api/recipes/nope" + secrets.token_hex(5), + "-H", f"Authorization: Bearer {tok}") + if code == "200": + say(" mealie: READBACK UNUSABLE — a slug that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/recipes/{t['slug']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["name"] in body + say(f" mealie: readback of the seeded recipe http={code} ok={ok}") + return ok + + +# ============================================================================================= +class N8n: + """n8n's own owner-setup API creates the first account; its own login proves it survived.""" + sub = "auto" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill" + secrets.token_hex(8) + "1" + rc, code, body = w.app_curl(sub, "/rest/owner/setup", "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "firstName": "drill", + "lastName": "drill", "password": pw}), + method="POST") + say(f" n8n: /rest/owner/setup http={code}") + if code not in ("200", "201"): + say(f" n8n: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return False + def login(p): + return w.app_curl(sub, "/rest/login", "-H", "Content-Type: application/json", + data=json.dumps({"emailOrLdapLoginId": t["email"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" n8n: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" and t["email"] in body + say(f" n8n: login as the seeded owner http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Zipline: + """Zipline's own setup/login API. Zipline 4 creates the first user through its own endpoint.""" + sub = "img" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/healthcheck", want=("200",), tries=90): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=30): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + for path in ("/api/auth/register", "/api/auth/setup"): + rc, code, body = w.app_curl(sub, path, "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), + method="POST") + say(f" zipline: {path} http={code} :: {body[:160]}") + if code in ("200", "201"): + return {"user": user, "pw": pw} + say(" zipline: neither register nor setup accepted a first user") + return None + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=60): + return False + def login(p): + return w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" zipline: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" + say(f" zipline: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vikunja: + """Vikunja's own REST API: register a user, log in, create a project, read the project back. + Four calls, all the app's own front door.""" + sub = "tasks" + + def _token(self, w, sub, t, pw=None): + rc, code, body = w.app_curl(sub, "/api/v1/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], + "password": pw or t["pw"]}), method="POST") + if code != "200": + return None, f"http={code} {body[:160]}" + try: + return json.loads(body)["token"], "" + except Exception: + return None, body[:160] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/v1/register", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw, + "email": f"{user}@gate.invalid"}), + method="POST") + say(f" vikunja: register http={code}") + if code not in ("200", "201"): + say(f" vikunja: refused {body[:220]}") + return None + t = {"user": user, "pw": pw} + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: could not log in after registering :: {why}") + return None + title = "drill-" + secrets.token_hex(5) + # Vikunja CREATES with PUT, not POST — a POST answers `405 Method Not Allowed`, which + # reads like a broken fixture and is really the wrong verb. Measured 2026-09-21. + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="PUT") + say(f" vikunja: create project http={code}") + if code not in ("200", "201"): + say(f" vikunja: project refused {body[:220]}") + return None + t["title"] = title + t["pid"] = json.loads(body).get("id") + return t + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return False + bad, why = self._token(w, sub, t, pw="wrong-" + secrets.token_hex(6)) + if bad: + say(" vikunja: READBACK UNUSABLE — a wrong password authenticated") + return False + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: the seeded account no longer authenticates :: {why}") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{t['pid']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["title"] in body + say(f" vikunja: readback of the seeded project http={code} ok={ok}") + return ok + + +# ============================================================================================= +class OpenGist: + """Opengist's own sign-up and sign-in FORMS. + + Two things had to be measured. Its sign-up is CSRF-protected: a bare POST answers 500 with an + HTML page, which reads like a broken app and is really a missing token — fetch the form, keep + its cookie, send its `_csrf` back. And its REST API refuses the account's own password + (`401 {"message":"Bad crendentials"}`) because it wants a token the app will not mint without a + browser. So the SEEDED DATA is the account itself and the READBACK is a real sign-in, which is + the same shape the docmost and navidrome fixtures use. + + LIMITATION, recorded rather than papered over: this is the DATABASE half. A gist's CONTENT is + not seeded, because that needs the API token above. + """ + sub = "gist" + + def _form(self, w, sub, path, jar, fields): + rc, code, html = w.app_curl(sub, path, "-b", jar, "-c", jar) + m = re.search(r'name="_csrf"[^>]*value="([^"]+)"', html or "") + if not m: + return None, f"no _csrf on {path} (http={code})" + body = "&".join([f"_csrf={m.group(1)}"] + [f"{k}={v}" for k, v in fields.items()]) + rc, code, out = w.app_curl(sub, path, "-b", jar, "-c", jar, + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + return code, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, out = self._form(w, sub, "/register", jar, {"username": user, "password": pw}) + say(f" opengist: /register (with its own _csrf) http={code}") + if code not in ("200", "302", "303"): + say(f" opengist: refused {str(out)[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + # Wait for the LOGIN FORM, not for the root page. Measured 2026-09-21: immediately after a + # successful update the root answers while /login does not yet carry its `_csrf`, so the + # sign-in silently fails and the app looks like it lost the account. It had not. + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" opengist: /login never came back after the update") + return False + for _ in range(24): + rc, code, html = w.app_curl(sub, "/login") + if code == "200" and '_csrf' in (html or ""): + break + time.sleep(5) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar, + {"username": t["user"], "password": "wrong-" + secrets.token_hex(5)}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar) + if t["user"] in (home or ""): + say(" opengist: READBACK UNUSABLE — a wrong password signed in") + return False + jar2 = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar2, {"username": t["user"], "password": t["pw"]}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar2) + ok = t["user"] in (home or "") + say(f" opengist: sign-in as the seeded account http={code} name_on_page={ok}") + return ok + + +# ============================================================================================= +class Papra: + """Papra's own e-mail sign-up and sign-in endpoints.""" + sub = "papra" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/auth/sign-up/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "password": pw, + "name": "drill"}), method="POST") + say(f" papra: sign-up http={code}") + if code not in ("200", "201"): + say(f" papra: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def signin(p): + return w.app_curl(sub, "/api/auth/sign-in/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": t["email"], "password": p}), method="POST") + rc, code, _ = signin("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" papra: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = signin(t["pw"]) + ok = code == "200" + say(f" papra: sign-in as the seeded account http={code} ok={ok}") + return ok + + +# ============================================================================================= +class HomeAssistant: + """Home Assistant's own onboarding API creates the owner account and hands back a code the + same API exchanges for a token. Both are the app's own documented non-browser route.""" + sub = "ha" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/onboarding/users", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "name": "drill", "username": user, + "password": pw, "language": "en"}), + method="POST") + say(f" home-assistant: /api/onboarding/users http={code}") + if code not in ("200", "201"): + say(f" home-assistant: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def _login(self, w, sub, user, pw): + """The app's own login flow: start it, then answer it. A 200 with a step_id of + `mfa`/`init` means the credentials were REFUSED; only `create_entry` is a pass.""" + rc, code, body = w.app_curl(sub, "/auth/login_flow", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "handler": ["homeassistant", None], + "redirect_uri": f"https://{sub}.felhom.invalid/", + "type": "authorize"}), method="POST") + if code not in ("200", "201"): + return None, f"flow start http={code} {body[:160]}" + try: + fid = json.loads(body)["flow_id"] + except Exception: + return None, body[:160] + rc, code, body = w.app_curl(sub, f"/auth/login_flow/{fid}", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "username": user, "password": pw}), + method="POST") + try: + j = json.loads(body) + except Exception: + return None, body[:160] + return (j.get("result") if j.get("type") == "create_entry" else None), body[:200] + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return False + bad, why = self._login(w, sub, t["user"], "wrong-" + secrets.token_hex(6)) + if bad: + say(" home-assistant: READBACK UNUSABLE — a wrong password authenticated") + return False + good, why = self._login(w, sub, t["user"], t["pw"]) + ok = bool(good) + say(f" home-assistant: login as the seeded owner ok={ok}") + if not ok: + say(f" home-assistant: {why}") + return ok + + +# ============================================================================================= +class Romm: + """RomM's own user API, driven the way RomM's own front end drives it. + + Three things had to be measured rather than guessed, and each one answered a 403 or a 422 that + looked like a different fault: RomM sets a **`romm_csrftoken` cookie** on any GET and requires + it back in an **`x-csrftoken` header** (a bare POST is `403 CSRF token verification failed`, + which reads like an auth problem); the fields go in the **JSON body**, not the query string (a + query-string POST is `422 Field required` for every field it was just given); and `email` is + required alongside username, password and role. + + On a fresh install with no admin the first `POST /api/users` is accepted unauthenticated; + afterwards it is not — which is what makes the readback (`POST /api/login` as that user) a real + authentication rather than a repeat of the seed. + + LIMITATION: this is the DATABASE half. RomM's other half is the ROM library on the drive, which + this does not populate. + """ + sub = "arcade" + + def _csrf(self, w, sub): + jar = f"/tmp/romm-{secrets.token_hex(4)}.jar" + w.app_curl(sub, "/api/heartbeat", "-c", jar) + out = w.sh(["bash", "-lc", f"grep -i csrf {jar} | awk '{{print $7}}'"]).stdout or "" + return jar, out.strip() + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + jar, tok = self._csrf(w, sub) + if not tok: + say(" romm: no romm_csrftoken cookie was set on /api/heartbeat") + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl( + sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": pw, "role": "admin"}), method="POST") + say(f" romm: POST /api/users http={code}") + if code not in ("200", "201"): + say(f" romm: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + return False + jar, tok = self._csrf(w, sub) + rc, code, _ = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:wrong-{secrets.token_hex(5)}", method="POST") + if code == "200": + say(" romm: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:{t['pw']}", method="POST") + ok = code == "200" + say(f" romm: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" romm: body {body[:200]}") + return ok + + +FIXTURES = { + "home-assistant": HomeAssistant(), + "romm": Romm(), + "vikunja": Vikunja(), + "opengist": OpenGist(), + "papra": Papra(), + "mealie": Mealie(), + "n8n": N8n(), + "zipline": Zipline(), + "grafana": Grafana(), + "audiobookshelf": AudiobookShelf(), + "actualbudget": ActualBudget(), + "nextcloud": Nextcloud(), + "adventurelog": Django("adventurelog", "travel", "/admin/login/"), + "tandoor": Django("tandoor", "recipes", "/accounts/login/", + python="/opt/recipes/venv/bin/python", workdir="/opt/recipes"), + "privatebin": PrivateBin(), + "docmost": Docmost(), + "bookstack": BookStack(), + "gitea": Gitea(), + "navidrome": Navidrome(), + "vaultwarden": Vaultwarden(), +} diff --git a/documentation/audits/cleanup-2026-09-23/live.py b/documentation/audits/cleanup-2026-09-23/live.py new file mode 100644 index 00000000..8a024100 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/live.py @@ -0,0 +1,194 @@ +#!/usr/bin/env python3 +"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes. + +Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/, +fetches the app page in both languages, and reads the seeds back through each app's own front door. +Seed C is written immediately before each Update, after every backup — so only the undo's own +last-second copy can bring it back. +""" +import json, re, sys, time, html as H +sys.path.insert(0, ".") +import walk as w +from spike import load, save, ts +import bakeoff as b + +HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"} + + +def seed_c(app, s): + return b.seed_b(app, s["sub"], s["seedA"]) + + +def page_lines(app): + out = {} + for lang in ("hu", "en"): + h = H.unescape(w.page(f"/apps/{app}?lang={lang}")) + m = re.search(r'data-update-undone="true">([^<]*)<', h) + hold = re.search(r'data-held="true">([^<]*)<', h) + out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None} + return out + + +def press(app, poll=0.5, on_phase=None): + code, d = w.ctl("POST", f"/api/stacks/{app}/update") + w.say(f" Update -> {code} {str(d)[:140]}") + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < 1500: + try: + st = w.stack(app) + except Exception as e: + st = {} + ph = st.get("update_phase") + if ph != seen and ph is not None: + seen = ph + phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label"))) + w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}") + if on_phase and on_phase(ph): + return phases, "interrupted" + if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2: + break + time.sleep(poll) + st = w.stack(app) + w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}") + return phases, st + + +def readback(app, s, with_c=True): + sub = s["sub"] + w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2) + A = b.FX[app].verify(w, sub, s["seedA"], w.say) + B = b.verify_b(app, sub, s["seedA"], s["seedB"]) + C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None + return {"A": A, "B": B, "C": C} + + +def stage_undo(app): + s = load(app) + w.say(f"=== {app}: live undo by the product") + s["seedC"] = seed_c(app, s); s["seedC_at"] = ts() + w.say(" seed C written right before the Update:", bool(s["seedC"])) + before = b.db_state(app); w.say(" db before:", before) + phases, st = press(app) + obs = w.observables(app) + w.say(" observables:", json.dumps(obs)) + rb = readback(app, s) + after = b.db_state(app) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})") + pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False)) + w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip()) + s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")}, + "readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs} + save(app, s) + + + + +def set_box_language(lang): + """POST /settings/language — the household's language switch (form + the dashboard's CSRF).""" + sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip() + r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", + "-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}", + f"{w.BASE}/settings/language"]) + w.say(f" box language -> {lang}: http {r.stdout.strip()}") + + +def stage_powercut(app): + """Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard + stop); boot it again and let the controller resume the undo.""" + import subprocess + s = load(app) + w.say(f"=== {app}: power cut DURING the undo") + s["seedC"] = seed_c(app, s) + cut = {} + + def on_phase(ph): + if ph == "undoing": + t = time.time() + r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120) + cut["at"] = ts(); cut["rc"] = r.returncode + w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s") + return True + return False + phases, _ = press(app, poll=0.3, on_phase=on_phase) + if not cut: + w.say(" the cut never landed in `undoing` — recorded as a MISS"); return + r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180) + w.say(f" guest started again: rc={r.returncode}") + for i in range(60): + time.sleep(5) + try: + w.login(); st = w.stack(app) + if st: + break + except SystemExit: + continue + w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6")) + t0 = time.time() + while time.time() - t0 < 600: + st = w.stack(app) + if not st.get("updating") and st.get("update_phase") in ("undone", "failed"): + break + time.sleep(3) + w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}") + rb = readback(app, s) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}") + w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip()) + s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb} + save(app, s) + + +def stage_cutoff(app): + """Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of + the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it + back and HOLD with the new prefix.""" + s = load(app) + w.say(f"=== {app}: a cut-off undo copy") + done = {} + + def on_phase(ph): + if ph == "verifying" and not done: + out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c" +docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""") + done["out"] = out + w.say(" >>> finished-marker removed from one copy:\n" + out) + return False + phases, st = press(app, poll=0.5, on_phase=on_phase) + before = b.db_state(app) + pl = page_lines(app) + w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False)) + set_box_language("en") + pl_en = page_lines(app) + w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False)) + set_box_language("hu") + w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}")) + s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done} + save(app, s) + + +def stage_fixprobe_and_press(app): + """After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work + and end the undone note.""" + s = load(app) + frm, to, p0, p1 = b.EDGE[app] + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f) + w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"]) + w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + w.sync_rescan(app, to) + time.sleep(20) + w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan") + w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)") + w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False)) + phases, st = press(app) + rb = readback(app, s) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}") + pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False)) + w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml")) + s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl} + save(app, s) + + +if __name__ == "__main__": + w.login() + {"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff, + "fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/cleanup-2026-09-23/press1b.py b/documentation/audits/cleanup-2026-09-23/press1b.py new file mode 100644 index 00000000..6e36ca3e --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/press1b.py @@ -0,0 +1,5 @@ +import time, walk as w +time.sleep(20) +w.login() +c, d = w.ctl("POST", "/api/backup/run") +w.say(f"WHOLE-BOX BACKUP pressed 20 s after the deploy was accepted -> {c} {str(d)[:160]}") diff --git a/documentation/audits/cleanup-2026-09-23/press5.py b/documentation/audits/cleanup-2026-09-23/press5.py new file mode 100644 index 00000000..8c4fecbf --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/press5.py @@ -0,0 +1,5 @@ +import time, walk as w +time.sleep(5) +w.login() +c, d = w.ctl("POST", "/api/backup/run") +w.say(f"WHOLE-BOX BACKUP pressed 5 s after the deploy was accepted -> {c} {str(d)[:160]}") diff --git a/documentation/audits/cleanup-2026-09-23/repoint.py b/documentation/audits/cleanup-2026-09-23/repoint.py new file mode 100644 index 00000000..a20587d0 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/repoint.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config. + +`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is +`controller.yaml.pre-cleanup` (NOT the older `.pre-28`, which a restore must never pick up). +""" +import re, sys, io +sys.path.insert(0, '.') +import walk as w + +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" + + +def creds(): + for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"): + m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l) + if m: + return m.group(1), m.group(2) + raise SystemExit("no admin credential") + + +def to_drill(): + u, t = creds() + print(w.guest(f""" +set -e +test -f {VOL}/controller.yaml.pre-cleanup || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-cleanup +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml" +s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +if not re.search(r'^update:', s, re.M): + s += "update:\\n health_timeout: 90s\\n" +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -A2 '^update:' {VOL}/controller.yaml +""")) + + +def restore(): + print(w.guest(f""" +set -e +cp -p {VOL}/controller.yaml.pre-cleanup {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -c '^update:' {VOL}/controller.yaml || true +""")) + + +if __name__ == "__main__": + to_drill() if sys.argv[1] == "drill" else restore() diff --git a/documentation/audits/cleanup-2026-09-23/spike.py b/documentation/audits/cleanup-2026-09-23/spike.py new file mode 100644 index 00000000..369d1f64 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/spike.py @@ -0,0 +1,160 @@ +#!/usr/bin/env python3 +"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202. + +EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup, +sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being +spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the +steps the product would take, each one timed. + +State between stages lives in state-.json so each stage can be run, read, and only then +followed by the next (the hand undo needs a person looking at what the previous step left). +""" +import json, os, sys, time, re +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +HERE = os.path.dirname(os.path.abspath(__file__)) +FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()} +SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"} + + +def st_path(app): + return os.path.join(HERE, f"state-{os.environ.get('GUEST', '9202')}-{app}.json") + + +def load(app): + return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {} + + +def save(app, s): + json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False) + + +def ts(): + return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + + +# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only +# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it +# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier. +def docmost_seed_b(sub, A): + jar = "/tmp/dm.jar" + w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + name = "drillB" + os.urandom(3).hex() + rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json", + data=json.dumps({"name": name, "slug": name.lower()}), method="POST") + w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}") + return {"space": name} if code in ("200", "201") else None + + +def docmost_verify_b(sub, A, B): + jar = "/tmp/dm.jar" + rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + if code not in ("200", "201"): + w.say(f" docmost B: cannot log in (http={code})"); return False + rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json", + data="{}", method="POST") + ok = code in ("200", "201") and B["space"] in out + neg = "drillBnever" in out + w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def break_edge(app, frm, to, port_from, port_to): + """The failing edge: a REAL migrating image move, plus — in the DRILL template only — the + health probe pointed at a port the app does not answer. Both in one drill commit.""" + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read() + m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f) + assert m, "probe port not found" + f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():] + open(fy, "w").write(f) + h = w.drill_bump(app, frm, to) + return h + + +def pg_state(container, db, user): + return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1") + + +def romm_seed_b(sub, A): + """A SECOND RomM user, created by the first (admin) one — written after the backup.""" + jar, tok = FX["romm"]._csrf(w, sub) + user = "drillb" + os.urandom(3).hex() + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}), + method="POST") + w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}") + return {"user": user} if code in ("200", "201") else None + + +def romm_verify_b(sub, A, B): + jar, tok = FX["romm"]._csrf(w, sub) + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}") + ok = code == "200" and B["user"] in body + neg = "drillbnever" in body + w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def tree(paths): + cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths) + return w.guest(cmd) + + +def my_state(container="romm-db", db="romm"): + # The root password is used INSIDE the container from its own env — it never leaves it. + q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '" + return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) " +echo -n "alembic=$({q("select version_num from alembic_version")}) " +echo -n "users=$({q("select count(*) from users")})" +""") + + +def vik_seed_b(sub, A): + """A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the + app writes into its files volume), uploaded through the app's own attachment API.""" + tok, why = FX["vikunja"]._token(w, sub, A) + title = "drillB-" + os.urandom(4).hex() + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT") + if code not in ("200", "201"): + w.say(f" vikunja B: project refused {code} {body[:120]}"); return None + pid = json.loads(body)["id"] + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT") + tid = json.loads(body)["id"] if code in ("200", "201") else None + content = "drill attachment " + os.urandom(8).hex() + fn = "/tmp/vik-att.txt"; open(fn, "w").write(content) + rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}", + "-F", f"files=@{fn}", method="PUT") + w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}") + return {"title": title, "pid": pid, "tid": tid, "att": content} + + +def vik_verify_b(sub, A, B): + tok, why = FX["vikunja"]._token(w, sub, A) + if not tok: + w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False} + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}") + proj = code == "200" and B["title"] in body + rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}") + att_ok = False + try: + atts = json.loads(body) + if atts: + aid = atts[0]["id"] + rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}") + att_ok = c3 == "200" and B["att"] in b3 + except Exception as e: + w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}") + w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}") + return {"project": proj, "attachment": att_ok} diff --git a/documentation/audits/cleanup-2026-09-23/walk.py b/documentation/audits/cleanup-2026-09-23/walk.py new file mode 100644 index 00000000..8ed18e65 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/walk.py @@ -0,0 +1,494 @@ +#!/usr/bin/env python3 +"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints. + +EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses: + POST /api/stacks//deploy · POST /api/sync · POST /api/stacks/rescan + POST /api/stacks//update · POST /api/stacks//remove +and reads GET /api/stacks/. No controller code exists for it. + +The walk, per `09` §6.4 and the update-night brief §4: + 1 deploy from the DRILL catalog at the LIVE pin + 2 seed through the app's OWN front door (R-156: never a volume, never SQL) + 3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing + 4 „Mentés most" + 5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages + 6 press the guarded Update, record every phase with timestamps + 7 read the seed back through the front door + 8 the four version observables side by side + 9 write the verdict record in `09`'s JSON shape + +`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`. +""" +import argparse, json, os, re, subprocess, sys, time +from datetime import datetime, timezone + +SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad" +EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/cleanup-2026-09-23" +DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest. +GUEST = os.environ.get("GUEST", "9202") +BASE = {"9202": "https://192.168.0.114", "9201": "https://192.168.0.138"}[GUEST] +DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu") +HOSTHDR = f"Host: felhom.{DOMAIN}" +HP = "demo-hp" + +LOG = [] + + +def say(*a): + line = " ".join(str(x) for x in a) + ts = datetime.now().strftime("%H:%M:%S") + print(f"{ts} {line}", flush=True) + LOG.append(f"{ts} {line}") + + +def sh(args, timeout=300, inp=None): + try: + return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp) + except (subprocess.TimeoutExpired, OSError) as e: + return subprocess.CompletedProcess(args, 124, "", f"{e}") + + +def guest(script, timeout=600): + """Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting).""" + r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP, + f"cat > /tmp/w{GUEST}.sh; pct push {GUEST} /tmp/w{GUEST}.sh /tmp/w.sh >/dev/null 2>&1; " + f"pct exec {GUEST} -- bash /tmp/w.sh; rm -f /tmp/w{GUEST}.sh"], + timeout=timeout, inp=script) + return r.stdout or "" + + +def login(): + pw = open(f"{SC}/.ctlpw").read().strip() + sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR, + "-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"]) + h = open(f"{SC}/hdr.txt").read() + m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I) + if not m: + sys.exit("login failed: no session cookie") + open(f"{SC}/sess.txt", "w").write(m.group(0)) + r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"]) + c = re.search(r'/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN. + + Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because + a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password + (grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only + reason this was safe to discover by running it (live-probes rule). + + A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on + the scratch drive first — the same act the drive browser performs for a household. + """ + code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields") + fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or [] + values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub} + made = [] + for f in fields: + ev, ty = f.get("env_var"), f.get("type") + if ev in values: + continue + # `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses + # when the caller sends none, deliberately ("the user needs to know their password"), + # while `.felhom.yml` declares `required: false` and the API serves that verbatim. A + # caller that trusts the contract gets a 400. Measured tonight on grafana; filed. + if not f.get("required") and ty != "password": + continue # the controller generates the optional secrets itself + if ty == "path": + p = f"{DRIVE}/{name}" + values[ev] = p + made.append(p) + elif ty in ("secret", "password"): + import secrets as _s + values[ev] = "Drill-" + _s.token_hex(12) + GENERATED.setdefault(name, {})[ev] = values[ev] + elif f.get("default"): + values[ev] = f["default"] + else: + values[ev] = f"drill-{name}" + if made: + guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made)) + say(f" [1] made the drive paths this app requires: {made}") + extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")] + if extra: + say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}") + return values + + +def deploy(name, sub, extra_values=None): + st = stack(name) + if st.get("deployed"): + say(f" [1] {name} already deployed — reusing") + return True + values = deploy_values(name, sub) + if extra_values: + values.update(extra_values) + code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values}) + say(f" [1] deploy -> {code} {str(d)[:120]}") + if code != "202": + return False + # WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the + # container `healthy` while the controller's own state read `unhealthy` — a gate on `running` + # alone therefore times out on an app that is up. The state is RECORDED rather than required; + # the real gate is the fixture's own `wait_app`, which asks whether the APP answers. + seen = None + for _ in range(90): + time.sleep(5) + st = stack(name) + seen = st.get("state") + # `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21: + # tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm + # read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the + # deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark + # (`runComposeDeploy` writes it), so that is what to wait for. + pins = (st.get("app_config") or {}).get("pinned_images") + if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"): + say(f" [1] deployed, controller state={seen}, " + f"pinned={(st.get('app_config') or {}).get('pinned_images')}") + if seen != "running": + say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, " + f"not treated as a failure; the fixture's front-door wait is the real gate") + return True + say(f" [1] never became deployed (last controller state={seen!r})") + return False + + +def backup_now(name): + """R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever. + + `POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and + restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product + has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup + exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one + app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its + phase list. A seed written "after the backup" is therefore written before the update's own backup + — the undo's last-second copy is still the one that must bring it back.""" + say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)") + return None + +def drill_bump(app, frm, to, service_hint=None): + """Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates). + + `frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in + two images (adventurelog's backend and frontend) moves both in one edge, while its engine + sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move + every service that carries the app's own version and no others. + """ + comp = f"{DRILL}/templates/{app}/docker-compose.yml" + fy = f"{DRILL}/templates/{app}/.felhom.yml" + s = open(comp).read() + froms = [x.strip() for x in frm.split(",") if x.strip()] + tos = [x.strip() for x in to.split(",") if x.strip()] + if len(froms) != len(tos): + say(f" [5] from/to lists differ in length: {froms} vs {tos}") + return None + for f1, t1 in zip(froms, tos): + if f"image: {f1}" not in s: + say(f" [5] FROM ref not found in compose: {f1}") + return None + s = s.replace(f"image: {f1}", f"image: {t1}") + open(comp, "w").write(s) + f = open(fy).read() + today = datetime.now().strftime("%Y-%m-%d") + f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M) + open(fy, "w").write(f) + sh(["git", "-C", DRILL, "add", "-A"]) + sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"]) + r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})") + return h + + +def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5): + """Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP. + + R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and + `catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not + enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the + Update that followed moved nothing and still reported "Frissitve". So when the caller knows + which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the + NUMBER R-607 asks for and has never had. + """ + t0 = time.time() + ctl("POST", "/api/sync") + time.sleep(2) + ctl("POST", "/api/stacks/rescan") + time.sleep(2) + if not expect_app or not expect_ref: + return None + for i in range(tries): + cat = stack(expect_app).get("catalog_images") or {} + if expect_ref in cat.values(): + waited = round(time.time() - t0, 1) + if i: + say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up " + f"to {expect_ref} — R-607's window, measured") + return waited + time.sleep(delay) + ctl("POST", "/api/sync") + time.sleep(1) + ctl("POST", "/api/stacks/rescan") + say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — " + f"catalog_images = {stack(expect_app).get('catalog_images')}") + return None + + +def badges(name): + out = {} + for lang, suffix in (("hu", ""), ("en", "?lang=en")): + h = page(f"/apps/{name}{suffix}") + m = re.findall(r']*title="([^"]*)"[^>]*>([^<]*)<', h) + out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3] + return out + + +def press_update(name, poll=1.0, cap_s=1800): + code, d = ctl("POST", f"/api/stacks/{name}/update") + say(f" [6] Update -> {code} {str(d)[:220]}") + if code not in ("202", "200"): + return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0} + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < cap_s: + st = stack(name) + ph = st.get("update_phase") + if ph != seen: + seen = ph + rec = {"t": round(time.time() - t0, 1), "phase": ph, + "label": st.get("update_phase_label"), "updating": st.get("updating"), + "error": st.get("update_error"), "hold": st.get("hold_reason")} + phases.append(rec) + say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} " + f"err={rec['error']} hold={rec['hold']}") + if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3: + break + time.sleep(poll) + st = stack(name) + return {"accepted": True, "http": code, "phases": phases, + "duration_s": round(time.time() - t0, 1), + "final_phase": st.get("update_phase"), "update_error": st.get("update_error"), + "hold_reason": st.get("hold_reason"), "state": st.get("state")} + + +def observables(name): + st = stack(name) + ac = st.get("app_config") or {} + live = guest(f""" +grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//' +echo '---inspect---' +for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do + echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}' +done +""") + a, _, b = live.partition("---inspect---") + return { + "pinned_images": ac.get("pinned_images"), + "installed_images": {k: (v.get("ref") if isinstance(v, dict) else v) + for k, v in (ac.get("installed_images") or {}).items()}, + "catalog_images": st.get("catalog_images"), + "live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()], + "docker_inspect": [x for x in b.strip().splitlines() if x.strip()], + } + + +def app_logs(name, lines=400): + """The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is + one string with escaped newlines — a scan over the envelope sees a single enormous line and + finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96 + rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines.""" + code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}") + if isinstance(d, dict): + data = d.get("data") + if isinstance(data, dict) and isinstance(data.get("logs"), str): + return data["logs"] + if isinstance(d.get("_raw"), str): + return d["_raw"] + return str(d) + + +def write_verdict(rec, appdir): + os.makedirs(appdir, exist_ok=True) + p = os.path.join(appdir, "verdict.json") + json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False) + say(f" [9] verdict {rec['verdict']} -> {p}") + + +def remove(name): + """Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint + refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up.""" + c1, d1 = ctl("POST", f"/api/stacks/{name}/stop") + say(f" [X] stop -> {c1} {str(d1)[:100]}") + for _ in range(24): + time.sleep(5) + if stack(name).get("state") != "running": + break + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": True, "remove_backups": True}) + say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}") + if code == "409": + # R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive + # path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202 + # `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits + # this. The household's other choice — remove the app, KEEP the data — is accepted, and the + # harness takes it, then tidies its own directory by name at teardown. + say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)" + " — removing the app and KEEPING the drive data instead") + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": False, "remove_backups": True}) + say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}") + time.sleep(5) + st = stack(name) + left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; " + f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'") + say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}") + return code + + +def app_env(name, key): + """Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the + app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller + shows them the same value. Data still goes in through the app's own front door.""" + out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1") + if ":" in out: + return out.split(":", 1)[1].strip().strip('"').strip("'") + return "" + + +def snapshots(name): + """The restorable copies the backups page offers for this app.""" + code, d = ctl("GET", f"/api/backup/snapshots?stack={name}") + data = d.get("data") if isinstance(d, dict) else None + if isinstance(data, dict): + for k in ("snapshots", "items", "restore_points"): + if isinstance(data.get(k), list): + return data[k] + return data if isinstance(data, list) else [] + + +def restore(name, snapshot_id=None, wait_s=1200): + """The household's own way out: the „Visszaállítás a mentésből" button on the backups page. + + A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`, + `snapshot_id` — because that is the button the sentence tells them to press. + """ + snaps = snapshots(name) + if snapshot_id is None: + if not snaps: + say(f" [R] no restorable copy offered for {name}") + return {"ok": False, "why": "no snapshot offered", "snapshots": snaps} + first = snaps[0] + snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id") + say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)") + sess = open(f"{SC}/sess.txt").read().strip() + csrf = open(f"{SC}/csrf.txt").read().strip() + r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}", + "-X", "POST", + "--data-urlencode", f"_csrf={csrf}", + "--data-urlencode", f"stack_name={name}", + "--data-urlencode", f"snapshot_id={snapshot_id}", + f"{BASE}/backup/restore"], timeout=180) + head = (r.stdout or "").split("\n")[0].strip() + loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")] + say(f" [R] POST /backup/restore -> {head} {loc[:1]}") + t0 = time.time() + last = None + while time.time() - t0 < wait_s: + code, d = ctl("GET", "/api/backup/restore-status") + dd = d.get("data") or {} + cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message")) + if cur != last: + say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}") + last = cur + if not dd.get("running", False) and time.time() - t0 > 5: + break + time.sleep(2) + st = stack(name) + say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} " + f"phase={st.get('update_phase')}") + return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps, + "http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1), + "state_after": st.get("state"), "hold_after": st.get("hold_reason"), + "observables_after": observables(name)} diff --git a/documentation/audits/cleanup-2026-09-23/watch.py b/documentation/audits/cleanup-2026-09-23/watch.py new file mode 100644 index 00000000..529dead4 --- /dev/null +++ b/documentation/audits/cleanup-2026-09-23/watch.py @@ -0,0 +1,26 @@ +"""Deploy one app (or none) and record every change of deployed/deploying/state with a timestamp.""" +import sys, time, json +import walk as w +app = sys.argv[1]; cap = int(sys.argv[2]) if len(sys.argv) > 2 else 900 +w.login() +if "--deploy" in sys.argv: + vals = w.deploy_values(app, app) + c, d = w.ctl("POST", f"/api/stacks/{app}/deploy", {"values": vals} if vals else {}) + w.say(f"deploy {app} -> {c} {str(d)[:160]}") +last = None; t0 = time.time() +while time.time() - t0 < cap: + st = w.stack(app) + cur = (st.get("deployed"), st.get("deploying"), st.get("state"), st.get("deploy_error")) + if cur != last: + w.say(f"+{time.time()-t0:6.1f}s {app}: deployed={cur[0]} deploying={cur[1]} state={cur[2]} err={str(cur[3])[:200]!r}") + last = cur + if "--until-settled" in sys.argv and cur[1] is False and time.time() - t0 > 30 and cur[2] in ("running", "not_deployed", "error", "stopped", "exited"): + # settled: keep watching 60 s more for late flips + end = time.time() + 60 + while time.time() < end: + st = w.stack(app); c2 = (st.get("deployed"), st.get("deploying"), st.get("state"), st.get("deploy_error")) + if c2 != last: w.say(f"+{time.time()-t0:6.1f}s {app}: deployed={c2[0]} deploying={c2[1]} state={c2[2]} err={str(c2[3])[:200]!r}"); last = c2 + time.sleep(3) + break + time.sleep(3) +w.say("final:", json.dumps({k: w.stack(app).get(k) for k in ("deployed","deploying","state","deploy_error","containers")}, default=str)[:600]) diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 258f3f5c..6088f1d5 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -348,3 +348,13 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-620** | **A disabled notifier dropped every event with no local trace (P3).** Closed in controller **v0.264.0**: one WARN per event type per process naming the dropped event, then DEBUG. Live on 9202: six types each WARNed once; `app_start_failed` WARN at 12:02:30Z then DEBUG at 12:03:00Z. | **CLOSED 2026-09-23 — PROVEN-LIVE (`24-9202-r620-warn-then-debug.txt`)** | as above | | **R-646** | **An app pinned before v0.263.2 had no record of its own `.felhom.yml` (P3).** Closed in controller **v0.264.0**: a startup pass records `applied-meta/` for every deployed, pinned app CURRENT with the catalog; a behind app is skipped by name (its file is gone — not backfillable by design). Live: 9202 recorded 3, 9201 recorded 10, skipped 0. | **CLOSED 2026-09-23 — PROVEN-LIVE (`28-9202-controller-log.txt`, `30-9201-before-and-upgrade.txt`)** | as above | +## 2026-09-23 (evening) — clean-up: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0) + +| Row | What | Closed | Full text | +|---|---|---|---| +| **R-634** | **An app could RUN while the controller recorded it as not deployed (P1).** The unremovable half closed in v0.262.0; **the mechanism** is now diagnosed and fixed in **v0.265.0** (`0054d4b`): the whole-box backup's app list read the in-memory `Deployed` flag, true from the moment a deploy is accepted, so the volume leg stopped a DEPLOYING app (`compose down`), dumped half-made volumes and ran a second `compose up -d` beside the deploy's own — both failed and the deploy recorded „not deployed" (reproduced on demand on 9202, `audits/cleanup-2026-09-23/12-*`, `13-*`). Fix: deploying apps leave every backup list; the volume leg re-asks before the stop; `StopStack`/`StartStack` refuse a deploying stack for every caller. `sparkyfitness` did not reproduce alone. The deploy's own-failure question → **R-649**. | **CLOSED 2026-09-23 — PROVEN-LIVE (`32-*`, `33-*`: backup across a live deploy stopped three other apps and never outline; deploy `deployed`)** | `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` | +| **R-625** | **A held app invited an update its button refused (P2).** v0.265.0: badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed" (`tag-error`, title = the hold's first sentence per reader), no Update button, 409 `held` unchanged. | **CLOSED 2026-09-23 — PROVEN-LIVE (`44-*`, all four box/reader language pairs)** | as above | +| **R-636** | **Six hours of OOM kills sent the same single warning as one hiccup (P2).** v0.265.0 + hub v0.121.0: the kernel `oom_kill` counter (the `OOMKilled` flag is sticky and cannot count); ≥ 20 kills in 30 min of one container run → ONE `app_oom_storm` (error, operator-only, per-app cooldown). | **CLOSED 2026-09-23 — PROVEN-LIVE on the controller (`45-*`, `46-*`: RomM at 320M, storm at 21 kills, still one at 49); hub side by unit tests** | as above | +| **R-647** | **Three leftovers of the update mail (P3).** v0.265.0: a held update's error is the key `update.error.held`, rendered per reader on both pages and the API; `copy_holds` travels as its key; the two log wordings fixed. | **CLOSED 2026-09-23 — PROVEN-LIVE for (1) (`43-*`, `44-*`); (2)(3) by red-proofed tests** | as above | +| **R-648** | **The drill's „Mentés most" was whole-box (P3).** No per-app backup endpoint exists; the harness (`audits/cleanup-2026-09-23/walk.py` `backup_now`) now presses nothing and the guarded update's own `backing-up` phase backs up the throwaway app alone. | **CLOSED 2026-09-23 — PROVEN-LIVE (`43-*`: phase `backing-up` for vikunja only)** | as above | + diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index a9b8e5a0..4a114995 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -792,7 +792,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-622** | **[P2-MEDIUM] `adventurelog v0.13.0` migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe.** MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. `v0.12.1 → v0.13.0` (backend AND frontend together, PostGIS held constant). The backend applied **nine migrations, every one `... OK`** — `adventures.0072_trail_wanderer_author_fields` through `integrations.0009_alter_endurainintegration_auth_method`, plus `billing.0001_initial` — and then the container's own healthcheck failed with `URLError: [Errno 111] Connection refused` on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. **THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING:** the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." **This is exactly the case `09` §4 exists for:** the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. **What it needs:** adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/`. | **READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails** | | **R-623** | **[P3-LOW] The unattended-update caller turned every SUCCESS into a `timeout`, and then refused to press that app again — the instrument, not the box.** FOUND 2026-09-21 (update night) by reading `unattended-caller.py` before relying on it for the Q4 hold measurement. Its `call()` returns the API **envelope** — `{"ok": true, "data": {…}}` — and `follow()` read `update_phase` and `updating` **off the envelope**, where neither exists. Both were therefore always `None`; the end test `not updating and phase in ("done","failed")` could never fire; every followed update ran the full **900-second** timeout and was recorded `timeout`, which the caller treats as terminal and adds to `never_again`. `main()` unwraps `data` for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. **The 2026-09-21 run did not catch it because the only pass that reached `follow()` was Scenario F, whose log was lost to a buffering `tail`** — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. **This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement**, and worse, it is a measurement that says the box behaved badly when the box behaved well. **FIXED in the same file 2026-09-21** (unwrap `data`, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. **What it does NOT invalidate:** the G-b no-retry proof, which is entirely in the refusal path and never reached `follow()`. **What it DOES qualify:** any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | **CLOSED 2026-09-21 — fixed in `audits/update-arc-gaps-2026-09-21/unattended-caller.py`** | | **R-624** | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. | **READY — rank P3-LOW; owner: CC (catalog harness)** | -| **R-625** | **[P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one.** MEASURED 2026-09-21 (update night, leg B6). `glance` was HELD by a genuine unattended failed update. The drill catalog then published a **fixed newer version** — a real forward route, the thing a household would hope for. Afterwards the app page read **„Frissítés elérhető — ma" / „Update available — today"**, in both languages, with the Update button offered; pressing it answered **`409 reason='held'`** and the hold sentence. **The BEHAVIOUR is correct and is DESIGN, not a defect:** `09` §6.1 says the hold is `settings.RestoreHold` and that **a successful unit restore lifts an update hold (only that kind)** — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. **The DEFECT is that the page says otherwise.** R-524 settled precisely this shape for the Ahead case — *"(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled"* — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. **What it needs:** the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while `RestoreHold` stands. One verdict, read by both surfaces, exactly as R-524 did it. **AND A QUESTION FOR `09` §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER:** the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: `audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-626** | **[P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed.** FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. `navidrome` was removed through the product: the `remove_hdd_data:true` call was correctly refused `409` (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the `remove_hdd_data:false` call returned **200** with `volumes_removed: ['navidrome_navidrome_data']`. **Two seconds later a container carrying `com.docker.compose.project=navidrome` was CREATED** (`.Created = 19:12:20Z`), and Docker's `restart: unless-stopped` started it again at the next guest boot (`.StartedAt = 20:19:33Z`, the B5 power cut). Seven hours later: **`app.yaml` absent, `deployed=false`, one volume back, and the controller happily probing it — `Health probe navidrome: API GET :4533/ping → 200`.** **The customer-visible shape is the bad one:** *"I deleted that app and it came back"* — and it came back **blank**, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on `deployed`, which is exactly why the teardown found it and nothing else did. **WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container.** The controller was restarted several times later in the night and its log no longer reaches that moment — **the second time in one night that a restart destroyed the evidence of the thing that mattered** (see R-621). **Needs:** reproduce with a loop that removes an app and watches `docker events` for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. **And one instrument lesson worth keeping:** this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: `audits/update-night-2026-09-21/26-removed-app-came-back.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-628** | **[P2-MEDIUM] An empty search of a mailbox I do not control was turned into a claim about what a THIRD PARTY had done, and it went into the register as fact.** FOUND 2026-09-22; the operator caught it within minutes by producing the thread. The `due_checks_gate` fired R-433 (*have Hetzner answered?*). Two searches of the felhom catch-all came back empty, and the emptiness was written into R-433 as **"Hetzner has not answered"** and **"there is no evidence the tickets were ever opened"**. Both were false: ticket **#2026090103040671** had been opened and answered. **THE PRECISE FAULT IS THE INFERENCE, NOT THE QUERY — and that distinction is the whole value of this row.** Re-run afterwards WITH a positive control: `from:monitoring@felhom.eu` returns **201 threads**, `in:anywhere … includeTrash` reaches SENT, TRASH and mail back to January — **the instrument works.** And the exact ticket number, the exact subject `Storage Box issue`, and `from:hetzner` each still return **nothing**. So the literal finding — *this correspondence is not in this mailbox* — was CORRECT. What was invented was the step from there to *Hetzner has not answered* and *the tickets were never opened*. **A mailbox I can read is not the only place a reply can be**, and the operator's own account is exactly where a support ticket he opened would land. An absent record in ONE place can never answer a question about what SOMEONE ELSE did. **This is R-96 rule 3 in a new surface, one step further out than R-607**: there the instrument reported a stale value as current; here a correct observation was promoted to a conclusion it could not carry. **It is sharper still because the same session, the night before, gave every fixture a negative control and every gate a red-proof — and then reached for a search with neither.** **THE RULE, in one line:** *an empty search may be reported as "absent from the place I looked", never as "it did not happen" — and only after a control query that MUST hit has been seen to hit.* **Done:** the rule is written into the Gmail-access memory, where the next session meets it before it searches rather than after. | **CLOSED 2026-09-22 — rule recorded; R-433 corrected with the real answers** | | **R-627** | **[P2-MEDIUM] Nothing checked that the register is a well-formed table, so an append that ate two rows' state cells went unnoticed until a person read the file — and one row had been broken the same way for 45 days.** FOUND 2026-09-22. The 2026-09-21 update night appended measured results to eight rows with a regex that matched each row's trailing state cell; on **R-446** and **R-458** it consumed the cell and did not restore it, the cell reappeared as a stray FOURTH cell on a DUPLICATED copy of **R-626** and **R-625**, and a blank line was left between each pair. The register then reported **317 rows for 315 findings**, two rows carried no state at all, and two findings existed twice with contradictory state cells. **Nothing caught it:** `one_register_gate.py` compares this file against ROADMAP and `closed_register_gate.py` forbids an id in BOTH files — neither asks whether the file is a well-formed table, and neither notices an id duplicated WITHIN it. **THE RED-PROOF THEN FOUND AN OLDER INSTANCE NOBODY HAD SEEN: R-254 lost its state cell on 2026-08-08 (commit `59527d0`) and had rendered without a State column for 45 days.** **Closed the same day:** `scripts/register_shape_gate.py`, registered in `repo_gates.py` as gate 14 and reached by the pre-push hook, refusing a row that does not end with `|` (an eaten state cell), a duplicated id, or a blank line splitting the table; four decoys in `test_gate_decoys.py` — three convicting on the exact damage shapes and one asserting a healthy register still passes. **TWO THINGS THE RED-PROOF CORRECTED IN THE GATE ITSELF, kept because they are the finding's real content:** a first draft counted CELLS and convicted **125 innocent rows** — register cells carry literal `|` inside prose and shell snippets (`owner: CC | …`), so a row cannot be split on `|`, and a count that cannot be computed is not a check; and it skipped malformed rows before counting ids, reporting **5 duplicates where there were 2**. **Repaired:** both state cells restored from the stray cells that carried them, the two duplicate rows deleted, R-254's verdict sentence given its cell back, and **15 blank lines that split the register into 12 separate markdown tables** removed — every row's text byte-identical afterwards, proven by diff. **And the gate immediately earned itself:** the very next row inserted in this session (R-628) left a blank line behind and the gate refused it. | **CLOSED 2026-09-22 — gate 14, four decoys, register repaired 317→315** | @@ -801,16 +800,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-631** | **[P3-LOW] Five templates cannot be judged by the probe gate at all, and one is correct only by accident.** FOUND 2026-09-22 when `check-probe-matches-compose.py` was run over all 53. The gate's oracle is the probed service's own compose `healthcheck.test`; where that dials no loopback URL the gate has nothing to compare and reports a WARNING rather than a pass. **`crafty-controller`, `mealie` and `uptime-kuma`** run their healthcheck through a python or script helper (`ssl._create_unverified_context()`, `socket.create_connection`, `extra/healthcheck`), so the port is inside code the gate does not execute. **`vikunja`** has no compose healthcheck on its probed service at all. **`home-assistant`** is the interesting one: its probe path `/api/` differs from the compose's `/manifest.json`, and it is healthy today **only because its check is `type: api` with no `expect` block, which `probeHTTP` treats as "any response is healthy"** — add `expect: {status: 200}` to that template, a change that looks like a tightening, and the app goes permanently unhealthy and every successful update of it starts stopping it. **The gate warns on it for exactly that reason and refuses to call it a pass.** **Needs:** a live probe reading for each of the five on a scratch guest — deploy, read `GET /api/stacks/`, compare with the front door — which is one rotation night's work and closes the last gap R-618 left. **CLOSED 2026-09-22 — all five read live on guest 9202, and all five probes are CORRECT.** Each app was deployed, its listening sockets read from inside the probed container, and the probe's own target dialled **on the compose network**, which is the call the controller makes. `mealie` `tcp 9000` → listens `0.0.0.0:9000`, dial 200. `uptime-kuma` `http 3001` → listens `*:3001`, dial 302 (and `http` calls any response healthy, so 302 passes and proves something answers). `vikunja` `api 3456 /api/v1/info` **with `expect: {status: 200}`** → dial **200** — the one that could have failed, because its expect block compares the code. `crafty-controller` `tcp 8443` → the controller's own log reads `Health probe crafty-controller: TCP :8443 -> ok (1ms)` twice, six minutes apart. **`home-assistant` is the one to carry forward:** `api 8123 /api/` with NO expect → the dial returns **401**, not 200. It reads healthy only because `probeHTTP` treats any response as healthy for that shape (`healthprobe.go:253-262`). **Add `expect: {status: 200}` to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy and every successful update of it starts stopping it.** That is R-618 one edit away, now a measured number rather than a caution. **So the gate's WARN list is not a backlog of suspects: it is four correct templates the gate honestly cannot prove, plus one correct by accident.** Evidence: `audits/the-28-2026-09-22/sidejobs/r631.json`. | **CLOSED 2026-09-22 — five live readings, five correct probes; home-assistant's fragility is now a number** | | **R-632** | **[P3-LOW] Twenty-eight of the 53 templates have never been deployed by any update drill, so nothing is known about whether their updates work.** COUNTED 2026-09-22 against the 2026-09-21 sweep, which is the widest one ever run. **20 apps have a verdict record** (14 proven, 3 failed, 3 inconclusive, plus tandoor re-walked to proven on 2026-09-22); **4 more were deployed as props in the bad-days legs with no edge walked** (`bentopdf`, `glance`, `uptime-kuma`, `wishlist`); **1 was deployed only to measure its probe** (`wger`); and **28 have never been deployed at all**: `calcom`, `calibre-web`, `claper`, `code-server`, `crafty-controller`, `emby`, `ghost`, `gokapi`, `gramps-web`, `homebox`, `homepage`, `immich`, `jellyfin`, `kimai`, `komga`, `onlyoffice`, `outline`, `paperless-ngx`, `plant-it`, `plex`, `radarr`, `rallly`, `recipe-importer`, `seerr`, `sonarr`, `sparkyfitness`, `termix`, `wanderer`. **THIS IS NOT A COMPLAINT ABOUT THE SWEEP** — it took one app-catalog-wide night to go from 3 apps ever measured to 21, and a night is the unit available. It is a record of what the catalog's update promise currently rests on: **for 28 of 53 apps, nothing.** **The list is the nightly rotation's queue**, smallest and least stateful first; `paperless-ngx` should be early because R-630 needs a live reading from it anyway, and `crafty-controller`, `mealie` and `uptime-kuma` should be early because R-631 needs one from each. Machine-readable copy: `audits/probe-fix-2026-09-22/not-judged.json`. **WORKED IN ONE NIGHT, 2026-09-22 — all 28 walked, so this row CLOSES and hands its findings to others.** Every one was installed on guest 9202 against the private drill catalog and taken through the same walk: deploy at the live pin, seed through the app's own front door, read it back, „Mentés most”, the guarded Update where a real within-a-major edge exists upstream, **restore from that copy and read the seed back a second time** (the half the update night skipped), then remove and a 60-second check that nothing came back. **26 of 28 deployed; 6 proven; the rest inconclusive, no-edge or refused.** **What the night produced that this row could not have predicted:** R-630 raised to P1 by measurement (a stack with no probe container has its working app STOPPED by a successful update), R-633 (a remove during a restore leaves an orphan with a live public route), R-634 (an app running and healthy while recorded as not deployed, and then unremovable), and one real upstream edge that HELD honestly (`outline 1.9.1 → 1.10.1`). **Also settled:** `plant-it` is `lifecycle: abandoned` and the product refuses to install it — the only lifecycle-gated template in the catalog, and its gate is now proven live. Full record: `audits/DRILL-the-28-2026-09-22.md`. | **CLOSED 2026-09-22 — all 28 walked in one night; the findings live in R-630, R-633, R-634** | | **R-633** | **[P2-MEDIUM] A remove sent while a restore is still running reports success, deletes the app's record, and leaves a container restarting forever with a live public route.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, during the twenty-eight walk. `gokapi` was restored from its own local copy at 11:34:07 and removed at 11:34:22. `POST /backup/restore` answers **302 and works in the background**; the remove tore down what existed and the restore's own `compose up` then RE-CREATED the container at **11:34:24**. **Both calls returned success.** Twenty-five minutes later: `GET /api/stacks/gokapi` reads **`deployed: false`**, and `docker ps -a` shows `gokapi` **`Restarting (1)`** with `RestartCount` climbing, carrying its full traefik label set — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid, generating new salt… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **WHY THIS IS A ROW AND NOT A HARNESS ARTEFACT:** a household can press exactly these two buttons in exactly this order, the product accepted both, and **the remove reported success while leaving the orphan**. Presence of a success message is not evidence of a result. **WHAT IT COSTS:** an app the household believes is gone keeps a container in a restart loop, keeps a route registered on the public reverse proxy, and is invisible to every product surface because the controller no longer records the stack. Nothing in the alarm ladder fires: `08` §4 keys on stacks the controller KNOWS about. **This is R-626's class with the mechanism finally visible** — that row recorded a removed `navidrome` coming back and could not diagnose it because the controller had restarted; here the window is 17 seconds and both halves are in the evidence. **Needs:** the remove path to refuse, or to wait, while a restore for the same stack is in flight — and, either way, to verify the teardown rather than report success without looking. Evidence: `audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt`, `apps/gokapi/log.txt`. **FIXED in v0.262.0, both halves.** **(a) The guard the product already had everywhere else:** `RemoveStack` now consults `UpdateGuards.Busy` (backup single-flight, restore status, app-data holds) and `IsUpdating`, and refuses with the app's own sentence in both languages. The product refused exactly this clash for `update` and for `restore` — the restore refusal even NAMES the blocking app — and `remove` was the one door with no lock on it. **(b) `down` returning 0 is a request, not a result:** the compose project is now watched for 25 s after `down`, anything carrying its label is removed BY NAME with its labels logged, and the API answer carries `verified: true/false`. That is the instrument R-626 asked for, in the product rather than in a drill script. **PROVEN LIVE ON 9202, and the live proof found a defect a test had not.** The guard fires: with `privatebin` STOPPED (so the pre-existing "still running" check could not answer first) and a box-wide backup in flight, the remove was refused and the controller said `RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)`, with the household's own sentence *„Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik.”* **But it answered HTTP 500.** `router.go`'s status mapping greps the error TEXT for `not deployed` / `still running` / `not found` / `protected`, and the busy sentence contains none of them, so it fell through to the default. **A 500 tells the UI something broke; this is “wait a moment”.** Fixed in **v0.262.1**: a typed `*stacks.RemoveBusyError` matched with `errors.As`, answered **409**, carrying both the Hungarian bytes and the key — and its test asserts the sentence contains none of the words the text mapping greps for, so the type is load-bearing rather than decorative. **The verification half is proven too:** a permitted remove answered `verified: true` with no reappearance, and no container carrying the project label existed 60 s later. **AND A FIRST ATTEMPT THAT PROVED NOTHING, recorded because it nearly went down as a pass:** the first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING running check, not this guard — the restore had already finished by then. A refusal from the wrong rule is not evidence for the new one. | **CLOSED 2026-09-22 — v0.262.0 + v0.262.1; refusal and verification both proven live, and the live proof corrected the status code — **re-proven on v0.262.1: HTTP 409 with the household's sentence**, `RemoveStack privatebin REFUSED (busy)`** | -| **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. **THE BOUNDED HALF IS FIXED in v0.262.0; the MECHANISM IS STILL NOT DIAGNOSED, and this row stays open for it.** `RemoveStack` refused on `!stack.Deployed` — a FLAG — while the machine plainly had containers, a compose file and an `app.yaml`. It now asks whether anything EXISTS (`halfStateEvidence`): containers, a compose file, or an `app.yaml` are each enough, and it removes what exists with every other fence unchanged (R-442 stays fail-closed). **The household must always be able to remove what the box shows them** — that is true whatever the cause of the bad record. Red-proof seen failing. **What is NOT done:** why `deployed` stays false while containers run. `runComposeDeploy`'s pin-write path was not read against a fresh reproduction in this session. **THE UNREMOVABLE HALF IS PROVEN LIVE ON 9202:** `sparkyfitness` rebuilt in its exact measured shape — `app.yaml` and a compose file on disk, no containers, `deployed: false`, `state: not_deployed` — and the remove answered **200** with `leftovers: NONE`. Under v0.261.0 the identical call answered `stack "sparkyfitness" is not deployed`. **THE MECHANISM, DIAGNOSED 2026-09-23 (evening) on 9202, controller v0.264.0, before any code** (`audits/cleanup-2026-09-23/12-*`, `13-*`). **Reproduced on demand:** deploy `outline`, press the whole-box backup 20 s later. The chain, at `file:line`: (1) `DeployStack` sets the in-memory `Deployed=true` at ACCEPT (`stacks/deploy.go:395`, the no-stale-button UX); (2) `stackAdapter.ListDeployedStacks` (`cmd/controller/main.go:2556`) filters on that flag only, so a DEPLOYING app is in the backup's list; (3) `runVolumeDumps` (`backup/backup.go:693`) calls `DumpAppVolumesSafe` on it — **`StopStack` runs `docker compose down` in the middle of the deploy** (`14:39:32 StopStack outline: current state=deploying deployed=true containers=0`), tars half-made volumes (1.5 KB each), and **`StartStack` runs a SECOND `compose up -d` while the deploy's own is still running** (`backup.go:918`, `14:39:33`); (4) both `compose up` calls fail at the same second (`14:39:48`, exit 1), and `runComposeDeploy`'s failure branch (`stacks/deploy.go:421-435`) writes `Deployed=false` to memory and disk without asking whether containers exist. **Which of the two `up` calls wins decides the end state:** on 2026-09-22 the backup's restart won → containers RUNNING under `deployed=false` (outline's 11:56:09 `installed_images`, recorded by `StartStack`'s `recordInstalledImages`); tonight both lost → containers `Created` under `deployed=false`. **Shape (i) of the brief is ruled out:** there is no deploy time limit (`composeExecCustomEnv` has no timeout). **`sparkyfitness` did NOT reproduce alone** at today's pin on v0.264.0: deployed in 47.3 s (`10-*`) — the 2026-09-22 "alone" reading is not reproduced and is most likely the same backup race inside the walk (the walk presses the whole-box backup). **Not a design choice:** the brief itself names the fix (a whole-box backup skips a deploying stack, as the update refuses `busy`). **One design question remains and goes to the operator:** when `compose up` fails for its OWN reasons and leaves containers behind, should the deploy remove what it started, or keep the record "failed" with the containers visible? | **OPEN — P1 for the MECHANISM only; the unremovable half is CLOSED in v0.262.0 and proven live** | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** FOUND 2026-09-22 17:37 **because the operator heard the fans**, which is the only reason it was found at all. `romm` was promoted `5.0.0 -> 5.3.0` on the live catalog that morning (`audits/PROBE-FIX-2026-09-22.md`) after the edge was PROVEN on scratch guest 9202, and the guarded Update was then pressed on demo-hp guest 9201 at **09:08:22Z**, reaching **`done` in 74.8 s** with the app `running`. **It ran clean for two hours.** The first worker kill is at **11:09:20Z**; by 15:38Z there had been **4,530** of them — `Worker (pid:…) was sent SIGKILL! Perhaps out of memory?` — with `docker inspect` reading **`OOMKilled: true`**, the container pinned at **457 MiB of its 512 MiB limit**, and `docker stats` showing **499.51% CPU**. The host's load average sat at **5.2 while otherwise idle**. `rq_cron` is killed and restarted every few seconds in a permanent storm. **THREE THINGS THIS ESTABLISHES, and the third is the one that changes how promotions are judged.** (1) The template's own comment says *`RAM: ~300MB (mem_limit: 1024M total — romm 512M + mariadb 384M + redis 128M)`*; **5.3.0 needs more than 512M and the template was not re-sized when the version moved.** A version move is not only an `image:` line. (2) **The update reported `done` and the app reads `running`**, because nginx answers `GET /` with 200 while the gunicorn workers behind it are being killed — a THIRD variant of the R-618/R-630 theme: the probe is right, the port is right, and the answer is still a false green. (3) **THE ALARM DID FIRE, AND MY FIRST WRITE-UP OF THIS ROW SAID IT DID NOT — CORRECTED 2026-09-22 BY THE OPERATOR, WHO PRODUCED THE MAILS.** The controller HAS an OOM detector (`main.go:821`, `[deadapp] romm: container romm was OOM-killed`, evaluated every 30 s), it emits **`app_oom`** (`notifier.go:726`), the hub allow-lists it (`dispatcher.go:636`) and dispatched it to the OPERATOR channel as **SENT** — `admin@felhom.eu` received **`[Felhom] ⚠️ demo-hp: app_oom`** at **11:09 CEST** and again at **17:48 CEST**, each carrying the app name, the Hungarian sentence and the dashboard link. The CUSTOMER channel is **SKIPPED**, which is correct. **I asserted an absence without looking at the instrument** — the hub's own Events and Notifications tabs show all of it — and I did it by reasoning from a memory note (`lxc-docker-oom-signals-unreliable`, R-528) instead of reading the hub. **That is R-628's shape exactly, four days on, and from the same hand.** **WHAT IS ACTUALLY WRONG, and it is narrower and real:** `notifier.go:715-726` keys the alarm on `container|startedAt` in an `oomSeen` map and emits **once per container lifetime**. So **six hours of continuous thrashing — 4,530 worker kills — produced exactly ONE mail**, severity `warning`, never escalating, while the app went on reading `running`. **A six-hour storm is indistinguishable from a single transient kill.** The hub's App Telemetry panel did carry the magnitude (RomM: **5,023 errors, 632 warnings**, peak 855 MB) but nothing turns that into a second, louder signal. So the fix worth having is not a detector — there is one — but an ESCALATION: a warning that repeats for hours should stop looking like a warning that happened once. **AND THE LESSON FOR R-462's METHOD, which is the real cost:** every `proven` verdict in the update night and in the twenty-eight measures the app for the **minutes of the walk**, not for a day of running. `romm` passed its walk, was seeded, read back and restored — and broke two hours later. **`proven` currently means "the update applied and the data survived", NOT "the new version runs".** Needs: decide between rolling the catalog back to 5.0.0 and raising romm's `mem_limit` (measured, not guessed); and a soak longer than a walk before any future promotion. **FIXED AND MEASURED 2026-09-22, in two steps, and the FIRST step was still a guess.** *Step 1 (operator's choice):* the limit was raised 512M → **768M** (`app-catalog-felhom.eu@886956d`). It slowed the kills from ~12/min to ~7/min and **stopped nothing** — 37 SIGKILLs in five minutes, `OOMKilled` still true, and the cgroup's own `memory.peak` read **exactly 768 MiB**: it hit the new ceiling and died there. *Step 2, from a MEASUREMENT instead:* the per-process RSS inside the container reads **~216 MiB per warm uvicorn worker**, so the image's default of four workers plus the master needs **~882 MiB** before nginx and the job runner — more than any sensible limit for this box. **The lever was in the image all along:** `/init:143` runs `--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a SERVER default on an appliance serving one household.** Setting **`WEB_SERVER_CONCURRENCY=2`** (`app-catalog-felhom.eu@f4eb94f`, limit left at 768M) and applying it through the product's own Update button fixed it. **PROVEN UNDER LOAD, not just at idle:** 6 concurrent callers driven at romm through the household's own route for 300 s — **26,645 requests** (9,687 × 200, 16,958 × 401 on the auth-gated endpoints), CPU a steady **~200%** (exactly two workers saturated, by design), memory oscillating **416–614 MiB against the 768 MiB limit and trending DOWN**, and **zero** SIGKILLs, `OOMKilled: false`, `RestartCount: 0` throughout. At idle afterwards: **1.64% CPU**, 610 MiB, host load falling from 5.2 to 2.4. **AND THE FIRST SOAK MEASURED NOTHING, which is worth more than the second one:** it was pointed at `arcade.enkisfelhom.hu` — the scratch-guest fixture's default subdomain — while this box deployed romm at `jatek`. Every request 404'd at traefik in 9 ms, romm sat idle at 0.64% CPU, and the counter cheerfully reported **14,026 successful requests**. It was caught only because 0.64% CPU under load is not believable. **A positive AND a negative control are now asserted before any load is driven** (`jatek` must not 404; a nonsense host must). R-96 rule 3, in a new surface. **WHAT STAYS OPEN, and it is the part that outlives romm:** 610 MiB of 768 MiB is **79%** — it works with ~158 MiB of headroom and the soak never exceeded 614 MiB, but it is not generous, and nothing watches it. **And the method lesson for R-462:** a version move is not only an `image:` line — the new version's SHAPE (worker counts, per-worker footprint) has to be measured too, and a walk lasting minutes cannot see a ceiling reached in two hours. Every `proven` verdict in the update night and in the twenty-eight means *"the update applied and the data survived"*, **not** *"the new version runs"*. Evidence: `audits/probe-fix-2026-09-22/romm-soak.json`, `romm-soak.out`. | **CLOSED 2026-09-22 — two workers, 768M, proven under 26,645 requests; the 79% headroom and the `proven`-means-minutes lesson are carried into R-462** | -| **R-636** | **[P2-MEDIUM] An app that has been out of memory for six hours sends the same single warning an app that hiccuped once sends.** FOUND 2026-09-22 when the operator produced the alarm mails I had wrongly written off as absent (R-635). **The detector is correct and works.** `main.go:821` re-checks every 30 s and logs `[deadapp] romm: container romm was OOM-killed`; `notifier.go:726` emits **`app_oom`** at severity `warning`; the hub allow-lists it (`dispatcher.go:636`) and delivered it to the OPERATOR channel — two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER is SKIPPED, correctly. **The defect is the SHAPE of the signal, not its absence.** `notifier.go:715-724` keys on `container|startedAt` in an `oomSeen` map and returns early on a repeat, so the alarm fires **once per container lifetime**. romm's 09:08 container produced **4,530 worker kills over six hours and exactly one mail**; the 15:47 container produced one more. Severity never escalates, the app goes on reading `running`, and `08` §4 rightly does not put it in `IsDownState`. **So a six-hour storm that pins five cores is indistinguishable, in the operator's inbox, from one transient kill at 3 a.m.** — and the operator, who had two correct mails, still found the fault by hearing the fans. **The once-per-lifetime rule is right in itself** (it is what stops a crash loop from mailing 4,530 times, which is R-629's lesson) — what is missing is the second, louder signal when the same key keeps re-firing. **The magnitude IS already collected:** the hub's App Telemetry panel read **RomM: 5,023 errors, 632 warnings, peak 855 MB against a 1280M limit** while the same panel showed every other app at 0. Nothing turns that into an event. **Candidate shapes, none chosen here:** escalate `app_oom` to `error` when the same `container|startedAt` key re-fires past a threshold; or a periodic digest for a key still firing after N minutes; or let App Telemetry raise its own event when an app's error count crosses a bound. All are hub/controller code. Evidence: the operator's screenshots of the Events, Notifications and App Telemetry tabs plus the two mails; `felhom-controller/controller/internal/notify/notifier.go:715-726`. | **OPEN — P2; owner: CC; product code, so not fixed unattended** | | **R-638** | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | | **R-640** | **[P2-MEDIUM] A TRUNCATED PostgreSQL copy loads with exit 0 into an EMPTY database — and `ValidateDump` would accept it.** MEASURED 2026-09-23 on 9202 in a scratch database beside docmost's: the first half of the 135 816-byte undo copy, loaded with `psql -v ON_ERROR_STOP=1 --single-transaction`, returned **rc 0** and left **42 tables and 48 migration rows — and 0 users, 0 spaces, 0 constraints, 0 indexes**; psql treats end-of-file inside a `COPY` as end of data and commits. The whole copy ends with `-- PostgreSQL database dump complete` (plus a `\unrestrict` line); the truncated one does not. `ValidateDump` (`appbackup/dbdump.go:415`, read from source) checks the header and a `CREATE TABLE` only. **An undo or a restore fed a truncated copy would report success, start the app on an empty database, and pass a health check.** MariaDB's truncated load failed (rc 1) but NON-atomically — half the tables already replaced; its copy carries `-- Dump completed`. Fix: check the engine's completion marker before any load. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-46`, `docmost-47`, `romm-44`. **-- NARROWED 2026-09-23:** the undo is not exposed (a folder copy, validated by its own finished-marker). The restore paths still load dumps through `ValidateDump`, which does not check the engine's completion marker. | **OPEN — P2, narrowed to the restore paths; owner: CC** | | **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** | | **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** | | **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | -| **R-647** | **[P3-LOW] Three small leftovers of the update mail and of R-606, found by the live proof of 2026-09-23.** (1) **A reader in the OTHER language than the box reads a held update's sentence in the box's language.** `fillHoldReason` and the stored `UpdateError` of a hold use the box language at hold time (the hold sentence is composed, not a key), so on an English box `?lang=hu` shows the English hold, and on a Hungarian box `UpdateErrorIn("en")` falls back to the stored Hungarian. The HOUSEHOLD always reads its own box language, which works — only the `?lang=` test switch and an operator reading a box in the other language see it. Fix shape: store a held update's error as a key the page can re-render (`RestoreHoldForLang`), and let the Hungarian `holdText` also call it. (2) **The `Note:` line of an English household mail prints the raw details, whose `copy_holds` is the Hungarian phrase** (`"copy_holds":"a beállításokat, …"`). Fix shape: send the phrase's key, or leave `copy_holds` out of the household copy. (3) **Two log wordings of my own:** the undone event logs `hold recorded: true` (meaningless for an undone update), and the disabled-notifier line for `health_change` prints the health STATUS where it says `severity` (`severity warn`) — the real events carry correct severities. Evidence: `audits/undo-fleet-2026-09-23/23-9202-held.txt`, `41-mails-as-received.md`, `24-9202-r620-warn-then-debug.txt`. | **OPEN — P3; owner: CC (controller)** | -| **R-648** | **[P3-LOW] The drill's „Mentés most" is a WHOLE-BOX backup, so every drill that seeds a throwaway app stops and restarts every standing app for its volume dump.** Found 2026-09-23 on 9201: two presses (one per throwaway app) restarted 9 of the 10 standing apps twice, a few seconds each (`backup.go` "Stopping for safe volume dump"); all came back healthy with byte-identical files, and the nightly does the same — but a brief that says "standing apps untouched" is not met. Fix shape: the harness backs up only the throwaway app (a per-app backup call, if the product has one; otherwise rely on the guarded update's own `backing-up` phase and skip the press). Evidence: `audits/undo-fleet-2026-09-23/49-9201-teardown.txt`. | **OPEN — P3; owner: CC (harness)** | +| **R-649** | **[P2-MEDIUM] OPERATOR QUESTION: when `docker compose up` fails for the deploy's OWN reasons and leaves containers behind, what should the record say?** Found while diagnosing R-634 (2026-09-23). The backup race is fixed (v0.265.0), but `runComposeDeploy`'s failure branch (`stacks/deploy.go:421-435`) still writes `Deployed=false` without asking whether containers exist — e.g. a dependency's healthcheck that times out after its containers started. Since v0.262.0 the household can REMOVE such an app (`halfStateEvidence`), so nothing is stranded; but the page says „not installed" over containers that may be running. **Options:** (a) the failed deploy runs `compose down` on what it started, so „failed" means nothing runs — cost: a slow-but-healthy start is thrown away, and the household must press Deploy again; (b) keep the containers and record the deploy as FAILED-WITH-CONTAINERS, a state the page shows with a Remove button and the compose error in both languages — cost: a new state every surface must learn; (c) leave it (today). **Recommendation: (a)** — it keeps the R-634 promise (never containers under „not deployed") with the least new surface, and a deploy that failed is re-pressed anyway. Not taken unattended: it changes what a failed deploy does to what the household started (rule 1). | **OPEN — P2; owner: operator (decision), then CC** | +| **R-650** | **[P3-LOW] A controller unit test that falls through to the REAL docker acts on the build host — DooPlex.** MEASURED 2026-09-23: my own first draft of `TestR634_VolumeLegNeverStopsADeployingApp` drove the real `DumpAppVolumesSafe` for its control case, and `docker run … tar` CREATED an empty volume `outline_outline_data` on DooPlex (removed by name, verified empty and unused; nothing else touched). The same draft of the stacks test ran a real `docker compose down` from a temp dir named `nextcloud` (no such project existed). Both tests now use seams (`dumpVolumesSafe`; `composeCmd` + an empty PATH). **The class is unguarded:** any test that reaches `exec.Command("docker", …)` on DooPlex acts on production Docker. Fix shape: a test-binary guard in `composeExec`/`execCommand` that refuses a real docker exec when running under `go test` unless an explicit env opt-in is set, plus a decoy test that proves it refuses. | **OPEN — P3; owner: CC (controller)** |