Clean-up evening: R-634 cause fixed, held badge, OOM storm; floor 0.265.0
gates / gates (push) Successful in 26s

R-634, R-625, R-636, R-647, R-648 closed; R-649 (operator question) and
R-650 opened. Open rows 334 -> 331. 08 §6.2 storm rung; 09 §6.4 parts
8-9 SHIPPED. Evidence: audits/cleanup-2026-09-23/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 17:40:01 +02:00
parent 3add9fa678
commit 210394ec6e
38 changed files with 3121 additions and 115 deletions
+22
View File
@@ -14,6 +14,28 @@
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## 2026-09-23 (late evening) — clean-up: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0); floor 0.265.0
**R-634 mechanism (reproduced on demand on 9202):** the in-memory `Deployed` flag is true from the moment a
deploy is ACCEPTED (`stacks/deploy.go:395`), and `stackAdapter.ListDeployedStacks` filtered on it alone, so
every backup leg saw a DEPLOYING app; the volume leg's `DumpAppVolumesSafe` ran `compose down` and a second
`compose up -d` in the middle of the deploy's own `up`; both failed and `runComposeDeploy`'s failure branch
wrote `Deployed=false` over whatever containers the race left. **Fix:** deploying apps leave the list; the
volume leg re-asks (`deployingReporter`, optional interface — no fake churn) right before the stop and SKIPs;
`StopStack`/`StartStack` return `ErrStackDeploying` for every caller. **Open:** R-649 (operator question: a
deploy that fails on its own and leaves containers).
**R-625:** held → badge `badge.update.held` (`tag-error`), no Update button (the list already hid it).
**R-636:** `ScanOOMKilled` reads `memory.events` `oom_kill` / `memory.max` / `memory.peak` by `docker exec
cat` for FLAGGED containers only (the controller's own cgroupns cannot see others); the notifier keeps 30 min
of samples per `container|startedAt` and sends `app_oom_storm` (error) once at ≥20; hub: allow-list +
`operatorOnlyEvents` + `perAppCooldownEvents`. `08` §6.2 has the rung.
**R-647:** web `readerFuncs(lang, prev)` apply to EVERY language (holdText / updateErrorText / held badge
title) — same bytes on a Hungarian box; held error key `stacks.UpdateErrorKeyHeld`; API uses
`Manager.UpdateErrorFor` / `HoldReasonFor` (guards' optional `HoldForLang`).
**R-648:** there is NO per-app backup endpoint; drills press nothing and rely on the update's `backing-up`.
**R-650:** a test that reaches real docker acts on DooPlex — measured once (my draft), unguarded.
Evidence: `audits/cleanup-2026-09-23/README.md`.
## 2026-09-23 (evening) — the household is told: controller v0.264.0 + hub v0.120.0 (`09` §6.4 parts 2–3); floor 0.264.0
**Events.** `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, sent by the
+82 -94
View File
@@ -1,11 +1,13 @@
# REPORT — the undo reaches the fleet, and the household is TOLD (controller v0.264.0, hub v0.120.0)
# REPORT — clean-up evening: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0)
2026-09-23 (evening). Repos touched: **felhom-controller** (v0.264.0, `bc27894`), **felhom.eu** (hub
v0.120.0 `d3b2848`, manifest `7caa24f`, docs, register, evidence), **admin/app-catalog-drill** (drill
commits, reset to live `main` at the end). **felhom-agent, app-catalog-felhom.eu: untouched.**
Architecture read and named: `documentation/architecture/09-update-architecture.md` (§3 decision 15,
§6.4). Evidence: `documentation/audits/undo-fleet-2026-09-23/README.md`. **Method: endpoint-level**
(the endpoints the UI invokes; mails read from the catch-all mailbox; no browser).
2026-09-23 (late evening). Baselines verified live: controller `0a3026180ae8` (v0.264.0), agent
`d9864a94bf62`, felhom.eu `21f17ed32bdc` (hub v0.120.0), catalog `cfcfe5278428` — as the brief said.
Repos touched: **felhom-controller** (v0.265.0 `0054d4b`), **felhom.eu** (hub v0.121.0, manifest, docs,
register, evidence), **admin/app-catalog-drill** (drill commits, reset to live `main`).
**felhom-agent, app-catalog-felhom.eu: untouched.** Architecture read and named:
`09-update-architecture.md` §6.1/§6.1a/§6.4, `08-alarm-ladder.md` §4/§6.2, `02-controller-module-map.md`
(stacks ↔ backup seam). Evidence: `documentation/audits/cleanup-2026-09-23/README.md`. **Method:
endpoint-level** (the endpoints the UI invokes; no browser).
---
@@ -13,104 +15,90 @@ Architecture read and named: `documentation/architecture/09-update-architecture.
| item | state |
|---|---|
| Part 0 — floor 0.263.2 | **done**, read back (`00-floor-0.263.2.txt`) |
| Part 1 — R-646 startup pass | **done**; live: 3 recorded on 9202, 10 on 9201, 0 skipped |
| Part 2 — two events, hub v0.120.0 first | **done**; hub deployed 11:49Z, before any box ran v0.264.0 |
| Part 3 — R-606 | **done, with leftovers → R-647** (a reader in the other language than the box sees a HELD sentence in the box's language; the English mail's raw `Note:` line keeps the Hungarian `copy_holds` phrase) |
| Part 4 — R-620 | **done** |
| Part 5 — live proof | **done**; all proofs passed → **floor raised to 0.264.0** |
| "standing apps untouched" | **NOT fully met.** The harness's „Mentés most" is a whole-box backup: on 9201 it stopped and restarted 9 of the 10 standing apps for their volume dumps, twice (a few seconds each). All healthy afterwards, files byte-identical. **R-648** filed. |
| the brief's Hungarian undone line | **changed:** it ends „semmi nem veszett el, nincs teendőd." — the brief's „teendője nincs" is the formal „ön" voice; the product speaks „te" everywhere and the voice gate ratchets the formal form (18). |
| the customer cooldown | **changed shape:** a NEW customer-side register (`perAppCustomerCooldownEvents`) instead of putting the customer leg of R-389's register to work — that would also have changed `app_start_failed`'s household grain, which nobody asked for (pinned by a test). |
| a box-side seed + two page toggles | **added, not in the brief:** the box pushes its own `EnabledEvents` to the hub on every notifications-page save, so the hub migration alone would be undone by the next save. |
| Part 1 — R-634 | **cause found in ~25 min (cap 2 h), fixed, proven live.** One design question left → **R-649** (operator) |
| Part 2 — R-625 | **done.** The Update button was ALREADY hidden for a held app on the list (since v0.238.0); the defect was the badge only. Now pinned by a test with a positive control. The badge applies to **both** hold kinds (a failed restore holds the app stopped too; the way back is also a restore) — the brief said `update_failed` only |
| Part 3 — R-636 | **done, and the brief's rule changed shape:** "the same key re-fires ≥ 20 times" cannot work — `OOMKilled` is sticky, so the key re-fires every 30 s after ONE kill (a hiccup would storm after 10 min). The controller now counts real kills (kernel `oom_kill`); the 20-in-30-min threshold is applied to THAT |
| Part 3 — hub live `curl` | **not done:** posting a synthetic event with a real customer's key would put a fabricated alarm in the operator's record; the hub side is proven by unit tests |
| Part 4 — R-647 | **done** |
| Part 5 — R-648 | **done — no per-app endpoint exists**, so the drill presses nothing |
| Part 6 — floor | **raised to 0.265.0** — every live proof passed |
| a mistake of mine | my first draft of the R-634 backup test drove the real dump and **created an empty volume `outline_outline_data` on DooPlex**; my inspection of it pulled `alpine:3.20`. Both verified new and unused, removed by name. Tests rewritten on seams. **R-650** filed (the class is unguarded) |
**Claims in the brief, checked:**
1. *"No app-update event exists today"* — **true for apps.** The only near-misses: `controller_updated` is
the controller's OWN self-update, and `update_available` exists only as a debug-page test option — never
emitted, and not in the hub's allow-list (it would 400).
2. *"`enabled_events` is the household's only mail gate"* — **wrong.** Four more gates stand in the way:
the `operatorOnlyEvents` register, a blocked customer, a missing e-mail address, and the cooldown —
keyed `customer:type`, which would have merged two apps' mails into one. And the box OVERWRITES
the hub's list on each page save (above).
3. *"9201 can follow the drill catalog without disturbing its standing apps"* — **true for the catalog,
false for the harness.** Measured: the 10 standing apps' compose + `.felhom.yml` were byte-identical
before, during and after (the drill differed from live only in the throwaway apps). The disturbance
came from the harness's whole-box backup press (R-648), not from the catalog.
1. *"sparkyfitness reproduces alone at today's pin"* — **wrong.** Deployed alone in 47.3 s on v0.264.0
(`10-*`). The 2026-09-22 "alone" walk pressed the whole-box backup itself — the same race.
2. *"a whole-box backup touches a DEPLOYING stack at all (never read)"* — **true, and it is THE cause.**
3. *"20 kills in 30 minutes separates a storm from a hiccup"* — **true only if kills are counted, and they
were not.** RomM's real rate: 4,530 kills in 6 h ≈ **375 per 30 min**; a hiccup is 1–3. 20 kept, applied
to the kernel's counter. Live: 21 kills 2.5 min after start → one storm.
4. *"a per-app backup endpoint exists"* — **wrong.** `router.go` has `/backup/run` and `/backup/tier2`
only; the per-app backup (`RunAppBackupNow`) is reachable only through the guarded update.
## 2. Floor raises, read back from the hub
## 2. R-634 — the mechanism, at `file:line`
| raise | read back | fleet |
1. `stacks/deploy.go:395` — `DeployStack` sets the in-memory `Deployed=true` at ACCEPT (no-stale-button UX).
2. `cmd/controller/main.go:2556` — `ListDeployedStacks` filtered on that flag only → a deploying app is in
every backup leg's list.
3. `backup/backup.go:693` → `DumpAppVolumesSafe` — `StopStack` (`compose down`) in the middle of the deploy,
tar of half-made volumes, `StartStack` (a SECOND `compose up -d`, `backup.go:918`).
4. Both `up` calls fail together; `stacks/deploy.go:421-435` writes `Deployed=false` without asking whether
containers exist. Which `up` wins decides the end: running under „not deployed" (2026-09-22) or
`Created` (tonight, 14:39:48). **No deploy time limit exists** — shape (i) ruled out.
**Fix:** deploying apps leave the list; the volume leg re-asks right before the stop (SKIP, never FAIL);
`StopStack`/`StartStack` refuse a deploying stack (`ErrStackDeploying`) — the backstop for quiesce,
restore, export, storage and the Stop button, which all had the same exposure.
**Live (`32-*`, `33-*`):** outline's image removed so the deploy must pull, backup at +5 s: the backup ran
15:23:07–15:23:31, stopped gokapi, paperless-ngx and privatebin, **never outline**; `deployed
successfully (took 48.9s)`. The first fix run is kept and labelled NOT A RACE (images cached, 6.3 s).
## 3. Red-proofs (each seen failing, mutation restored, `=== RUN` checked)
| # | mutation | test → failure |
|---|---|---|
| 0.263.2 / MinAgent 0.131.0 | `min_controller_version=0.263.2`, `min_agent=0.131.0` | `00-floor-0.263.2.txt` |
| **0.264.0 / MinAgent 0.131.0** (14:35:08) | `min_controller_version=0.264.0`, `min_agent=0.131.0`; hub log `managed floor SERVED for demo-felhom: floor 0.264.0 … from declared` | demo-felhom 0.263.2 → **0.264.0 in 20 s**; demo-hp 9201 already 0.264.0 (`50-floor-0.264.0.txt`) |
| 1 | backup skip removed | `TestR634_VolumeLegNeverStopsADeployingApp` — `dumped=[outline]` |
| 2 | StopStack guard removed | `TestR634_StopAndStartRefuseADeployingStack` — `exit code -1` instead of `ErrStackDeploying` |
| 3 | held badge removed (v0.264.0's shape) | `TestR625_AHeldAppSaysStoppedRestoreNeeded` — badge missing, „Frissítés elérhető" present |
| 4 | reader funcs off | `TestR647_AReaderInTheOtherLanguageReadsTheHoldInTheirs` — English hold for the Hungarian reader |
| 5 | Hungarian phrase wired back | `TestR647_HeldEventCarriesTheCopyHoldsKey` |
| 6 | health status passed as severity | `TestR647_DisabledHealthChangeNamesTheSeverity` — `severity warn` |
| 7 | escalation removed | `TestR636_TwentyKillsInThirtyMinutesIsOneStorm` — `kills=20 … storms=0, want 1` |
| 8 | counter never read | `TestR636_ScanReadsTheCounterOfFlaggedContainersOnly` — `Kills:-1` |
| 9 | hub: not operator-only | `TestAppOOMStormIsAllowlistedAndOperatorOnly` — "must be operator-only" |
| 10 | hub: not allow-listed | same test — "must be in allowedEventTypes" |
| 11 | hub: no per-app cooldown | `TestR389_TheAllowListHasExactlyOneMember` |
## 3. The hub migration — rows changed (one-time, add-only)
Also pinned without a mutation run: 19 kills → no storm, 200 → one, one hiccup over 6 h → none, 20 kills
over 3 h → none, an unreadable counter → never. Gates: controller `go test ./...` rc=0 +
`controller_gates.py` rc=0; hub `go test ./...` rc=0 + `repo_gates.py` rc=0. Parity: `app_info_deployed`
and `stacks_full` regenerated — measured diff **exactly one line each, the new held badge**.
| customer | before | after |
|---|---|---|
| `demo-felhom` | 10 types | the same 10 + `app_update_undone`, `app_update_held` |
| `demo-hp` | `null` (nothing on) | `["app_update_undone","app_update_held"]` |
| `_resend-rotation-test` | `[]` | `["app_update_undone","app_update_held"]` |
## 4. Live proofs in both languages (9202)
Guard row `seed_app_update_events_v1 = 2026-09-23T11:49:03Z`. E-mail column never selected.
- **R-625 + R-647 (1)** (`44-*`), every box/reader pair: hu reader „Megállítva — visszaállítás szükséges",
title „A frissítés nem sikerült, és az automatikus visszaállítás sem."; en reader "Stopped — restore
needed", title "The update did not succeed, and the automatic undo did not either."; with the box in
ENGLISH the Hungarian reader reads Hungarian (page and API). No vikunja Update button; other apps keep
theirs (control). `POST …/update` → **409 `held`**.
- **R-636** (`45-*`, `46-*`): romm at 320M — kills 8, 13, **21 → `OOM STORM — 21 kills in 30 min (limit
320M, peak 320M)`** + `DROPPED event app_oom_storm (severity error)`; one storm line at 49 kills.
- **R-648** (`43-*`): vikunja's update ran its own `backing-up` phase.
## 4. The mails (9201, customer demo-hp) and their `notification_log` rows
## 5. Rows
| row | time (UTC) | type | leg | subject as received |
|---|---|---|---|---|
| 937 / 938 | 12:15 | undone | op / household | op `⚠️ demo-hp: app_update_undone` · hh **`Figyelmeztetés: vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut`** |
| 939 / 940 | 12:16 | held | op / household | op `🔴 demo-hp: app_update_held` · hh **`Hiba: vikunja: az alkalmazás leállítva, visszaállítás szükséges`** |
| 942 / 943 | 12:29 | undone | op / household | hh **`Warning: glance: the update did not work; the app runs on its previous version`** |
| 944 / 946 | 12:31 | held | op / household | hh **`Error: glance: the app is stopped and needs a restore`** |
**Closed (5):** R-634, R-625, R-636, R-647, R-648. **Opened (2):** R-649 (operator question — a deploy that
fails on its own), R-650 (tests that reach real docker act on DooPlex). **Open rows 334 → 331.**
`09` §6.4 parts 8 and 9 marked SHIPPED; `08` §6.2 has the storm rung and the two per-app grains.
All 8 `sent`. The box language was switched to English once, between the two rounds (the hub had two
reports in English before the second round). Household bodies carry the box's own sentence (the undone
line / the hold sentence) in the household's language; the operator's carry the Hungarian wire text.
Verbatim bodies: `41-mails-as-received.md`. **Per-app cooldown, live:** row 943 went out 14 minutes after
row 938, inside the 6-hour customer cooldown.
## 6. Floor
## 5. The pages in both languages (9202, then 9201)
0.265.0 / MinAgent 0.131.0, read back from the hub (`min_controller_version=0.265.0`,
`min_agent=0.131.0`; `managed floor SERVED` for demo-felhom and demo-hp); both on 0.265.0 in 20 s.
- undone: hu *„A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan
visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en *"The update of vikunja
at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically —
nothing was lost."*
- held, box hu: `?lang=hu` → the Hungarian hold sentence; `?lang=en` → *"The update did not succeed, and the
automatic undo did not either. … own drive, 2026-09-23 13:58 — this copy holds the settings, the database
and the data volumes."* — the whole sentence English now (before v0.264.0 only the prefix was).
- `GET /api/stacks/vikunja?lang=en`: label *"The update did not succeed"*, error and hold in English.
- **Leftover (R-647):** box en + `?lang=hu` showed the English hold.
## 7. Teardown — three layers
## 6. Red-proofs (each seen failing, then restored)
**Controller (9):** backfill stores nothing → `TestR646_…` *"the current app must have its record"*;
`dropped()` not called → `TestR620_…` *"want exactly TWO WARN lines, got 0"*; undone event not emitted →
*"exactly ONE app_update_undone event, got []"*; held event not emitted → *"…got []"*; page
`updateErrorText` not localised → `TestR606_UpdateSentencesFollowTheReader`; hold sentence ignores the
language → `undo_hold_test.go:62`; `UpdateErrorIn` returns the Hungarian → *"English = „A frissítés nem
indult el…""*; box seed not run → *"got [backup_failed offbox_enlarge_blocked]"*; sink not wired in main →
*"SetUpdateEventSink is never called"* (first mutation did not compile; retried with a compiling one).
**Hub (5):** seed call does nothing → *"c1: enabled_events = [backup_failed], want […]"*; customer per-app
suffix removed → *"want 3 household mails …, got 2"*; allow-list entry removed → *"must be in
allowedEventTypes"*; app-named headline removed → subject *"%s: a frissítés…"*; seed default removed →
*"defaultSeedEvents must include app_update_undone"* (first attempt matched no test; retried with a new
test that names both types). Every `-run` was checked for `=== RUN`.
Gates: controller `go test ./...` rc=0, `controller_gates.py` rc=0; hub `go build/vet/test` rc=0,
`repo_gates.py` rc=0. The R-389 fence test was widened on purpose, with its reason; the settings-page
parity fixture was regenerated (measured diff: exactly the two new toggles).
## 7. Rows
**Closed (3, moved to `CLOSED-ITEMS.md`):** R-606, R-620, R-646. **Opened (2):** R-647 (the three
leftovers), R-648 (the whole-box backup press). **Open rows 335 → 334.** `09` §6.4 parts 2 and 3 marked
SHIPPED; the capability map's undo row now carries the mail; STATUS rewritten.
## 8. Teardown — three layers
- **machine:** 9202 and 9201 — throwaway apps removed through the product (0 volumes, 0 undo copies),
test images removed, `controller.yaml` identical to the saved copy, catalog on live `cfcfe52`, 9201's
language back to `hu`, standing apps deployed and healthy. Both on 0.264.0 (= the floor).
- **host:** nothing provisioned on demo-hp.
- **hub:** v0.120.0 (planned); floor 0.264.0. Scratchpad password files deleted; hub DB copies deleted.
- **drill repo:** reset to live `main`, read back from the remote.
- **machine (9202):** all four test apps removed through the product; romm's drive folder by name (R-442
fence, as always); 0 volumes, 0 undo copies; test images removed by name where unused; config identical to
the saved copy; live catalog; language `hu`; recorders stopped. 9201 not touched.
- **host:** nothing provisioned. **hub:** v0.121.0 + floor (planned). **DooPlex:** the volume and image
above, removed. **drill repo:** reset to live `main`. Password files deleted from the scratchpad.
+16 -14
View File
@@ -1,23 +1,25 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-23 (evening) — the undo is on every demo machine now, and the household gets a mail when an update fails. The floor is raised.**
**Updated 2026-09-23 (late evening) — clean-up evening. The "runs but not installed" fault has its cause found and fixed. The fleet has the new version.**
**Decisions I took on my own: none.** One small deviation from your brief: the Hungarian mail line ends "nincs teendőd", not "teendője nincs", because the product speaks to the household as "te" everywhere.
**Decisions I took on my own: none.** One question for you is below.
**What the household gets now.** When an update fails and the machine puts the app back, the household gets one mail: "docmost: the update did not work; the app runs on its previous version". When the undo fails too, they get one mail that says the app is stopped and needs a restore, and which backup to use. The mail is in the household's own language. The operator gets the same two events. Two apps on the same night give two mails, not one.
**The top fault: an app could run while the box said "not installed".** Found and fixed. The cause: the "back up now" button backs up every app. It also caught an app that was still installing. It stopped that app in the middle of the install and started it again. The install then failed and wrote "not installed", while the app's containers sometimes kept running. I made this happen on purpose on the scratch machine, twice. With the fix, the backup leaves an installing app alone. I proved that on the same machine: the backup stopped three other apps and did not touch the one installing, and the install finished normally. Sparkyfitness did not fail on its own this time. It installed in 47 seconds.
**Proven on real machines, through the same buttons the page uses:**
- On the demo-hp machine: 4 mails to the household and 4 to the operator, first in Hungarian, then in English after one language switch. All arrived. The hub log shows all 8 as sent.
- On the scratch machine: the update messages on both app pages show in Hungarian and English. A machine with no hub now writes one warning line per kind of lost message.
- At start, the new version stores the health-check file of each up-to-date app, so its first undo asks the right question. It did this for 3 apps on one machine and 10 on the other.
- The floor is 0.264.0. The N100 demo machine updated itself within 20 seconds.
**A stopped app now says so.** When an update failed and the app is stopped, the badge says "Stopped — restore needed", in Hungarian or English. There is no Update button. Proven on the scratch machine in both languages.
**What was not clean.**
- My test "back up now" button backs up the whole machine. On demo-hp it stopped and restarted 9 of the 10 standing apps for a few seconds, twice. All came back healthy, with unchanged files. So "standing apps untouched" is not fully true. I filed a row to fix the test method.
- Three small leftovers, filed as one row: a reader who forces the other language sees a stopped app's sentence in the machine's language; the raw detail line of an English mail has one Hungarian phrase; and two of my log lines use unclear words.
**A long memory problem is now louder.** The box now counts real memory kills. If one app is killed 20 times in 30 minutes, you get one louder alarm ("error"), once. Households do not get it. Proven with RomM on the scratch machine: the alarm came at 21 kills and stayed one alarm at 49.
**Rows.** Three closed, two new. The list went from 335 to 334.
**The three small mail and log leftovers are fixed.** A reader who picks the other language now reads a stopped app's message in their own language.
**What needs you: nothing.** If you do nothing, the fleet keeps the new version. The leftovers are small and wait for a free evening.
**The test method is fixed.** My tests no longer press "back up now". An app's update makes its own backup instead.
**Nothing on your own machine, Peti's machine or the off-site box was touched, except the hub update. The test apps are gone from both demo machines. Both are back on the real catalogue.**
**Rows.** Five closed, two new. The list went from 334 to 331. The floor is 0.265.0, and both demo machines updated themselves within 20 seconds.
**What was not clean.** While I wrote one test, it created an empty storage volume on your own machine. I saw it at once, checked it was new and unused, and removed it. Nothing else was touched. I filed a row so tests cannot do this again.
**What needs you — one question.** When an install fails for its own reasons, some of its containers can stay running while the box says "not installed". The household can remove them, so nothing is stuck. What should happen?
- **Remove what the failed install started (my pick):** "failed" then always means nothing runs. The household presses Install again.
- **Leave it as it is:** nothing changes. The page can say "not installed" over running containers until the household removes them.
**Nothing on Peti's machine or the off-site box was touched. On your own machine, only the hub was updated, plus the test mistake above. The demo-hp machine was not touched by hand. The scratch machine is back to its three standing apps and the real catalogue.**
@@ -211,6 +211,22 @@ minutes later (90 m instead of 60). **One value, everywhere:** both staleness ch
hardcoded 30 m / 1 h until then; pinned by `TestControllerStatus_FollowsConfiguredThreshold`). A running
hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`).
**An out-of-memory storm gets its own, louder rung (R-636, controller v0.265.0 / hub v0.121.0, 2026-09-23).**
`app_oom` stays exactly as it was: `warning`, operator-only, ONCE per container run (R-514) — the once
is what stops a crash loop from mailing thousands of times (R-629). Beside it, `app_oom_storm`:
| event | severity | who | minted by | why that audience |
|---|---|---|---|---|
| `app_oom_storm` | **error** | **operator only** | controller v0.265.0, when the kernel's `oom_kill` counter of the SAME container run rises by **≥ 20 within 30 min**; once per run | raw container names and memory figures; the household's side is the dashboard tag |
**Why a counter and not the flag:** Docker's `OOMKilled` is sticky — true for the whole run after ONE
kill — so "the key re-fired N times" measures only how long ago the first kill was. The kernel's
`memory.events` `oom_kill` counts kills. **Why 20 in 30 minutes:** RomM's measured rate on 2026-09-22
was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202 (2026-09-23): RomM at 320M
reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:**
the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose
`OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms.
**Two event types added 2026-09-17, with who receives them:**
| event | severity | who | minted by | why that audience |
@@ -221,6 +237,8 @@ hub prints it at startup (`node_stale after 45m0s, node_down after 1h30m0s`).
| Family | Grain | Key carries | Why |
|---|---|---|---|
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
| update outcome (`app_update_undone`, `app_update_held`) | **per APP**, both legs | `…:<stack_name>` | hub v0.120.0 — one mail per app per outcome; the household leg has its own register (`perAppCustomerCooldownEvents`) |
| OOM storm (`app_oom_storm`) | **per APP**, operator | `…:<stack_name>` | hub v0.121.0 — no digest; the controller already sends it at most once per container run |
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
| everything else, incl. `crossdrive_failed` and `backup_integrity_failed` | per TYPE, per hour | — | coarse **on purpose** |
@@ -1049,8 +1049,8 @@ what the part can do to a household's data if it is wrong, not how likely that i
| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
| **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
| **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **8** | **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |
@@ -0,0 +1,8 @@
16:35:59 deploy sparkyfitness -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
16:35:59 + 0.0s sparkyfitness: deployed=True deploying=True state=not_deployed err='None'
16:36:08 + 9.1s sparkyfitness: deployed=True deploying=True state=deploying err='None'
16:36:29 + 30.3s sparkyfitness: deployed=True deploying=True state=degraded err='None'
16:36:48 + 48.4s sparkyfitness: deployed=True deploying=False state=starting err='None'
16:37:00 + 60.5s sparkyfitness: deployed=True deploying=False state=running err='None'
16:38:00 final: {"deployed": true, "deploying": false, "state": "running", "deploy_error": null, "containers": [{"name": "sparkyfitness", "image": "codewithcj/sparkyfitness:v0.17.3", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "sparkyfitness-server", "image": "codewithcj/sparkyfitness_server:v0.17.3", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "sparkyfitness-db", "image": "postgres:15-alpine", "state": "running", "status": "Up About a minute (healthy)"}]}
rc=0
@@ -0,0 +1,155 @@
== controller log (sparkyfitness lines + deploy/stop/start)
2026/09/23 14:35:59 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness/deploy-fields
2026/09/23 14:35:59 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness/deploy-fields (path=/stacks/sparkyfitness/deploy-fields)
2026/09/23 14:35:59 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/sparkyfitness/deploy
2026/09/23 14:35:59 router.go:82: [DEBUG] [api] POST /api/stacks/sparkyfitness/deploy (path=/stacks/sparkyfitness/deploy)
2026/09/23 14:35:59 router.go:438: [INFO] [api] Deploy requested for stack: sparkyfitness
2026/09/23 14:35:59 router.go:82: [DEBUG] [api] deployStack: name=sparkyfitness contentLength=70
2026/09/23 14:35:59 deploy.go:253: [DEBUG] Deploy sparkyfitness: received 2 user values
2026/09/23 14:35:59 deploy.go:258: [DEBUG] SUBDOMAIN = "sparkyfitness"
2026/09/23 14:35:59 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 4 encrypted, 4 sensitive fields
2026/09/23 14:35:59 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness
2026/09/23 14:35:59 deploy.go:383: [INFO] [stacks] Deploying stack sparkyfitness with 6 env vars: [APP_DB_PASSWORD, API_ENCRYPTION_KEY, BETTER_AUTH_SECRET, DOMAIN, SUBDOMAIN, DB_PASSWORD]
2026/09/23 14:35:59 manager.go:1533: [INFO] [stacks] Deploying stack sparkyfitness — checking 3 images...
2026/09/23 14:35:59 manager.go:1537: [DEBUG] codewithcj/sparkyfitness_server:v0.17.3 — not found locally, will pull
2026/09/23 14:35:59 manager.go:1537: [DEBUG] codewithcj/sparkyfitness:v0.17.3 — not found locally, will pull
2026/09/23 14:35:59 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=sparkyfitness, waiting 5s for state refresh
2026/09/23 14:35:59 manager.go:1414: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/sparkyfitness)
2026/09/23 14:35:59 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:35:59 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:02 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:02 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:04 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=sparkyfitness integrationsFound=0
2026/09/23 14:36:05 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:05 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:08 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:08 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:11 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:11 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:14 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:14 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:17 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml
2026/09/23 14:36:17 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:17 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:20 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:20 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:23 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:23 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:26 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:26 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:27 manager.go:815: [DEBUG] [stacks] restart-policy of down member "sparkyfitness" = "unless-stopped"
2026/09/23 14:36:29 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:29 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:32 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:32 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:35 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:35 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:37 healthprobe.go:99: [DEBUG] [stacks] RunHealthProbes: sparkyfitness has a health check but no container to probe — candidates: [sparkyfitness(stopped) sparkyfitness-server(stopped) sparkyfitness-db(starting)]
2026/09/23 14:36:38 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:38 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:41 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:41 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:45 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:45 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:46 deploy.go:451: [INFO] [stacks] Stack sparkyfitness deployed successfully (took 47.3s)
2026/09/23 14:36:46 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 4 encrypted, 4 sensitive fields
2026/09/23 14:36:46 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness
2026/09/23 14:36:47 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 0 encrypted, 4 sensitive fields
2026/09/23 14:36:47 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness
2026/09/23 14:36:47 installed.go:408: [INFO] [stacks] installed-images sparkyfitness: recorded 3 service(s) (sparkyfitness-db=postgres:15-alpine (sha256:f7d23353e1b1…), sparkyfitness-frontend=codewithcj/sparkyfitness:v0.17.3 (sha256:46d90e46bd87…), sparkyfitness-server=codewithcj/sparkyfitness_server:v0.17.3 (sha256:6aa7d9832324…))
2026/09/23 14:36:47 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/sparkyfitness — 6 env vars, 0 encrypted, 4 sensitive fields
2026/09/23 14:36:47 [INFO] [stacks] SaveAppConfig: saved config for sparkyfitness
2026/09/23 14:36:47 pin.go:93: [INFO] [stacks] pin sparkyfitness: sparkyfitness-db=postgres:15-alpine, sparkyfitness-frontend=codewithcj/sparkyfitness:v0.17.3, sparkyfitness-server=codewithcj/sparkyfitness_server:v0.17.3
2026/09/23 14:36:48 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:48 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:50 manager.go:1414: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/sparkyfitness)
2026/09/23 14:36:50 manager.go:1494: [INFO] [stacks] Stack sparkyfitness post-start status:
2026/09/23 14:36:50 manager.go:1497: [INFO] [stacks] sparkyfitness codewithcj/sparkyfitness:v0.17.3 running Up 3 seconds (health: starting)
2026/09/23 14:36:50 manager.go:1497: [INFO] [stacks] sparkyfitness-db postgres:15-alpine running Up 24 seconds (healthy)
2026/09/23 14:36:50 manager.go:1497: [INFO] [stacks] sparkyfitness-server codewithcj/sparkyfitness_server:v0.17.3 running Up 19 seconds (healthy)
2026/09/23 14:36:51 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:51 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:54 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:54 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:36:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:36:57 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:37:00 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:37:00 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:37:00 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:37:00 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
2026/09/23 14:37:03 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/sparkyfitness
2026/09/23 14:37:03 router.go:82: [DEBUG] [api] GET /api/stacks/sparkyfitness (path=/stacks/sparkyfitness)
== docker events (sparkyfitness)
1790174166 image pull codewithcj/sparkyfitness
1790174183 image pull codewithcj/sparkyfitness_server
1790174183 network create sparkyfitness_sparkyfitness-internal
1790174185 container create sparkyfitness-db
1790174185 container create sparkyfitness-server
1790174185 container create sparkyfitness
1790174185 network connect sparkyfitness_sparkyfitness-internal
1790174185 container start sparkyfitness-db
1790174190 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174190 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174190 container exec_die sparkyfitness-db
1790174190 container health_status: healthy sparkyfitness-db
1790174191 network connect sparkyfitness_sparkyfitness-internal
1790174191 container start sparkyfitness-server
1790174196 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174196 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174196 container exec_die sparkyfitness-server
1790174200 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174200 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174200 container exec_die sparkyfitness-db
1790174201 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174201 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174201 container exec_die sparkyfitness-server
1790174206 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174206 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174206 container exec_die sparkyfitness-server
1790174206 container health_status: healthy sparkyfitness-server
1790174206 network connect sparkyfitness_sparkyfitness-internal
1790174206 container start sparkyfitness
1790174210 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174210 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174210 container exec_die sparkyfitness-db
1790174211 container exec_create: wget --spider -q http://127.0.0.1:80/ sparkyfitness
1790174211 container exec_start: wget --spider -q http://127.0.0.1:80/ sparkyfitness
1790174211 container exec_die sparkyfitness
1790174211 container health_status: healthy sparkyfitness
1790174220 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174220 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174220 container exec_die sparkyfitness-db
1790174230 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174230 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174230 container exec_die sparkyfitness-db
1790174236 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174236 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174236 container exec_die sparkyfitness-server
1790174240 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174240 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174240 container exec_die sparkyfitness-db
1790174241 container exec_create: wget --spider -q http://127.0.0.1:80/ sparkyfitness
1790174241 container exec_start: wget --spider -q http://127.0.0.1:80/ sparkyfitness
1790174242 container exec_die sparkyfitness
1790174250 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174250 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174250 container exec_die sparkyfitness-db
1790174260 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174260 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174260 container exec_die sparkyfitness-db
1790174266 container exec_create: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174266 container exec_start: node -e require('http').get('http://127.0.0.1:3010/api/health',r=>process.exit(r.statusCode<400?0:1)).on('error',()=>process.exit(1)) sparkyfitness-server
1790174266 container exec_die sparkyfitness-server
1790174270 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174270 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174271 container exec_die sparkyfitness-db
1790174272 container exec_create: wget --spider -q http://127.0.0.1:80/ sparkyfitness
1790174272 container exec_start: wget --spider -q http://127.0.0.1:80/ sparkyfitness
1790174272 container exec_die sparkyfitness
1790174281 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174281 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174281 container exec_die sparkyfitness-db
1790174291 container exec_create: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174291 container exec_start: /bin/sh -c pg_isready -U sparky -d sparkyfitness_db sparkyfitness-db
1790174291 container exec_die sparkyfitness-db
grep: write error: Broken pipe
@@ -0,0 +1,7 @@
16:39:08 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
16:39:08 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None'
16:39:17 + 9.1s outline: deployed=True deploying=True state=deploying err='None'
16:39:51 + 42.4s outline: deployed=False deploying=False state=deploying err='exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021'
16:39:57 + 48.4s outline: deployed=False deploying=False state=stopped err='exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021'
16:40:57 final: {"deployed": false, "deploying": false, "state": "stopped", "deploy_error": "exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021cb Pulling fs layer 0B\n f25fb26220a8 Pulling fs layer 0B\n 4d027270bfe6 Pulling fs layer 0B\n d256567eb8f6 Pulling fs layer 0B\n bc77c6bc12ba Pulling fs layer 0B\n 167140e5ddd9 Pulling fs layer 0B\n b6b3029da283 Pulling fs layer 0B\n 5d5585bcfc4f Pulling fs layer 0B\n 5726d50585da Pulling fs layer 0B\n 4a58d711ba36
rc=0
@@ -0,0 +1 @@
16:39:28 WHOLE-BOX BACKUP pressed 20 s after the deploy was accepted -> 200 {'ok': True, 'message': 'Mentés elindítva'}
@@ -0,0 +1,132 @@
== controller log: outline + backup legs
2026/09/23 14:39:08 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline/deploy-fields
2026/09/23 14:39:08 router.go:82: [DEBUG] [api] GET /api/stacks/outline/deploy-fields (path=/stacks/outline/deploy-fields)
2026/09/23 14:39:08 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/outline/deploy
2026/09/23 14:39:08 router.go:82: [DEBUG] [api] POST /api/stacks/outline/deploy (path=/stacks/outline/deploy)
2026/09/23 14:39:08 router.go:438: [INFO] [api] Deploy requested for stack: outline
2026/09/23 14:39:08 router.go:82: [DEBUG] [api] deployStack: name=outline contentLength=64
2026/09/23 14:39:08 deploy.go:253: [DEBUG] Deploy outline: received 2 user values
2026/09/23 14:39:08 deploy.go:258: [DEBUG] SUBDOMAIN = "outline"
2026/09/23 14:39:08 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields
2026/09/23 14:39:08 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 14:39:08 deploy.go:383: [INFO] [stacks] Deploying stack outline with 5 env vars: [DOMAIN, SUBDOMAIN, SECRET_KEY, UTILS_SECRET, DB_PASSWORD]
2026/09/23 14:39:08 manager.go:1533: [INFO] [stacks] Deploying stack outline — checking 3 images...
2026/09/23 14:39:08 manager.go:1537: [DEBUG] outlinewiki/outline:1.9.1 — not found locally, will pull
2026/09/23 14:39:08 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=outline, waiting 5s for state refresh
2026/09/23 14:39:08 manager.go:1414: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline)
2026/09/23 14:39:08 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:08 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:11 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:11 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:13 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=outline integrationsFound=0
2026/09/23 14:39:14 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:14 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:17 backup.go:465: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [gokapi, outline, privatebin]
2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml
2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml
2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml
2026/09/23 14:39:17 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/outline/docker-compose.yml
2026/09/23 14:39:17 recovery_unit.go:224: [INFO] [backup] Recovery unit captured for outline → /mnt/sys_drive/felhom-data/backups/primary/outline (images=3, secrets-referenced=3, data_keys=0, portable-carried=3/3, withheld=0)
2026/09/23 14:39:17 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:17 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:20 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:20 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:23 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:23 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:26 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:26 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:28 backup.go:520: [INFO] [backup] Starting database dump run
2026/09/23 14:39:29 backup.go:908: [INFO] [backup] Stopping gokapi for safe volume dump
2026/09/23 14:39:29 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:29 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:30 backup.go:809: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_config → 3.0 KB
2026/09/23 14:39:31 backup.go:809: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_data → 99.5 KB
2026/09/23 14:39:31 backup.go:918: [INFO] [backup] Restarting gokapi after volume dump
2026/09/23 14:39:32 backup.go:908: [INFO] [backup] Stopping outline for safe volume dump
2026/09/23 14:39:32 manager.go:1234: [DEBUG] [stacks] StopStack outline: current state=deploying deployed=true containers=0
2026/09/23 14:39:32 manager.go:1237: [INFO] [stacks] Stopping stack: outline
2026/09/23 14:39:32 manager.go:1414: [DEBUG] Running: docker compose down (in /opt/docker/stacks/outline)
2026/09/23 14:39:32 manager.go:1246: [INFO] [stacks] Stack outline stopped successfully (took 0.1s)
2026/09/23 14:39:32 backup.go:784: [DEBUG] [backup] Dumping volume outline_outline_data for outline
2026/09/23 14:39:32 backup.go:809: [INFO] [backup] Volume dump: outline/outline_outline_data → 1.5 KB
2026/09/23 14:39:32 backup.go:784: [DEBUG] [backup] Dumping volume outline_outline_postgres_data for outline
2026/09/23 14:39:32 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:32 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:33 backup.go:809: [INFO] [backup] Volume dump: outline/outline_outline_postgres_data → 1.5 KB
2026/09/23 14:39:33 backup.go:784: [DEBUG] [backup] Dumping volume outline_outline_redis_data for outline
2026/09/23 14:39:33 backup.go:809: [INFO] [backup] Volume dump: outline/outline_outline_redis_data → 1.5 KB
2026/09/23 14:39:33 backup.go:918: [INFO] [backup] Restarting outline after volume dump
2026/09/23 14:39:33 manager.go:1152: [DEBUG] [stacks] StartStack outline: current state=deploying deployed=true
2026/09/23 14:39:33 manager.go:1155: [INFO] [stacks] Starting stack: outline
2026/09/23 14:39:33 manager.go:1162: [DEBUG] [stacks] StartStack outline: prepared 11 env vars for compose
2026/09/23 14:39:33 manager.go:1414: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline)
2026/09/23 14:39:36 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:36 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:39 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:39 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:42 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:42 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:45 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:45 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:48 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:48 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:48 manager.go:1423: [ERROR] [stacks] Command failed: docker compose up -d (in /opt/docker/stacks/outline) — exit code 1 (took 40.0s)
2026/09/23 14:39:48 manager.go:1429: [ERROR] [stacks] stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:39:48 deploy.go:424: [ERROR] [stacks] Stack outline deploy failed after 40.0s: exit code 1
stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:39:48 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields
2026/09/23 14:39:48 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 14:39:48 manager.go:1423: [ERROR] [stacks] Command failed: docker compose up -d (in /opt/docker/stacks/outline) — exit code 1 (took 15.3s)
2026/09/23 14:39:48 manager.go:1429: [ERROR] [stacks] stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:39:48 manager.go:1166: [ERROR] [stacks] Stack outline start failed after 15.3s: exit code 1
stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:39:48 backup.go:921: [ERROR] [backup] Failed to restart outline after volume dump: starting stack outline: exit code 1
stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:39:48 backup.go:740: [ERROR] [backup] Volume dump failed for outline: volume dump OK but restart failed for outline: starting stack outline: exit code 1
stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:39:48 backup.go:908: [INFO] [backup] Stopping paperless-ngx for safe volume dump
2026/09/23 14:39:51 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:51 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:54 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:54 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:56 backup.go:809: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 543.5 KB
2026/09/23 14:39:56 backup.go:809: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 67.8 MB
2026/09/23 14:39:57 backup.go:809: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 490.0 KB
2026/09/23 14:39:57 backup.go:918: [INFO] [backup] Restarting paperless-ngx after volume dump
2026/09/23 14:39:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:57 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:39:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:39:57 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:00 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:00 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:03 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:03 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:06 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:06 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:08 backup.go:908: [INFO] [backup] Stopping privatebin for safe volume dump
2026/09/23 14:40:09 backup.go:809: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB
2026/09/23 14:40:09 backup.go:918: [INFO] [backup] Restarting privatebin after volume dump
2026/09/23 14:40:09 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:09 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:09 backup.go:638: [WARN] [backup] some backup steps failed (FAIL outline volumes: volume dump OK but restart failed for outline: starting stack outline: exit code 1
stderr: Image outlinewiki/outline:1.9.1 Pulling
2026/09/23 14:40:12 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:12 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:15 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:15 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:17 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml
2026/09/23 14:40:18 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:18 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
2026/09/23 14:40:21 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/outline
2026/09/23 14:40:21 router.go:82: [DEBUG] [api] GET /api/stacks/outline (path=/stacks/outline)
== docker events: outline
1790174388 image pull outlinewiki/outline
1790174388 image pull outlinewiki/outline
1790174388 network create outline_outline-internal
1790174388 container create outline-redis
1790174388 container create outline-postgres
== now
outline-postgres Created
outline-redis Created
deployed: false
desired_state: running
@@ -0,0 +1,3 @@
# hub v0.121.0 deploy 2026-09-23T15:18:56Z
gitea.dooplex.hu/admin/felhom-hub:0.121.0
sync=Synced health=Healthy rev=3add9fa678b89973462a4b6ae3a3c35d2e642b9d
@@ -0,0 +1,10 @@
17:20:26 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
17:20:26 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None'
17:20:29 + 3.0s outline: deployed=True deploying=True state=stopped err='None'
17:20:35 + 9.1s outline: deployed=True deploying=False state=starting err='None'
17:20:56 + 30.3s outline: deployed=True deploying=False state=stopped err='None'
17:20:56 + 30.3s outline: deployed=True deploying=False state=degraded err='None'
17:21:03 + 36.4s outline: deployed=True deploying=False state=starting err='None'
17:21:18 + 51.5s outline: deployed=True deploying=False state=running err='None'
17:21:57 final: {"deployed": true, "deploying": false, "state": "running", "deploy_error": null, "containers": [{"name": "outline", "image": "outlinewiki/outline:1.9.1", "state": "running", "status": "Up 54 seconds (healthy)"}, {"name": "outline-postgres", "image": "postgres:16-alpine", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "outline-redis", "image": "redis:7-alpine", "state": "running", "status": "Up About a minute (healthy)"}]}
rc=0
@@ -0,0 +1 @@
17:20:46 WHOLE-BOX BACKUP pressed 20 s after the deploy was accepted -> 200 {'ok': True, 'message': 'Mentés elindítva'}
@@ -0,0 +1,67 @@
2026/09/23 15:19:56 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml
2026/09/23 15:19:56 backup.go:1187: [INFO] [backup] Backup status cache refreshed
2026/09/23 15:19:57 sync.go:416: [DEBUG] [sync] outline/docker-compose.yml: hash match, skipped
2026/09/23 15:19:57 sync.go:416: [DEBUG] [sync] outline/.felhom.yml: hash match, skipped
2026/09/23 15:20:26 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/outline/deploy
2026/09/23 15:20:26 router.go:81: [DEBUG] [api] POST /api/stacks/outline/deploy (path=/stacks/outline/deploy)
2026/09/23 15:20:26 router.go:431: [INFO] [api] Deploy requested for stack: outline
2026/09/23 15:20:26 router.go:81: [DEBUG] [api] deployStack: name=outline contentLength=64
2026/09/23 15:20:26 deploy.go:253: [DEBUG] Deploy outline: received 2 user values
2026/09/23 15:20:26 deploy.go:258: [DEBUG] SUBDOMAIN = "outline"
2026/09/23 15:20:26 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields
2026/09/23 15:20:26 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:20:26 deploy.go:383: [INFO] [stacks] Deploying stack outline with 5 env vars: [DOMAIN, SUBDOMAIN, SECRET_KEY, UTILS_SECRET, DB_PASSWORD]
2026/09/23 15:20:26 manager.go:1546: [INFO] [stacks] Deploying stack outline — checking 3 images...
2026/09/23 15:20:26 manager.go:1552: [DEBUG] outlinewiki/outline:1.9.1 — found locally
2026/09/23 15:20:26 manager.go:1427: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline)
2026/09/23 15:20:26 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=outline, waiting 5s for state refresh
2026/09/23 15:20:31 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=outline integrationsFound=0
2026/09/23 15:20:32 deploy.go:451: [INFO] [stacks] Stack outline deployed successfully (took 6.3s)
2026/09/23 15:20:32 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields
2026/09/23 15:20:32 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:20:33 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields
2026/09/23 15:20:33 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:20:33 installed.go:408: [INFO] [stacks] installed-images outline: recorded 3 service(s) (outline=outlinewiki/outline:1.9.1 (sha256:9fe2cbdcecce…), outline-postgres=postgres:16-alpine (sha256:721873c34ceb…), outline-redis=redis:7-alpine (sha256:858f009f9709…))
2026/09/23 15:20:33 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields
2026/09/23 15:20:33 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:20:33 pin.go:93: [INFO] [stacks] pin outline: outline=outlinewiki/outline:1.9.1, outline-postgres=postgres:16-alpine, outline-redis=redis:7-alpine
2026/09/23 15:20:36 manager.go:1427: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/outline)
2026/09/23 15:20:36 manager.go:1507: [INFO] [stacks] Stack outline post-start status:
2026/09/23 15:20:36 manager.go:1510: [INFO] [stacks] outline outlinewiki/outline:1.9.1 running Up 3 seconds (health: starting)
2026/09/23 15:20:36 manager.go:1510: [INFO] [stacks] outline-postgres postgres:16-alpine running Up 9 seconds (healthy)
2026/09/23 15:20:36 manager.go:1510: [INFO] [stacks] outline-redis redis:7-alpine running Up 9 seconds (healthy)
2026/09/23 15:20:46 dbdump.go:147: [DEBUG] DiscoverDatabases: docker ps output: a0affd38f0c9 outline outline outlinewiki/outline:1.9.1
10d80fc75f8c outline-postgres outline postgres:16-alpine
fdd9eaa9a3bd outline-redis outline redis:7-alpine
2026/09/23 15:20:46 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container outline (image=outlinewiki/outline:1.9.1, not a database)
2026/09/23 15:20:46 dbdump.go:176: [DEBUG] DiscoverDatabases: found postgres container: outline-postgres (id=10d80fc75f8c)
2026/09/23 15:20:46 dbdump.go:196: [DEBUG] DiscoverDatabases: outline-postgres → stack=outline, dbUser=outline, dbName=outline
2026/09/23 15:20:46 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container outline-redis (image=redis:7-alpine, not a database)
2026/09/23 15:20:46 backup.go:548: [INFO] [backup] Discovered 2 database(s): outline-postgres(postgres), paperless-postgres(postgres)
2026/09/23 15:20:46 dbdump.go:264: [DEBUG] DumpOne: starting dump for container=outline-postgres, stack=outline, dbType=postgres, dest=/mnt/sys_drive/felhom-data/backups/primary/outline/db-dumps/outline-postgres.sql
2026/09/23 15:20:46 dbdump.go:301: [DEBUG] DumpOne: pg_dump command: docker exec 10d80fc75f8c pg_dump -U outline -d outline --clean --if-exists --no-owner --no-privileges
2026/09/23 15:20:46 [WARN] [backup] ValidateDump: /mnt/sys_drive/felhom-data/backups/primary/outline/db-dumps/outline-postgres.sql is structurally valid (32 tables) but its accounts table has NO rows — the dump may predate the customer's data
2026/09/23 15:20:46 dbdump.go:403: [DEBUG] DumpOne: completed outline-postgres → outline-postgres.sql (size=98.6 KB, valid=true, tables=32, duration=254ms)
2026/09/23 15:20:46 dbdump.go:408: [INFO] [backup] DB dump: outline-postgres → outline-postgres.sql (98.6 KB, 254ms, 32 tables)
2026/09/23 15:20:48 backup.go:918: [INFO] [backup] Stopping outline for safe volume dump
2026/09/23 15:20:48 manager.go:1247: [DEBUG] [stacks] StopStack outline: current state=starting deployed=true containers=3
2026/09/23 15:20:48 manager.go:1250: [INFO] [stacks] Stopping stack: outline
2026/09/23 15:20:48 manager.go:1427: [DEBUG] Running: docker compose down (in /opt/docker/stacks/outline)
2026/09/23 15:20:55 manager.go:1259: [INFO] [stacks] Stack outline stopped successfully (took 6.2s)
2026/09/23 15:20:55 backup.go:794: [DEBUG] [backup] Dumping volume outline_outline_redis_data for outline
2026/09/23 15:20:55 backup.go:819: [INFO] [backup] Volume dump: outline/outline_outline_redis_data → 2.5 KB
2026/09/23 15:20:55 backup.go:794: [DEBUG] [backup] Dumping volume outline_outline_data for outline
2026/09/23 15:20:55 backup.go:819: [INFO] [backup] Volume dump: outline/outline_outline_data → 1.5 KB
2026/09/23 15:20:55 backup.go:794: [DEBUG] [backup] Dumping volume outline_outline_postgres_data for outline
2026/09/23 15:20:56 backup.go:819: [INFO] [backup] Volume dump: outline/outline_outline_postgres_data → 49.3 MB
2026/09/23 15:20:56 backup.go:928: [INFO] [backup] Restarting outline after volume dump
2026/09/23 15:20:56 manager.go:1158: [DEBUG] [stacks] StartStack outline: current state=stopped deployed=true
2026/09/23 15:20:56 manager.go:1161: [INFO] [stacks] Starting stack: outline
2026/09/23 15:20:56 manager.go:1168: [DEBUG] [stacks] StartStack outline: prepared 11 env vars for compose
== now
outline Up 57 seconds (healthy)
outline-postgres Up About a minute (healthy)
outline-redis Up About a minute (healthy)
deployed: true
desired_state: running
grep: write error: Broken pipe
@@ -0,0 +1,8 @@
17:23:02 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
17:23:02 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None'
17:23:08 + 6.1s outline: deployed=True deploying=True state=deploying err='None'
17:23:47 + 45.4s outline: deployed=True deploying=True state=degraded err='None'
17:23:54 + 51.4s outline: deployed=True deploying=False state=starting err='None'
17:24:27 + 84.8s outline: deployed=True deploying=False state=running err='None'
17:25:27 final: {"deployed": true, "deploying": false, "state": "running", "deploy_error": null, "containers": [{"name": "outline", "image": "outlinewiki/outline:1.9.1", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "outline-postgres", "image": "postgres:16-alpine", "state": "running", "status": "Up About a minute (healthy)"}, {"name": "outline-redis", "image": "redis:7-alpine", "state": "running", "status": "Up About a minute (healthy)"}]}
rc=0
@@ -0,0 +1 @@
17:23:07 WHOLE-BOX BACKUP pressed 5 s after the deploy was accepted -> 200 {'ok': True, 'message': 'Mentés elindítva'}
@@ -0,0 +1,46 @@
10d80fc75f8c outline-postgres outline postgres:16-alpine
fdd9eaa9a3bd outline-redis outline redis:7-alpine
2026/09/23 15:23:02 router.go:81: [DEBUG] [api] POST /api/stacks/outline/deploy (path=/stacks/outline/deploy)
2026/09/23 15:23:02 router.go:431: [INFO] [api] Deploy requested for stack: outline
2026/09/23 15:23:02 router.go:81: [DEBUG] [api] deployStack: name=outline contentLength=64
2026/09/23 15:23:02 deploy.go:253: [DEBUG] Deploy outline: received 2 user values
2026/09/23 15:23:02 deploy.go:258: [DEBUG] SUBDOMAIN = "outline"
2026/09/23 15:23:02 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields
2026/09/23 15:23:02 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:23:02 deploy.go:383: [INFO] [stacks] Deploying stack outline with 5 env vars: [UTILS_SECRET, DB_PASSWORD, DOMAIN, SUBDOMAIN, SECRET_KEY]
2026/09/23 15:23:02 manager.go:1546: [INFO] [stacks] Deploying stack outline — checking 3 images...
2026/09/23 15:23:02 manager.go:1550: [DEBUG] outlinewiki/outline:1.9.1 — not found locally, will pull
2026/09/23 15:23:02 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=outline, waiting 5s for state refresh
2026/09/23 15:23:02 manager.go:1427: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/outline)
2026/09/23 15:23:07 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=outline integrationsFound=0
2026/09/23 15:23:07 backup.go:918: [INFO] [backup] Stopping gokapi for safe volume dump
2026/09/23 15:23:08 backup.go:819: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_config → 3.0 KB
2026/09/23 15:23:09 backup.go:819: [INFO] [backup] Volume dump: gokapi/gokapi_gokapi_data → 99.5 KB
2026/09/23 15:23:10 backup.go:918: [INFO] [backup] Stopping paperless-ngx for safe volume dump
2026/09/23 15:23:17 backup.go:819: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 580.0 KB
2026/09/23 15:23:18 backup.go:819: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 67.8 MB
2026/09/23 15:23:18 backup.go:819: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 512.5 KB
2026/09/23 15:23:30 backup.go:918: [INFO] [backup] Stopping privatebin for safe volume dump
2026/09/23 15:23:31 backup.go:819: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB
2026/09/23 15:23:46 manager.go:815: [DEBUG] [stacks] restart-policy of down member "outline" = "unless-stopped"
2026/09/23 15:23:51 deploy.go:451: [INFO] [stacks] Stack outline deployed successfully (took 48.9s)
2026/09/23 15:23:51 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 3 encrypted, 3 sensitive fields
2026/09/23 15:23:51 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:23:51 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields
2026/09/23 15:23:51 installed.go:408: [INFO] [stacks] installed-images outline: recorded 3 service(s) (outline=outlinewiki/outline:1.9.1 (sha256:9fe2cbdcecce…), outline-postgres=postgres:16-alpine (sha256:721873c34ceb…), outline-redis=redis:7-alpine (sha256:858f009f9709…))
2026/09/23 15:23:51 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:23:51 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, 0 encrypted, 3 sensitive fields
2026/09/23 15:23:51 pin.go:93: [INFO] [stacks] pin outline: outline=outlinewiki/outline:1.9.1, outline-postgres=postgres:16-alpine, outline-redis=redis:7-alpine
2026/09/23 15:23:51 [INFO] [stacks] SaveAppConfig: saved config for outline
2026/09/23 15:23:54 manager.go:1427: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/outline)
2026/09/23 15:23:54 manager.go:1507: [INFO] [stacks] Stack outline post-start status:
2026/09/23 15:23:54 manager.go:1510: [INFO] [stacks] outline outlinewiki/outline:1.9.1 running Up 3 seconds (health: starting)
2026/09/23 15:23:54 manager.go:1510: [INFO] [stacks] outline-postgres postgres:16-alpine running Up 9 seconds (healthy)
2026/09/23 15:23:54 manager.go:1510: [INFO] [stacks] outline-redis redis:7-alpine running Up 9 seconds (healthy)
2026/09/23 15:23:56 manager.go:584: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=true composePath=/opt/docker/stacks/outline/docker-compose.yml
== now
outline Up About a minute (healthy)
outline-postgres Up About a minute (healthy)
outline-redis Up About a minute (healthy)
deployed: true
desired_state: running
@@ -0,0 +1,14 @@
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
17:27:51 catalog cache: d4eeff6 DRILL vikunja: vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.3.0
17:27:53 romm drill memory: 86: memory: 320M # DRILL (R-636 live proof): too small on purpose — workers are OOM-killed in a loop
@@ -0,0 +1,5 @@
17:28:14 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm']
17:28:14 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH']
17:28:14 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
17:29:34 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}
17:29:34 romm deployed: True
@@ -0,0 +1,13 @@
17:29:52 === prep vikunja: deploy at the drill FROM pin
17:29:52 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
17:29:57 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
17:29:57 deployed: True
17:30:03 vikunja: register http=200
17:30:03 vikunja: create project http=201
17:30:03 vikunja: readback of the seeded project http=200 ok=True
17:30:03 C1 A: True
17:30:03 [4] backup press SKIPPED for vikunja (R-648: whole-box only; the update's backing-up phase backs up vikunja alone)
17:30:03 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill1a7c22","created":"2026-09
17:30:04 vikunja B: project readback=True attachment content readback=True
17:30:04 B reads back: True
17:30:07 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
@@ -0,0 +1,29 @@
17:30:19 [5] drill commit 21b56e4891e3: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
17:30:24 badge caught up after 4.5 s
17:30:24 === vikunja: a cut-off undo copy
17:30:24 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
17:30:24 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt…
17:30:26 + 2.1s phase=safety-dump label=Adatbázis pillanatkép…
17:30:27 + 3.2s phase=pulling label=Új verzió letöltése…
17:30:32 + 7.9s phase=copying label=Az adatok másolása a frissítés előtt…
17:30:33 + 9.5s phase=starting label=Indítás az új verzióval…
17:30:34 + 10.0s phase=verifying label=Működés ellenőrzése…
17:30:37 >>> finished-marker removed from one copy:
copy: vikunja_vikunja_data.pre-update-20260923T153032Z
total 12
drwxr-xr-x 3 root root 4096 Sep 23 15:30 .
drwxr-xr-x 1 root root 4096 Sep 23 15:30 ..
drwxr-xr-x 2 root root 4096 Sep 23 15:30 data
-rw-r--r-- 1 root root 0 Sep 23 15:30 felhom-undo-complete
data
17:32:05 + 100.9s phase=undoing label=Visszaállítás az előző változatra…
17:32:05 + 101.5s phase=failed label=A frissítés nem sikerült
17:32:05 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
17:32:09 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 17:32 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: second drive, 2026-09-23 17:30 — this copy holds the settings, the database and the data volumes."}}
17:32:09 box language -> en: http 302
17:32:09 PAGE (box language en): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 17:32-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: második meghajtó, 2026-09-23 17:30 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 17:32 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: second drive, 2026-09-23 17:30 — this copy holds the settings, the database and the data volumes."}}
17:32:09 box language -> hu: http 302
17:32:11 copies kept: vikunja_vikunja_data.pre-update-20260923T153032Z
vikunja_vikunja_db.pre-update-20260923T153032Z
@@ -0,0 +1,16 @@
17:32:35 box language -> hu: http 302
17:32:36 box=hu reader=hu: badge=('Megállítva — visszaállítás szükséges', 'A frissítés nem sikerült, és az automatikus visszaállítás sem.')
17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True
17:32:36 API: phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új v'
17:32:36 box=hu reader=en: badge=('Stopped — restore needed', 'The update did not succeed, and the automatic undo did not either.')
17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True
17:32:36 API: phase=failed err='The update did not succeed, and the automatic undo did not either. The data is a'
17:32:36 box language -> en: http 302
17:32:36 box=en reader=hu: badge=('Megállítva — visszaállítás szükséges', 'A frissítés nem sikerült, és az automatikus visszaállítás sem.')
17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True
17:32:36 API: phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új v'
17:32:36 box=en reader=en: badge=('Stopped — restore needed', 'The update did not succeed, and the automatic undo did not either.')
17:32:36 invite words on the app page: False | vikunja Update button on the list: False | control, another app's button: True
17:32:36 API: phase=failed err='The update did not succeed, and the automatic undo did not either. The data is a'
17:32:36 the button's endpoint still refuses: 409 {'ok': False, 'data': {'reason': 'held'}, 'error': 'The update did not succeed, and the automatic undo did not either. The data is as the new version left it. T
17:32:36 box language -> hu: http 302
@@ -0,0 +1,13 @@
15:32:53
== the scan's own lines (first 4)
2026/09/23 15:30:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 8, limit 320M, peak 320M)
2026/09/23 15:30:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 13, limit 320M, peak 320M)
2026/09/23 15:31:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 21, limit 320M, peak 320M)
2026/09/23 15:31:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 28, limit 320M, peak 320M)
== storm + dropped lines
2026/09/23 15:30:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom (severity warning) — further app_oom events are logged at DEBUG only
2026/09/23 15:31:10 notifier.go:738: [ERROR] [notify] romm: container romm OOM STORM — 21 kills in 30 min (limit 320M, peak 320M)
2026/09/23 15:31:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom_storm (severity error) — further app_oom_storm events are logged at DEBUG only
== kernel counter now
oom_kill 45
started=2026-09-23T15:28:39.049958056Z restarts=0 oom=true
@@ -0,0 +1,104 @@
== 1b (v0.264.0) docker events for outline
1790174388 image pull outlinewiki/outline
1790174388 image pull outlinewiki/outline
1790174388 network create outline_outline-internal
1790174388 container create outline-redis
1790174388 container create outline-postgres
1790174712 container destroy outline-postgres
1790174712 container destroy outline-redis
1790174712 network destroy outline_outline-internal
== fix run (v0.265.0) docker events for outline
1790176826 network create outline_outline-internal
1790176826 container create outline-redis
1790176826 container create outline-postgres
1790176826 container create outline
1790176827 network connect outline_outline-internal
1790176827 network connect outline_outline-internal
1790176827 container start outline-redis
1790176827 container start outline-postgres
1790176832 container exec_create: redis-cli ping outline-redis
1790176832 container exec_start: redis-cli ping outline-redis
1790176832 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176832 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176832 container exec_die outline-redis
1790176832 container health_status: healthy outline-redis
1790176832 container exec_die outline-postgres
1790176832 container health_status: healthy outline-postgres
1790176832 network connect outline_outline-internal
1790176832 container start outline
1790176837 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176837 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176837 container exec_die outline
1790176842 container exec_create: redis-cli ping outline-redis
1790176842 container exec_start: redis-cli ping outline-redis
1790176842 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176842 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176842 container exec_die outline-redis
1790176842 container exec_die outline-postgres
1790176842 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176842 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176843 container exec_die outline
1790176846 container exec_create: pg_dump -U outline -d outline --clean --if-exists --no-owner --no-privileges outline-postgres
1790176846 container exec_start: pg_dump -U outline -d outline --clean --if-exists --no-owner --no-privileges outline-postgres
1790176846 container exec_die outline-postgres
1790176848 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176848 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176848 container exec_die outline
1790176848 container kill outline
1790176852 container exec_create: redis-cli ping outline-redis
1790176852 container exec_start: redis-cli ping outline-redis
1790176852 container exec_die outline-redis
1790176852 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176852 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176852 container exec_die outline-postgres
1790176853 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176853 container exec_start: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
1790176853 container exec_die outline
1790176854 network disconnect outline_outline-internal
1790176854 container stop outline
1790176854 container die outline
1790176854 container destroy outline
1790176854 container kill outline-redis
1790176854 container kill outline-postgres
1790176854 network disconnect outline_outline-internal
1790176854 container stop outline-redis
1790176854 container die outline-redis
1790176854 container destroy outline-redis
1790176854 network disconnect outline_outline-internal
1790176854 container stop outline-postgres
1790176854 container die outline-postgres
1790176854 container destroy outline-postgres
1790176855 network destroy outline_outline-internal
1790176856 network create outline_outline-internal
1790176856 container create outline-redis
1790176856 container create outline-postgres
1790176856 container create outline
1790176856 network connect outline_outline-internal
1790176856 network connect outline_outline-internal
1790176856 container start outline-postgres
1790176856 container start outline-redis
1790176861 container exec_create: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176861 container exec_start: /bin/sh -c pg_isready -U outline -d outline outline-postgres
1790176861 container exec_create: redis-cli ping outline-redis
1790176861 container exec_start: redis-cli ping outline-redis
1790176861 container exec_die outline-postgres
1790176861 container health_status: healthy outline-postgres
1790176861 container exec_die outline-redis
1790176861 container health_status: healthy outline-redis
1790176862 network connect outline_outline-internal
1790176862 container start outline
1790176867 container exec_create: node -e const http = require('http'); http.get('http://127.0.0.1:3000/_health', (r) => { process.exit(r.statusCode === 200 ? 0 : 1) }).on('error', () => process.exit(1)) outline
== storm log: every OOM/STORM line
2026/09/23 15:30:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 8, limit 320M, peak 320M)
2026/09/23 15:30:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom (severity warning) — further app_oom events are logged at DEBUG only
2026/09/23 15:30:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 13, limit 320M, peak 320M)
2026/09/23 15:31:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 21, limit 320M, peak 320M)
2026/09/23 15:31:10 notifier.go:738: [ERROR] [notify] romm: container romm OOM STORM — 21 kills in 30 min (limit 320M, peak 320M)
2026/09/23 15:31:10 notifier.go:1255: [WARN] notifier disabled (no hub configured): DROPPED event app_oom_storm (severity error) — further app_oom_storm events are logged at DEBUG only
2026/09/23 15:31:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 28, limit 320M, peak 320M)
2026/09/23 15:32:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 34, limit 320M, peak 320M)
2026/09/23 15:32:40 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 40, limit 320M, peak 320M)
2026/09/23 15:33:10 main.go:854: [WARN] [deadapp] romm: container romm was OOM-killed (started 2026-09-23T15:28:39.049958056Z; kills 49, limit 320M, peak 320M)
== kernel counter at teardown
oom_kill 49
recorders left: 0
@@ -0,0 +1,52 @@
17:33:26 === remove vikunja through the product
17:33:26 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
17:33:58 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
17:34:06 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
17:34:06 === remove romm through the product
17:34:17 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'}
17:34:22 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
17:34:22 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
17:34:49 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat
17:34:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm'
17:34:58 undo copies left: 0
17:35:00 volumes left: 0
17:35:02 romm drive folder: removed-by-name
17:35:10 rmi vikunja/vikunja:2.3.0
rmi vikunja/vikunja:2.6.0
rmi rommapp/romm:5.3.0
rmi mariadb:11.4
rmi outlinewiki/outline:1.9.1
KEPT postgres:16-alpine (in use by 1)
rmi codewithcj/sparkyfitness:v0.17.3
rmi codewithcj/sparkyfitness_server:v0.17.3
rmi postgres:15-alpine
rmi gitea.dooplex.hu/admin/felhom-controller:0.264.0
r634-scratch-removed
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
17:36:00 catalog cache: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
17:36:02 controller.yaml vs saved: IDENTICAL
17:36:04 language: hu
17:36:04 deployed: ['gokapi', 'paperless-ngx', 'privatebin']
17:36:06 felhom-controller Up 54 seconds (healthy)
filebrowser Up 5 hours (healthy)
gokapi Restarting (1) 9 seconds ago
paperless-postgres Up 12 minutes (healthy)
paperless-redis Up 12 minutes (healthy)
paperless-webserver Up 12 minutes (healthy)
privatebin Up 12 minutes (healthy)
traefik Up 5 hours
gitea.dooplex.hu/admin/felhom-controller:0.265.0
@@ -0,0 +1,16 @@
=== FLOOR -> 0.265.0, MinAgent 0.131.0 declared (2026-09-23 17:36:28)
POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set
--- read back from the hub:
"min_controller_version" value="0.265.0"
"min_agent" value="0.131.0"
"min_agent" value="0.131.0"
--- watching the two demo boxes (up to 3 min)
+10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up 3 hours (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up 3 hours (healthy)
+20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.265.0 Up 7 seconds (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.265.0 Up 10 seconds (healthy)
--- hub log:
2026/09/23 17:36:31 [INFO] managed floor SERVED for demo-felhom: floor 0.265.0, agent requirement "0.131.0" from declared (golden 0.258.0)
2026/09/23 17:36:31 [INFO] managed floor SERVED for demo-hp: floor 0.265.0, agent requirement "0.131.0" from declared (golden 0.258.0)
--- hosts page:
demo-felhom-8363b5
demo-hp-bb76ea
drill-r50-0a4f9a
@@ -0,0 +1,50 @@
# Clean-up evening — R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0), 2026-09-23
**Method: endpoint-level** on scratch guest 9202 (the endpoints the UI invokes; pages as HTML with
`?lang=hu|en`); `docker events` + the controller log recorded inside the guest from before the first
call (`setsid docker events/logs -f` to a file — never piped through a buffering command). No browser.
Harness copied from `undo-fleet-2026-09-23/`; **`walk.py backup_now` presses nothing now (R-648)**.
## Part 1 — R-634, diagnosed before any code
| # | what | result | evidence |
|---|---|---|---|
| 1a | `sparkyfitness` alone, live pin, v0.264.0 | **did NOT reproduce**: deployed in 47.3 s, running | `10-*`, `11-*` |
| 1b | `outline` + whole-box backup at +20 s, v0.264.0 | **reproduced**: `StopStack outline: state=deploying deployed=true containers=0` (14:39:32) → `Restarting outline after volume dump` (14:39:33) → both `compose up` fail at 14:39:48 → `deployed=false`, containers `Created` | `12-*`, `13-*` |
| fix, try 1 | same, v0.265.0 | **not a race** — images cached, deploy done in 6.3 s before the backup came; kept as a record, not as proof | `30-fix-attempt1-*` |
| fix, try 2 | outline image removed by name, backup at +5 s, v0.265.0 | **the backup ran 15:23:07–15:23:31 across the deploy, stopped gokapi, paperless-ngx, privatebin and never outline**; deploy `successfully (took 48.9s)`, `deployed=true`, running | `32-*`, `33-*` |
Mechanism at `file:line`: `stacks/deploy.go:395` (in-memory `Deployed=true` at accept) →
`cmd/controller/main.go:2556` `ListDeployedStacks` (flag only) → `backup/backup.go:693` `runVolumeDumps` →
`DumpAppVolumesSafe` `StopStack` + `StartStack` (`backup.go:908/918`) → `stacks/deploy.go:421-435` failure
branch writes `Deployed=false`. There is no deploy time limit (the brief's shape (i) is ruled out).
## Parts 2–4 (live, 9202, drill catalog: romm 320M/4 workers, vikunja at 2.3.0)
| proof | result | evidence |
|---|---|---|
| R-625 badge, 4 box/reader pairs | hu reader: „Megállítva — visszaállítás szükséges", title „A frissítés nem sikerült, és az automatikus visszaállítás sem."; en reader: "Stopped — restore needed", title "The update did not succeed, and the automatic undo did not either."; no „Frissítés elérhető"/"Update available"; **no vikunja Update button** on the list while other apps keep theirs (control); API `POST …/update` → **409 `held`** | `44-*` |
| R-647 (1) | box **en**, `?lang=hu`: the hold in Hungarian on the app page, and `update_error` Hungarian in the API | `43-*`, `44-*` |
| R-648 | vikunja's update ran its own `backing-up` phase; no whole-box press | `42-*`, `43-*` |
| R-636 | romm: kills 8 → 13 → **21 → `OOM STORM — 21 kills in 30 min (limit 320M, peak 320M)`** + `DROPPED event app_oom_storm (severity error)` (R-620 witness, 9202 has no hub); still ONE storm at 49 kills | `45-*`, `46-*` |
Hub v0.121.0 deployed first (`20-*`). The hub side of R-636 is proven by unit tests only: posting a
synthetic event with a real customer's key would put a fabricated alarm into the operator's record.
## Floor
0.265.0 / MinAgent 0.131.0, read back; demo-hp and demo-felhom both on 0.265.0 within 20 s (`50-*`).
## Teardown — three layers
- **machine (9202):** outline, sparkyfitness, vikunja, romm removed through the product (romm's
with-data remove refused 409 on the drive path, R-442 as in every drill — its drive folder removed by
name); 0 volumes, 0 undo copies; test images removed by name only where no container used them
(`postgres:16-alpine` kept — in use); `controller.yaml` identical to `.pre-cleanup`; catalog on live
`cfcfe52`; language `hu`; recorders stopped and `/root/r634` removed; standing apps as at the start.
- **host:** nothing provisioned on demo-hp. **9201 not touched** (no backup press; it reached 0.265.0 by
the floor).
- **hub:** v0.121.0 (planned); floor 0.265.0.
- **DooPlex:** one empty volume `outline_outline_data` and one `alpine:3.20` image were created by MY first
test draft / my inspection of it; both verified unused and removed by name (R-650).
- **drill repo:** reset to live `main`, read back from the remote.
@@ -0,0 +1,266 @@
#!/usr/bin/env python3
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
drill catalog.
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
Start supplies the app's secrets (they are never decrypted here).
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
"""
import json, os, sys, time
sys.path.insert(0, ".")
import walk as w
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
romm_verify_b, vik_seed_b, vik_verify_b
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
SUFFIX = ".pre-undo"
def seed_b(app, sub, A):
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
def verify_b(app, sub, A, B):
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
return all(r.values()) if isinstance(r, dict) else r
def db_state(app):
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
if app == "docmost":
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
if app == "romm":
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
alembic = q("select version_num from alembic_version")
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
return w.guest("""python3 - <<'PY'
import sqlite3
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
m=c.execute("select count(*), max(id) from migration").fetchone()
print(f"tables={t} ledger={m[0]} newest={m[1]}")
PY""").strip()
def volumes(app):
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
if not v.endswith(SUFFIX)]
def containers(app):
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
def stage_prep(app):
s = {"app": app, "sub": SUB[app]}
w.say(f"=== prep {app}: deploy at the drill FROM pin")
w.say("deployed:", w.deploy(app, s["sub"]))
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
w.backup_now(app)
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
s["volumes"] = volumes(app)
w.say("named volumes:", s["volumes"])
save(app, s)
def stage_fcopy(app):
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
s = load(app)
s["state_before_update"] = db_state(app)
w.say("db state before the update:", s["state_before_update"])
vols = s["volumes"]
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
script = f"""
set -u
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
t0=$(date +%s.%N)
docker stop $cs >/dev/null
t1=$(date +%s.%N)
for v in {' '.join(vols)}; do
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
docker volume create "$v{SUFFIX}" >/dev/null
a=$(date +%s.%N)
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
b=$(date +%s.%N)
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
mkdir -p /root/bk; c=$(date +%s.%N)
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
done
t2=$(date +%s.%N)
docker start $cs >/dev/null
t3=$(date +%s.%N)
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
"""
out = w.guest(script, timeout=1800)
w.say(out)
t0 = time.time()
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
s["fcopy"] = out
s["fcopy_at"] = ts()
save(app, s)
def stage_break(app):
s = load(app)
frm, to, p0, p1 = EDGE[app]
s["break_commit"] = break_edge(app, frm, to, p0, p1)
w.say("badge caught up after", w.sync_rescan(app, to), "s")
save(app, s)
def stage_update(app):
s = load(app)
r = w.press_update(app, poll=1.0)
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
s.setdefault("updates", []).append(r)
save(app, s)
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
docker restart felhom-controller >/dev/null; sleep 15"""
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
def lift_and_pinback(app):
frm, to, _, _ = EDGE[app]
t = time.time()
w.say(w.guest(LIFT.format(app=app)))
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
w.login()
return round(time.time() - t, 1)
def stage_undoF(app):
s = load(app)
w.say(f"=== undo by F: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
vols = s["volumes"]
script = f"""
t0=$(date +%s.%N)
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
for v in {' '.join(vols)}; do
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
done
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
"""
w.say(w.guest(script, timeout=1800))
t0 = time.time()
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
w.say("product start ->", c)
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_undoD(app):
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
s = load(app)
w.say(f"=== undo by D: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
if app == "vikunja":
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
return
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
time.sleep(8)
if app == "docmost":
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
marker = "-- PostgreSQL database dump complete"
else:
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
marker = "-- Dump completed"
script = f"""
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
t0=$(date +%s.%N)
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
"""
w.say(w.guest(script, timeout=900))
t0 = time.time()
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_cutoff(app):
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
s = load(app)
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
script = f"""
v={big}
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
rm -f /root/cut.sql
else echo "D: no safety dump exists for this app (R-641)"; fi
"""
w.say(f"=== cut-off copy: {app} (largest volume {big})")
out = w.guest(script, timeout=600)
w.say(out)
s["cutoff"] = out
save(app, s)
if __name__ == "__main__":
w.login()
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,194 @@
#!/usr/bin/env python3
"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes.
Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/<n>,
fetches the app page in both languages, and reads the seeds back through each app's own front door.
Seed C is written immediately before each Update, after every backup — so only the undo's own
last-second copy can bring it back.
"""
import json, re, sys, time, html as H
sys.path.insert(0, ".")
import walk as w
from spike import load, save, ts
import bakeoff as b
HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}
def seed_c(app, s):
return b.seed_b(app, s["sub"], s["seedA"])
def page_lines(app):
out = {}
for lang in ("hu", "en"):
h = H.unescape(w.page(f"/apps/{app}?lang={lang}"))
m = re.search(r'data-update-undone="true">([^<]*)<', h)
hold = re.search(r'data-held="true">([^<]*)<', h)
out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None}
return out
def press(app, poll=0.5, on_phase=None):
code, d = w.ctl("POST", f"/api/stacks/{app}/update")
w.say(f" Update -> {code} {str(d)[:140]}")
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < 1500:
try:
st = w.stack(app)
except Exception as e:
st = {}
ph = st.get("update_phase")
if ph != seen and ph is not None:
seen = ph
phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label")))
w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}")
if on_phase and on_phase(ph):
return phases, "interrupted"
if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2:
break
time.sleep(poll)
st = w.stack(app)
w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}")
return phases, st
def readback(app, s, with_c=True):
sub = s["sub"]
w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2)
A = b.FX[app].verify(w, sub, s["seedA"], w.say)
B = b.verify_b(app, sub, s["seedA"], s["seedB"])
C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None
return {"A": A, "B": B, "C": C}
def stage_undo(app):
s = load(app)
w.say(f"=== {app}: live undo by the product")
s["seedC"] = seed_c(app, s); s["seedC_at"] = ts()
w.say(" seed C written right before the Update:", bool(s["seedC"]))
before = b.db_state(app); w.say(" db before:", before)
phases, st = press(app)
obs = w.observables(app)
w.say(" observables:", json.dumps(obs))
rb = readback(app, s)
after = b.db_state(app)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})")
pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False))
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")},
"readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs}
save(app, s)
def set_box_language(lang):
"""POST /settings/language — the household's language switch (form + the dashboard's CSRF)."""
sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}",
"-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}",
f"{w.BASE}/settings/language"])
w.say(f" box language -> {lang}: http {r.stdout.strip()}")
def stage_powercut(app):
"""Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard
stop); boot it again and let the controller resume the undo."""
import subprocess
s = load(app)
w.say(f"=== {app}: power cut DURING the undo")
s["seedC"] = seed_c(app, s)
cut = {}
def on_phase(ph):
if ph == "undoing":
t = time.time()
r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120)
cut["at"] = ts(); cut["rc"] = r.returncode
w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s")
return True
return False
phases, _ = press(app, poll=0.3, on_phase=on_phase)
if not cut:
w.say(" the cut never landed in `undoing` — recorded as a MISS"); return
r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180)
w.say(f" guest started again: rc={r.returncode}")
for i in range(60):
time.sleep(5)
try:
w.login(); st = w.stack(app)
if st:
break
except SystemExit:
continue
w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6"))
t0 = time.time()
while time.time() - t0 < 600:
st = w.stack(app)
if not st.get("updating") and st.get("update_phase") in ("undone", "failed"):
break
time.sleep(3)
w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}")
rb = readback(app, s)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}")
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb}
save(app, s)
def stage_cutoff(app):
"""Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of
the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it
back and HOLD with the new prefix."""
s = load(app)
w.say(f"=== {app}: a cut-off undo copy")
done = {}
def on_phase(ph):
if ph == "verifying" and not done:
out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c"
docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""")
done["out"] = out
w.say(" >>> finished-marker removed from one copy:\n" + out)
return False
phases, st = press(app, poll=0.5, on_phase=on_phase)
before = b.db_state(app)
pl = page_lines(app)
w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False))
set_box_language("en")
pl_en = page_lines(app)
w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False))
set_box_language("hu")
w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}"))
s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done}
save(app, s)
def stage_fixprobe_and_press(app):
"""After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work
and end the undone note."""
s = load(app)
frm, to, p0, p1 = b.EDGE[app]
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f)
w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"])
w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120)
w.sync_rescan(app, to)
time.sleep(20)
w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan")
w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)")
w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False))
phases, st = press(app)
rb = readback(app, s)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}")
pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False))
w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml"))
s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl}
save(app, s)
if __name__ == "__main__":
w.login()
{"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff,
"fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2])
@@ -0,0 +1,5 @@
import time, walk as w
time.sleep(20)
w.login()
c, d = w.ctl("POST", "/api/backup/run")
w.say(f"WHOLE-BOX BACKUP pressed 20 s after the deploy was accepted -> {c} {str(d)[:160]}")
@@ -0,0 +1,5 @@
import time, walk as w
time.sleep(5)
w.login()
c, d = w.ctl("POST", "/api/backup/run")
w.say(f"WHOLE-BOX BACKUP pressed 5 s after the deploy was accepted -> {c} {str(d)[:160]}")
@@ -0,0 +1,60 @@
#!/usr/bin/env python3
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
`controller.yaml.pre-cleanup` (NOT the older `.pre-28`, which a restore must never pick up).
"""
import re, sys, io
sys.path.insert(0, '.')
import walk as w
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
def creds():
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
if m:
return m.group(1), m.group(2)
raise SystemExit("no admin credential")
def to_drill():
u, t = creds()
print(w.guest(f"""
set -e
test -f {VOL}/controller.yaml.pre-cleanup || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-cleanup
python3 - <<'PY'
import re
p = "{VOL}/controller.yaml"
s = open(p).read()
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
if not re.search(r'^update:', s, re.M):
s += "update:\\n health_timeout: 90s\\n"
open(p, "w").write(s)
PY
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -A2 '^update:' {VOL}/controller.yaml
"""))
def restore():
print(w.guest(f"""
set -e
cp -p {VOL}/controller.yaml.pre-cleanup {VOL}/controller.yaml
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -c '^update:' {VOL}/controller.yaml || true
"""))
if __name__ == "__main__":
to_drill() if sys.argv[1] == "drill" else restore()
@@ -0,0 +1,160 @@
#!/usr/bin/env python3
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
steps the product would take, each one timed.
State between stages lives in state-<app>.json so each stage can be run, read, and only then
followed by the next (the hand undo needs a person looking at what the previous step left).
"""
import json, os, sys, time, re
sys.path.insert(0, ".")
import walk as w
import fixtures as fx
HERE = os.path.dirname(os.path.abspath(__file__))
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
def st_path(app):
return os.path.join(HERE, f"state-{os.environ.get('GUEST', '9202')}-{app}.json")
def load(app):
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
def save(app, s):
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
def ts():
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
def docmost_seed_b(sub, A):
jar = "/tmp/dm.jar"
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
name = "drillB" + os.urandom(3).hex()
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
return {"space": name} if code in ("200", "201") else None
def docmost_verify_b(sub, A, B):
jar = "/tmp/dm.jar"
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
if code not in ("200", "201"):
w.say(f" docmost B: cannot log in (http={code})"); return False
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
data="{}", method="POST")
ok = code in ("200", "201") and B["space"] in out
neg = "drillBnever" in out
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
return ok and not neg
def break_edge(app, frm, to, port_from, port_to):
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
health probe pointed at a port the app does not answer. Both in one drill commit."""
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read()
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
assert m, "probe port not found"
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
open(fy, "w").write(f)
h = w.drill_bump(app, frm, to)
return h
def pg_state(container, db, user):
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
def romm_seed_b(sub, A):
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
jar, tok = FX["romm"]._csrf(w, sub)
user = "drillb" + os.urandom(3).hex()
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
method="POST")
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
return {"user": user} if code in ("200", "201") else None
def romm_verify_b(sub, A, B):
jar, tok = FX["romm"]._csrf(w, sub)
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}")
ok = code == "200" and B["user"] in body
neg = "drillbnever" in body
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
return ok and not neg
def tree(paths):
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
return w.guest(cmd)
def my_state(container="romm-db", db="romm"):
# The root password is used INSIDE the container from its own env — it never leaves it.
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
echo -n "alembic=$({q("select version_num from alembic_version")}) "
echo -n "users=$({q("select count(*) from users")})"
""")
def vik_seed_b(sub, A):
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
app writes into its files volume), uploaded through the app's own attachment API."""
tok, why = FX["vikunja"]._token(w, sub, A)
title = "drillB-" + os.urandom(4).hex()
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
if code not in ("200", "201"):
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
pid = json.loads(body)["id"]
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
tid = json.loads(body)["id"] if code in ("200", "201") else None
content = "drill attachment " + os.urandom(8).hex()
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
"-F", f"files=@{fn}", method="PUT")
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
return {"title": title, "pid": pid, "tid": tid, "att": content}
def vik_verify_b(sub, A, B):
tok, why = FX["vikunja"]._token(w, sub, A)
if not tok:
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
proj = code == "200" and B["title"] in body
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
att_ok = False
try:
atts = json.loads(body)
if atts:
aid = atts[0]["id"]
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
att_ok = c3 == "200" and B["att"] in b3
except Exception as e:
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
return {"project": proj, "attachment": att_ok}
@@ -0,0 +1,494 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/cleanup-2026-09-23"
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = {"9202": "https://192.168.0.114", "9201": "https://192.168.0.138"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"cat > /tmp/w{GUEST}.sh; pct push {GUEST} /tmp/w{GUEST}.sh /tmp/w.sh >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash /tmp/w.sh; rm -f /tmp/w{GUEST}.sh"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
@@ -0,0 +1,26 @@
"""Deploy one app (or none) and record every change of deployed/deploying/state with a timestamp."""
import sys, time, json
import walk as w
app = sys.argv[1]; cap = int(sys.argv[2]) if len(sys.argv) > 2 else 900
w.login()
if "--deploy" in sys.argv:
vals = w.deploy_values(app, app)
c, d = w.ctl("POST", f"/api/stacks/{app}/deploy", {"values": vals} if vals else {})
w.say(f"deploy {app} -> {c} {str(d)[:160]}")
last = None; t0 = time.time()
while time.time() - t0 < cap:
st = w.stack(app)
cur = (st.get("deployed"), st.get("deploying"), st.get("state"), st.get("deploy_error"))
if cur != last:
w.say(f"+{time.time()-t0:6.1f}s {app}: deployed={cur[0]} deploying={cur[1]} state={cur[2]} err={str(cur[3])[:200]!r}")
last = cur
if "--until-settled" in sys.argv and cur[1] is False and time.time() - t0 > 30 and cur[2] in ("running", "not_deployed", "error", "stopped", "exited"):
# settled: keep watching 60 s more for late flips
end = time.time() + 60
while time.time() < end:
st = w.stack(app); c2 = (st.get("deployed"), st.get("deploying"), st.get("state"), st.get("deploy_error"))
if c2 != last: w.say(f"+{time.time()-t0:6.1f}s {app}: deployed={c2[0]} deploying={c2[1]} state={c2[2]} err={str(c2[3])[:200]!r}"); last = c2
time.sleep(3)
break
time.sleep(3)
w.say("final:", json.dumps({k: w.stack(app).get(k) for k in ("deployed","deploying","state","deploy_error","containers")}, default=str)[:600])
+10
View File
@@ -348,3 +348,13 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-620** | **A disabled notifier dropped every event with no local trace (P3).** Closed in controller **v0.264.0**: one WARN per event type per process naming the dropped event, then DEBUG. Live on 9202: six types each WARNed once; `app_start_failed` WARN at 12:02:30Z then DEBUG at 12:03:00Z. | **CLOSED 2026-09-23 — PROVEN-LIVE (`24-9202-r620-warn-then-debug.txt`)** | as above |
| **R-646** | **An app pinned before v0.263.2 had no record of its own `.felhom.yml` (P3).** Closed in controller **v0.264.0**: a startup pass records `applied-meta/` for every deployed, pinned app CURRENT with the catalog; a behind app is skipped by name (its file is gone — not backfillable by design). Live: 9202 recorded 3, 9201 recorded 10, skipped 0. | **CLOSED 2026-09-23 — PROVEN-LIVE (`28-9202-controller-log.txt`, `30-9201-before-and-upgrade.txt`)** | as above |
## 2026-09-23 (evening) — clean-up: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0)
| Row | What | Closed | Full text |
|---|---|---|---|
| **R-634** | **An app could RUN while the controller recorded it as not deployed (P1).** The unremovable half closed in v0.262.0; **the mechanism** is now diagnosed and fixed in **v0.265.0** (`0054d4b`): the whole-box backup's app list read the in-memory `Deployed` flag, true from the moment a deploy is accepted, so the volume leg stopped a DEPLOYING app (`compose down`), dumped half-made volumes and ran a second `compose up -d` beside the deploy's own — both failed and the deploy recorded „not deployed" (reproduced on demand on 9202, `audits/cleanup-2026-09-23/12-*`, `13-*`). Fix: deploying apps leave every backup list; the volume leg re-asks before the stop; `StopStack`/`StartStack` refuse a deploying stack for every caller. `sparkyfitness` did not reproduce alone. The deploy's own-failure question → **R-649**. | **CLOSED 2026-09-23 — PROVEN-LIVE (`32-*`, `33-*`: backup across a live deploy stopped three other apps and never outline; deploy `deployed`)** | `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
| **R-625** | **A held app invited an update its button refused (P2).** v0.265.0: badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed" (`tag-error`, title = the hold's first sentence per reader), no Update button, 409 `held` unchanged. | **CLOSED 2026-09-23 — PROVEN-LIVE (`44-*`, all four box/reader language pairs)** | as above |
| **R-636** | **Six hours of OOM kills sent the same single warning as one hiccup (P2).** v0.265.0 + hub v0.121.0: the kernel `oom_kill` counter (the `OOMKilled` flag is sticky and cannot count); ≥ 20 kills in 30 min of one container run → ONE `app_oom_storm` (error, operator-only, per-app cooldown). | **CLOSED 2026-09-23 — PROVEN-LIVE on the controller (`45-*`, `46-*`: RomM at 320M, storm at 21 kills, still one at 49); hub side by unit tests** | as above |
| **R-647** | **Three leftovers of the update mail (P3).** v0.265.0: a held update's error is the key `update.error.held`, rendered per reader on both pages and the API; `copy_holds` travels as its key; the two log wordings fixed. | **CLOSED 2026-09-23 — PROVEN-LIVE for (1) (`43-*`, `44-*`); (2)(3) by red-proofed tests** | as above |
| **R-648** | **The drill's „Mentés most" was whole-box (P3).** No per-app backup endpoint exists; the harness (`audits/cleanup-2026-09-23/walk.py` `backup_now`) now presses nothing and the guarded update's own `backing-up` phase backs up the throwaway app alone. | **CLOSED 2026-09-23 — PROVEN-LIVE (`43-*`: phase `backing-up` for vikunja only)** | as above |
File diff suppressed because one or more lines are too long