The household is told: controller v0.264.0 + hub v0.120.0 proven live, floor 0.264.0
gates / gates (push) Successful in 28s

09 §6.4 parts 2-3 SHIPPED. R-606, R-620, R-646 closed; R-647 (three
leftovers) and R-648 (whole-box backup press in the harness) opened.
Open rows 335 -> 334. Evidence: audits/undo-fleet-2026-09-23/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 14:39:31 +02:00
parent 7caa24ffe1
commit 21f17ed32b
38 changed files with 3269 additions and 85 deletions
+22
View File
@@ -14,6 +14,28 @@
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## 2026-09-23 (evening) — the household is told: controller v0.264.0 + hub v0.120.0 (`09` §6.4 parts 2–3); floor 0.264.0
**Events.** `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, sent by the
controller through `Manager.SetUpdateEventSink` (wired in main.go). Details `{app, stack_name, from, to, at,
copy_tier, copy_date, copy_holds}`. The held mail's message line is the hold sentence rendered in the
household's language (`pushEventBoth`); the operator gets the Hungarian wire text.
**Hub.** Allow-listed, NOT operator-only, in `defaultSeedEvents`; `mail.event.*` hu+en name the app via the
new register `appNamedMailEvents` (reads `details.stack_name`). **Existing households:** a one-time add-only
migration `Store.SeedEventTypesOnce` guarded by `hub_settings.seed_app_update_events_v1` — it changed all
three rows (`demo-felhom`, `demo-hp`, `_resend-rotation-test`; before/after in
`audits/undo-fleet-2026-09-23/10-*`, `11-*`). **Cooldown:** per app on both legs — operator via R-389's
`perAppCooldownEvents` (widened on purpose), household via a NEW sibling `perAppCustomerCooldownEvents`
(the customer key was `customer:type`). `app_start_failed`'s household grain is unchanged (pinned).
**Why the box also seeds:** the controller pushes its own `EnabledEvents` to the hub on every notifications
page save, so a hub-only migration would be undone by the next save; the box seeds once
(`app_update_events_seeded`) and the page has the two toggles.
**R-606:** update sentences are keys + args, rendered per reader; leftovers R-647. **R-620:** a disabled
notifier WARNs once per type. **R-646:** startup `BackfillAppliedMeta` for apps current with the catalog.
**Voice:** the undone line ends „nincs teendőd" (the product's „te"), not the brief's „teendője nincs".
**Harness lesson (R-648):** the drill's „Mentés most" is whole-box and restarts standing apps.
Evidence: `audits/undo-fleet-2026-09-23/README.md`.
## 2026-09-23 (afternoon) — the undo is built: controller v0.263.2 (decisions 15, 19, 20)
**Operator rulings 19 and 20** recorded in `09` §3: the copy method is chosen by a bake-off; on update
+97 -62
View File
@@ -1,11 +1,11 @@
# REPORT — the undo: bake-off, build (controller v0.263.0 → v0.263.2), live proof
# REPORT — the undo reaches the fleet, and the household is TOLD (controller v0.264.0, hub v0.120.0)
2026-09-23 (afternoon). Repos touched: **felhom-controller** (v0.263.0, v0.263.1, v0.263.2),
**felhom.eu** (docs, register, capability map, STATUS, CONTEXT, evidence), **admin/app-catalog-drill**
(drill commits, reset to live `main` at the end). **app-catalog-felhom.eu, felhom-agent, hub: untouched.**
Architecture read first and named: `documentation/architecture/09-update-architecture.md` (§3 decisions
11–20, §4, §6.1, §6.1a, §6.4). Baselines verified live: controller `b9deec19077b`, agent
`d9864a94bf62`, felhom.eu `4c92beab8fdb`, catalog `cfcfe5278428` — all as the brief said.
2026-09-23 (evening). Repos touched: **felhom-controller** (v0.264.0, `bc27894`), **felhom.eu** (hub
v0.120.0 `d3b2848`, manifest `7caa24f`, docs, register, evidence), **admin/app-catalog-drill** (drill
commits, reset to live `main` at the end). **felhom-agent, app-catalog-felhom.eu: untouched.**
Architecture read and named: `documentation/architecture/09-update-architecture.md` (§3 decision 15,
§6.4). Evidence: `documentation/audits/undo-fleet-2026-09-23/README.md`. **Method: endpoint-level**
(the endpoints the UI invokes; mails read from the catch-all mailbox; no browser).
---
@@ -13,69 +13,104 @@ Architecture read first and named: `documentation/architecture/09-update-archite
| item | state |
|---|---|
| rulings 19–20 | **recorded** (`09` §3), commit `5a349d9` |
| Part 1 — bake-off | **done in ~40 min** of the 2 h cap; folder copy chosen |
| Part 2 — build | **done — but in THREE releases, not one.** v0.263.0 failed its first two live proofs honestly (HELD, data put back). Two defects only the live box could show; each fixed + red-proofed + released: **v0.263.1** (the undo's probe never ran on an app the current probe held `unhealthy`), **v0.263.2** (the "old" `.felhom.yml` was already the new one — it flows in on every catalog sync; now recorded at pin time). 0.263.0/0.263.1 ran only on 9202 and were removed from it. |
| the household mail + operator event of decision 15 | **not built** — `09` §6.4 part 2 (R-606). The page line and the hold sentence are built. |
| the floor | **not raised** — the operator's question (STATUS) |
| immich/nextcloud rate test | **nextcloud deployed and was used** (185 MB MariaDB volume, 300 files through WebDAV); immich not tried |
| cut-off copy on vikunja in the bake-off | **could not be cut** (2.9 MB finishes before a kill lands); proven on docmost and romm instead; in the LIVE proof the marker was removed from a vikunja copy mid-update |
| Part 0 — floor 0.263.2 | **done**, read back (`00-floor-0.263.2.txt`) |
| Part 1 — R-646 startup pass | **done**; live: 3 recorded on 9202, 10 on 9201, 0 skipped |
| Part 2 — two events, hub v0.120.0 first | **done**; hub deployed 11:49Z, before any box ran v0.264.0 |
| Part 3 — R-606 | **done, with leftovers → R-647** (a reader in the other language than the box sees a HELD sentence in the box's language; the English mail's raw `Note:` line keeps the Hungarian `copy_holds` phrase) |
| Part 4 — R-620 | **done** |
| Part 5 — live proof | **done**; all proofs passed → **floor raised to 0.264.0** |
| "standing apps untouched" | **NOT fully met.** The harness's „Mentés most" is a whole-box backup: on 9201 it stopped and restarted 9 of the 10 standing apps for their volume dumps, twice (a few seconds each). All healthy afterwards, files byte-identical. **R-648** filed. |
| the brief's Hungarian undone line | **changed:** it ends „semmi nem veszett el, nincs teendőd." — the brief's „teendője nincs" is the formal „ön" voice; the product speaks „te" everywhere and the voice gate ratchets the formal form (18). |
| the customer cooldown | **changed shape:** a NEW customer-side register (`perAppCustomerCooldownEvents`) instead of putting the customer leg of R-389's register to work — that would also have changed `app_start_failed`'s household grain, which nobody asked for (pinned by a test). |
| a box-side seed + two page toggles | **added, not in the brief:** the box pushes its own `EnabledEvents` to the hub on every notifications-page save, so the hub migration alone would be undone by the next save. |
**Claims in the brief that turned out wrong or unmeasured:**
1. *"All three apps keep their data only in named volumes"* — true from compose AND on disk for these
three; but romm also has two BIND folders (roms, resources), empty throughout — never copied, by rule.
2. *"A volume copy with the containers stopped is consistent for both engines"* — **measured true**:
docmost (PostgreSQL) and romm (MariaDB) came back with ledgers equal, three times each.
3. *"Immich or Nextcloud deploys on 9202"* — Nextcloud did; Immich was not tried.
4. **Not in the brief, and the most important:** *"the old `.felhom.yml`"* does not exist at update
time. It is replaced on every catalog sync (`09` §5.4). The build now records it when a version is
pinned. R-646 for apps pinned before that.
**Claims in the brief, checked:**
1. *"No app-update event exists today"* — **true for apps.** The only near-misses: `controller_updated` is
the controller's OWN self-update, and `update_available` exists only as a debug-page test option — never
emitted, and not in the hub's allow-list (it would 400).
2. *"`enabled_events` is the household's only mail gate"* — **wrong.** Four more gates stand in the way:
the `operatorOnlyEvents` register, a blocked customer, a missing e-mail address, and the cooldown —
keyed `customer:type`, which would have merged two apps' mails into one. And the box OVERWRITES
the hub's list on each page save (above).
3. *"9201 can follow the drill catalog without disturbing its standing apps"* — **true for the catalog,
false for the harness.** Measured: the 10 standing apps' compose + `.felhom.yml` were byte-identical
before, during and after (the drill differed from live only in the throwaway apps). The disturbance
came from the harness's whole-box backup press (R-648), not from the catalog.
## 2. The bake-off (`audits/undo-bakeoff-2026-09-23/README.md`)
## 2. Floor raises, read back from the hub
Both methods passed every row on docmost, romm, vikunja (seeds A and B back, ledgers equal, cut-off
detected before anything moves, ≤ 5.3 s extra downtime). **Folder copy chosen (decision 19):** an app
with no database server gets no dump, so dump-and-load would need the folder copy anyway. Rate: 185 MB
in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s (warm cache); 5 GB ≈ 12 s here, 25–50 s on a cold or spinning
disk. Found on the way: the undo must never read the recovery unit (R-645), and killing `docker run`
does not stop the copy container.
## 3. Red-proofs (each seen failing, then restored)
| # | mutation | test that failed |
| raise | read back | fleet |
|---|---|---|
| 1 | the undo call removed (v0.262.1's shape) | `TestUndo_FailedUpdateIsPutBackWithItsData` — `held=true phase="failed"` |
| 2 | copy validation removed | `TestUndo_CutOffCopyIsRefusedBeforeAnythingMoves` — state `half` |
| 3 | old-probe rule removed | `TestUndo_UsesTheOldProbe` — `not_started` |
| 3b | the old probe read from the stack dir (v0.263.1's shape) | `TestUndo_UsesTheOldProbe` |
| 4 | the `undoing` recovery arm removed | `TestUndo_PowerCutDuringTheUndoResumesIt` — `resumed=[]` |
| 5a | `last_update_undone` not recorded | `TestUndo_FailedUpdateIsPutBackWithItsData` — `got <nil>` |
| 5b | the page line removed from the handler | `TestUndo_PageSaysTheBoxPutTheAppBack` (hu and en) |
| 6 | the hold prefix ignored | `TestUndo_HoldSentenceSaysTheUndoWasTriedAndTheDataState` |
| 7 | R-642: "completed" back | `TestR642_StartIsNeverReportedCompleted` |
| 8 | a successful update keeps the undone note | `TestUndo_ASuccessfulUpdateEndsTheUndoneNote` |
| 9 | the undo probe switched off on an `unhealthy` app (v0.263.0's shape) | `TestUndo_OldProbeRunsOnAnAppTheCurrentProbeMarkedUnhealthy` — `not healthy within 1s (last: state unhealthy)`, the live message verbatim |
| 10 | the pin advance records no probe | `TestUndo_PinningRecordsThatVersionsProbe` |
| 0.263.2 / MinAgent 0.131.0 | `min_controller_version=0.263.2`, `min_agent=0.131.0` | `00-floor-0.263.2.txt` |
| **0.264.0 / MinAgent 0.131.0** (14:35:08) | `min_controller_version=0.264.0`, `min_agent=0.131.0`; hub log `managed floor SERVED for demo-felhom: floor 0.264.0 … from declared` | demo-felhom 0.263.2 → **0.264.0 in 20 s**; demo-hp 9201 already 0.264.0 (`50-floor-0.264.0.txt`) |
`go build ./... && go vet ./... && go test ./...` rc=0 before each of the three commits; controller
gates rc=0 (the `docker -v` gate needed four named-volume mounts allowlisted with their why).
## 3. The hub migration — rows changed (one-time, add-only)
## 4. Live proof (`audits/undo-live-2026-09-23/README.md`) — endpoint-level, both languages
| customer | before | after |
|---|---|---|
| `demo-felhom` | 10 types | the same 10 + `app_update_undone`, `app_update_held` |
| `demo-hp` | `null` (nothing on) | `["app_update_undone","app_update_held"]` |
| `_resend-rotation-test` | `[]` | `["app_update_undone","app_update_held"]` |
Three apps undone by the product with seeds A, B, C back and ledgers equal (30–52 s of undo); page
line hu/en; cut-off copy → HOLD `untouched` (prefix in the box's language, hu and en quoted); power cut
during `undoing` → resumed and undone; a person's press after an undo → `done`, note cleared; removal
deleted kept copies; R-642 Start answer.
Guard row `seed_app_update_events_v1 = 2026-09-23T11:49:03Z`. E-mail column never selected.
## 5. Rows
## 4. The mails (9201, customer demo-hp) and their `notification_log` rows
**Closed (4, moved to `CLOSED-ITEMS.md`):** R-637, R-639, R-641, R-642. **Narrowed:** R-638, R-640 (to
the restore paths). **Ruled:** R-643 (decision 20). **Noted:** R-645. **Opened:** R-645 (earlier this
session, filed with the bake-off) and R-646 (apps pinned before v0.263.2). **Open rows 337 → 335.**
| row | time (UTC) | type | leg | subject as received |
|---|---|---|---|---|
| 937 / 938 | 12:15 | undone | op / household | op `⚠️ demo-hp: app_update_undone` · hh **`Figyelmeztetés: vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut`** |
| 939 / 940 | 12:16 | held | op / household | op `🔴 demo-hp: app_update_held` · hh **`Hiba: vikunja: az alkalmazás leállítva, visszaállítás szükséges`** |
| 942 / 943 | 12:29 | undone | op / household | hh **`Warning: glance: the update did not work; the app runs on its previous version`** |
| 944 / 946 | 12:31 | held | op / household | hh **`Error: glance: the app is stopped and needs a restore`** |
## 6. Teardown — three layers
All 8 `sent`. The box language was switched to English once, between the two rounds (the hub had two
reports in English before the second round). Household bodies carry the box's own sentence (the undone
line / the hold sentence) in the household's language; the operator's carry the Hungarian wire text.
Verbatim bodies: `41-mails-as-received.md`. **Per-app cooldown, live:** row 943 went out 14 minutes after
row 938, inside the 6-hour customer cooldown.
Machine (9202): the three apps and every copy removed through the product; test images removed by
name; `controller.yaml` restored and read back; catalog cache on live `cfcfe52`; **9202 stays on
controller 0.263.2** (self-update off, fleet floor untouched). Host: nothing provisioned; 9202 stopped
and started once for the power cut. Hub: untouched. Drill repo: reset to live `main`.
## 5. The pages in both languages (9202, then 9201)
- undone: hu *„A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan
visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en *"The update of vikunja
at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically —
nothing was lost."*
- held, box hu: `?lang=hu` → the Hungarian hold sentence; `?lang=en` → *"The update did not succeed, and the
automatic undo did not either. … own drive, 2026-09-23 13:58 — this copy holds the settings, the database
and the data volumes."* — the whole sentence English now (before v0.264.0 only the prefix was).
- `GET /api/stacks/vikunja?lang=en`: label *"The update did not succeed"*, error and hold in English.
- **Leftover (R-647):** box en + `?lang=hu` showed the English hold.
## 6. Red-proofs (each seen failing, then restored)
**Controller (9):** backfill stores nothing → `TestR646_…` *"the current app must have its record"*;
`dropped()` not called → `TestR620_…` *"want exactly TWO WARN lines, got 0"*; undone event not emitted →
*"exactly ONE app_update_undone event, got []"*; held event not emitted → *"…got []"*; page
`updateErrorText` not localised → `TestR606_UpdateSentencesFollowTheReader`; hold sentence ignores the
language → `undo_hold_test.go:62`; `UpdateErrorIn` returns the Hungarian → *"English = „A frissítés nem
indult el…""*; box seed not run → *"got [backup_failed offbox_enlarge_blocked]"*; sink not wired in main →
*"SetUpdateEventSink is never called"* (first mutation did not compile; retried with a compiling one).
**Hub (5):** seed call does nothing → *"c1: enabled_events = [backup_failed], want […]"*; customer per-app
suffix removed → *"want 3 household mails …, got 2"*; allow-list entry removed → *"must be in
allowedEventTypes"*; app-named headline removed → subject *"%s: a frissítés…"*; seed default removed →
*"defaultSeedEvents must include app_update_undone"* (first attempt matched no test; retried with a new
test that names both types). Every `-run` was checked for `=== RUN`.
Gates: controller `go test ./...` rc=0, `controller_gates.py` rc=0; hub `go build/vet/test` rc=0,
`repo_gates.py` rc=0. The R-389 fence test was widened on purpose, with its reason; the settings-page
parity fixture was regenerated (measured diff: exactly the two new toggles).
## 7. Rows
**Closed (3, moved to `CLOSED-ITEMS.md`):** R-606, R-620, R-646. **Opened (2):** R-647 (the three
leftovers), R-648 (the whole-box backup press). **Open rows 335 → 334.** `09` §6.4 parts 2 and 3 marked
SHIPPED; the capability map's undo row now carries the mail; STATUS rewritten.
## 8. Teardown — three layers
- **machine:** 9202 and 9201 — throwaway apps removed through the product (0 volumes, 0 undo copies),
test images removed, `controller.yaml` identical to the saved copy, catalog on live `cfcfe52`, 9201's
language back to `hu`, standing apps deployed and healthy. Both on 0.264.0 (= the floor).
- **host:** nothing provisioned on demo-hp.
- **hub:** v0.120.0 (planned); floor 0.264.0. Scratchpad password files deleted; hub DB copies deleted.
- **drill repo:** reset to live `main`, read back from the remote.
+14 -17
View File
@@ -1,26 +1,23 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-23 (afternoon) — a failed update now puts the app back by itself. Built, and proven on the scratch machine. One question for you: whether the fleet gets it.**
**Updated 2026-09-23 (evening) — the undo is on every demo machine now, and the household gets a mail when an update fails. The floor is raised.**
**Decisions I took on my own: none.** Your two afternoon rulings are written down: the bake-off picks the copy method, and on update nights the full-system backup waits for the updates.
**Decisions I took on my own: none.** One small deviation from your brief: the Hungarian mail line ends "nincs teendőd", not "teendője nincs", because the product speaks to the household as "te" everywhere.
**The bake-off.** Both ways of keeping the last-second copy passed every test on three apps. Copying the app's data folders won, because one of the apps has no database server and so gets no database copy at all. The extra downtime was 1 to 5 seconds.
**What the household gets now.** When an update fails and the machine puts the app back, the household gets one mail: "docmost: the update did not work; the app runs on its previous version". When the undo fails too, they get one mail that says the app is stopped and needs a restore, and which backup to use. The mail is in the household's own language. The operator gets the same two events. Two apps on the same night give two mails, not one.
**What the machine does now.** When an update fails its health check, the machine puts back the previous version and the data exactly as it was seconds before the update. It stops the app only if that undo fails too, and then it says so. It never touches the household's own folders (photos, documents).
**Proven on real machines, through the same buttons the page uses:**
- On the demo-hp machine: 4 mails to the household and 4 to the operator, first in Hungarian, then in English after one language switch. All arrived. The hub log shows all 8 as sent.
- On the scratch machine: the update messages on both app pages show in Hungarian and English. A machine with no hub now writes one warning line per kind of lost message.
- At start, the new version stores the health-check file of each up-to-date app, so its first undo asks the right question. It did this for 3 apps on one machine and 10 on the other.
- The floor is 0.264.0. The N100 demo machine updated itself within 20 seconds.
**Proven on the scratch machine, through the same buttons the page uses:**
- Three apps undone by the machine. Data written before the backup, after it, and seconds before the update all came back. It took 30 to 52 seconds.
- The app page shows one line in Hungarian or English: the update failed, the machine put the app back, nothing was lost.
- A damaged copy is caught before anything is poured back, and the app is held with an honest sentence.
- A power cut in the middle of the undo: after restart, the machine finished the undo.
- A person pressed Update again after the catalogue was fixed, and it worked.
**What was not clean.**
- My test "back up now" button backs up the whole machine. On demo-hp it stopped and restarted 9 of the 10 standing apps for a few seconds, twice. All came back healthy, with unchanged files. So "standing apps untouched" is not fully true. I filed a row to fix the test method.
- Three small leftovers, filed as one row: a reader who forces the other language sees a stopped app's sentence in the machine's language; the raw detail line of an English mail has one Hungarian phrase; and two of my log lines use unclear words.
**What went wrong on my side.** My first build failed its own live test twice. Both times the app stayed stopped with an honest sentence, and the data was safe. The first fault: the machine never asked the old version the right health question. The second: it kept the wrong copy of that question. My unit tests had passed both times. I fixed both the same afternoon and proved the fixes. The released version is the third build.
**Rows.** Three closed, two new. The list went from 335 to 334.
**Rows.** Four closed, two narrowed, one new. The list went from 337 to 335.
**What needs you: nothing.** If you do nothing, the fleet keeps the new version. The leftovers are small and wait for a free evening.
**What needs you — one question.** Should the demo machines and the fleet get this version now?
- **Yes (my pick):** I raise the floor, and both demo machines update themselves within a minute. A failed update then puts the app back instead of stopping it.
- **Not yet:** nothing changes. A failed update keeps stopping the app until someone restores it.
**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. The scratch machine runs the new version and is back on the real catalogue.**
**Nothing on your own machine, Peti's machine or the off-site box was touched, except the hub update. The test apps are gone from both demo machines. Both are back on the real catalogue.**
File diff suppressed because one or more lines are too long
@@ -1043,8 +1043,8 @@ what the part can do to a household's data if it is wrong, not how likely that i
| # | part | rulings / rows | cost | depends on | risk to customer data |
|---|---|---|---|---|---|
| **1** | **SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23** (`audits/undo-live-2026-09-23/`). **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none |
| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
| **2** | **SHIPPED — controller v0.264.0 + hub v0.120.0, proven live on 9202 and 9201 2026-09-23** (`audits/undo-fleet-2026-09-23/`). **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held: events `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, on by default, mailed in the household's language with the app named in the subject; per-app cooldown on both legs. Leftovers: R-647. | R-606, 15 | **1** | — | none |
| **3** | **SHIPPED — controller v0.264.0, proven live on 9202 2026-09-23.** **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
| **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
| **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
@@ -0,0 +1,16 @@
=== FLOOR -> 0.263.2, MinAgent 0.131.0 declared (2026-09-23 13:07:31)
POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set
--- read back from the hub:
"min_controller_version" value="0.263.2"
"min_agent" value="0.131.0"
"min_agent" value="0.131.0"
--- watching the two demo boxes (up to 3 min)
+10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.262.1 Up 15 hours (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.262.1 Up 15 hours (healthy)
+20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 5 seconds (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 8 seconds (healthy)
--- hub log:
2026/09/23 13:07:34 [INFO] managed floor SERVED for demo-felhom: floor 0.263.2, agent requirement "0.131.0" from declared (golden 0.258.0)
2026/09/23 13:07:34 [INFO] managed floor SERVED for demo-hp: floor 0.263.2, agent requirement "0.131.0" from declared (golden 0.258.0)
--- hosts page:
demo-felhom-8363b5
demo-hp-bb76ea
drill-r50-0a4f9a
@@ -0,0 +1,5 @@
# customer_notifications BEFORE hub v0.120.0 (read-only, 2026-09-23T11:47:29Z); e-mail column deliberately not selected
_resend-rotation-test|[]
demo-felhom|["backup_failed","db_dump_failed","offbox_enlarge_blocked","storage_disconnected","node_down","health_critical","disk_warning","disk_critical","expected_backup_missed","expected_dbdump_missed"]
demo-hp|null
# guard row:
@@ -0,0 +1,17 @@
# hub v0.120.0 rollout 2026-09-23T11:49:50Z pod hub-b6ff59c47-n9vp8
gitea.dooplex.hu/admin/felhom-hub:0.120.0
# startup log lines (migration):
2026/09/23 13:49:03 [INFO] [store] app-update event types added to 3 household(s)' notification prefs (one-time, add-only): [demo-felhom _resend-rotation-test demo-hp]
2026/09/23 13:49:03 [INFO] Default controller-version floor: 0.120.0
2026/09/23 13:49:04 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000
2026/09/23 13:49:04 [INFO] Registry version checker started (every 6h)
2026/09/23 13:49:04 [DEBUG] Registry version check: latest = 0.263.2
2026/09/23 13:49:35 [INFO] Listening on :8080
# customer_notifications AFTER (read-only; e-mail column not selected)
# guard row:
# customer_notifications AFTER (copy of hub.db+wal+shm taken 2026-09-23T11:50:35Z; e-mail column not selected)
_resend-rotation-test|["app_update_undone","app_update_held"]
demo-felhom|["backup_failed","db_dump_failed","offbox_enlarge_blocked","storage_disconnected","node_down","health_critical","disk_warning","disk_critical","expected_backup_missed","expected_dbdump_missed","app_update_undone","app_update_held"]
demo-hp|["app_update_undone","app_update_held"]
# guard row:
seed_app_update_events_v1|2026-09-23T11:49:03Z
@@ -0,0 +1,17 @@
13:54:56 === 9202 setup: drill vikunja at the FROM pin, 9202 onto the drill catalog
13:54:57 [5] drill commit 847d544b8064: vikunja vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.3.0 (push rc=0)
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
13:57:23 [sync] the badge NEVER caught up to vikunja/vikunja:2.3.0 in 127.6s — catalog_images = None
13:57:23 badge caught up after None s
13:57:25 catalog cache: 847d544 DRILL vikunja: vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.3.0
https://<cred>@gitea.dooplex.hu/admin/app-catalog-drill.git
@@ -0,0 +1,16 @@
13:57:31 === prep vikunja: deploy at the drill FROM pin
13:57:31 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
13:57:36 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
13:57:36 deployed: True
13:57:41 vikunja: register http=200
13:57:41 vikunja: create project http=201
13:57:42 vikunja: readback of the seeded project http=200 ok=True
13:57:42 C1 A: True
13:57:42 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
13:58:07 [4] backup idle; last=None
13:58:07 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill73bafb","created":"2026-09
13:58:07 vikunja B: project readback=True attachment content readback=True
13:58:07 B reads back: True
13:58:10 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
port: 3456
@@ -0,0 +1,34 @@
13:58:21 [5] drill commit 073bc953e182: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
13:58:25 badge caught up after 4.5 s
13:58:25 === vikunja: live undo by the product
13:58:25 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drill73bafb","created":"2026-09
13:58:25 seed C written right before the Update: True
13:58:29 db before: tables=36 ledger=117 newest=SCHEMA_INIT
13:58:29 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
13:58:29 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
13:58:30 + 1.1s phase=pulling label=Új verzió letöltése…
13:58:33 + 4.2s phase=copying label=Az adatok másolása a frissítés előtt…
13:58:35 + 5.8s phase=starting label=Indítás az új verzióval…
13:58:35 + 6.3s phase=verifying label=Működés ellenőrzése…
14:00:06 + 96.7s phase=undoing label=Visszaállítás az előző változatra…
14:00:08 + 99.3s phase=undone label=Visszaállítva az előző változatra
14:00:08 END phase=undone err=None hold=None
14:00:11 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]}
14:00:11 vikunja: readback of the seeded project http=200 ok=True
14:00:12 vikunja B: project readback=True attachment content readback=True
14:00:12 vikunja B: project readback=True attachment content readback=True
14:00:14 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT)
14:00:14 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
14:00:16 leftover copies: 0
14:00:16 GET /api/stacks/vikunja: phase='undone' label='Visszaállítva az előző változatra' err=None
14:00:16 GET /api/stacks/vikunja?lang=en: label='Put back to the previous version' err=None
14:00:16 --- controller log (positive observables):
14:00:18 2026/09/23 11:52:40 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only
2026/09/23 11:55:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only
2026/09/23 11:57:31 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deploy_started (severity info) — further app_deploy_started events are logged at DEBUG only
2026/09/23 11:57:34 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deployed (severity info) — further app_deployed events are logged at DEBUG only
2026/09/23 12:00:00 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event health_change (severity warn) — further health_change events are logged at DEBUG only
2026/09/23 12:00:08 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true)
2026/09/23 12:00:08 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only
2026/09/23 12:00:08 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 3s — the previous version is running on the data from before the update (the app's health check passed)
@@ -0,0 +1,34 @@
14:00:30 === vikunja: a cut-off undo copy
14:00:31 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
14:00:31 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
14:00:32 + 1.1s phase=pulling label=Új verzió letöltése…
14:00:33 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt…
14:00:34 + 3.7s phase=verifying label=Működés ellenőrzése…
14:00:37 >>> finished-marker removed from one copy:
copy: vikunja_vikunja_data.pre-update-20260923T120033Z
total 12
drwxr-xr-x 3 root root 4096 Sep 23 12:00 .
drwxr-xr-x 1 root root 4096 Sep 23 12:00 ..
drwxr-xr-x 2 root root 4096 Sep 23 12:00 data
-rw-r--r-- 1 root root 0 Sep 23 12:00 felhom-undo-complete
data
14:02:05 + 94.7s phase=undoing label=Visszaállítás az előző változatra…
14:02:06 + 95.2s phase=failed label=A frissítés nem sikerült
14:02:06 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
14:02:08 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."}}
14:02:08 box language -> en: http 302
14:02:08 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."}}
14:02:08 box language -> hu: http 302
14:02:10 copies kept: vikunja_vikunja_data.pre-update-20260923T120033Z
vikunja_vikunja_db.pre-update-20260923T120033Z
14:02:10 GET /api/stacks/vikunja: phase='failed' label='A frissítés nem sikerült' err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
14:02:10 GET /api/stacks/vikunja?lang=en: label='The update did not succeed' err='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes.' hold='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes.'
14:02:10 --- controller log:
14:02:13 2026/09/23 12:00:08 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only
2026/09/23 12:02:05 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T120033Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T120033Z}]
2026/09/23 12:02:05 update_guard.go:553: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T11:58:06Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched")
2026/09/23 12:02:05 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true)
2026/09/23 12:02:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only
@@ -0,0 +1,10 @@
http 200 type list n 2000
--- every DROPPED line in the debug ring (WARN once per type, then DEBUG):
{"timestamp": "2026-09-23T11:57:31Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_deploy_started (severity info) — further app_deploy_started events are logged at DEBUG only", "source": "notifier.go:1163"}
{"timestamp": "2026-09-23T11:57:34Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_deployed (severity info) — further app_deployed events are logged at DEBUG only", "source": "notifier.go:1163"}
{"timestamp": "2026-09-23T12:00:00Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event health_change (severity warn) — further health_change events are logged at DEBUG only", "source": "notifier.go:1163"}
{"timestamp": "2026-09-23T12:00:08Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only", "source": "notifier.go:1163"}
{"timestamp": "2026-09-23T12:02:05Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only", "source": "notifier.go:1163"}
{"timestamp": "2026-09-23T12:02:30Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_start_failed (severity warning) — further app_start_failed events are logged at DEBUG only", "source": "notifier.go:1163"}
{"timestamp": "2026-09-23T12:03:00Z", "level": "DEBUG", "message": "notifier disabled: dropped event app_start_failed (severity warning)", "source": "notifier.go:1167"}
--- counts per (type, level): {('app_deploy_started', 'WARN'): 1, ('app_deployed', 'WARN'): 1, ('health_change', 'WARN'): 1, ('app_update_undone', 'WARN'): 1, ('app_update_held', 'WARN'): 1}
@@ -0,0 +1,57 @@
2026/09/23 11:52:35 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 3 [gokapi paperless-ngx privatebin]; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
2026/09/23 11:52:40 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only
2026/09/23 11:55:00 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
2026/09/23 11:55:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only
2026/09/23 11:57:31 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deploy_started (severity info) — further app_deploy_started events are logged at DEBUG only
2026/09/23 11:57:34 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deployed (severity info) — further app_deployed events are logged at DEBUG only
2026/09/23 11:58:29 update.go:488: [INFO] [stacks] update vikunja: accepted — guarded update started
2026/09/23 11:58:29 update.go:1145: [INFO] [stacks] update vikunja: phase checking
2026/09/23 11:58:29 update.go:665: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T11:58:06Z (0s old, limit 24h0m0s)
2026/09/23 11:58:29 update.go:1145: [INFO] [stacks] update vikunja: phase safety-dump
2026/09/23 11:58:29 update.go:697: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) []
2026/09/23 11:58:30 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.5 MiB
2026/09/23 11:58:30 update.go:1145: [INFO] [stacks] update vikunja: phase pinning
2026/09/23 11:58:30 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0)
2026/09/23 11:58:30 update.go:1145: [INFO] [stacks] update vikunja: phase pulling
2026/09/23 11:58:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying
2026/09/23 11:58:34 update.go:1145: [INFO] [stacks] update vikunja: phase copying
2026/09/23 11:58:34 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T115834Z in 471ms
2026/09/23 11:58:34 update.go:1145: [INFO] [stacks] update vikunja: phase copying
2026/09/23 11:58:35 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T115834Z in 431ms
2026/09/23 11:58:35 update.go:1145: [INFO] [stacks] update vikunja: phase starting
2026/09/23 11:58:35 update.go:1145: [INFO] [stacks] update vikunja: phase verifying
2026/09/23 12:00:00 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event health_change (severity warn) — further health_change events are logged at DEBUG only
2026/09/23 12:00:06 update.go:859: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy)
2026/09/23 12:00:06 update.go:851: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T120006Z/compose-logs.txt before stopping it (R-621)
2026/09/23 12:00:06 update.go:1145: [INFO] [stacks] update vikunja: phase undoing
2026/09/23 12:00:06 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy))
2026/09/23 12:00:08 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true)
2026/09/23 12:00:08 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only
2026/09/23 12:00:08 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 3s — the previous version is running on the data from before the update (the app's health check passed)
2026/09/23 12:00:31 update.go:488: [INFO] [stacks] update vikunja: accepted — guarded update started
2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase checking
2026/09/23 12:00:31 update.go:665: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T11:58:06Z (2m0s old, limit 24h0m0s)
2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase safety-dump
2026/09/23 12:00:31 update.go:697: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) []
2026/09/23 12:00:31 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB
2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase pinning
2026/09/23 12:00:31 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0)
2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase pulling
2026/09/23 12:00:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying
2026/09/23 12:00:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying
2026/09/23 12:00:33 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T120033Z in 476ms
2026/09/23 12:00:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying
2026/09/23 12:00:34 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T120033Z in 413ms
2026/09/23 12:00:34 update.go:1145: [INFO] [stacks] update vikunja: phase starting
2026/09/23 12:00:34 update.go:1145: [INFO] [stacks] update vikunja: phase verifying
2026/09/23 12:02:05 update.go:859: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy)
2026/09/23 12:02:05 update.go:851: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T120205Z/compose-logs.txt before stopping it (R-621)
2026/09/23 12:02:05 update.go:1145: [INFO] [stacks] update vikunja: phase undoing
2026/09/23 12:02:05 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy))
2026/09/23 12:02:05 undo.go:363: [ERROR] [stacks] update vikunja: the undo copy vikunja_vikunja_data.pre-update-20260923T120033Z has no finished-marker — it is cut off or missing; NOTHING is put back
2026/09/23 12:02:05 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T120033Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T120033Z}]
2026/09/23 12:02:05 update_guard.go:553: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T11:58:06Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched")
2026/09/23 12:02:05 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true)
2026/09/23 12:02:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only
2026/09/23 12:02:30 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_start_failed (severity warning) — further app_start_failed events are logged at DEBUG only
2026/09/23 12:03:00 notifier.go:1167: [DEBUG] notifier disabled: dropped event app_start_failed (severity warning)
@@ -0,0 +1,27 @@
14:03:26 --- evidence off the machine FIRST (R-320): the controller log of this phase
14:03:28 saved 57 lines to 28-9202-controller-log.txt
14:03:29 === teardown 9202: vikunja removed through the product
14:03:29 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
14:04:00 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
14:04:08 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
14:04:10 undo copies left: 0
14:04:12 vikunja volumes left: 0
14:04:14 images: Deleted: sha256:8711aa5bb344152bb7043c3c8a1b1013379492383d68414b3150fe3ae192ec51
Deleted: sha256:3c889b99e3ccd594794c0c3738db5c8150535d0894b764667dfa041bb51c8dfe
Deleted: sha256:cf151044fdc4de3c2d630d3aea2fe854f70c0be91239a085c55e4fa8f16bef03
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
14:05:04 catalog cache: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
14:05:06 controller.yaml vs saved copy: IDENTICAL
14:05:06 deployed apps now: ['gokapi', 'paperless-ngx', 'privatebin']
@@ -0,0 +1,25 @@
14:06:24 === 9201 BEFORE: standing apps, their compose + .felhom.yml checksums, image, controller
14:06:27 adventurelog 947bcad8f388 54cbd75330e3 Up 7 hours (healthy)
bentopdf 39679e28cdd6 3f63855ee03f Up 7 hours (healthy)
bookstack 190e20714347 7b57e5c06fb7 Up 7 hours (healthy)
calibre-web 295cee174968 8c3b9bec9c3a Up 7 hours (healthy)
docmost 9d01fd6fac49 da820f01b95c Up 7 hours (healthy)
kimai c17a7027190c 1fae6f4d5478 Up 7 hours (healthy)
opengist a44845f6db19 17d8ac371c07 Up 7 hours (healthy)
paperless-ngx ac1dd1354afb 19d1960ff453 Up 7 hours (healthy)
privatebin 79e7877dc3e8 eb4a2c7cc42a Up 7 hours (healthy)
romm cab4100bf0b0 2671d4102d5f Up 7 hours (healthy)
gitea.dooplex.hu/admin/felhom-controller:0.263.2
14:06:27 === upgrade 9201 to 0.264.0 (hub v0.120.0 is already live)
14:06:56 gitea.dooplex.hu/admin/felhom-controller:0.264.0
gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up Less than a second (health: starting)
2026/09/23 12:06:31 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 10 [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin romm]; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
"app_update_events_seeded":
seeded: True | language: hu
enabled_events: None
top keys with notif: ['notifications', 'app_update_events_seeded']
{'cooldown_hours': 6}
@@ -0,0 +1,11 @@
adventurelog 947bcad8f388 54cbd75330e3 Up 7 hours (healthy)
bentopdf 39679e28cdd6 3f63855ee03f Up 7 hours (healthy)
bookstack 190e20714347 7b57e5c06fb7 Up 7 hours (healthy)
calibre-web 295cee174968 8c3b9bec9c3a Up 7 hours (healthy)
docmost 9d01fd6fac49 da820f01b95c Up 7 hours (healthy)
kimai c17a7027190c 1fae6f4d5478 Up 7 hours (healthy)
opengist a44845f6db19 17d8ac371c07 Up 7 hours (healthy)
paperless-ngx ac1dd1354afb 19d1960ff453 Up 7 hours (healthy)
privatebin 79e7877dc3e8 eb4a2c7cc42a Up 7 hours (healthy)
romm cab4100bf0b0 2671d4102d5f Up 7 hours (healthy)
gitea.dooplex.hu/admin/felhom-controller:0.263.2
@@ -0,0 +1,28 @@
14:07:37 === 9201 onto the drill catalog (saved: controller.yaml.pre-undofleet)
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
14:08:23 catalog cache: a8d8da9 DRILL glance: glanceapp/glance:v0.8.5 -> glanceapp/glance:v0.8.4
https://<cred>@gitea.dooplex.hu/admin/app-catalog-drill.git
14:08:25 control 1 — standing apps' compose + .felhom.yml unchanged by the drill sync: True (10 apps)
14:08:25 adventurelog 947bcad8f388 54cbd75330e3 Up 7 hours (healthy)
bentopdf 39679e28cdd6 3f63855ee03f Up 7 hours (healthy)
bookstack 190e20714347 7b57e5c06fb7 Up 7 hours (healthy)
calibre-web 295cee174968 8c3b9bec9c3a Up 7 hours (healthy)
docmost 9d01fd6fac49 da820f01b95c Up 7 hours (healthy)
kimai c17a7027190c 1fae6f4d5478 Up 7 hours (healthy)
opengist a44845f6db19 17d8ac371c07 Up 7 hours (healthy)
paperless-ngx ac1dd1354afb 19d1960ff453 Up 7 hours (healthy)
privatebin 79e7877dc3e8 eb4a2c7cc42a Up 7 hours (healthy)
romm cab4100bf0b0 2671d4102d5f Up 7 hours (healthy)
14:08:25 control 2 — standing apps' badges: {'adventurelog': (None, None), 'bentopdf': (None, None), 'bookstack': (None, None), 'calibre-web': (None, None), 'docmost': (None, None), 'kimai': (None, None), 'opengist': (None, None), 'paperless-ngx': (None, None), 'privatebin': (None, None), 'romm': (None, None)}
14:08:25 control 3 — live catalog main: cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main
@@ -0,0 +1,24 @@
14:08:40 === prep vikunja: deploy at the drill FROM pin
14:08:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
14:08:45 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
14:08:45 deployed: True
14:08:46 vikunja: register http=200
14:08:46 vikunja: create project http=201
14:08:46 vikunja: readback of the seeded project http=200 ok=True
14:08:46 C1 A: True
14:08:46 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
14:10:42 [4] backup idle; last=None
14:10:42 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill229ad3","created":"2026-09
14:10:42 vikunja B: project readback=True attachment content readback=True
14:10:42 B reads back: True
14:10:46 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
14:10:46 === glance: deploy at the drill FROM pin (v0.8.4)
14:10:46 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
14:11:06 [1] deployed, controller state=running, pinned={'glance': 'glanceapp/glance:v0.8.4'}
14:11:06 deployed: True
14:11:06 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
14:13:07 [4] backup idle; last=None
14:13:07 glance answers: True
14:13:09 vikunja probe: port: 3456
glance probe: port: 8080
@@ -0,0 +1,54 @@
14:13:18 [5] drill commit 88b959602f91: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
14:13:23 badge caught up after 4.5 s
14:13:23 === vikunja: live undo by the product
14:13:23 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drill229ad3","created":"2026-09
14:13:23 seed C written right before the Update: True
14:13:25 db before: tables=36 ledger=117 newest=SCHEMA_INIT
14:13:26 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
14:13:26 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
14:13:27 + 1.1s phase=pulling label=Új verzió letöltése…
14:13:31 + 4.8s phase=copying label=Az adatok másolása a frissítés előtt…
14:13:33 + 6.9s phase=starting label=Indítás az új verzióval…
14:13:33 + 7.4s phase=verifying label=Működés ellenőrzése…
14:15:03 + 97.7s phase=undoing label=Visszaállítás az előző változatra…
14:15:06 + 100.4s phase=undone label=Visszaállítva az előző változatra
14:15:06 END phase=undone err=None hold=None
14:15:09 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]}
14:15:09 vikunja: readback of the seeded project http=200 ok=True
14:15:09 vikunja B: project readback=True attachment content readback=True
14:15:09 vikunja B: project readback=True attachment content readback=True
14:15:11 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT)
14:15:11 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 14:15 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
14:15:13 leftover copies: 0
14:15:13 === held: the same bad step pressed again, with one undo copy cut off
14:15:13 === vikunja: a cut-off undo copy
14:15:14 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
14:15:14 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
14:15:15 + 1.1s phase=pulling label=Új verzió letöltése…
14:15:16 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt…
14:15:17 + 3.2s phase=starting label=Indítás az új verzióval…
14:15:17 + 3.7s phase=verifying label=Működés ellenőrzése…
14:15:20 >>> finished-marker removed from one copy:
copy: vikunja_vikunja_data.pre-update-20260923T121516Z
total 12
drwxr-xr-x 3 root root 4096 Sep 23 12:15 .
drwxr-xr-x 1 root root 4096 Sep 23 12:15 ..
drwxr-xr-x 2 root root 4096 Sep 23 12:15 data
-rw-r--r-- 1 root root 0 Sep 23 12:15 felhom-undo-complete
data
14:16:48 + 94.3s phase=undoing label=Visszaállítás az előző változatra…
14:16:49 + 95.3s phase=failed label=A frissítés nem sikerült
14:16:49 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
14:16:51 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."}}
14:16:51 box language -> en: http 302
14:16:51 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."}}
14:16:51 box language -> hu: http 302
14:16:53 copies kept: vikunja_vikunja_data.pre-update-20260923T121516Z
vikunja_vikunja_db.pre-update-20260923T121516Z
14:16:56 2026/09/23 12:15:06 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true)
2026/09/23 12:15:06 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 2s — the previous version is running on the data from before the update (the app's health check passed)
2026/09/23 12:16:48 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T121516Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T121516Z}]
2026/09/23 12:16:48 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true)
@@ -0,0 +1,4 @@
14:17:15 box language -> en: http 302
language = en
@@ -0,0 +1,35 @@
14:27:49 === glance (box language en): the failing edge v0.8.4 -> v0.8.5, probe 8080 -> 8999 (drill only)
14:27:49 [5] drill commit d36a5fc5294b: glance glanceapp/glance:v0.8.4 -> glanceapp/glance:v0.8.5 (push rc=0)
14:27:54 badge caught up after 4.4 s
14:27:54 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'The update has started – the card shows how it is going'}
14:27:54 + 0.0s phase=safety-dump label=Database snapshot…
14:27:55 + 1.1s phase=pulling label=Downloading the new version…
14:27:56 + 2.1s phase=copying label=Copying the data before the update…
14:27:57 + 2.7s phase=starting label=Starting the new version…
14:27:57 + 3.2s phase=verifying label=Checking that it works…
14:29:28 + 93.5s phase=undoing label=Putting the previous version back…
14:29:35 + 100.4s phase=undone label=Put back to the previous version
14:29:35 END phase=undone err=None hold=None
14:29:35 PAGE: {"hu": {"undone_line": "A(z) glance frissítése 2026-09-23 14:29-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of glance at 2026-09-23 14:29 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
14:29:37 leftover copies: 0
14:29:37 === glance held: pressed again, one undo copy cut off in `verifying`
14:29:37 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'The update has started – the card shows how it is going'}
14:29:37 + 0.0s phase=safety-dump label=Database snapshot…
14:29:38 + 0.5s phase=pulling label=Downloading the new version…
14:29:39 + 2.1s phase=copying label=Copying the data before the update…
14:29:40 + 2.6s phase=starting label=Starting the new version…
14:29:41 + 3.2s phase=verifying label=Checking that it works…
14:29:43 >>> finished-marker removed from one copy:
copy: glance_glance_config.pre-update-20260923T122939Z
data
14:31:11 + 93.8s phase=undoing label=Putting the previous version back…
14:31:12 + 94.3s phase=failed label=The update did not succeed
14:31:12 END phase=failed err='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.' hold='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.'
14:31:12 PAGE: {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes."}}
14:31:12 GET /api/stacks/glance (box en): label='The update did not succeed' err='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.'
14:31:14 2026/09/23 12:15:06 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true)
2026/09/23 12:16:48 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true)
2026/09/23 12:29:34 undo.go:474: [INFO] [stacks] update glance: event app_update_undone (hold recorded: true)
2026/09/23 12:31:11 undo.go:474: [INFO] [stacks] update glance: event app_update_held (hold recorded: true)
@@ -0,0 +1,11 @@
# notification_log, customer demo-hp, app_update_* (copy of hub.db+wal+shm 2026-09-23T12:31:46Z); message cut to 90 chars
name id customer_id event_type severity message status error_message created_at channel
id|created_at|event_type|severity|channel|status|error_message|message
937|2026-09-23 12:15:06|app_update_undone|warning|operator|sent||A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaáll
938|2026-09-23 12:15:07|app_update_undone|warning|customer|sent||A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaáll
939|2026-09-23 12:16:49|app_update_held|error|operator|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál
940|2026-09-23 12:16:49|app_update_held|error|customer|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál
942|2026-09-23 12:29:35|app_update_undone|warning|operator|sent||A(z) glance frissítése 2026-09-23 14:29-kor nem sikerült. A doboz automatikusan visszaállí
943|2026-09-23 12:29:35|app_update_undone|warning|customer|sent||A(z) glance frissítése 2026-09-23 14:29-kor nem sikerült. A doboz automatikusan visszaállí
944|2026-09-23 12:31:12|app_update_held|error|operator|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál
946|2026-09-23 12:31:12|app_update_held|error|customer|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál
@@ -0,0 +1,55 @@
# The mails as received — 2026-09-23, guest 9201 (customer demo-hp), read with the Gmail connector
Household address = the demo household's catch-all address (`drill@`); operator = `admin@`. Both land in the
@felhom.eu catch-all mailbox. Bodies quoted verbatim (plain-text part). The `notification_log` rows are in
`40-notification-log.txt` (ids 937–946, all `sent`).
## Round 1 — box language Hungarian, app vikunja
**Household, undone** (12:15:08Z, log row 938)
Subject: `[Felhom] Figyelmeztetés: vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut`
```
vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut
- Szint: Figyelmeztetés
- Típus: app_update_undone
- Üzenet: A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el, nincs teendőd.
- Megjegyzés: {"app":"vikunja","stack_name":"vikunja","from":{"vikunja":"vikunja/vikunja:2.3.0"},"to":{"vikunja":"vikunja/vikunja:2.6.0"},"at":"2026-09-23T12:15:06Z"}
```
**Operator, undone** (12:15:07Z, row 937) Subject: `[Felhom] ⚠️ demo-hp: app_update_undone` — Message: the same Hungarian sentence.
**Household, held** (12:16:50Z, row 940)
Subject: `[Felhom] Hiba: vikunja: az alkalmazás leállítva, visszaállítás szükséges`
```
- Üzenet: A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
- Megjegyzés: {...,"copy_tier":1,"copy_date":"2026-09-23T12:13:01Z","copy_holds":"a beállításokat, az adatbázist és az adatköteteket tartalmazza"}
```
**Operator, held** (12:16:49Z, row 939) Subject: `[Felhom] 🔴 demo-hp: app_update_held` — Message: the same Hungarian hold sentence.
## Box language switched to English once (14:17:15 local); hub reports received 14:17:18 and 14:22:42
## Round 2 — box language English, app glance
**Household, undone** (12:29:35Z, row 943)
Subject: `[Felhom] Warning: glance: the update did not work; the app runs on its previous version`
```
glance: the update did not work; the app runs on its previous version
- Level: Warning
- Type: app_update_undone
- Message: The update of glance at 2026-09-23 14:29 did not work. The box put back the previous version and its data automatically — nothing was lost, and there is nothing you need to do.
- Note: {"app":"glance","stack_name":"glance","from":{"glance":"glanceapp/glance:v0.8.4"},"to":{"glance":"glanceapp/glance:v0.8.5"},"at":"2026-09-23T12:29:34Z"}
```
**Operator, undone** (12:29:35Z, row 942) — Message in Hungarian (the operator's wire text), as designed.
**Household, held** (12:31:12Z, row 946)
Subject: `[Felhom] Error: glance: the app is stopped and needs a restore`
```
- Message: The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.
- Note: {...,"copy_holds":"a beállításokat, az adatbázist és az adatköteteket tartalmazza"} <- Hungarian in the raw Note line (row filed)
```
**Operator, held** (12:31:12Z, row 944) Subject: `[Felhom] 🔴 demo-hp: app_update_held` — Hungarian hold sentence.
## What else arrived
- The operator also received `app_start_failed` for each held app (12:17:12Z vikunja, 12:31:12Z glance) — pre-existing
behaviour: a held app is a stopped app. The household did not (the type is not in its list).
- **Per-app cooldown, live:** glance's undone mail (row 943) went out 14 minutes after vikunja's (row 938) inside
the 6-hour customer cooldown. Before hub v0.120.0 the key was `customer:type`, and it would have been swallowed.
@@ -0,0 +1,12 @@
2026/09/23 12:06:31 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 10 [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin romm]; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
2026/09/23 12:07:41 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
2026/09/23 12:15:06 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true)
2026/09/23 12:15:06 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 2s — the previous version is running on the data from before the update (the app's health check passed)
2026/09/23 12:16:48 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T121516Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T121516Z}]
2026/09/23 12:16:48 update_guard.go:553: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T12:13:01Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched")
2026/09/23 12:16:48 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true)
2026/09/23 12:29:34 undo.go:474: [INFO] [stacks] update glance: event app_update_undone (hold recorded: true)
2026/09/23 12:29:34 undo.go:405: [INFO] [stacks] update glance: UNDONE in 7s — the previous version is running on the data from before the update (the app's health check passed)
2026/09/23 12:31:11 update.go:871: [ERROR] [stacks] update glance: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{glance_glance_config glance_glance_config.pre-update-20260923T122939Z}]
2026/09/23 12:31:11 update_guard.go:553: [WARN] [backup] glance is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T12:12:42Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched")
2026/09/23 12:31:11 undo.go:474: [INFO] [stacks] update glance: event app_update_held (hold recorded: true)
@@ -0,0 +1,43 @@
14:32:19 evidence first: 12 log lines saved
14:32:19 === remove vikunja through the product
14:32:19 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
14:32:51 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
14:32:59 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
14:32:59 === remove glance through the product
14:32:59 [X] stop -> 200 {'ok': True, 'message': 'Stack glance stop completed'}
14:33:30 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'glance', 'volumes_removed': ['glance_glance_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alka
14:33:38 [X] after remove: deployed=False leftovers='/opt/docker/stacks/glance'
14:33:40 undo copies left: 0
14:33:42 volumes left: 0
14:33:44 images: 18
14:33:44 box language -> hu: http 302
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
14:34:35 catalog cache: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
14:34:37 controller.yaml vs saved: IDENTICAL
14:34:39 language: hu
14:34:41 standing apps' files identical to before: True
14:34:41 adventurelog 947bcad8f388 54cbd75330e3 Up 23 minutes (healt
bentopdf 39679e28cdd6 3f63855ee03f Up 8 hours (healthy)
bookstack 190e20714347 7b57e5c06fb7 Up 23 minutes (healt
calibre-web 295cee174968 8c3b9bec9c3a Up 22 minutes (healt
docmost 9d01fd6fac49 da820f01b95c Up 22 minutes (healt
kimai c17a7027190c 1fae6f4d5478 Up 22 minutes (healt
opengist a44845f6db19 17d8ac371c07 Up 22 minutes (healt
paperless-ngx ac1dd1354afb 19d1960ff453 Up 22 minutes (healt
privatebin 79e7877dc3e8 eb4a2c7cc42a Up 22 minutes (healt
romm cab4100bf0b0 2671d4102d5f Up 21 minutes (healt
14:34:41 deployed: ['adventurelog', 'bentopdf', 'bookstack', 'calibre-web', 'docmost', 'kimai', 'opengist', 'paperless-ngx', 'privatebin', 'romm']
@@ -0,0 +1,15 @@
=== FLOOR -> 0.264.0, MinAgent 0.131.0 declared (2026-09-23 14:35:08)
POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set
--- read back from the hub:
"min_controller_version" value="0.264.0"
"min_agent" value="0.131.0"
"min_agent" value="0.131.0"
--- watching the two demo boxes (up to 3 min)
+10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up About a minute (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up About an hour (healthy)
+20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up About a minute (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up 10 seconds (healthy)
--- hub log:
2026/09/23 14:35:11 [INFO] managed floor SERVED for demo-felhom: floor 0.264.0, agent requirement "0.131.0" from declared (golden 0.258.0)
--- hosts page:
demo-felhom-8363b5
demo-hp-bb76ea
drill-r50-0a4f9a
@@ -0,0 +1,50 @@
# The undo reaches the fleet, and the household is told — controller v0.264.0 + hub v0.120.0, 2026-09-23
**Method: endpoint-level.** Every act is the endpoint the UI invokes (deploy, backup, sync, rescan, update,
remove, the language switch); pages are fetched as HTML (`/apps/<app>?lang=hu|en`); the mails are read from
the @felhom.eu catch-all mailbox with the Gmail connector; the hub's rows from a copy of `hub.db`+`-wal`+`-shm`.
No browser. The product performs every undo and every hold; the harness (`walk.py`, `live.py`, `bakeoff.py`,
copied from `undo-live-2026-09-23/`, now selectable by `GUEST=9201|9202`) only presses and reads.
## Order of events
| # | what | evidence |
|---|---|---|
| 0 | floor 0.263.2 (Part 0) | `00-floor-0.263.2.txt` |
| 1 | hub v0.120.0 deployed by manifest bump + ArgoCD, **before any box ran v0.264.0**; one-time add-only migration changed all 3 `enabled_events` rows | `10-enabled-events-before.txt`, `11-hub-0.120.0-deploy-and-migration.txt` |
| 2 | 9202 (no hub): v0.264.0; R-646 backfill recorded 3; vikunja undone then held; pages hu/en; WARN once per type, then DEBUG | `20-*` … `29-*` |
| 3 | 9201 (hub on): v0.264.0; backfill recorded 10; drill repoint with three controls; vikunja undone+held (box hu); language → en; glance undone+held (box en) | `30-*` … `35-*` |
| 4 | 8 mails as received + 8 `notification_log` rows | `40-notification-log.txt`, `41-mails-as-received.md` |
| 5 | teardown of 9201, floor 0.264.0 read back | `48-*`, `49-*`, `50-floor-0.264.0.txt` |
## Results
| proof | result |
|---|---|
| migration (hub) | `demo-felhom` kept its 10 types and gained 2; `demo-hp` and `_resend-rotation-test` (`null` / `[]` = nothing on) gained the 2; guard row set 11:49:03Z; log names all three |
| R-646 | `applied-meta backfill (R-646): recorded 3 [gokapi paperless-ngx privatebin]` (9202); `recorded 10 […]` (9201); skipped 0 on both |
| undone (9202, vikunja) | `undone` at +99 s; seeds A, B, C back; ledger equal; page hu *„A(z) vikunja frissítése … — semmi nem veszett el."* / en *"The update of vikunja … — nothing was lost."*; API `?lang=en` label *"Put back to the previous version"* |
| held (9202, cut-off copy) | `failed` 1 s into `undoing`; hold sentence hu on the hu page and **en on the en page (R-606)**; both copies kept |
| R-620 | WARN once each for `controller_started`, `app_deploy_started`, `app_deployed`, `health_change`, `app_update_undone`, `app_update_held`, `app_start_failed`; the second `app_start_failed` at 12:03:00Z is DEBUG |
| mails (9201) | 4 household + 4 operator, rows 937–946 all `sent`; subjects name the app; household text in the box language (hu for vikunja, en for glance); operator text Hungarian. Quoted in `41-mails-as-received.md` |
| per-app cooldown | glance's undone mail went 14 min after vikunja's, inside the 6 h customer cooldown |
| standing apps (9201) | compose + `.felhom.yml` byte-identical before/after (10 apps), catalog back on live `cfcfe52`, `controller.yaml` identical to the saved copy. **But not untouched:** the harness's two „Mentés most" presses (whole-box backups) stopped and restarted 9 of 10 standing apps for their volume dumps — R-648 |
## Found
- **R-647** — (1) a reader in the other language than the box reads a HELD sentence in the box's language
(seen: 9202 box en, `?lang=hu` page showed English); (2) the English mail's raw `Note:` line carries the
Hungarian `copy_holds` phrase; (3) two log wordings (`hold recorded: true` on an undone event; `severity warn`
printing a health status).
- **R-648** — the drill's backup press is whole-box.
## Teardown — three layers
- **machine:** 9202 — vikunja removed through the product (0 volumes, 0 undo copies), test images removed,
`controller.yaml` identical to `controller.yaml.pre-undofleet`, catalog on live `cfcfe52`, apps as at the
start. 9201 — vikunja and glance removed through the product (0 volumes, 0 copies), test images removed,
language back to `hu`, config identical to the saved copy, catalog on live `cfcfe52`, the 10 standing apps
deployed and healthy. Both stay on controller 0.264.0 (= the floor).
- **host:** nothing provisioned on demo-hp.
- **hub:** v0.120.0 deployed (the one planned change); floor 0.264.0 / MinAgent 0.131.0; no customer reset.
- **drill repo:** reset to live `main` (`cfcfe52`), read back from the remote.
@@ -0,0 +1,266 @@
#!/usr/bin/env python3
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
drill catalog.
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
Start supplies the app's secrets (they are never decrypted here).
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
"""
import json, os, sys, time
sys.path.insert(0, ".")
import walk as w
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
romm_verify_b, vik_seed_b, vik_verify_b
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
SUFFIX = ".pre-undo"
def seed_b(app, sub, A):
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
def verify_b(app, sub, A, B):
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
return all(r.values()) if isinstance(r, dict) else r
def db_state(app):
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
if app == "docmost":
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
if app == "romm":
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
alembic = q("select version_num from alembic_version")
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
return w.guest("""python3 - <<'PY'
import sqlite3
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
m=c.execute("select count(*), max(id) from migration").fetchone()
print(f"tables={t} ledger={m[0]} newest={m[1]}")
PY""").strip()
def volumes(app):
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
if not v.endswith(SUFFIX)]
def containers(app):
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
def stage_prep(app):
s = {"app": app, "sub": SUB[app]}
w.say(f"=== prep {app}: deploy at the drill FROM pin")
w.say("deployed:", w.deploy(app, s["sub"]))
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
w.backup_now(app)
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
s["volumes"] = volumes(app)
w.say("named volumes:", s["volumes"])
save(app, s)
def stage_fcopy(app):
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
s = load(app)
s["state_before_update"] = db_state(app)
w.say("db state before the update:", s["state_before_update"])
vols = s["volumes"]
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
script = f"""
set -u
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
t0=$(date +%s.%N)
docker stop $cs >/dev/null
t1=$(date +%s.%N)
for v in {' '.join(vols)}; do
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
docker volume create "$v{SUFFIX}" >/dev/null
a=$(date +%s.%N)
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
b=$(date +%s.%N)
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
mkdir -p /root/bk; c=$(date +%s.%N)
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
done
t2=$(date +%s.%N)
docker start $cs >/dev/null
t3=$(date +%s.%N)
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
"""
out = w.guest(script, timeout=1800)
w.say(out)
t0 = time.time()
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
s["fcopy"] = out
s["fcopy_at"] = ts()
save(app, s)
def stage_break(app):
s = load(app)
frm, to, p0, p1 = EDGE[app]
s["break_commit"] = break_edge(app, frm, to, p0, p1)
w.say("badge caught up after", w.sync_rescan(app, to), "s")
save(app, s)
def stage_update(app):
s = load(app)
r = w.press_update(app, poll=1.0)
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
s.setdefault("updates", []).append(r)
save(app, s)
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
docker restart felhom-controller >/dev/null; sleep 15"""
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
def lift_and_pinback(app):
frm, to, _, _ = EDGE[app]
t = time.time()
w.say(w.guest(LIFT.format(app=app)))
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
w.login()
return round(time.time() - t, 1)
def stage_undoF(app):
s = load(app)
w.say(f"=== undo by F: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
vols = s["volumes"]
script = f"""
t0=$(date +%s.%N)
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
for v in {' '.join(vols)}; do
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
done
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
"""
w.say(w.guest(script, timeout=1800))
t0 = time.time()
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
w.say("product start ->", c)
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_undoD(app):
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
s = load(app)
w.say(f"=== undo by D: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
if app == "vikunja":
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
return
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
time.sleep(8)
if app == "docmost":
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
marker = "-- PostgreSQL database dump complete"
else:
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
marker = "-- Dump completed"
script = f"""
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
t0=$(date +%s.%N)
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
"""
w.say(w.guest(script, timeout=900))
t0 = time.time()
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_cutoff(app):
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
s = load(app)
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
script = f"""
v={big}
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
rm -f /root/cut.sql
else echo "D: no safety dump exists for this app (R-641)"; fi
"""
w.say(f"=== cut-off copy: {app} (largest volume {big})")
out = w.guest(script, timeout=600)
w.say(out)
s["cutoff"] = out
save(app, s)
if __name__ == "__main__":
w.login()
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,194 @@
#!/usr/bin/env python3
"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes.
Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/<n>,
fetches the app page in both languages, and reads the seeds back through each app's own front door.
Seed C is written immediately before each Update, after every backup — so only the undo's own
last-second copy can bring it back.
"""
import json, re, sys, time, html as H
sys.path.insert(0, ".")
import walk as w
from spike import load, save, ts
import bakeoff as b
HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}
def seed_c(app, s):
return b.seed_b(app, s["sub"], s["seedA"])
def page_lines(app):
out = {}
for lang in ("hu", "en"):
h = H.unescape(w.page(f"/apps/{app}?lang={lang}"))
m = re.search(r'data-update-undone="true">([^<]*)<', h)
hold = re.search(r'data-held="true">([^<]*)<', h)
out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None}
return out
def press(app, poll=0.5, on_phase=None):
code, d = w.ctl("POST", f"/api/stacks/{app}/update")
w.say(f" Update -> {code} {str(d)[:140]}")
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < 1500:
try:
st = w.stack(app)
except Exception as e:
st = {}
ph = st.get("update_phase")
if ph != seen and ph is not None:
seen = ph
phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label")))
w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}")
if on_phase and on_phase(ph):
return phases, "interrupted"
if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2:
break
time.sleep(poll)
st = w.stack(app)
w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}")
return phases, st
def readback(app, s, with_c=True):
sub = s["sub"]
w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2)
A = b.FX[app].verify(w, sub, s["seedA"], w.say)
B = b.verify_b(app, sub, s["seedA"], s["seedB"])
C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None
return {"A": A, "B": B, "C": C}
def stage_undo(app):
s = load(app)
w.say(f"=== {app}: live undo by the product")
s["seedC"] = seed_c(app, s); s["seedC_at"] = ts()
w.say(" seed C written right before the Update:", bool(s["seedC"]))
before = b.db_state(app); w.say(" db before:", before)
phases, st = press(app)
obs = w.observables(app)
w.say(" observables:", json.dumps(obs))
rb = readback(app, s)
after = b.db_state(app)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})")
pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False))
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")},
"readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs}
save(app, s)
def set_box_language(lang):
"""POST /settings/language — the household's language switch (form + the dashboard's CSRF)."""
sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}",
"-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}",
f"{w.BASE}/settings/language"])
w.say(f" box language -> {lang}: http {r.stdout.strip()}")
def stage_powercut(app):
"""Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard
stop); boot it again and let the controller resume the undo."""
import subprocess
s = load(app)
w.say(f"=== {app}: power cut DURING the undo")
s["seedC"] = seed_c(app, s)
cut = {}
def on_phase(ph):
if ph == "undoing":
t = time.time()
r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120)
cut["at"] = ts(); cut["rc"] = r.returncode
w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s")
return True
return False
phases, _ = press(app, poll=0.3, on_phase=on_phase)
if not cut:
w.say(" the cut never landed in `undoing` — recorded as a MISS"); return
r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180)
w.say(f" guest started again: rc={r.returncode}")
for i in range(60):
time.sleep(5)
try:
w.login(); st = w.stack(app)
if st:
break
except SystemExit:
continue
w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6"))
t0 = time.time()
while time.time() - t0 < 600:
st = w.stack(app)
if not st.get("updating") and st.get("update_phase") in ("undone", "failed"):
break
time.sleep(3)
w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}")
rb = readback(app, s)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}")
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb}
save(app, s)
def stage_cutoff(app):
"""Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of
the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it
back and HOLD with the new prefix."""
s = load(app)
w.say(f"=== {app}: a cut-off undo copy")
done = {}
def on_phase(ph):
if ph == "verifying" and not done:
out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c"
docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""")
done["out"] = out
w.say(" >>> finished-marker removed from one copy:\n" + out)
return False
phases, st = press(app, poll=0.5, on_phase=on_phase)
before = b.db_state(app)
pl = page_lines(app)
w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False))
set_box_language("en")
pl_en = page_lines(app)
w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False))
set_box_language("hu")
w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}"))
s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done}
save(app, s)
def stage_fixprobe_and_press(app):
"""After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work
and end the undone note."""
s = load(app)
frm, to, p0, p1 = b.EDGE[app]
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f)
w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"])
w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120)
w.sync_rescan(app, to)
time.sleep(20)
w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan")
w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)")
w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False))
phases, st = press(app)
rb = readback(app, s)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}")
pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False))
w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml"))
s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl}
save(app, s)
if __name__ == "__main__":
w.login()
{"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff,
"fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2])
@@ -0,0 +1,60 @@
#!/usr/bin/env python3
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
`controller.yaml.pre-undofleet` (NOT the older `.pre-28`, which a restore must never pick up).
"""
import re, sys, io
sys.path.insert(0, '.')
import walk as w
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
def creds():
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
if m:
return m.group(1), m.group(2)
raise SystemExit("no admin credential")
def to_drill():
u, t = creds()
print(w.guest(f"""
set -e
test -f {VOL}/controller.yaml.pre-undofleet || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-undofleet
python3 - <<'PY'
import re
p = "{VOL}/controller.yaml"
s = open(p).read()
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
if not re.search(r'^update:', s, re.M):
s += "update:\\n health_timeout: 90s\\n"
open(p, "w").write(s)
PY
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -A2 '^update:' {VOL}/controller.yaml
"""))
def restore():
print(w.guest(f"""
set -e
cp -p {VOL}/controller.yaml.pre-undofleet {VOL}/controller.yaml
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -c '^update:' {VOL}/controller.yaml || true
"""))
if __name__ == "__main__":
to_drill() if sys.argv[1] == "drill" else restore()
@@ -0,0 +1,160 @@
#!/usr/bin/env python3
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
steps the product would take, each one timed.
State between stages lives in state-<app>.json so each stage can be run, read, and only then
followed by the next (the hand undo needs a person looking at what the previous step left).
"""
import json, os, sys, time, re
sys.path.insert(0, ".")
import walk as w
import fixtures as fx
HERE = os.path.dirname(os.path.abspath(__file__))
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
def st_path(app):
return os.path.join(HERE, f"state-{os.environ.get('GUEST', '9202')}-{app}.json")
def load(app):
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
def save(app, s):
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
def ts():
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
def docmost_seed_b(sub, A):
jar = "/tmp/dm.jar"
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
name = "drillB" + os.urandom(3).hex()
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
return {"space": name} if code in ("200", "201") else None
def docmost_verify_b(sub, A, B):
jar = "/tmp/dm.jar"
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
if code not in ("200", "201"):
w.say(f" docmost B: cannot log in (http={code})"); return False
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
data="{}", method="POST")
ok = code in ("200", "201") and B["space"] in out
neg = "drillBnever" in out
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
return ok and not neg
def break_edge(app, frm, to, port_from, port_to):
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
health probe pointed at a port the app does not answer. Both in one drill commit."""
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read()
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
assert m, "probe port not found"
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
open(fy, "w").write(f)
h = w.drill_bump(app, frm, to)
return h
def pg_state(container, db, user):
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
def romm_seed_b(sub, A):
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
jar, tok = FX["romm"]._csrf(w, sub)
user = "drillb" + os.urandom(3).hex()
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
method="POST")
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
return {"user": user} if code in ("200", "201") else None
def romm_verify_b(sub, A, B):
jar, tok = FX["romm"]._csrf(w, sub)
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}")
ok = code == "200" and B["user"] in body
neg = "drillbnever" in body
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
return ok and not neg
def tree(paths):
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
return w.guest(cmd)
def my_state(container="romm-db", db="romm"):
# The root password is used INSIDE the container from its own env — it never leaves it.
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
echo -n "alembic=$({q("select version_num from alembic_version")}) "
echo -n "users=$({q("select count(*) from users")})"
""")
def vik_seed_b(sub, A):
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
app writes into its files volume), uploaded through the app's own attachment API."""
tok, why = FX["vikunja"]._token(w, sub, A)
title = "drillB-" + os.urandom(4).hex()
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
if code not in ("200", "201"):
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
pid = json.loads(body)["id"]
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
tid = json.loads(body)["id"] if code in ("200", "201") else None
content = "drill attachment " + os.urandom(8).hex()
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
"-F", f"files=@{fn}", method="PUT")
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
return {"title": title, "pid": pid, "tid": tid, "att": content}
def vik_verify_b(sub, A, B):
tok, why = FX["vikunja"]._token(w, sub, A)
if not tok:
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
proj = code == "200" and B["title"] in body
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
att_ok = False
try:
atts = json.loads(body)
if atts:
aid = atts[0]["id"]
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
att_ok = c3 == "200" and B["att"] in b3
except Exception as e:
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
return {"project": proj, "attachment": att_ok}
@@ -0,0 +1,172 @@
{
"app": "vikunja",
"sub": "tasks",
"seedA": {
"user": "drill229ad3",
"pw": "<not recorded>",
"title": "drill-251d05a1ca",
"pid": 2
},
"seedB": {
"title": "drillB-f2e2cc24",
"pid": 3,
"tid": 1,
"att": "drill attachment c2abe62af7869525"
},
"volumes": [
"vikunja_vikunja_data",
"vikunja_vikunja_db"
],
"break_commit": "88b959602f91",
"seedC": {
"title": "drillB-935bb593",
"pid": 4,
"tid": 2,
"att": "drill attachment 2b1d7877c1929b80"
},
"seedC_at": "2026-09-23T12:13:23Z",
"live_undo": {
"phases": [
[
0.0,
"safety-dump",
"Adatbázis pillanatkép…"
],
[
1.1,
"pulling",
"Új verzió letöltése…"
],
[
4.8,
"copying",
"Az adatok másolása a frissítés előtt…"
],
[
6.9,
"starting",
"Indítás az új verzióval…"
],
[
7.4,
"verifying",
"Működés ellenőrzése…"
],
[
97.7,
"undoing",
"Visszaállítás az előző változatra…"
],
[
100.4,
"undone",
"Visszaállítva az előző változatra"
]
],
"end": {
"update_phase": "undone",
"update_error": null,
"hold_reason": null
},
"readback": {
"A": true,
"B": true,
"C": true
},
"db_before": "tables=36 ledger=117 newest=SCHEMA_INIT",
"db_after": "tables=36 ledger=117 newest=SCHEMA_INIT",
"page": {
"hu": {
"undone_line": "A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.",
"hold": null
},
"en": {
"undone_line": "The update of vikunja at 2026-09-23 14:15 did not succeed. The box put back the previous version and its data automatically — nothing was lost.",
"hold": null
}
},
"obs": {
"pinned_images": {
"vikunja": "vikunja/vikunja:2.3.0"
},
"installed_images": {
"vikunja": "vikunja/vikunja:2.3.0"
},
"catalog_images": {
"vikunja": "vikunja/vikunja:2.6.0"
},
"live_compose_image_lines": [
"image: vikunja/vikunja:2.3.0"
],
"docker_inspect": [
"vikunja vikunja/vikunja:2.3.0 running=true restarts=0"
]
}
},
"cutoff_live": {
"phases": [
[
0.0,
"safety-dump",
"Adatbázis pillanatkép…"
],
[
1.1,
"pulling",
"Új verzió letöltése…"
],
[
2.1,
"copying",
"Az adatok másolása a frissítés előtt…"
],
[
3.2,
"starting",
"Indítás az új verzióval…"
],
[
3.7,
"verifying",
"Működés ellenőrzése…"
],
[
94.3,
"undoing",
"Visszaállítás az előző változatra…"
],
[
95.3,
"failed",
"A frissítés nem sikerült"
]
],
"end": {
"update_phase": "failed",
"hold_reason": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
},
"page_hu_box": {
"hu": {
"undone_line": null,
"hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
},
"en": {
"undone_line": null,
"hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."
}
},
"page_en_box": {
"hu": {
"undone_line": null,
"hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."
},
"en": {
"undone_line": null,
"hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."
}
},
"marker": {
"out": "copy: vikunja_vikunja_data.pre-update-20260923T121516Z\ntotal 12\ndrwxr-xr-x 3 root root 4096 Sep 23 12:15 .\ndrwxr-xr-x 1 root root 4096 Sep 23 12:15 ..\ndrwxr-xr-x 2 root root 4096 Sep 23 12:15 data\n-rw-r--r-- 1 root root 0 Sep 23 12:15 felhom-undo-complete\ndata\n"
}
}
}
@@ -0,0 +1,167 @@
{
"app": "vikunja",
"sub": "tasks",
"seedA": {
"user": "drill73bafb",
"pw": "<not recorded>",
"title": "drill-94f9145a94",
"pid": 2
},
"seedB": {
"title": "drillB-b6e6bb28",
"pid": 3,
"tid": 1,
"att": "drill attachment 12d38cc6ade81d67"
},
"volumes": [
"vikunja_vikunja_data",
"vikunja_vikunja_db"
],
"break_commit": "073bc953e182",
"seedC": {
"title": "drillB-6d4934fb",
"pid": 4,
"tid": 2,
"att": "drill attachment d4f2ee5a03683645"
},
"seedC_at": "2026-09-23T11:58:25Z",
"live_undo": {
"phases": [
[
0.0,
"safety-dump",
"Adatbázis pillanatkép…"
],
[
1.1,
"pulling",
"Új verzió letöltése…"
],
[
4.2,
"copying",
"Az adatok másolása a frissítés előtt…"
],
[
5.8,
"starting",
"Indítás az új verzióval…"
],
[
6.3,
"verifying",
"Működés ellenőrzése…"
],
[
96.7,
"undoing",
"Visszaállítás az előző változatra…"
],
[
99.3,
"undone",
"Visszaállítva az előző változatra"
]
],
"end": {
"update_phase": "undone",
"update_error": null,
"hold_reason": null
},
"readback": {
"A": true,
"B": true,
"C": true
},
"db_before": "tables=36 ledger=117 newest=SCHEMA_INIT",
"db_after": "tables=36 ledger=117 newest=SCHEMA_INIT",
"page": {
"hu": {
"undone_line": "A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.",
"hold": null
},
"en": {
"undone_line": "The update of vikunja at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically — nothing was lost.",
"hold": null
}
},
"obs": {
"pinned_images": {
"vikunja": "vikunja/vikunja:2.3.0"
},
"installed_images": {
"vikunja": "vikunja/vikunja:2.3.0"
},
"catalog_images": {
"vikunja": "vikunja/vikunja:2.6.0"
},
"live_compose_image_lines": [
"image: vikunja/vikunja:2.3.0"
],
"docker_inspect": [
"vikunja vikunja/vikunja:2.3.0 running=true restarts=0"
]
}
},
"cutoff_live": {
"phases": [
[
0.0,
"safety-dump",
"Adatbázis pillanatkép…"
],
[
1.1,
"pulling",
"Új verzió letöltése…"
],
[
2.1,
"copying",
"Az adatok másolása a frissítés előtt…"
],
[
3.7,
"verifying",
"Működés ellenőrzése…"
],
[
94.7,
"undoing",
"Visszaállítás az előző változatra…"
],
[
95.2,
"failed",
"A frissítés nem sikerült"
]
],
"end": {
"update_phase": "failed",
"hold_reason": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
},
"page_hu_box": {
"hu": {
"undone_line": null,
"hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
},
"en": {
"undone_line": null,
"hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."
}
},
"page_en_box": {
"hu": {
"undone_line": null,
"hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."
},
"en": {
"undone_line": null,
"hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."
}
},
"marker": {
"out": "copy: vikunja_vikunja_data.pre-update-20260923T120033Z\ntotal 12\ndrwxr-xr-x 3 root root 4096 Sep 23 12:00 .\ndrwxr-xr-x 1 root root 4096 Sep 23 12:00 ..\ndrwxr-xr-x 2 root root 4096 Sep 23 12:00 data\n-rw-r--r-- 1 root root 0 Sep 23 12:00 felhom-undo-complete\ndata\n"
}
}
}
@@ -0,0 +1,495 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-fleet-2026-09-23"
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = {"9202": "https://192.168.0.114", "9201": "https://192.168.0.138"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"cat > /tmp/w{GUEST}.sh; pct push {GUEST} /tmp/w{GUEST}.sh /tmp/w.sh >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash /tmp/w.sh; rm -f /tmp/w{GUEST}.sh"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
code, d = ctl("POST", "/api/backup/run")
say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}")
for _ in range(90):
time.sleep(5)
c2, s = ctl("GET", "/api/backup/status")
dd = s.get("data") or {}
if not dd.get("running", False):
say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}")
return True
say(" [4] backup still running after 7.5 min — carrying on")
return False
def drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
+9
View File
@@ -339,3 +339,12 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-639** | **After a held update the previous definition survived only in the recovery unit (P3).** Closed in controller **v0.263.0/v0.263.2**: the pre-update compose, applied definition, pin and the pinned version's `.felhom.yml` are kept until the undo is over; the undo never reads the unit (R-645). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — SHIPPED** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
| **R-641** | **An app with no database server had no last-second copy (P2).** Closed in controller **v0.263.0**: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
| **R-642** | **Start answered 200 "start completed" over a crash loop (P3).** Closed in controller **v0.263.0**: start/restart answer `requested — state now: <state>`, never "completed" (`startAnswer`, pinned by `TestR642_*`, red-proofed); live on 9202: `Stack romm start requested — state now: running`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
## 2026-09-23 — the household is told when an update is undone or held (controller v0.264.0, hub v0.120.0)
| Row | What | Closed | Full text |
|---|---|---|---|
| **R-606** | **Every sentence the update path shows a household was Hungarian-only (P2).** Closed in controller **v0.264.0** (`bc27894`): `UpdateError` is stored as a key + args; phase labels, refusals, failure lines, the undone line, the hold sentence and its prefix and `UpdateCopyHolds` render in the reader's language on both pages and in `GET /api/stacks/<n>`; the stored Hungarian is byte-identical (parity gate green). Live: 9202 and 9201, both pages, both languages. Three leftovers → **R-647**. | **CLOSED 2026-09-23 — PROVEN-LIVE (endpoint-level; `audits/undo-fleet-2026-09-23/22-*`, `23-*`, `35-*`)** | `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
| **R-620** | **A disabled notifier dropped every event with no local trace (P3).** Closed in controller **v0.264.0**: one WARN per event type per process naming the dropped event, then DEBUG. Live on 9202: six types each WARNed once; `app_start_failed` WARN at 12:02:30Z then DEBUG at 12:03:00Z. | **CLOSED 2026-09-23 — PROVEN-LIVE (`24-9202-r620-warn-then-debug.txt`)** | as above |
| **R-646** | **An app pinned before v0.263.2 had no record of its own `.felhom.yml` (P3).** Closed in controller **v0.264.0**: a startup pass records `applied-meta/` for every deployed, pinned app CURRENT with the catalog; a behind app is skipped by name (its file is gone — not backfillable by design). Live: 9202 recorded 3, 9201 recorded 10, skipped 0. | **CLOSED 2026-09-23 — PROVEN-LIVE (`28-9202-controller-log.txt`, `30-9201-before-and-upgrade.txt`)** | as above |
File diff suppressed because one or more lines are too long