SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s

THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
2026-09-01 21:35:32 +02:00
parent ac079b8c43
commit 56c7e373a3
43 changed files with 2026 additions and 287 deletions
@@ -99,7 +99,10 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Scenario | Components | Status | Evidence | Gap / roadmap |
|---|---|---|---|---|
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3` | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller | **PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
@@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap
### `sync/`
| File | Class | Reason | Risk |
|---|---|---|---|
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). | clean |
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean |
### `system/` — split per-function (not per-file)
| File | Class | Reason | Risk |
@@ -503,3 +503,77 @@ own; every caller that is not the customer must decide for itself whether the ap
- **S6:** §5(3) self-restore-test → **status-display only**; the agent owns orchestration.
- **Self-update resolved (03 §11):** `updater.go` → **DELETE(→agent)**, `state.go` →
DELETE(obsolete), `version.go` KEEP; §6 + §5(2) updated (bulk = `backup=0` mountpoint recipe).
---
## The app-definition seam — what happens to a deployed app when its compose file is rewritten
> **STATUS: MEASURED BEHAVIOUR, NOT A RECORDED DESIGN.** This section describes what the system was
> observed to do on 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`), with controls. It is
> deliberately **not** marked `[DESIGN]`, because the operator has not ruled on whether the behaviour
> was intended. **Do not read this as an endorsement, and do not spec against it as though it were
> settled.** The open questions are R-438 and R-441.
Until this was measured, no architecture document said what happens here, and the gap itself is
R-438. The three facts below are the ones a reader needs before touching any of it.
### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
`Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies
`docker-compose.yml` and `.felhom.yml` into the matching stack folder. **There is no test of whether
the app is deployed.** The only guard is a sha256 content compare (`copyIfChanged`) and the only
exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus
one immediate sync at controller start (`sync.go:98`).
**It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing
else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree,
and they stay that way until something else acts.
### 2. Every lifecycle path resolves that disagreement, silently, by upgrading
`StartStack`, `RestartStack` and `UpdateStack` (`stacks/manager.go:1029 / 1133 / 1170`) all end in
`docker compose up -d`. `up -d` makes the container match the file, **and pulls the image itself if it
is absent** — measured at 18.3 s with a pull versus 0.5 s without, against a negative control
(unchanged file) that did not even recreate the container.
`RestartStack` carries an explicit in-source comment saying this is deliberate — *"so that … any
template changes (new images, healthchecks) are picked up"*. **That is a fact about the source, not an
operator ruling**, and it speaks only for the customer-pressed restart.
**Thirteen non-API call sites across nine files** reach `up -d` without anyone pressing anything. The
full table is §8 of the spike doc; the three that matter most are:
- `bootrecon.Reconciler.Run` → `StartStack` — `bootrecon/bootrecon.go:269`
- `backup.AppStopGuard.Recover` → `StartStack` — `backup/appstop_marker.go:283`
- the drive-return gate, `web.Server.restartStacks` → `StartStack` — `web/intermediary.go:222`
**A plain power cut does NOT trigger this.** Docker's `restart: unless-stopped` restores the existing
containers on the old image, the reconciler finds no orphan, and it logs so
(`no boot-orphaned apps (nothing to start)`). The unattended upgrade needs the narrower precondition
*"and the app did not come back"* — which was measured, and does upgrade.
### 3. Nothing takes a copy first, and the tag cannot be put back afterwards
`UpdateStack` runs `compose pull` then `compose up -d --remove-orphans` and does nothing else — no
dump, no copy, no hold. The R-361 safety dump (`backup/offbox_reconstitute.go:207`) is **database-only**
and is not on this path at all; an app with no database gets nothing from it even on the paths where it
does run.
**And restoring the old tag is not a rollback.** Once an app has migrated its data, the old image
refuses to start — measured on Nextcloud: *"the version of the data (32.0.9.2) is higher than the
docker image version (31.0.14.1) and downgrading is not supported"*. The data itself survives; only the
downgrade is blocked. So the only route back is a **data** restore from a copy taken before the
update.
**And that route fights this seam** (R-441): `stackAdapter.RecreateStackDefinitionFromUnit`
(`cmd/controller/main.go:2570`) writes the recovery unit's captured compose — with the OLD pin — into
the live stack dir, and `copyIfChanged` overwrites it again on the next tick. The overwrite is
measured; that the restore writes to that path is read, not measured.
### What a change here must not break
- The syncer's overwrite is also the **repair** path: a locally-corrupted compose is replaced by the
catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
- `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to
`docker compose restart` would silently stop doing that.
@@ -0,0 +1,626 @@
# SPIKE — what an app update actually does, and which other paths do it too (2026-09-01)
> **THE ANSWER TO PHASE 1, IN ONE SENTENCE: YES — the Restart button upgrades the app, and so does
> the box's own repair when an app fails to come back, because every one of these paths ends in
> `docker compose up -d`, and `up -d` makes the container match whatever the file now says, pulling
> the image itself if it is missing.**
>
> **AND THE SECOND ANSWER, WHICH THE OPERATOR PAGE GOT WRONG IN THE HELPFUL DIRECTION: a plain power
> cut does NOT do this.** Docker's own `restart: unless-stopped` puts the existing containers back on
> the OLD image, so the boot reconciler finds nothing to repair and never runs `up -d`. The unattended
> upgrade happens only in the narrower case where an app does **not** come back by itself — and there
> it happens with nobody pressing anything.
>
> **AND THE THIRD, WHICH DECIDES THE VOCABULARY OF THIS WHOLE ARC: app data CANNOT be rolled back.**
> Measured on Nextcloud: once a migration has actually run, putting the old image tag back produces a
> container that refuses to start — *"the version of the data (32.0.9.2) is higher than the docker
> image version (31.0.14.1) and downgrading is not supported"*. **So "rollback" is the wrong word and
> should be struck before anyone specs against it.** The shape that is actually available is
> *pre-update copy plus refuse-and-explain*.
**Scope.** A spike. **No production code was written, in any repo.** The controller was read only:
`felhom-controller` is at `960d29b0612c` before and after, working tree clean, and
`go build ./... && go vet ./... && go test ./...` is green (28 packages, 0 FAIL) — run at the end
precisely to prove the tree was left untouched.
**Method.** Endpoint-level throughout, which is the standard method here because there is no browser on
DooPlex: every action was the exact endpoint the UI invokes (`POST /api/stacks/{name}/restart`,
`/update`, `/start`, `/stop`, `/deploy`, `/remove`), driven over HTTPS with the customer's own session
cookie and CSRF token. **No raw agent attach, no hand-set state — the F9 bypass was not used.** The two
places where state was staged by hand are named as such at the point of use (Phase 1's compose edits,
and Phase 1c-ii's `docker rm -f`), and each says exactly what was staged and why.
**Target.** All destructive work ran on **`demo-hp` (192.168.0.104, `ssh hp`; guest 9201 =
192.168.0.138)**, which `runbooks/target-selection.md` classes **Tier 0 — disposable**. DooPlex,
`demo-felhom`, `ep0` and Peti's box were not mutated. Peti's box was not touched at all.
**Evidence.** `audits/evidence-spike-app-update-2026-09-01/` — written straight onto DooPlex as each
phase produced it, before every revert, so R-96 rule 5 is satisfied continuously rather than at the
end. Nothing was lost and nothing had to be reproduced.
---
## 1. Confirmed baselines — all four matched §1 of the task, none had moved
| Repo | `main` @ commit at start | at end | note |
|---|---|---|---|
| felhom-controller | `960d29b0612c` | `960d29b0612c` | **read only — unchanged** |
| felhom.eu | `1d59353df437` | moved (this spike's documents) | Phase 0 + Phase 7 |
| app-catalog-felhom.eu | `29edad9c5bf4` | moved, then **reverted to the same tree** | `214d448` → `30bd892` → `5d8f25f` |
| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched |
Live controller on demo-hp: **0.232.0**, matching the newest release. Sync interval **15m**
(`internal/config/config.go:351`, the default, no override on the box).
---
## 2. Phase 1 — THE GATE. Does `compose up -d` upgrade an app whose compose file already moved?
App: **`bentopdf`** — one container, no database, no volume, no data of any kind, and deployed on
`demo-hp` only. Chosen so the gate could be measured with zero data risk anywhere.
Baseline: container `9473b6e4a18f`, `ghcr.io/alam00000/bentopdf:v2.8.6`, digest
`sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3`, started `04:05:50Z`.
| # | Variant | Result | Discriminator |
|---|---|---|---|
| **1d** | **negative control** — file NOT edited, then `restart` | **no change** | container id, image id, digest and `StartedAt` all IDENTICAL. `up -d` saw no change and did not even recreate. |
| **1a** | `restart`, target image **ABSENT** from the local store | **UPGRADED, and it pulled** | new container `93e9db74e97f`, image `v2.8.5`, digest `sha256:2d867aac…`, and `v2.8.5` **appeared in `docker images`** where it had not been. **18.3 s.** |
| **1b** | `restart`, target image **ALREADY PRESENT** | **UPGRADED, no pull** | new container `222f8178dbd0`, image id unchanged from baseline `4baadf01bd68`. **0.5 s.** |
| **1c** | **boot recovery** — hard guest reset, app left running | **NO upgrade** | SAME container `222f8178dbd0` restored by Docker at `17:47:45Z` on `v2.8.6` while the file said `v2.8.5`. |
| **1c-ii** | boot recovery where the app did **not** come back | **UPGRADED, unattended** | new container `aa443836a4cd` on `v2.8.5` at `17:55:44Z`. Nobody pressed anything. |
**The 18.3 s vs 0.5 s split is the quantitative discriminator** and it is worth more than the tags:
the restart path contains no pull step in the source, yet variant 1a spent 18 seconds and ended with a
new image in the local store. `up -d` pulled it.
Quoted live, variant 1a (`phase1-1a-controller-log.txt`):
```
17:35:45 router.go:566: [INFO] [api] restart requested for stack: bentopdf
17:35:45 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
17:36:03 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 18.3s)
17:36:07 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running
```
Quoted live, variant 1c-ii — **the unattended upgrade, attributed to the exact symbol**
(`phase1-1cii-orphan.txt`):
```
17:55:44 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 10s — sweeping
17:55:44 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]
17:55:44 manager.go:1039: [INFO] [stacks] Starting stack: bentopdf
17:55:44 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
17:55:48 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running
```
**1c's negative result is a POSITIVE observable, not an absent log line** — the reconciler says so in
its own words, which is exactly what it was built to do:
```
17:48:34 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
```
### What 1c-ii staged, stated plainly
`docker rm -f bentopdf` removed the container so it could not be restored by Docker's restart policy,
which is the shape a power cut leaves when an app does not come back. `desired_state: running` was left
untouched in `app.yaml`. The controller was then restarted, because that is what runs the boot
reconciler. **Nothing else was staged**, and the reconciler selected the app on its own.
### Why 1c could not use a hand-edited file, and how that was handled
At controller start the initial catalog sync runs immediately (`sync.go:98`, `17:47:49Z`) while the
boot reconciler waits out `bootReconcileSettle` = 5 s plus a settle window (`17:48:34Z`). **The sync
wins by ~45 s**, so a hand-edited compose would have been overwritten before the reconciler ever saw
it, and 1c would have measured nothing. 1c and 1c-ii were therefore run inside Phase 2's window, where
the **catalog itself** carried the new tag — which is also the faithful production shape.
### The design intent is already stated in the source, and it is not hidden
`Manager.RestartStack` (`internal/stacks/manager.go:1133`) carries this comment:
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
> changes (new images, healthchecks) are picked up."*
**So the restart behaviour was chosen, deliberately, and written down.** What is NOT written down
anywhere is the consequence once the catalog syncer moves the file underneath a deployed app. That gap
is R-438, and whether the choice should extend to the unattended paths is the operator's ruling.
---
## 3. Phase 2 — the sync overwrite, confirmed live
A real tag change (`bentopdf` `v2.8.6` → `v2.8.5`) was pushed to the catalog `main` at **17:38:22Z**
(`214d448`) and left to travel the real 15-minute cycle. No hand-edit; the catalog was the only input.
| time (UTC) | file on demo-hp | running container |
|---|---|---|
| 17:38:35 (pre-sync) | `v2.8.6` | `v2.8.6` |
| 17:45:15 | `v2.8.6` | `v2.8.6` |
| **17:45:37 (post-sync)** | **`v2.8.5`** | **`v2.8.6`** |
The sync's own line, and the file's mtime, agree to the second:
```
17:45:17 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
mtime: 2026-09-01 17:45:17.646017366 +0000 /opt/docker/stacks/bentopdf/docker-compose.yml
container started: 2026-09-01T17:36:35Z (unchanged)
```
`Syncer.copyTemplates` (`internal/sync/sync.go:319`) has **no deployed check of any kind**. Its only
guard is a sha256 content compare in `copyIfChanged`, and its only exclusion is `app.yaml`. The
post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the sync does not
restart anything**, which is why the file and the container can disagree indefinitely.
### Was the customer told? **No — and the page does not even show a version.**
Searched with ASCII fragments and both controls, per the standing rule.
| page | `BentoPDF` (positive control) | `zzz-never-present` (negative control) | `v2.8.5` | `v2.8.6` |
|---|---|---|---|---|
| `/apps/bentopdf` (40 360 B) | 4 | 0 | **0** | **0** |
| `/stacks` (140 883 B) | 1 | 0 | **0** | **0** |
**A correction to my own method, made here rather than buried:** the first pass used
`grep -o "2.8.6"`, where `.` is a regex wildcard, and it reported 2 hits. Re-run with `grep -F` the
count is **0**. The controls are what exposed it. The customer's pages carry **no version string at
all** — not the running one, not the pending one.
No event, no notification and no email were emitted by the sync. The only `Friss…` strings on the
pages are the catalog-sync toast and the `Frissítés` button label.
### The revert, verified on the box
`30bd892` reverted the pin; the sync landed it at **18:10:29Z**
(`[INFO] [sync] Updated bentopdf/docker-compose.yml`, `Sablonok frissítve — frissítve: bentopdf`) and
the file read `v2.8.6` again. **The catalog tree is byte-identical to `29edad9c5bf4`.**
**An unlooked-for second result:** that same sync also overwrote a hand-broken compose file (Phase 3's
`alpine:3.20` edit) with the catalog's version. **So the syncer is also the repair path for a broken
app definition — within 15 minutes, automatically.** The container stays broken until something runs
`up -d`, but the definition heals itself.
---
## 4. Phase 3 — what a failed update looks like to the customer
### 3a — the pull fails. **Loud, honest, and harmless.**
Compose pointed at `ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist`, then `POST .../update`:
```
HTTP 500
{"ok":false,"error":"pulling images for bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ... Error manifest unknown\nError response from daemon: manifest unknown"}
```
**The app was untouched:** same container `aa443836a4cd`, `v2.8.5`, `status=running`, `RestartCount=0`.
`UpdateStack` returns after the failed `pull` and never reaches `up -d`.
**3a-ii, the landmine that was not one.** The failed update leaves the file pointing at an image that
does not exist, while the page shows a green healthy app with a Restart button beside the Update
button. Pressing Restart in that state was measured: **HTTP 500, and the app still ran.** `compose up
-d` resolves every image before it touches a container, so a bad reference fails before anything stops.
**This is the reassuring half of Phase 3 and it should be said as plainly as the bad half.**
**What the customer is shown, though, is raw English Docker output** — `exit code 1`, `stderr:`,
`manifest unknown`, `Error response from daemon` — on a product whose every other error string is
Hungarian.
### 3b — the pull succeeds and the app does not. **This is the one that matters.**
Compose pointed at `alpine:3.20` — a real image that pulls cleanly and then exits at once, so
`restart: unless-stopped` puts it in a crash loop. This is the shape of a real upstream image whose
configuration the template can no longer supply, and the catalog's own history contains exactly that
case (`7350cd9`, *"wger: revert 2.6 -> 2.3 (2.6 needs a full DB config the template cannot supply)"*).
```
POST /api/stacks/bentopdf/update
HTTP 200
{"ok":true,"message":"Stack bentopdf update completed"}
18:00:47 manager.go:1195: [INFO] [stacks] Stack bentopdf updated successfully (took 3.5s)
```
Reality 37 seconds later: `status=restarting`, `RestartCount=9`, `alpine:3.20`.
**The controller's own post-start line told the truth** — and it ran *after* the API had already
answered `ok:true`:
```
18:00:50 manager.go:1403: [INFO] [stacks] bentopdf alpine:3.20 restarting Restarting (0) Less than a second ago
```
**What the customer's page said** (`/stacks`, live HTML):
> `BentoPDF` · `pdf.enkisfelhom.hu` · **„URL nem elérhető – útvonal nincs publikálva"** ·
> badge **„Újraindítás…"** · `Restarting (0) 15 seconds ago` ·
> buttons: **`Frissítés` `Újraindítás` `Leállítás` `Naplók` `Részletek`**
**„Újraindítás…" reads as transient, not as failure.** Nothing on the page says the update broke the
app. `isOperationalState` (`internal/web/funcmap.go:90`) counts `StateRestarting` and `StateDegraded`
as operational, so the full button row — including the green `Frissítés` — is rendered over a
crash-looping app.
**Is there a route back? Not from the page.** Every button offered re-runs the same broken definition
or stops the app. The compose file is not customer-editable. The routes that exist are (a) the catalog
being corrected, which then heals the file within 15 minutes, or (b) a restore, which has its own
problem — see §7.
### The alarm DOES fire — 5 minutes 16 seconds later, and by a different road
This was measured to a positive observable rather than inferred from silence:
```
18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
```
Update was at `18:00:43`; the alarm at `18:05:59`. The delay is `crashLoopAfter = 5 * time.Minute`
(`internal/stacks/manager.go`), and it is deliberate and well-argued in its own comment — a shorter
threshold would alarm on every routine deploy. **The severity word is `warning`, which is inside the
hub's exact vocabulary**, so this one really does reach the customer (R-328/R-329 class avoided).
**So the honest summary of 3b is not "silently broken".** It is: *the button lied at the moment it was
pressed, the page then described a failure as a restart, and the truth arrived five minutes later
through the dead-app alarm rather than through the update the customer actually performed.*
---
## 5. Phase 4 — how far behind is the real fleet
**Read from the running container, never from the file** — the file is the thing that has already
moved. Both tag and digest recorded.
### demo-hp — **zero drift by tag.** Nine deployed apps, every running tag equal to the catalog pin.
This is not luck: the box was reinstalled 2026-08-21 and its apps were deployed ~13 h before this run.
It is a *young* box, and it is the reason Phase 5 could not be priced on it (§6).
### demo-felhom — one deployed app, `opengist`, running `ghcr.io/thomiceli/opengist:1.13` = catalog pin.
### But "up to date by tag" is not up to date. **Two floating pins have already moved upstream.**
Running digests compared against what the registry serves for the same tag today, with fully-pinned
tags as the control:
| image | kind | verdict |
|---|---|---|
| `postgres:16-alpine` | floating | SAME |
| `redis:7-alpine` | floating | SAME |
| `mariadb:11.6` | floating | SAME |
| `ghcr.io/thomiceli/opengist:1.13` | floating | SAME |
| **`mariadb:11.4`** | **floating** | **MOVED** — running `sha256:4f1d8d20…`, upstream now `sha256:611a2fcc…` |
| **`mariadb:12.3`** | **floating** | **MOVED** — running `sha256:a02fe89c…`, upstream now `sha256:dd9b303a…` |
| `rommapp/romm:5.0.0` | **pinned (CONTROL)** | SAME |
| `privatebin/pdo:2.0.5` | **pinned (CONTROL)** | SAME |
**Both controls held and both positives are floating tags.** So on a box with no visible drift at all,
pressing `Frissítés` today would silently swap the **database engine build** under `romm` (MariaDB
11.4) and `bookstack` (MariaDB 12.3) — with no catalog change, no version change on any screen, and no
record anywhere of what it was before. That is R-440, no longer as an argument but as a measurement.
### Peti's box — **NOT measurable, and the task's premise here was wrong**
The task asks Phase 4 to inspect three boxes. **`runbooks/target-selection.md:161` states plainly:
*"Currently DOWN, no enrolled host. No access route from DooPlex, and nothing here needs one."*** The
hub agrees: the Hosts page lists exactly **two** enrolled hosts (`demo-felhom-8363b5`,
`demo-hp-bb76ea`), and the `peti-felhom` customer shows **status DOWN, last report 48 d ago, controller
0.115.0**.
So there is no running container to read, and **the hub does not record image tags at all** — the
report's container payload carries name, state, CPU and memory, and no image field. The row is
therefore recorded as **UNKNOWN**, not guessed.
**What IS knowable read-only, and it matters:** the box last reported ~2026-07-15 running `rallly` +
`rallly-postgres`. On **2026-07-18 — three days after it went quiet — the catalog moved
`rallly` `3.11.2` → `4.11.1` [MAJOR]** (`e3f3a81`). So a one-major upgrade is queued behind that box's
next boot. **On this spike's own measurements that upgrade will NOT fire on the power-on itself**
(1c: Docker restores the containers on the old image), **but it will fire the moment any app fails to
come back, or anyone presses Restart or Update.** Nothing was touched on that box to establish this;
it is the catalog's git history plus the hub's own record.
---
## 6. Phase 5 — what a pre-update copy would cost
### The existing safety machinery is DATABASE-ONLY, and that is the headline
`Manager.writeSafetyDump` (`internal/backup/offbox_reconstitute.go:207`) discovers the app's databases
and dumps each one. **An app with no database gets nothing at all** — `len(mine) == 0` returns an empty
set. There is no file-level safety copy on the R-361 path.
### The database half is nearly free. Measured on demo-hp:
Every `.sql` dump on the box, including the real `pre-restore-*` undo copies from the August restore
work: **48 KB – 395 KB**. Largest is `paperless-ngx-postgres.sql` at 395 065 bytes.
### The file half is the cost, and demo-hp cannot price it. Stated, not papered over.
The box holds a few MB of app *content*; the bulk is engine data directories:
| app | biggest components | total |
|---|---|---|
| kimai | `kimai_db_data` 166 M + `kimai_var` 56 M | **≈ 222 MB** |
| romm | `romm_db_data` 166 M + `romm_redis_data` 25 M + roms 1.1 M | **≈ 192 MB** |
| bookstack | `bookstack_db_data` 166 M + `bookstack_config` 6.8 M | **≈ 173 MB** |
**These are not customer-scale numbers** — 166 MB is a fresh MariaDB's preallocated files, not
anyone's data. So rather than over-claim from a young box, the arc's already-measured figures are
cited:
> **`CAMPAIGN-10-two-storage-soak-2026-07-31.md` §6, 66 restores plus an M-band point 327× larger:
> backup ≈ 29 s + 17.4 s/GB, and a DB-backed app's recovery unit is 1.90× its data**
> (21.1 GB of data produced a 40.2 GB unit).
Applying that to demo-hp's three heaviest: **≈ 32–33 s each**, unit ≈ 330–420 MB. Trivial.
**And that is exactly what makes the ceiling the real finding.** For a large catalog app — Immich,
Nextcloud, Plex — the same arithmetic gives, for 100 GB of data, **≈ 29 minutes and a ~190 GB copy**,
against a default appliance whose `/mnt/sys_drive` ships at **20 GB**. **A pre-update copy of a large
app does not fit on a default box, and no amount of tuning the copy changes that.** Whatever is
designed here has to answer that before it answers anything else. Where the copy lives, and whether it
is a copy at all or a snapshot, is a design decision and is not proposed here.
---
## 7. Phase 6 — can app data be rolled back at all? **NO.**
Run on operator confirmation, on a throwaway Nextcloud on `demo-hp`. Nothing else on any box was
involved; the app was created, used and destroyed inside this phase.
**Seeded on 31.0.14.1** (`occ status`), two independent markers — a database-backed system config value
and a real file on disk:
```
System config value spike_marker set to string SPIKE-SEEDED-ON-31.0.14-2026-09-01
/var/www/html/data/admin/files/spike-marker.txt = SPIKE-FILE-CONTENT-31.0.14-2026-09-01
occ files:scan admin → 5 Folders, 53 Files, 2 Updated, 0 Errors
```
### 6a — the 3-major jump (31.0.14 → 34.0.1). **Refused by the app, reported as success by the product.**
```
POST /api/stacks/nextcloud/update → HTTP 200 {"ok":true,"message":"Stack nextcloud update completed"}
```
The container then crash-looped, saying:
> **`Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.`**
> **`It is only possible to upgrade one major version at a time.`**
**This is R-40, measured.** The catalog really does carry this jump today: `5e2c1ae`, 2026-07-18,
*"nextcloud: 31.0.14-apache -> 34.0.1-apache [MAJOR]"*.
**It was fully recoverable** — putting `31.0.14` back gave a healthy app in 9 seconds with both markers
intact. **But that is only because nothing migrated.** The refusal is Nextcloud's own safety net doing
its job, and it is why 6a is the *easy* case.
### 6b — a migration that actually runs (31.0.14 → 32.0.9, one major). **It ran.**
```
Initializing nextcloud 32.0.9.2 ...
Upgrading nextcloud from 31.0.14.1 ...
Updated database
Updated <dav> to 1.34.2 · <files> to 2.4.0 · <files_sharing> to 1.24.1 · … (20+ apps)
```
Healthy in 31 s on `32.0.9.2`, both markers still readable.
### 6c — THE ROLLBACK ATTEMPT. **Refused. The old image will not start on migrated data.**
```
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker
image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the
newest image version?
```
Crash loop, indefinitely.
### 6d — positive control: **the data is not destroyed, only the downgrade is blocked**
Putting `32.0.9` back gave a healthy app in 46 s, `occ status` `32.0.9.2`, and **both markers read back
byte-identical to what was seeded on 31**. So the failure in 6c is a refusal, not corruption — which is
the distinction that decides the remedy.
### What this settles
**Putting the old image tag back is not a rollback and must not be described as one.** The only route
back from a migration that has run is **restoring the DATA from a copy taken before the update** —
which is precisely the thing the update path does not take (§2 of the task, confirmed by reading
`Manager.UpdateStack`, and confirmed live: no dump, no copy, no hold, in any of the six updates run
here).
---
## 8. The exact symbols — found by reading, as required
Every path that ends in `compose up -d` on a customer's stack folder, with the enclosing function.
| what brings the app back | symbol | line |
|---|---|---|
| the four compose callers | `Manager.StartStack` / `StopStack` / `RestartStack` / `UpdateStack` | `internal/stacks/manager.go:1029 / 1106 / 1133 / 1170` |
| partial start (R-47 DB-only window) | `Manager.StartStackServices` | `internal/stacks/manager.go:1082` |
| **the boot reconciler** | `Reconciler.Run` → `r.stacks.StartStack(name)` | `internal/bootrecon/bootrecon.go:223` → **`:269`** |
| its scheduler | `runBootReconcile` → `bootReconcileFn` | `cmd/controller/main.go:2127` → `:2172`, called at `:450` |
| **the app-stop guard** | `AppStopGuard.Recover` → `g.starter.StartStack(name)` | `internal/backup/appstop_marker.go:264` → **`:283`** |
| its hold-aware wrapper | `gatedAppStopStarter.StartStack` | `cmd/controller/main.go:2008` |
| **the drive-return gate** | `Server.restartStacks` → `s.stackMgr.StartStack(name)` | `internal/web/intermediary.go:220` → **`:222`** |
| guest-boot change handler | `Server.processGuestBootChange` | `internal/web/intermediary.go:395` → `:458` |
| quiesce restart-after-backup | `Loop.restartAll` | `internal/quiesce/quiesce.go:730` → `:733` |
| off-site reconstitution | `Manager.ReconstituteFromOffsite` | `internal/backup/offbox_reconstitute.go:480` → `:692` |
| restore from unit / local / tier-2 | `restore_unit.go:381`, `restore.go:76`, `tier2_restore.go:418`, `backup.go:895` | — |
| `.fab` export / import | `appexport/export.go:286`, `appexport/restore.go:461` | — |
| integrations | `onlyoffice_filebrowser.go:61, :96` (`RestartStack`) | — |
**CORRECTION TO THE TASK: not five other paths — thirteen.** Excluding the three API actions
(`start`/`restart`/`update`) and excluding interface declarations and adapters, there are **13 call
sites across 9 files** that call `StartStack` or `RestartStack`, and every one of them ends in
`docker compose up -d` against the live compose file.
---
## 9. R-439, confirmed by reading and refined
`Router.actionStack` (`internal/api/router.go:565`) checks the hold under
`if action == "start" || action == "restart"` — **`update` is absent** and falls through to
`UpdateStack`. The design intent is stated in `RestoreHoldFor`'s own comment
(`internal/backup/offbox_reconstitute.go:323`):
> *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the
> boot reconciler — because a hold that only one path honours is not a hold."*
**The severity argument in the task is correct, but for a narrower reason than it states.** The task
says the UI hides `Frissítés` unless the app is operational, and a held app is stopped. That is right —
but `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded`
as operational too**, which was observed live in Phase 3b: the green `Frissítés` button was rendered
over a crash-looping app. So the button is hidden specifically because a held app is `StateStopped`,
not because broken apps hide it. **The conclusion (LOW, not customer-reachable) survives; the reason
needs stating precisely, and the fix needs a test pinning it or the comment stays a wish.**
---
## 10. Every claim in the task that turned out to be wrong, named
1. **"a power cut on a sleeping customer's box is an unattended three-major-version upgrade"** —
**NOT AS STATED.** Measured: a hard guest reset upgraded nothing, because Docker restored the
containers itself and the reconciler had no orphan. The unattended upgrade is real but needs the
narrower precondition *"and the app did not come back"*. **The exposure is smaller than the operator
page claims, and saying so is more useful than leaving the scarier version standing.**
2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites across nine
files (§8).
3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and has **no
access route from DooPlex** by the project's own runbook. Two boxes were measured live; Peti's row
is UNKNOWN, with what *is* knowable recorded from the hub and the catalog history (§5).
4. **"R-440 — 23 catalog image pins float"** — **CONFIRMED EXACTLY** (79 `image:` lines, 53 apps, 66
distinct; 23 with no patch component). One arguable 24th is recorded rather than rounded away:
`ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly and
leaves the PostgreSQL patch floating.
5. **The catalog history numbers** — spot-checked and **all correct**: 153 commits touching
`templates/`, 53 apps, 0 `.felhom.yml` carrying any upgrade metadata (the only `upgrade` matches are
prose about STARTTLS), and all four multi-major bumps confirmed on 2026-07-18 —
nextcloud `5e2c1ae`, grafana `b789acc`, calcom `147cee7`, vikunja `3fa63cd`.
6. **"The Update button … takes no safety copy, cannot undo itself, does not stop the app if it goes
wrong"** — **CONFIRMED**, by reading and across six live updates.
7. **A methodological correction of my own, not the task's:** the first customer-page search used
`grep -o "2.8.6"` and the unescaped `.` produced two false hits. Re-run with `grep -F`: zero. The
controls caught it (§3).
---
## 11. What was NOT measured
- **Whether the drive-return gate and `AppStopGuard.Recover` upgrade in practice.** Both were located
by reading (§8) and both call `StartStack`, which is the same function variant 1c-ii measured
upgrading an app. **The mechanism is measured; these two specific entry points were not exercised
live.** Naming them as read-not-measured is deliberate.
- **Whether `RecreateStackDefinitionFromUnit`'s rollback is really undone by the sync in a live
restore** — see §12, which states which half is measured and which is read.
- **Any behaviour on Peti's box** (§5).
---
## 12. Observations — noticed, documented, not acted on
**O1 — the restore path and the catalog sync disagree about the image, and the sync wins.**
`stackAdapter.RecreateStackDefinitionFromUnit` (`cmd/controller/main.go:2570`) writes the recovery
unit's captured `docker-compose.yml` — **carrying the OLD image pin** — straight into the live stack
dir (`os.WriteFile(filepath.Join(stackDir, fname), data, 0644)`), and `restore_unit.go:317` says so:
*"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But
`copyIfChanged` overwrites any file whose content differs from the catalog, on the next 15-minute tick.
**MEASURED:** a locally-modified compose (Phase 3's `alpine:3.20`) was overwritten by the sync at
`18:10:29Z`. **READ, not measured:** that the restore writes to that same path. So a restore's
image-level rollback has a **≤15-minute half-life**, and then the next `up -d` from any source
re-applies the catalog pin. Filed as **R-441**.
**O2 — `remove_hdd_data: true` is inert on this box, and the response says neither removed nor
preserved.** Removing the Phase 6 Nextcloud with `{"remove_hdd_data":true,"remove_backups":true}`
returned `HTTP 200` with `"hdd_paths_removed":null,"hdd_paths_preserved":null`, and left **128 MB** at
`/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **Root cause, with controls:** `Paths.HDDPath`
(`internal/config/config.go:117`) has **no default** — only an env override at `:403` — and demo-hp's
`controller.yaml` `paths:` block contains only `data_dir`, `stacks_dir`, `system_data_path`. The
container has no `FELHOM_PATHS_*` variable at all (measured, count 0). So `cfg.Paths.HDDPath == ""` and
`ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its first line —
`[INFO] found 0 HDD mounts` — for a compose that plainly contains
`- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal also
no-op'd:** `[WARN] Refusing to remove backup path outside expected directory:
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. Filed as **R-442**.
**O3 — an update can report success over a broken app, and the truth arrives by another road 5 minutes
later.** §4. Filed as **R-443**.
**O4 — demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** This run
added ~1.05 GiB that `local-lvm` did not reclaim; `fstrim` inside the unprivileged container is refused
(`FITRIM ioctl failed: Operation not permitted`), and `pct fstrim 9201` from the host then trimmed
30.2 GiB + 57 GiB and took `local-lvm` from **70.91% → 26.78%** — i.e. **23.8 GB below this run's own
starting point**. Nothing runs `pct fstrim` on the fleet. Filed as **R-444**.
**O5 — the hub keeps app telemetry for an app that no longer exists anywhere.** The throwaway Nextcloud
now sets a **fleet-wide** `Suggested Limit (P95×1.2) = 352 MB` for Nextcloud, from ~15 minutes of a
crash-looping instance, plus three MariaDB `io_uring` "Known Issues" attributed to demo-hp. **Retained
deliberately, not cleared** — see §13. Filed as **R-445**.
**O6 — NOT-A-FINDING: the sync's debug hash line cannot show what changed.** `logFileHashes`
(`internal/sync/sync.go:386`) reads the destination *after* the write, so it prints
`src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word "changed".
Harmless (DEBUG only, and the `Updated <app>/<file>` INFO line above it carries the fact), but it
cannot serve the purpose its name implies. Not filed; recorded here so the next person does not trust
it.
---
## 13. Teardown — all three layers
**Layer 1 — the machine.**
- `bentopdf` **restored to its catalog tag** `v2.8.6`, digest `sha256:eaeea1e447205a79…` — **byte-identical
to the run's baseline**. Container `02c80375fcba`, running, healthy.
- The throwaway `nextcloud` stack **removed** via `POST /api/stacks/nextcloud/remove` (after the
required stop): containers gone, all three named volumes gone, `app.yaml` gone. The stack dir holds
only the catalog template (`docker-compose.yml`, `.felhom.yml`), i.e. not deployed.
- The **128 MB the product did not remove** (O2) was deleted by hand, together with the nextcloud
backup dirs on both drives. `find /mnt -iname "*nextcloud*"` returns nothing.
- **Four images this run pulled were removed by targeted `docker rmi`** — `nextcloud:31.0.14-apache`,
`32.0.9-apache`, `34.0.1-apache`, `bentopdf:v2.8.5`, plus `alpine:3.20`. **No `prune` of any kind was
run, anywhere.**
- **No guest was created.** Guest 9201 was hard-reset once, by design (variant 1c), and came back with
all nine apps.
**Layer 2 — the host.** `pvesm status`, before → after:
| pool | before | after trim | note |
|---|---|---|---|
| `local` | 44.42% | 44.44% | unchanged in substance |
| `local-lvm` | **68.97%** (38 959 729 KiB) | **26.78%** (15 127 469 KiB) | **the run's ~1.05 GiB was returned, and `pct fstrim 9201` reclaimed 23.8 GB more that predated this run** (O4) |
Guest: `/` 957 M used (baseline 957 M), `/mnt/sys_drive` 12 G / 18%, `/mnt/felhom-drives/hdd_1` 5.5 G —
all back to their pre-run values.
**Layer 3 — the hub. This run provisioned NOTHING.**
- **No customer record and no appliance record was created.** The customers list is unchanged at five
rows: `demo-felhom`, `demo-hp`, `drill-r50`, `peti-felhom`, `tester-1` — identical to the list read at
the start of the run. The existing `demo-hp` customer was used throughout.
- **What the run DID create is events** — `app_start_failed` (warning) for BentoPDF at `18:05:59Z`,
plus deploy/remove events for the throwaway Nextcloud. **These are retained deliberately:** the event
log is an append-only record and deleting from it to tidy up a test would damage the very surface
this project relies on for history.
- **One piece of residue is retained rather than cleared, and the reason is stated:** the Nextcloud app
telemetry row (O5). The hub offers `POST /apps/nextcloud/reset-telemetry`, whose own confirm text is
**"Delete all telemetry data for nextcloud? This cannot be undone."** That is an irreversible write on
the operator's own surface, and the operator authorised Phase 6, not this. **The exact command is
recorded here so it is a one-line decision:**
`curl -u ":$HUB_PW" -X POST http://<hub-clusterIP>:8080/apps/nextcloud/reset-telemetry`
**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified.** Peti's box was never contacted.
---
## 14. What goes to the operator
**One decision, and it is not a bug report.** §2 shows the restart behaviour was *chosen* and is stated
in the source. §7 shows the word "rollback" does not describe anything this product can do. The two
together mean the safety question is not "fix the Update button" — it is **where the safety belongs**,
and that is in `STATUS.md`, phrased as one answerable question with what happens if nothing is done.
**No design is proposed here, deliberately.** Four production designs in this project were specced
against unvalidated mechanisms and all four were wrong; this spike exists so the fifth is not.
@@ -0,0 +1,12 @@
### date: 2026-09-01T17:35:06Z
### container:
id=9473b6e4a18f5f6faf0640357c923a347a478029de38b979d0443332a4b9901a image=ghcr.io/alam00000/bentopdf:v2.8.6 imageid=sha256:4baadf01bd68b6fe61fd2c74612e800209907e2e1c16e58b770479af2cc3ae22 started=2026-09-01T04:05:50.905652046Z
### repodigest:
ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
### compose image line:
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
### compose sha256:
39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1 /opt/docker/stacks/bentopdf/docker-compose.yml
### local bentopdf images:
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68 2 months ago
@@ -0,0 +1,13 @@
2026/09/01 17:35:45 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/bentopdf/restart
2026/09/01 17:35:45 router.go:80: [DEBUG] [api] POST /api/stacks/bentopdf/restart (path=/stacks/bentopdf/restart)
2026/09/01 17:35:45 router.go:566: [INFO] [api] restart requested for stack: bentopdf
2026/09/01 17:35:45 router.go:80: [DEBUG] [api] actionStack: action=restart name=bentopdf
2026/09/01 17:35:45 manager.go:1140: [DEBUG] [stacks] RestartStack bentopdf: current state=running deployed=true containers=1
2026/09/01 17:35:45 manager.go:1143: [INFO] [stacks] Restarting stack: bentopdf
2026/09/01 17:35:45 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
2026/09/01 17:36:03 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 18.3s)
2026/09/01 17:36:04 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=bentopdf, waiting 5s for state refresh
2026/09/01 17:36:06 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
2026/09/01 17:36:07 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
2026/09/01 17:36:07 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running Up 3 seconds (health: starting)
2026/09/01 17:36:09 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=bentopdf integrationsFound=0
@@ -0,0 +1,18 @@
### 1a — hand-edit compose to v2.8.5 (ABSENT from the local docker store), then RESTART
file now:
11: image: ghcr.io/alam00000/bentopdf:v2.8.5
local images (v2.8.5 must be ABSENT):
ghcr.io/alam00000/bentopdf:v2.8.6
container BEFORE:
id=9473b6e4a18f5f6faf0640357c923a347a478029de38b979d0443332a4b9901a image=ghcr.io/alam00000/bentopdf:v2.8.6 started=2026-09-01T04:05:50.905652046Z
t0=2026-09-01T17:35:45Z
### POST /api/stacks/bentopdf/restart
{"ok":true,"message":"Stack bentopdf restart completed"}
HTTP=200
### AFTER 1a, date: 2026-09-01T17:36:16Z
id=93e9db74e97f1bb354108cc553b9438f52f7a4a36201bc8c7d74d24838ae1d3d image=ghcr.io/alam00000/bentopdf:v2.8.5 imageid=sha256:ab6eeed172e49347fa0b9d8471772ffd0b3df648ceaa83d3054f3d80edf3fae9 started=2026-09-01T17:36:03.775567517Z status=running
digest=ghcr.io/alam00000/bentopdf@sha256:2d867aacb8ab5b196d00ee86944b1899d09d72df355384c5e15cf974737963a0
local images now:
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68 2 months ago
ghcr.io/alam00000/bentopdf:v2.8.5 ab6eeed172e4 3 months ago
@@ -0,0 +1,18 @@
### 1b — hand-edit compose to v2.8.6 (image ALREADY PRESENT locally), then RESTART
file now:
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
local images (v2.8.6 must be PRESENT):
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68
ghcr.io/alam00000/bentopdf:v2.8.5 ab6eeed172e4
container BEFORE:
id=93e9db74e97f1bb354108cc553b9438f52f7a4a36201bc8c7d74d24838ae1d3d image=ghcr.io/alam00000/bentopdf:v2.8.5 started=2026-09-01T17:36:03.775567517Z
t0=2026-09-01T17:36:35Z
### POST /api/stacks/bentopdf/restart
{"ok":true,"message":"Stack bentopdf restart completed"}
HTTP=200
### AFTER 1b, date: 2026-09-01T17:36:42Z
id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 image=ghcr.io/alam00000/bentopdf:v2.8.6 imageid=sha256:4baadf01bd68b6fe61fd2c74612e800209907e2e1c16e58b770479af2cc3ae22 started=2026-09-01T17:36:35.916981126Z status=running
digest=ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68
ghcr.io/alam00000/bentopdf:v2.8.5 ab6eeed172e4
@@ -0,0 +1,50 @@
### 1c PRE-RESET state (catalog=v2.8.5, file on box=v2.8.5, container=v2.8.6)
2026-09-01T17:47:20Z
image: ghcr.io/alam00000/bentopdf:v2.8.5
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:36:35.916981126Z
bentopdf bookstack bookstack-db calibre-web cloudflared docmost docmost-postgres docmost-redis felhom-controller filebrowser kimai kimai-db opengist paperless-postgres paperless-redis paperless-webserver privatebin romm romm-db romm-redis traefik
### 1c — HARD RESET of guest 9201 on demo-hp (pct stop = ungraceful, then start)
2026-09-01T17:47:27Z
stop_rc=0
start_rc=0
VMID Status Lock Name
9201 running demo-hp
2026-09-01T17:47:41Z
2026-09-01T17:47:48Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:48:05Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:48:22Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:48:39Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:48:56Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:49:12Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:49:28Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:49:45Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:50:01Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:50:18Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:50:34Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:50:51Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:51:07Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:51:23Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:51:40Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:51:56Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:52:13Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:52:29Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:52:46Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:53:02Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:53:18Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:53:35Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:53:51Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:54:08Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
2026-09-01T17:54:24Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
### 1c — the boot reconciler's OWN verdict line (positive observable, not an absent log)
2026/09/01 17:47:59 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
2026/09/01 17:48:04 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
2026/09/01 17:48:09 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
2026/09/01 17:48:14 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
2026/09/01 17:48:24 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
2026/09/01 17:48:34 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 40s (3 identical samples 5s apart) — sweeping
2026/09/01 17:48:34 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
### did the initial sync run before it, and did it rewrite bentopdf?
2026/09/01 17:47:49 sync.go:98: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
2026/09/01 17:47:49 sync.go:189: [INFO] [sync] Starting catalog sync
2026/09/01 17:47:50 sync.go:103: [INFO] [sync] Initial sync: Sablonok naprakészek — nincs változás
@@ -0,0 +1,46 @@
### 1c-ii — SIMULATED boot orphan. Staged explicitly: 'docker rm -f bentopdf' removes the
### container so it does NOT come back on its own (the shape a power cut leaves when docker
### cannot restore an app). Desired state stays 'running'. Then the CONTROLLER is restarted,
### which is what runs the boot reconciler. Nothing else is staged.
2026-09-01T17:55:18Z
file says:
image: ghcr.io/alam00000/bentopdf:v2.8.5
container BEFORE:
ghcr.io/alam00000/bentopdf:v2.8.6 222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2
desired state on disk:
desired_state: running
REMOVED bentopdf container
(empty above = gone)
### restarting the controller so the boot reconciler runs
2026-09-01T17:55:27Z
felhom-controller
2026-09-01T17:55:30Z bentopdf: STILL ABSENT
2026-09-01T17:55:46Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:56:02Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:56:19Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:56:35Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:56:52Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:57:08Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:57:25Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:57:41Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:57:57Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:58:14Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
2026-09-01T17:58:30Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
### which component started bentopdf at 17:55:44?
2026/09/01 17:55:29 deploy.go:1051: [DEBUG] [stacks] InjectMissingFields: checking stack bentopdf — 2 deploy fields, 2 existing env vars
2026/09/01 17:55:29 backup.go:450: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin]
2026/09/01 17:55:29 sync.go:371: [DEBUG] [sync] bentopdf/docker-compose.yml: hash match, skipped
2026/09/01 17:55:29 sync.go:371: [DEBUG] [sync] bentopdf/.felhom.yml: hash match, skipped
2026/09/01 17:55:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bentopdf/docker-compose.yml
2026/09/01 17:55:44 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 10s (3 identical samples 5s apart) — sweeping
2026/09/01 17:55:44 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf] — up to 2 attempt(s)
2026/09/01 17:55:44 manager.go:1036: [DEBUG] [stacks] StartStack bentopdf: current state=stopped deployed=true
2026/09/01 17:55:44 manager.go:1039: [INFO] [stacks] Starting stack: bentopdf
2026/09/01 17:55:44 manager.go:1046: [DEBUG] [stacks] StartStack bentopdf: prepared 8 env vars for compose
2026/09/01 17:55:44 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
2026/09/01 17:55:45 manager.go:1054: [INFO] [stacks] Stack bentopdf started successfully (took 0.3s)
2026/09/01 17:55:45 bootrecon.go:274: [INFO] [bootrecon] Boot reconciliation attempt 1/2: started "bentopdf" (took 0.4s)
2026/09/01 17:55:45 bootrecon.go:303: [INFO] [bootrecon] Boot reconciliation complete: 1 app(s) recovered in 1 attempt(s): [bentopdf]
2026/09/01 17:55:48 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
2026/09/01 17:55:48 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
2026/09/01 17:55:48 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running Up 3 seconds (health: starting)
@@ -0,0 +1,11 @@
=== file immediately BEFORE the POST ===
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1 /opt/docker/stacks/bentopdf/docker-compose.yml
=== POST /api/stacks/bentopdf/restart ===
{"ok":true,"message":"Stack bentopdf restart completed"}
HTTP=200
### AFTER 1d, date: 2026-09-01T17:35:31Z
id=9473b6e4a18f5f6faf0640357c923a347a478029de38b979d0443332a4b9901a image=ghcr.io/alam00000/bentopdf:v2.8.6 imageid=sha256:4baadf01bd68b6fe61fd2c74612e800209907e2e1c16e58b770479af2cc3ae22 started=2026-09-01T04:05:50.905652046Z
digest=ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68
@@ -0,0 +1,40 @@
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=true composePath=/opt/docker/stacks/romm/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=false composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml
2026/09/01 17:36:27 healthprobe.go:153: [DEBUG] Health probe bentopdf: HTTP GET :8080/ → 200 (5ms)
2026/09/01 17:36:35 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/bentopdf/restart
2026/09/01 17:36:35 router.go:80: [DEBUG] [api] POST /api/stacks/bentopdf/restart (path=/stacks/bentopdf/restart)
2026/09/01 17:36:35 router.go:566: [INFO] [api] restart requested for stack: bentopdf
2026/09/01 17:36:35 router.go:80: [DEBUG] [api] actionStack: action=restart name=bentopdf
2026/09/01 17:36:35 manager.go:1140: [DEBUG] [stacks] RestartStack bentopdf: current state=running deployed=true containers=1
2026/09/01 17:36:35 manager.go:1143: [INFO] [stacks] Restarting stack: bentopdf
2026/09/01 17:36:35 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, SUBDOMAIN, IMPORT_PATH] (8 app + 0 system)
2026/09/01 17:36:35 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
2026/09/01 17:36:36 manager.go:1340: [DEBUG] Command completed: docker compose up -d (took 0.5s)
2026/09/01 17:36:36 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 0.5s)
2026/09/01 17:36:36 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=bentopdf, waiting 5s for state refresh
2026/09/01 17:36:39 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, SUBDOMAIN, IMPORT_PATH] (8 app + 0 system)
2026/09/01 17:36:39 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
2026/09/01 17:36:39 manager.go:1340: [DEBUG] Command completed: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (took 0.1s)
2026/09/01 17:36:39 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
2026/09/01 17:36:39 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.6 running Up 3 seconds (health: starting)
2026/09/01 17:36:41 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=bentopdf integrationsFound=0
2026/09/01 17:36:47 healthprobe.go:153: [DEBUG] Health probe bentopdf: HTTP GET :8080/ → 200 (6ms)
2026/09/01 17:36:57 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bentopdf — last check 10s ago, effective interval 5m0s, healthy=true
@@ -0,0 +1,4 @@
pre-sync 2026-09-01T17:38:35Z
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1 /opt/docker/stacks/bentopdf/docker-compose.yml
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2
@@ -0,0 +1,12 @@
2026-09-01T17:42:02Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:42:24Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:42:45Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:43:07Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:43:28Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:43:50Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:44:11Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:44:32Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:44:54Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:45:15Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
2026-09-01T17:45:37Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6
>>> SYNC LANDED
@@ -0,0 +1,46 @@
### controller log around the 17:45:17 sync
2026/09/01 17:45:07 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting
2026/09/01 17:45:07 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting
2026/09/01 17:45:07 manager.go:621: [INFO] [stacks] Status refresh: 20 containers across 56 stacks
2026/09/01 17:45:17 sync.go:189: [INFO] [sync] Starting catalog sync
2026/09/01 17:45:17 sync.go:277: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
2026/09/01 17:45:17 sync.go:279: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache
2026/09/01 17:45:17 sync.go:449: [DEBUG] [sync] Running: git fetch --depth 1 origin main
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈4456MB avail≈21442MB
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25429852 → total=25898MB avail=21442MB used=4456MB (17.2%)
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.7GB used=11.9GB avail=53.3GB (17.3%)
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.5GB avail=884.6GB (0.6%)
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.23 1.00 0.81 7/963 133176" → 1m=1.23 5m=1.00 15m=0.81
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 54.2°C (hwmon2)
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo done in 74ms — mem=4456MB/25898MB (17.2%), rootDisk=11.9GB/68.7GB (17.3%), load=1.23/1.00/0.81, temp=54.2°C (hwmon2), cpu=2.7%
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: backup-cache
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job offsite-credential-retry: execution starting
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job backup-cache: execution starting
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: hub-report
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job hub-report: execution starting
2026/09/01 17:45:17 builder.go:37: [INFO] [report] Building system report
2026/09/01 17:45:17 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.232.0, storagePaths=1
2026/09/01 17:45:17 scheduler.go:363: [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
2026/09/01 17:45:17 builder.go:62: [DEBUG] [report] BuildReport: configHash=042b71a3123d... (2989 bytes)
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: system-health
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job system-health: execution starting
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting
2026/09/01 17:45:17 manager.go:621: [INFO] [stacks] Status refresh: 20 containers across 56 stacks
2026/09/01 17:45:17 backup.go:450: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin]
2026/09/01 17:45:17 backup.go:450: [DEBUG] groupStacksByDrive: /mnt/felhom-drives/hdd_1 → [calibre-web, paperless-ngx, romm]
2026/09/01 17:45:17 backup.go:1083: [INFO] [backup] Found 13 DB dump files across drives
### file state on the box
2ebbbda3765b2b216f41ea6203fc7419cabee2fdf9dbdd60f46e2dd97839b30a /opt/docker/stacks/bentopdf/docker-compose.yml
2026-09-01 17:45:17.646017366 +0000 /opt/docker/stacks/bentopdf/docker-compose.yml
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:36:35.916981126Z
@@ -0,0 +1,26 @@
### the sync's own lines (grep: Updated / frissitve / event / notify)
2026/09/01 17:45:17 sync.go:189: [INFO] [sync] Starting catalog sync
2026/09/01 17:45:17 sync.go:277: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
2026/09/01 17:45:17 sync.go:279: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache
2026/09/01 17:45:17 sync.go:449: [DEBUG] [sync] Running: git fetch --depth 1 origin main
2026/09/01 17:45:17 sync.go:449: [DEBUG] [sync] Running: git reset --hard origin/main
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] actualbudget/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] actualbudget/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] adventurelog/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] adventurelog/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] audiobookshelf/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] audiobookshelf/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
2026/09/01 17:45:17 sync.go:398: [DEBUG] [sync] bentopdf/docker-compose.yml: src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] bentopdf/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] bookstack/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] bookstack/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calcom/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calcom/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calibre-web/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calibre-web/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] claper/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] claper/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] code-server/docker-compose.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] code-server/.felhom.yml: hash match, skipped
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] crafty-controller/docker-compose.yml: hash match, skipped
@@ -0,0 +1,21 @@
### LITERAL (grep -F) search — my earlier grep used '.' as a wildcard and over-counted. Redone.
--- page_apps_bentopdf.html ---
BentoPDF hits=4
zzz-never-present hits=0
v2.8.5 hits=0
v2.8.6 hits=0
2.8.5 hits=0
2.8.6 hits=0
Friss hits=1
bentopdf/update hits=0
bentopdf/restart hits=0
--- page_stacks.html ---
BentoPDF hits=1
zzz-never-present hits=0
v2.8.5 hits=0
v2.8.6 hits=0
2.8.5 hits=0
2.8.6 hits=0
Friss hits=10
bentopdf/update hits=0
bentopdf/restart hits=0
@@ -0,0 +1,2 @@
### context of every 2.8.6 / Friss hit on /apps/bentopdf
--- Friss ---
@@ -0,0 +1,9 @@
### Phase 2 revert verification + Phase 3 teardown
2026-09-01T18:12:52Z
### did the catalog revert reach the box's file?
image: ghcr.io/alam00000/bentopdf:v2.8.6
container=alpine:3.20 status=restarting
2026/09/01 18:10:29 sync.go:189: [INFO] [sync] Starting catalog sync
2026/09/01 18:10:29 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
2026/09/01 18:10:29 sync.go:244: [INFO] [sync] Catalog sync complete
2026/09/01 18:10:29 sync.go:116: [INFO] [sync] Periodic sync: Sablonok frissítve — frissítve: bentopdf
@@ -0,0 +1,12 @@
### 3a-ii — the LANDMINE: the update failed and left the file pointing at a tag that
### does not exist. The page still shows a green healthy app with a Restart button.
### What does that Restart button do?
2026-09-01T18:00:23Z
### POST /api/stacks/bentopdf/restart:
{"ok":false,"error":"restarting stack bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Error manifest unknown\nError response from daemon: manifest unknown"}
HTTP=500
### AFTER:
2026-09-01T18:00:29Z
image=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running
bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 Up 4 minutes (healthy)
@@ -0,0 +1,18 @@
### 3a — the PULL FAILS. Compose pointed at a tag that does not exist.
2026-09-01T17:59:50Z
file now:
image: ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist
container BEFORE:
ghcr.io/alam00000/bentopdf:v2.8.5 aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running
### POST /api/stacks/bentopdf/update — EXACT status and body:
{"ok":false,"error":"pulling images for bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Error manifest unknown\nError response from daemon: manifest unknown"}
HTTP=500
### 3a AFTER — is the app still running?
2026-09-01T18:00:05Z
image=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running restarts=0
image: ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist
### what the customer's app-list row says now (state badge, buttons)
http=200
|Pi kompatibilis||||Telepítés||Részletek|||||||||Audiobookshelf||||Nincs telepítve|||Hangoskönyv és podcast kezelő szerver|||~100M||Pi kompatibilis||HDD szükséges||||Telepítés||Részletek|||||||||BentoPDF||pdf.enkisfelhom.hu ↗||||Fut|||Adatvédelmi fókuszú PDF eszköztár|||~100M||Pi kompatibilis|||||bentopdf||Up 4 minutes (healthy)|||||Frissítés||Újraindítás||Leállítás||Naplók||Részletek|||||||<div
@@ -0,0 +1,4 @@
### does anything ALARM on a crash-looping app after a 'successful' update?
2026-09-01T18:01:48Z
2026/09/01 18:00:59 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
2026/09/01 18:01:29 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
@@ -0,0 +1,49 @@
--- 2026-09-01T18:02:31Z ---
--- 2026-09-01T18:03:17Z ---
--- 2026-09-01T18:04:04Z ---
--- 2026-09-01T18:04:50Z ---
--- 2026-09-01T18:05:37Z ---
--- 2026-09-01T18:06:23Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
--- 2026-09-01T18:07:09Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:07:56Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:08:42Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:09:29Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:10:15Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:11:02Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:11:48Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
--- 2026-09-01T18:12:35Z ---
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
@@ -0,0 +1,37 @@
### 3b — the PULL SUCCEEDS AND THE APP DOES NOT. The compose is pointed at alpine:3.20:
### a real image that pulls cleanly, then exits at once, so restart:unless-stopped puts it
### into a crash-loop. This is the shape of a real upstream image whose config the template
### can no longer supply — the catalog's own 2026-07-18 wger 2.6 revert was exactly that.
2026-09-01T18:00:42Z
file now:
image: alpine:3.20
alpine present locally?
0
container BEFORE:
ghcr.io/alam00000/bentopdf:v2.8.5 aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running
### POST /api/stacks/bentopdf/update — EXACT status and body:
{"ok":true,"message":"Stack bentopdf update completed"}
HTTP=200
### 3b AFTER — what the customer actually has
2026-09-01T18:01:20Z
image=alpine:3.20 id=436a8ae4fe639581dd05febd183f0476ad4aff91c8180063582ad9162ecacb3e status=restarting restarts=9 exit=0
bentopdf | alpine:3.20 | Restarting (0) 5 seconds ago
### the controller's own post-start status line (what it logged about the result)
2026/09/01 18:00:43 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/bentopdf/update
2026/09/01 18:00:43 router.go:80: [DEBUG] [api] POST /api/stacks/bentopdf/update (path=/stacks/bentopdf/update)
2026/09/01 18:00:43 router.go:566: [INFO] [api] update requested for stack: bentopdf
2026/09/01 18:00:43 router.go:80: [DEBUG] [api] actionStack: action=update name=bentopdf
2026/09/01 18:00:43 manager.go:1176: [INFO] [stacks] Updating stack: bentopdf
2026/09/01 18:00:43 manager.go:1439: [INFO] [stacks] Deploying stack bentopdf — checking 1 images...
2026/09/01 18:00:43 manager.go:1320: [DEBUG] Running: docker compose pull (in /opt/docker/stacks/bentopdf)
2026/09/01 18:00:46 manager.go:1320: [DEBUG] Running: docker compose up -d --remove-orphans (in /opt/docker/stacks/bentopdf)
2026/09/01 18:00:47 manager.go:1195: [INFO] [stacks] Stack bentopdf updated successfully (took 3.5s)
2026/09/01 18:00:50 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
2026/09/01 18:00:50 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
2026/09/01 18:00:50 manager.go:1403: [INFO] [stacks] bentopdf alpine:3.20 restarting Restarting (0) Less than a second ago
### 3b — the EXACT customer-facing row while the app is crash-looping
ataset.fallback&&this.dataset.step==='1'){this.dataset.step='2';this.src=this.dataset.fallback;}else{this.onerror=null;this.style.visibility='hidden';}">||BentoPDF||pdf.enkisfelhom.hu ↗||URL nem elérhető – útvonal nincs publikálva||||Újraindítás...|||Adatvédelmi fókuszú PDF eszköztár|||~100M||Pi kompatibilis|||||bentopdf||Restarting (0) 15 seconds ago|||||Frissítés||Újraindítás||Leállítás||Naplók||Részletek||||||<img class="stack-logo-lg" src="/static/assets/bookstack-logo.svg" alt="" data-fallback="/static/app-placeholder.svg"
onerror="if(!thi
@@ -0,0 +1,5 @@
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.232.0 Up 9 hours (healthy)
opengist ghcr.io/thomiceli/opengist:1.13 Up 9 hours (healthy)
filebrowser gtstef/filebrowser:1.3.3-stable Up 3 weeks (healthy)
cloudflared cloudflare/cloudflared:2026.6.0 Up 3 weeks
traefik traefik:v3.6.7 Up 3 weeks
@@ -0,0 +1,49 @@
=== demo-hp 2026-09-01T17:40:38Z ===
bentopdf|ghcr.io/alam00000/bentopdf:v2.8.6|ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
felhom-controller|gitea.dooplex.hu/admin/felhom-controller:0.232.0|gitea.dooplex.hu/admin/felhom-controller@sha256:51b2403520e425e0a7d4bc00cfd395d3b06078d79cf7927bee1035d662814b8c
romm|rommapp/romm:5.0.0|rommapp/romm@sha256:91f6611eca5a4dafc4f4a1d72a1ed7dd66a11375d939f28410dc1d1de0b80b1b
romm-db|mariadb:11.4|mariadb@sha256:4f1d8d202fcf7bcb3902f63af09f9c1a050c2922a89652f22abaec0d4f015e83
romm-redis|redis:7-alpine|redis@sha256:ff02b58f971e7d7d156a1267e283fcbbeee91773b6aa36c49dac28ecfe28eadf
privatebin|privatebin/pdo:2.0.5|privatebin/pdo@sha256:8a2cac16eff6caed4dc622e7b3ebd0c0ffcb28dfc0eba0b9c335636552eadac1
paperless-webserver|ghcr.io/paperless-ngx/paperless-ngx:2.20.15|ghcr.io/paperless-ngx/paperless-ngx@sha256:6c86cad803970ea782683a8e80e7403444c5bf3cf70de63b4d3c8e87500db92f
paperless-redis|redis:7-alpine|redis@sha256:ff02b58f971e7d7d156a1267e283fcbbeee91773b6aa36c49dac28ecfe28eadf
paperless-postgres|postgres:16-alpine|postgres@sha256:cf78e76683b9ca8c5733cbbdce6c9262b45b6767934dd0a95e671f9a0fc20685
opengist|ghcr.io/thomiceli/opengist:1.13|ghcr.io/thomiceli/opengist@sha256:dddc26031d1320ebb4bc5b913b3c42a9cb84c7528192d387f99ddcbbe57b0085
kimai|kimai/kimai2:apache-2.57.0|kimai/kimai2@sha256:efa66c5eadf948b94d8031e9436feae491d90b2ceaa3d5e45e4d7c1a59dd8c38
kimai-db|mariadb:11.6|mariadb@sha256:bfb1298c06cd15f446f1c59600b3a856dae861705d1a2bd2a00edbd6c74ba748
docmost|docmost/docmost:0.95.0|docmost/docmost@sha256:41c8d777cf23c74e78f94e676aec328b7d7856f48df5e573543dac68d371e37c
docmost-postgres|postgres:16-alpine|postgres@sha256:cf78e76683b9ca8c5733cbbdce6c9262b45b6767934dd0a95e671f9a0fc20685
docmost-redis|redis:7-alpine|redis@sha256:ff02b58f971e7d7d156a1267e283fcbbeee91773b6aa36c49dac28ecfe28eadf
calibre-web|crocodilestick/calibre-web-automated:v4.0.6|crocodilestick/calibre-web-automated@sha256:c31a738b6d5ec6982c050063dd3f063b6943eb1051fc81144789f840d9093a8d
bookstack|lscr.io/linuxserver/bookstack:26.05.2|lscr.io/linuxserver/bookstack@sha256:3db259db582808ab498d49ae96b0a63f935d9cf3635c9d5bd8b8815c6ff1f8a1
bookstack-db|mariadb:12.3|mariadb@sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
filebrowser|gtstef/filebrowser:1.3.3-stable|gtstef/filebrowser@sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
cloudflared|cloudflare/cloudflared:2026.6.0|cloudflare/cloudflared@sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
traefik|traefik:v3.6.7|traefik@sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
=== compose file image lines for DEPLOYED stacks ===
--- bentopdf
image: ghcr.io/alam00000/bentopdf:v2.8.6
--- bookstack
image: lscr.io/linuxserver/bookstack:26.05.2
image: mariadb:12.3
--- calibre-web
image: crocodilestick/calibre-web-automated:v4.0.6
--- docmost
image: docmost/docmost:0.95.0
image: postgres:16-alpine
image: redis:7-alpine
--- kimai
image: kimai/kimai2:apache-2.57.0
image: mariadb:11.6
--- opengist
image: ghcr.io/thomiceli/opengist:1.13
--- paperless-ngx
image: ghcr.io/paperless-ngx/paperless-ngx:2.20.15
image: postgres:16-alpine
image: redis:7-alpine
--- privatebin
image: privatebin/pdo:2.0.5
--- romm
image: rommapp/romm:5.0.0
image: mariadb:11.4
image: redis:7-alpine
@@ -0,0 +1,14 @@
image ref floating? verdict
----------------------------------------------------------------------------------------------------
docker.io/library/postgres:16-alpine FLOATING SAME
docker.io/library/redis:7-alpine FLOATING SAME
docker.io/library/mariadb:11.4 FLOATING MOVED
running sha256:4f1d8d202fcf7bcb3902f63af09f9c1a050c2922a89652f22abaec0d4f015e83
upstream sha256:611a2fcc5fa7c6ceb8644c6f74b25ede004ff6c3a6b38c8f8c23d3bbf6c26430
docker.io/library/mariadb:11.6 FLOATING SAME
docker.io/library/mariadb:12.3 FLOATING MOVED
running sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
upstream sha256:dd9b303aed4f4890ed09f766d8ca9ddfd176c0c6f6267feff53b3192ec65a979
ghcr.io/thomiceli/opengist:1.13 FLOATING SAME
docker.io/rommapp/romm:5.0.0 pinned SAME
docker.io/privatebin/pdo:2.0.5 pinned SAME
@@ -0,0 +1,17 @@
### Phase 4 — Peti's box: NOT measurable live (down, not enrolled, no access route).
### What IS knowable read-only: it last reported 48d ago (2026-07-15) running rallly.
### So: how far has the CATALOG's rallly pin moved since then?
current catalog pin:
image: lukevella/rallly:4.11.1
image: postgres:16-alpine
commits touching rallly since 2026-07-15:
d7ffcbe 2026-07-19 rallly: raise app memory 256M -> 768M (OOM-killed at 256M)
e3f3a81 2026-07-18 rallly: 3.11.2 -> 4.11.1 [MAJOR]
(empty above = the pin has NOT moved since Peti's box last reported)
### demo-felhom (opengist, the only deployed app there)
catalog:
image: ghcr.io/thomiceli/opengist:1.13
running: ghcr.io/thomiceli/opengist:1.13 (from phase4-demo-felhom-running.txt)
@@ -0,0 +1,72 @@
### Phase 5 — what a pre-update copy would cost, measured on demo-hp 2026-09-01T18:13:33Z
### 1. app data on disk, per deployed app (the thing a FILE copy would have to move)
56K /mnt/felhom-drives/hdd_1/appdata/
5.9M /mnt/felhom-drives/hdd_1/userdata/
742M /mnt/felhom-drives/hdd_1/backups/
748M /mnt/sys_drive/felhom-data/
### 2. the EXISTING safety machinery is DATABASE-ONLY (writeSafetyDump -> DumpOne).
### Existing dumps on the box, with size and age — the cost proxy for a DB copy:
62270 2026-08-21 21:02 /mnt/felhom-drives/hdd_1/backups/primary/romm/db-dumps/pre-restore-20260821T210246Z-romm-mariadb.sql
62270 2026-09-01 02:15 /mnt/felhom-drives/hdd_1/backups/primary/romm/db-dumps/romm-mariadb.sql
58775 2026-09-01 00:30 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/bookstack-mariadb.sql
58775 2026-08-22 14:24 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/pre-restore-20260822T142418Z-bookstack-mariadb.sql
58775 2026-08-22 16:25 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/pre-restore-20260822T162555Z-bookstack-mariadb.sql
58775 2026-08-22 21:55 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/pre-restore-20260822T215540Z-bookstack-mariadb.sql
142277 2026-09-01 00:30 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql
141363 2026-08-22 16:23 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
141363 2026-08-22 16:27 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
141363 2026-08-22 21:54 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
48217 2026-09-01 00:30 /mnt/felhom-drives/hdd_1/backups/secondary/kimai/recovery-unit/db-dumps/kimai-mariadb.sql
58775 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/bookstack-mariadb.sql
58775 2026-08-22 14:24 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/pre-restore-20260822T142418Z-bookstack-mariadb.sql
58775 2026-08-22 16:25 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/pre-restore-20260822T162555Z-bookstack-mariadb.sql
58775 2026-08-22 21:55 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/pre-restore-20260822T215540Z-bookstack-mariadb.sql
142277 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql
141363 2026-08-22 16:23 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
141363 2026-08-22 16:27 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
141363 2026-08-22 21:54 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
48217 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/primary/kimai/db-dumps/kimai-mariadb.sql
312381 2026-08-22 07:38 /mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql
395065 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx/recovery-unit/db-dumps/paperless-ngx-postgres.sql
312957 2026-08-22 07:56 /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx/recovery-unit/db-dumps/pre-restore-20260822T075658Z-paperless-ngx-postgres.sql
62270 2026-08-21 21:02 /mnt/sys_drive/felhom-data/backups/secondary/romm/recovery-unit/db-dumps/pre-restore-20260821T210246Z-romm-mariadb.sql
62270 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/secondary/romm/recovery-unit/db-dumps/romm-mariadb.sql
### per-app FILE data on demo-hp (this is what a pre-update copy would have to move)
== /mnt/sys_drive/felhom-data/userdata
12K /mnt/sys_drive/felhom-data/userdata/import/
== /mnt/felhom-drives/hdd_1/appdata
12K /mnt/felhom-drives/hdd_1/appdata/romm/
40K /mnt/felhom-drives/hdd_1/appdata/paperless/
== /mnt/felhom-drives/hdd_1/userdata
4.0K /mnt/felhom-drives/hdd_1/userdata/documents/
4.0K /mnt/felhom-drives/hdd_1/userdata/downloads/
12K /mnt/felhom-drives/hdd_1/userdata/import/
1.1M /mnt/felhom-drives/hdd_1/userdata/roms/
4.8M /mnt/felhom-drives/hdd_1/userdata/media/
### docker named volumes (the other place app data lives)
76K /var/lib/docker/volumes/ee40750a8284cf0481f03f3c8c357d2fdc9ab0c68fafd25321b7256c70218bae/
76K /var/lib/docker/volumes/f55b60c2fecb068286c83c306f07695557a5524a66542c548316a76535338ebe/
76K /var/lib/docker/volumes/f98ca17ff5c961ae51783dee774bd48f7e98ccc21a172d819183e73f187fa62e/
76K /var/lib/docker/volumes/fc61c5e41e1aad481c18ec4f5c983e8c095d3ad24c427b1d6cec7044a5be4bbc/
92K /var/lib/docker/volumes/filebrowser_filebrowser_data/
228K /var/lib/docker/volumes/opengist_opengist_data/
960K /var/lib/docker/volumes/calibre-web_calibre_web_config/
2.1M /var/lib/docker/volumes/privatebin_privatebin_data/
5.0M /var/lib/docker/volumes/paperless-ngx_paperless_data/
6.4M /var/lib/docker/volumes/paperless-ngx_paperless_redis_data/
6.8M /var/lib/docker/volumes/bookstack_bookstack_config/
25M /var/lib/docker/volumes/romm_romm_redis_data/
40M /var/lib/docker/volumes/felhom-controller-data/
55M /var/lib/docker/volumes/docmost_docmost_redis_data/
56M /var/lib/docker/volumes/kimai_kimai_var/
67M /var/lib/docker/volumes/docmost_docmost_postgres_data/
69M /var/lib/docker/volumes/paperless-ngx_paperless_postgres_data/
166M /var/lib/docker/volumes/bookstack_bookstack_db_data/
166M /var/lib/docker/volumes/kimai_kimai_db_data/
166M /var/lib/docker/volumes/romm_romm_db_data/
### LIMITATION, stated rather than papered over: demo-hp was reinstalled 2026-08-21 and its
### apps were deployed ~13 h before this run, so it holds a few MB of app data in total.
### It CANNOT price a pre-update copy at customer scale. Prior measured numbers are cited instead.
@@ -0,0 +1,19 @@
### Phase 6 — pin the stack to nextcloud:31.0.14-apache BEFORE deploying
2026-09-01T18:58:18Z
image: nextcloud:31.0.14-apache
image: mariadb:11.6
image: redis:7-alpine
### POST /api/stacks/nextcloud/deploy (file verified at nextcloud:31.0.14-apache above)
2026-09-01T18:58:40Z
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
HTTP=202 time=0.205833s
2026-09-01T18:58:40Z
2026-09-01T18:58:48Z
2026-09-01T18:59:10Z
2026-09-01T18:59:31Z nextcloud|nextcloud:31.0.14-apache|Created nextcloud-db|mariadb:11.6|Up 9 seconds (health: starting) nextcloud-redis|redis:7-alpine|Up 9 seconds (healthy)
>>> HEALTHY
2026-09-01T18:59:39Z nextcloud:31.0.14-apache|Up 6 seconds (health: starting)
2026-09-01T19:00:00Z nextcloud:31.0.14-apache|Up 28 seconds (health: starting)
2026-09-01T19:00:22Z nextcloud:31.0.14-apache|Up 49 seconds (healthy)
>>> nextcloud HEALTHY
@@ -0,0 +1,22 @@
### Phase 6 — SEED identifiable data on nextcloud 31.0.14-apache
2026-09-01T19:00:36Z
- installed: true
- version: 31.0.14.1
- versionstring: 31.0.14
- edition:
- maintenance: false
- needsDbUpgrade: false
- productname: Nextcloud
- extendedSupport: false
--- set a system config marker ---
System config value spike_marker set to string SPIKE-SEEDED-ON-31.0.14-2026-09-01
--- write a real file into the admin user data dir ---
Starting scan for user 1 out of 1 (admin)
+---------+-------+-----+---------+---------+--------+--------------+
| Folders | Files | New | Updated | Removed | Errors | Elapsed time |
+---------+-------+-----+---------+---------+--------+--------------+
| 5 | 53 | 0 | 2 | 0 | 0 | 00:00:00 |
+---------+-------+-----+---------+---------+--------+--------------+
--- read both back ---
SPIKE-SEEDED-ON-31.0.14-2026-09-01
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
@@ -0,0 +1,60 @@
### Phase 6 — THE 3-MAJOR JUMP: 31.0.14-apache -> 34.0.1-apache, via the real Update button
2026-09-01T19:00:53Z
file now:
image: nextcloud:34.0.1-apache
container BEFORE:
nextcloud:31.0.14-apache status=running
### POST /api/stacks/nextcloud/update:
{"ok":true,"message":"Stack nextcloud update completed"}
HTTP=200
2026-09-01T19:01:40Z
### what the app itself says after the 3-major jump
2026-09-01T19:01:50Z nextcloud:34.0.1-apache|Restarting (1) Less than a second ago
2026-09-01T19:02:06Z nextcloud:34.0.1-apache|Restarting (1) 10 seconds ago
2026-09-01T19:02:23Z nextcloud:34.0.1-apache|Restarting (1) 13 seconds ago
2026-09-01T19:02:39Z nextcloud:34.0.1-apache|Restarting (1) 4 seconds ago
2026-09-01T19:02:56Z nextcloud:34.0.1-apache|Restarting (1) 20 seconds ago
2026-09-01T19:03:12Z nextcloud:34.0.1-apache|Restarting (1) 36 seconds ago
### THE EXACT MESSAGE — container logs since the update
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
=> Configuring PHP session handler...
==> Using Redis as PHP session handler...
Initializing nextcloud 34.0.1.2 ...
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
@@ -0,0 +1,36 @@
### Phase 6 — CAN IT GO BACK? put the old tag back and press the button
2026-09-01T19:03:43Z
file now:
image: nextcloud:31.0.14-apache
### POST /api/stacks/nextcloud/restart:
{"ok":true,"message":"Stack nextcloud restart completed"}
HTTP=200
2026-09-01T19:03:46Z
2026-09-01T19:03:55Z nextcloud:31.0.14-apache|Up 9 seconds (healthy)
2026-09-01T19:04:11Z nextcloud:31.0.14-apache|Up 26 seconds (healthy)
2026-09-01T19:04:28Z nextcloud:31.0.14-apache|Up 42 seconds (healthy)
2026-09-01T19:04:44Z nextcloud:31.0.14-apache|Up 59 seconds (healthy)
2026-09-01T19:05:01Z nextcloud:31.0.14-apache|Up About a minute (healthy)
2026-09-01T19:05:17Z nextcloud:31.0.14-apache|Up About a minute (healthy)
### READ THE SEEDED DATA BACK
- installed: true
- version: 31.0.14.1
- versionstring: 31.0.14
- edition:
- maintenance: false
- needsDbUpgrade: false
- productname: Nextcloud
- extendedSupport: false
--- system config marker ---
SPIKE-SEEDED-ON-31.0.14-2026-09-01
--- the seeded file ---
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
--- is the file still known to the DB? ---
Starting scan for user 1 out of 1 (admin)
+---------+-------+-----+---------+---------+--------+--------------+
| Folders | Files | New | Updated | Removed | Errors | Elapsed time |
+---------+-------+-----+---------+---------+--------+--------------+
| 5 | 53 | 0 | 0 | 0 | 0 | 00:00:00 |
+---------+-------+-----+---------+---------+--------+--------------+
@@ -0,0 +1,58 @@
### Phase 6b — THE DECISIVE HALF: a migration that ACTUALLY RUNS. 31.0.14 -> 32.0.9 (one major).
### The refused 3-major jump above was recoverable precisely because nothing migrated.
2026-09-01T19:05:59Z
file now:
image: nextcloud:32.0.9-apache
### POST /api/stacks/nextcloud/update:
{"ok":true,"message":"Stack nextcloud update completed"}
HTTP=200
2026-09-01T19:06:44Z
2026-09-01T19:06:54Z nextcloud:32.0.9-apache|Up 10 seconds (health: starting)
2026-09-01T19:07:16Z nextcloud:32.0.9-apache|Up 31 seconds (healthy)
2026-09-01T19:07:37Z nextcloud:32.0.9-apache|Up 53 seconds (healthy)
2026-09-01T19:07:58Z nextcloud:32.0.9-apache|Up About a minute (healthy)
2026-09-01T19:08:20Z nextcloud:32.0.9-apache|Up About a minute (healthy)
2026-09-01T19:08:41Z nextcloud:32.0.9-apache|Up About a minute (healthy)
2026-09-01T19:09:03Z nextcloud:32.0.9-apache|Up 2 minutes (healthy)
2026-09-01T19:09:24Z nextcloud:32.0.9-apache|Up 2 minutes (healthy)
2026-09-01T19:09:46Z nextcloud:32.0.9-apache|Up 3 minutes (healthy)
2026-09-01T19:10:07Z nextcloud:32.0.9-apache|Up 3 minutes (healthy)
### did the migration RUN? (occ status + upgrade log)
- installed: true
- version: 32.0.9.2
- versionstring: 32.0.9
- edition:
- maintenance: false
- needsDbUpgrade: false
- productname: Nextcloud
- extendedSupport: false
--- markers ---
SPIKE-SEEDED-ON-31.0.14-2026-09-01
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
### the upgrade evidence in the container log
Initializing nextcloud 32.0.9.2 ...
Upgrading nextcloud from 31.0.14.1 ...
=> Searching for hook scripts (*.sh) to run, located in the folder "/docker-entrypoint-hooks.d/pre-upgrade"
==> Skipped: the "pre-upgrade" folder is empty (or does not exist)
Updated database
Updated <federation> to 1.22.0
Updated <lookup_server_connector> to 1.20.0
Updated <oauth2> to 1.20.0
Updated <password_policy> to 4.0.0
Updated <photos> to 5.0.0
Updated <activity> to 5.0.0
Updated <circles> to 32.0.0
Updated <cloud_federation_api> to 1.16.0
Updated <dav> to 1.34.2
Updated <files> to 2.4.0
Updated <files_sharing> to 1.24.1
Updated <files_trashbin> to 1.22.0
Updated <files_versions> to 1.25.0
Updated <sharebymail> to 1.22.0
Updated <webhook_listeners> to 1.3.0
Updated <workflowengine> to 2.14.0
Updated <comments> to 1.22.0
Updated <logreader> to 5.0.0
Updated <nextcloud_announcements> to 4.0.0
Updated <notifications> to 5.0.0
@@ -0,0 +1,41 @@
### Phase 6c — THE ROLLBACK ATTEMPT: put 31.0.14 back after a migration that RAN
2026-09-01T19:10:45Z
file BEFORE my edit (the sync may have flipped it):
image: nextcloud:34.0.1-apache
file now:
image: nextcloud:31.0.14-apache
### POST /api/stacks/nextcloud/restart:
{"ok":true,"message":"Stack nextcloud restart completed"}
HTTP=200
2026-09-01T19:10:49Z
2026-09-01T19:10:58Z nextcloud:31.0.14-apache|Restarting (1) Less than a second ago
2026-09-01T19:11:15Z nextcloud:31.0.14-apache|Restarting (1) 10 seconds ago
2026-09-01T19:11:31Z nextcloud:31.0.14-apache|Restarting (1) 13 seconds ago
2026-09-01T19:11:48Z nextcloud:31.0.14-apache|Restarting (1) 4 seconds ago
2026-09-01T19:12:04Z nextcloud:31.0.14-apache|Restarting (1) 20 seconds ago
2026-09-01T19:12:21Z nextcloud:31.0.14-apache|Restarting (1) 37 seconds ago
2026-09-01T19:12:37Z nextcloud:31.0.14-apache|Restarting (1) 2 seconds ago
2026-09-01T19:12:53Z nextcloud:31.0.14-apache|Restarting (1) 18 seconds ago
### THE EXACT MESSAGE from the app when the old version meets migrated data
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
Configuring Redis as session handler
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
@@ -0,0 +1,13 @@
### Phase 6d — POSITIVE CONTROL: the data is NOT destroyed, only the downgrade is refused.
### Put 32.0.9 back and read the markers.
2026-09-01T19:13:27Z
image: nextcloud:32.0.9-apache
{"ok":true,"message":"Stack nextcloud restart completed"}
HTTP=200
nextcloud:32.0.9-apache|Up 46 seconds (healthy)
- installed: true
- version: 32.0.9.2
- versionstring: 32.0.9
SPIKE-SEEDED-ON-31.0.14-2026-09-01
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
@@ -0,0 +1,50 @@
### OBSERVATION — remove_hdd_data:true, yet the HDD data dir is still there and the response
### reported neither removed nor preserved (hdd_paths_removed:null, hdd_paths_preserved:null).
total 88592
drwxrwx--- 4 www-data www-data 4096 Sep 1 19:07 .
drwxr-xr-x 5 root root 4096 Sep 1 18:59 ..
-rw-rw-r-- 1 www-data www-data 542 Sep 1 19:06 .htaccess
-rw-rw-r-- 1 www-data www-data 52 Sep 1 19:06 .ncdata
drwxr-xr-x 3 www-data www-data 4096 Sep 1 18:59 admin
drwxr-xr-x 5 www-data www-data 4096 Sep 1 19:07 appdata_ocrus79ehwsx
-rw-rw-r-- 1 www-data www-data 0 Sep 1 19:06 index.html
-rw-r----- 1 www-data www-data 90692753 Sep 1 19:07 nextcloud.log
128M /mnt/felhom-drives/hdd_1/appdata/nextcloud/
--- for contrast, the OTHER apps own theirs legitimately ---
128M /mnt/felhom-drives/hdd_1/appdata/nextcloud/
40K /mnt/felhom-drives/hdd_1/appdata/paperless/
12K /mnt/felhom-drives/hdd_1/appdata/romm/
2026/09/01 19:14:56 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, SUBDOMAIN, DB_PASSWORD, DOMAIN, HDD_PATH, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_PASSWORD, NEXTCLOUD_ADMIN_USER, USERDATA_PATH, IMPORT_PATH] (14 app + 0 system)
2026/09/01 19:15:06 router.go:80: [DEBUG] [api] removeStack: name=nextcloud
2026/09/01 19:15:06 router.go:80: [DEBUG] [api] removeStack: name=nextcloud removeHDDData=true removeBackups=true
2026/09/01 19:15:06 delete.go:289: [DEBUG] [stacks] RemoveStack called: name="nextcloud", removeHDDData=true, backupPathsToRemove=1
2026/09/01 19:15:06 delete.go:303: [DEBUG] [stacks] RemoveStack nextcloud: state=stopped, deployed=true, orphaned=false, deploying=false
2026/09/01 19:15:06 delete.go:327: [INFO] Removing deployed stack: nextcloud (removeHDDData=true, backupPaths=1)
2026/09/01 19:15:06 delete.go:337: [DEBUG] [stacks] RemoveStack nextcloud: found 0 HDD mounts from compose file
2026/09/01 19:15:06 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DB_PASSWORD, DOMAIN, HDD_PATH, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_PASSWORD, NEXTCLOUD_ADMIN_USER, SUBDOMAIN, USERDATA_PATH, IMPORT_PATH] (14 app + 0 system)
2026/09/01 19:15:07 delete.go:347: [DEBUG] [stacks] RemoveStack nextcloud: compose down output:
2026/09/01 19:15:07 delete.go:402: [DEBUG] [stacks] RemoveStack nextcloud: processing 1 backup paths for removal (base=backups)
2026/09/01 19:15:07 delete.go:408: [WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps
2026/09/01 19:15:07 delete.go:426: [DEBUG] [stacks] RemoveStack nextcloud: removing app.yaml at /opt/docker/stacks/nextcloud/app.yaml
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.6GB avail=884.5GB (0.6%)
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/hdd_1" total=937.8 GB used=5.6 GB avail=884.5GB (0.6%)
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.6GB avail=884.5GB (0.6%)
2026/09/01 19:15:09 healthcheck.go:42: [DEBUG] [monitor] Raw values: disk=23.6%, hdd=0.6% (configured=true), mem=15.9% (4106MB/25898MB), cpu=4.1%, temp=52.9°C (hwmon1)
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/hdd_1" total=937.8 GB used=5.6 GB avail=884.5GB (0.6%)
2026/09/01 19:15:29 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
### ROOT CAUSE of the leftover — positive + negative control
### POSITIVE control: the env override IS the only way HDDPath can be set, and it is:
0
(0 = the env override is ABSENT)
(no FELHOM_PATHS_* var at all)
### and controller.yaml paths: has no hdd_path key (read above).
### config.go:117 HDDPath has NO default — only envStr at :403. So cfg.Paths.HDDPath == "".
### ParseComposeHDDMounts (delete.go:600-603) returns nil on the FIRST line when hddPath == "".
### NEGATIVE control — this is not nextcloud-specific. The same 0-mounts line was logged
### for bentopdf earlier in this session, which genuinely has no HDD mount:
2026/09/01 17:55:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bentopdf/docker-compose.yml
### ParseComposeHDDMounts (delete.go:600-603) returns nil on the FIRST line when hddPath == "".
@@ -0,0 +1,8 @@
### TEARDOWN — restore bentopdf to its catalog tag via the product's own restart path
{"ok":true,"message":"Stack bentopdf restart completed"}
HTTP=200
2026-09-01T18:13:18Z
image: ghcr.io/alam00000/bentopdf:v2.8.6
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=02c80375fcba293c6bd06ee826e6fc851f030bf8f5a2d021394c54964507c69a status=running
digest=ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
@@ -0,0 +1,89 @@
### TEARDOWN LAYER 3 — the hub. This run created NO customer and NO appliance record.
### It used the EXISTING demo-hp customer. What it DID create is events. Listing them:
404 page not found
### customers list unchanged (5 rows, same as at the start of the run):
demo-felhom
demo-hp
drill-r50
peti-felhom
tester-1
### hub-side residue from the run — the nextcloud telemetry row and the events it pushed
nextcloud — Felhom Hub
Felhom |Hub
Dashboard
Customers
Apps
Hosts
Offsite
Configuration
← Apps
24h
7d
30d
Nextcloud
Reset Telemetry
App Name
nextcloud
Deployments
Catalog Estimate
256M
Catalog Limit
1024M
Suggested Limit (P95×1.2)
352 MB
Avg Memory
208 MB
P95 Memory
280 MB
Avg CPU
1%
Memory Trend
Known Issues |— filtered: demo-hp
Clear filter (fleet view)
Show dismissed
Severity
Message
Occurrences (all customers)
Affected Customers
First Seen
Last Seen
warn
0 [Warning] InnoDB: liburing disabled: falling back to innodb_use_native_aio=OFF
9 min ago
9 min ago
Details
fingerprint: |0 [warning] innodb: liburing disabled: falling back to innodb_use_native_aio=off
· severity: warn
· first seen: 2026-09-01 19:10:30
· last seen: 2026-09-01 19:10:30
affected customers:
demo-hp
Full message:
Copy
0 [Warning] InnoDB: liburing disabled: falling back to innodb_use_native_aio=OFF
No context captured (pre-v0.111 report or warn-severity issue).
warn
0 [Warning] mariadbd: io_uring_queue_init() failed with errno 0
9 min ago
9 min ago
Details
fingerprint: |0 [warning] mariadbd: io_uring_queue_init() failed with errno 0
· severity: warn
· first seen: 2026-09-01 19:10:30
· last seen: 2026-09-01 19:10:30
affected customers:
demo-hp
Full message:
Copy
0 [Warning] mariadbd: io_uring_queue_init() failed with errno 0
No context captured (pre-v0.111 report or warn-severity issue).
warn
0 [Warning] mariadbd: io_uring_queue_init() failed with errno 2
9 min ago
9 min ago
Details
fingerprint: |0 [warning]
### current containers demo-hp reports (nextcloud must be ABSENT):
3
@@ -0,0 +1,25 @@
### remove the one other image this spike pulled: bentopdf:v2.8.5
removed
removed alpine:3.20
bentopdf images now:
ghcr.io/alam00000/bentopdf:v2.8.6
### TEARDOWN LAYER 2 — the host: space returned
Filesystem Size Used Avail Use% Mounted on
/dev/mapper/pve-vm--9201--disk--0 32G 957M 29G 4% /
/dev/mapper/pve-vm--9201--disk--1 69G 12G 54G 18% /mnt/sys_drive
/dev/nvme0n1 938G 5.5G 885G 1% /mnt/felhom-drives/hdd_1
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 17977988 20388272 44.44%
local-lvm lvmthin active 56487936 40055595 16432340 70.91%
### the 1.05 GiB the guest freed did NOT return to the host thin pool. Trying fstrim.
before: local-lvm used 40055595 KiB (baseline before this spike: 38959729 KiB)
fstrim: /mnt/felhom-drives/hdd_1: FITRIM ioctl failed: Operation not permitted
fstrim: /var/lib/felhom: FITRIM ioctl failed: Operation not permitted
fstrim: /: FITRIM ioctl failed: Operation not permitted
local-lvm lvmthin active 56487936 40055595 16432340 70.91%
/var/lib/lxc/9201/rootfs/: 30.2 GiB (32480210944 bytes) trimmed
/var/lib/lxc/9201/rootfs/var/lib/felhom: 57 GiB (61255426048 bytes) trimmed
rc=0
local-lvm lvmthin active 56487936 15127469 41360466 26.78%
@@ -0,0 +1,61 @@
### TEARDOWN LAYER 1 — the machine: remove the throwaway nextcloud stack this spike created
2026-09-01T19:14:36Z
{"ok":false,"error":"stack \"nextcloud\" is still running — stop it first before removing"}
HTTP=409
### verify gone
nextcloud nextcloud:32.0.9-apache Up About a minute (healthy)
nextcloud-db mariadb:11.6 Up 15 minutes (healthy)
nextcloud-redis redis:7-alpine Up 15 minutes (healthy)
(empty above = containers gone)
nextcloud_nextcloud_db_data
nextcloud_nextcloud_html
nextcloud_nextcloud_redis_data
(empty above = volumes gone)
app.yaml
docker-compose.yml
nextcloud
paperless
romm
### stop first, then remove
{"ok":true,"message":"Stack nextcloud stop completed"}
HTTP=200
{"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":null,"hdd_paths_preserved":null},"message":"Stack nextcloud removed"}
HTTP=200
### verify gone
containers:
volumes:
hdd appdata:
nextcloud
paperless
romm
stack dir:
docker-compose.yml
### removing the spike's own leftover (128 MB the product did not remove)
2026-09-01T19:17:28Z
128M /mnt/felhom-drives/hdd_1/appdata/nextcloud/
after:
paperless
romm
any nextcloud left anywhere under /mnt:
stack dir (template only, app.yaml gone = not deployed):
docker-compose.yml
images:
nextcloud:34.0.1-apache
nextcloud:32.0.9-apache
nextcloud:31.0.14-apache
### remove ONLY the three nextcloud images this spike pulled (targeted rmi, never a prune)
/dev/mapper/pve-vm--9201--disk--0 32G 958M 29G 4% /
removed nextcloud:31.0.14-apache
removed nextcloud:32.0.9-apache
removed nextcloud:34.0.1-apache
remaining nextcloud images:
(none)
/dev/mapper/pve-vm--9201--disk--0 32G 958M 29G 4% /
stack dir now:
.
..
.felhom.yml
docker-compose.yml
+8 -3
View File
@@ -675,9 +675,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). | **OPEN — rank P3-LOW; owner: CC** |
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** |
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.