SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has already moved, and the Restart button does it — 18.3s with a network pull when the target image is absent, 0.5s when present, against a negative control that did not even recreate the container. The boot reconciler does the same thing unattended when an app fails to come back (bootrecon.go:269 -> StartStack). AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades nothing. Docker restores the old containers and the reconciler logs 'no boot-orphaned apps (nothing to start)'. AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run, the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported'. 'Rollback' is the wrong word for this arc and is struck. Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator confirmation. Peti's box was never contacted. No production code written in any repo: felhom-controller is at 960d29b0612c before and after, tree clean, and build/vet/test are green — run at the end to prove exactly that. Register: R-438/439/440 updated with live evidence; R-441..R-445 opened (restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB left behind; update reports success over a broken app; no fleet fstrim; hub telemetry outlives the app). Capability map gains three measured rows. The mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR, not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision. Five claims in the brief are named as wrong, including two of my own method.
This commit is contained in:
@@ -1,311 +1,214 @@
|
||||
# REPORT — the line under the backup arc, and one alarm that was telling the operator something untrue
|
||||
# REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01)
|
||||
|
||||
**hub v0.111.0 → v0.111.1 · controller v0.232.0 UNTOUCHED · 2026-09-01**
|
||||
**Task class: Spike.** No production code was written in any repo. Output is a findings doc, register
|
||||
rows, two architecture updates and one operator decision.
|
||||
|
||||
**No controller release. No agent release. No golden owed, no floor change.** The only behaviour
|
||||
change in this session is one sentence in one alarm. Everything else is register and documentation.
|
||||
**Findings doc:** `documentation/audits/SPIKE-app-update-2026-09-01.md`
|
||||
**Evidence:** `documentation/audits/evidence-spike-app-update-2026-09-01/` (19 files)
|
||||
|
||||
---
|
||||
|
||||
## 1. The alarm's new text, and proof the old promise is gone
|
||||
## 1. The Phase 1 answer, stated unambiguously
|
||||
|
||||
**Quoted from the RUNNING binary** — `kubectl cp` out of pod `hub-857678f9b4-v95s6`, byte-grepped:
|
||||
**YES — `docker compose up -d` upgrades an app whose compose file has already moved, and the Restart
|
||||
button does it.**
|
||||
|
||||
```
|
||||
Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than
|
||||
retention can explain. The daily Storage Box snapshots are read-only and still hold the older
|
||||
copy. The route back out of them is not yet established, so treat this as neither confirmed
|
||||
data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a
|
||||
deletion ran on the box.
|
||||
```
|
||||
Produced by **variants 1a and 1b, and independently by 1c-ii**:
|
||||
|
||||
**Every shape of the withdrawn promise is absent from the deployed bytes:**
|
||||
- **1a** (target image ABSENT locally): the restart took **18.3 s**, ended with the container on the
|
||||
new tag, and **left the new image in the local store where it had not been** — so `up -d` performed
|
||||
a network pull, in a code path that contains no pull step.
|
||||
- **1b** (target image PRESENT): **0.5 s**, container recreated onto the new tag, image id unchanged.
|
||||
- **1d, the negative control** (file NOT edited): container id, image id, digest and `StartedAt` all
|
||||
identical — `up -d` did not even recreate. **The method can show "no change" when nothing changed.**
|
||||
- **1c-ii** (unattended): the boot reconciler brought a failed app back on the **new** version with
|
||||
nobody pressing anything.
|
||||
|
||||
| fragment | in the deployed binary |
|
||||
**And the answer to 1c is NO, which narrows the exposure:** a hard guest reset upgraded nothing.
|
||||
Docker's `restart: unless-stopped` restored the existing containers on the old image and the
|
||||
reconciler logged its own verdict — `Boot reconciliation: no boot-orphaned apps (nothing to start)`.
|
||||
|
||||
**Per the task's own rule, no fix is proposed.** The behaviour may have been chosen — `RestartStack`
|
||||
says so in a comment — so it goes to the operator as a decision.
|
||||
|
||||
---
|
||||
|
||||
## 2. Confirmed baselines actually used
|
||||
|
||||
| Repo | at §1 of the task | measured at start | end | moved? |
|
||||
|---|---|---|---|---|
|
||||
| felhom-controller | `960d29b0612c` | `960d29b0612c` | `960d29b0612c` | **no — read only** |
|
||||
| felhom.eu | `1d59353df437` | `1d59353df437` | this session's documents | as planned |
|
||||
| app-catalog-felhom.eu | `29edad9c5bf4` | `29edad9c5bf4` | `5d8f25f`, **tree identical to `29edad9c5bf4`** | reverted |
|
||||
| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched |
|
||||
|
||||
**None had moved since the task was written.** Live controller on demo-hp: 0.232.0.
|
||||
|
||||
---
|
||||
|
||||
## 3. Per-phase results, each with its control
|
||||
|
||||
| Phase | Result | Control |
|
||||
|---|---|---|
|
||||
| **0** | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) |
|
||||
| **1** | **THE GATE: YES.** 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | **1d** — unchanged file, no change at all |
|
||||
| **2** | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (`BentoPDF`, 4 hits) + negative (`zzz-never-present`, 0) on both customer pages |
|
||||
| **3a** | Pull failure: HTTP 500, **app untouched and still running** | re-run as a Restart: also 500, app still ran |
|
||||
| **3b** | **HTTP 200 "update completed" over a crash-looping app**; alarm fired 5m16s later | the alarm is a positive observable (`Event pushed: app_start_failed`), plus the detector heartbeat `1 currently down` |
|
||||
| **4** | demo-hp and demo-felhom: **zero drift by tag**. But **2 of 6 floating pins have already MOVED upstream** | both fully-pinned images (`romm:5.0.0`, `pdo:2.0.5`) were SAME |
|
||||
| **5** | Existing safety dump is **database-only**; DB dumps are 48 KB–395 KB; the **file** half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead |
|
||||
| **6** | **App data CANNOT be rolled back.** Old image refuses to start on migrated data | returning to 32.0.9 restored **both** seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked |
|
||||
|
||||
---
|
||||
|
||||
## 4. The exact symbols that bring an app back — found by reading
|
||||
|
||||
| path | `file:symbol` |
|
||||
|---|---|
|
||||
| `recoverable file-by-file` | **absent** |
|
||||
| `recoverable file by file` | **absent** |
|
||||
| `so this is recoverable` | **absent** |
|
||||
| `it is NOT confirmed data loss` | **absent** |
|
||||
| *(negative control)* `zzz-no-such-string-r434` | absent — so the search can report absence |
|
||||
| *(positive control)* `offsite_snapshots_dropped` | present, 2 occurrences |
|
||||
| boot reconciler | `felhom-controller/controller/internal/bootrecon/bootrecon.go:269` — `Reconciler.Run` → `StartStack` |
|
||||
| its scheduler | `controller/cmd/controller/main.go:2127` — `runBootReconcile`, called at `:450` |
|
||||
| app-stop guard | `controller/internal/backup/appstop_marker.go:283` — `AppStopGuard.Recover` → `StartStack` |
|
||||
| its hold-aware wrapper | `controller/cmd/controller/main.go:2008` — `gatedAppStopStarter.StartStack` |
|
||||
| drive-return gate | `controller/internal/web/intermediary.go:222` — `Server.restartStacks` → `StartStack` |
|
||||
| guest-boot change | `controller/internal/web/intermediary.go:458` — `Server.processGuestBootChange` |
|
||||
| quiesce restart | `controller/internal/quiesce/quiesce.go:733` — `Loop.restartAll` |
|
||||
| off-site reconstitution | `controller/internal/backup/offbox_reconstitute.go:692` — `Manager.ReconstituteFromOffsite` |
|
||||
|
||||
All seven fragments of the new sentence were checked individually and are present.
|
||||
Full table of all 13 non-API call sites: findings doc §8.
|
||||
|
||||
**THE FIX IS A DELETION, NOT A REPLACEMENT — and that is what unblocked it.** R-434's own row said
|
||||
the fix was *"blocked on R-433"*, i.e. on first establishing what IS true. **That verdict was mine
|
||||
and it was wrong.** A sentence that asserts **neither** loss **nor** recovery is true under every
|
||||
possible answer to the provider questions, so it never needs a second rewrite. A *replacement* would
|
||||
have been blocked; a *withdrawal* is not. The reasoning is recorded in R-434's closing cell and in
|
||||
the function's doc comment, because the distinction is the transferable part.
|
||||
---
|
||||
|
||||
**It does not swing the other way either.** `TestR434_AlarmStillDoesNotClaimDataLoss` pins that:
|
||||
"your backups are gone" is still usually false — the snapshots exist and hold the older copy; what is
|
||||
unproven is our route to them. Clause (a) of the 2026-09-01 measurement — the box **cannot write**
|
||||
into the snapshot area — stands and is re-confirmed.
|
||||
## 5. Register rows opened, updated or re-ranked
|
||||
|
||||
### Tests, and the red-proof
|
||||
|
||||
Three tests in `hub/internal/monitor/offsite_r434_test.go`, all driving the production path
|
||||
(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator
|
||||
**receives** rather than the function that formats it. ASCII-only fragments, with a positive control
|
||||
(a phrase in every version of the alarm) and a negative control (a phrase that cannot exist).
|
||||
|
||||
**RED-PROOF (run before the fix was restored):** the v0.111.0 sentence was put back and **all three
|
||||
FAILED** — on `recoverable file-by-file`, on `so this is recoverable`, on all three required
|
||||
fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the offending
|
||||
sentence printed in the failure message. Fix restored, `git diff` clean, full hub suite green.
|
||||
|
||||
**One existing test was edited, and it caught this fix correctly.**
|
||||
`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the
|
||||
new wording. **The fragment was REMOVED, not updated** — the wording now has one home
|
||||
(`offsite_r434_test.go`), because duplicating it creates the second source that makes the next
|
||||
correction land in one file and not the other. Its signal fragments (`69`, `4`, `read-only`) are
|
||||
unchanged.
|
||||
|
||||
### R-435 written into the alarm's own documentation, no threshold changed
|
||||
|
||||
The comment above `snapshotDropFraction` now states what the detector does **not** see: **a mass
|
||||
deletion, yes; one app being wiped, no.** demo-hp's baseline is 69 across 9 apps, so ~35 must go
|
||||
before it speaks and one app's tag is ~9; `offbox.go:1388` runs `forget --prune` grouped by
|
||||
`host,tags`, so the blind spot sits on the most likely single-app failure. The comment says
|
||||
explicitly that the numbers must **not** be lowered to "fix" this, and that per-app detection needs a
|
||||
**second** signal keyed on the per-tag count.
|
||||
|
||||
## 2. The stopping line, as it now reads in all three places
|
||||
|
||||
**Register — `documentation/backlog/OPEN-ITEMS.md`**, a new section in the voice this file uses for a
|
||||
settled decision:
|
||||
|
||||
> **DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01)** … **CLOSED FOR BETA at
|
||||
> controller v0.232.0 / hub v0.111.1.** … **What is finished, and proven live:** everything a
|
||||
> customer does for themselves … **Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of `07` §8 — every one PROVEN.**
|
||||
> … **THIS REOPENS IF:** a customer-facing recovery path is found broken; **or** Hetzner's answers
|
||||
> change what the snapshots are worth …; **or** a real customer's data is at stake in one of the
|
||||
> deferred rows.
|
||||
|
||||
**Architecture — `documentation/architecture/07-backup-architecture.md`, at the head of §8**, so a
|
||||
reader of the matrix meets it before the blanks:
|
||||
|
||||
> **THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01)** … **The blanks in
|
||||
> the RTO column below are now blank ON PURPOSE, and that is the whole difference.** … **NO STATUS
|
||||
> MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day. A stopping line that
|
||||
> promotes a row is a stopping line that lies.**
|
||||
|
||||
**Operator page — `STATUS.md`**, under *Decided — and what would reopen each*:
|
||||
|
||||
> **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.** … **What is parked until
|
||||
> after beta:** everything **only I do, with you** — rebuilding a machine as itself, losing a whole
|
||||
> box, recovering from ransomware, restoring the hub, and losing Hetzner. **Six of these have never
|
||||
> been timed, and the hub has never been restored.** They are written down, they are real, and **none
|
||||
> of them stops a beta customer.**
|
||||
|
||||
## 3. The deferred set — row numbers, not a description
|
||||
|
||||
`07` §8 rows **4, 8, 9, 10, 11 (and 11b), 12**, each tagged **`[BETA-DEFERRED]`** in its status cell.
|
||||
|
||||
| row | failure | status today |
|
||||
| row | what | owner |
|
||||
|---|---|---|
|
||||
| **4** | primary drive dies — the drive-loss **journey** | `PARTIAL` |
|
||||
| **8** | host dies, drives intact — a host rebuilt as itself | `IMPLEMENTED`, never executed |
|
||||
| **9** | whole box lost (fire/theft) | `IMPLEMENTED / UNPROVEN` |
|
||||
| **10** | ransomware / malicious deletion | `PARTIAL` |
|
||||
| **11** (+**11b**) | hub lost — a hub restore | `UNPROVEN`, never performed |
|
||||
| **12** | off-site provider lost (Hetzner) | `[FACT]` only |
|
||||
| **R-438** | opened P1-HIGH, then **updated with the live evidence** (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | **VIKTOR rules, CC measures** |
|
||||
| **R-439** | opened P3-LOW, then **updated — the severity survives but its stated reason was imprecise** (`isOperationalState` counts `restarting`/`degraded` as operational) | CC |
|
||||
| **R-440** | opened P2-MEDIUM, then **updated: two floating pins have ALREADY moved, with a passing control** | CC |
|
||||
| **R-441** | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules |
|
||||
| **R-442** | NEW — **P1-HIGH**: `remove_hdd_data:true` is inert with no `paths.hdd_path`; 128 MB left, API said neither removed nor preserved | CC |
|
||||
| **R-443** | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules |
|
||||
| **R-444** | NEW — nothing runs `pct fstrim`; demo-hp's thin pool held ~23.8 GB of freed blocks | CC |
|
||||
| **R-445** | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements |
|
||||
|
||||
`grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md` returns **eight**
|
||||
lines — the seven tagged rows plus the one line in §8's header that defines the marker. That is
|
||||
stated in the register rather than left for the reader to trip over.
|
||||
**Nothing was closed and nothing was re-ranked.** The ranking of R-438 relative to existing rows is
|
||||
Viktor's, and I have not moved anything.
|
||||
|
||||
**A NUMBER IN THE BRIEF WAS WRONG AND IS CORRECTED IN PLACE.** The brief said *"six rows of §8 still
|
||||
have no measured time"*. **Six rows are DEFERRED; ELEVEN carry a blank RTO** — counted, not
|
||||
estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not
|
||||
deferred work, and collapsing them into one number is how a blank stops meaning anything:
|
||||
---
|
||||
|
||||
* **row 5** — `PROVEN` by construction; a derived copy, so there is no recovery to time.
|
||||
* **row 13** — `NONE for host-loss` **by design**; R exists in zero system copies.
|
||||
* **row 14** — `PROVEN`; the break-glass route works and has simply never been stopwatched.
|
||||
* **row 11b** — a consequences note attached to row 11, not a recovery row.
|
||||
* **row 15** — an **open DEFECT** (R-104, the stale-lock path). **It is NOT inside the stopping line**
|
||||
and must not be read as parked by it. This one matters most: parking a live defect by accident is
|
||||
the failure mode a stopping line invites.
|
||||
## 6. Claims in the task that turned out to be wrong, named
|
||||
|
||||
**No status was moved.** Nothing was proven today.
|
||||
1. **"a power cut … is an unattended three-major-version upgrade"** — not as stated. A plain power cut
|
||||
upgraded nothing (measured). The unattended upgrade needs *"and the app did not come back"*.
|
||||
**The exposure is smaller than the operator page claims.**
|
||||
2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites, nine files.
|
||||
3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and
|
||||
`runbooks/target-selection.md:161` says *"No access route from DooPlex"*. Two boxes measured live;
|
||||
Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history.
|
||||
4. **R-439's severity reason** — right conclusion, imprecise reason (see §5).
|
||||
5. **R-440's "23 pins"** — **exactly right** (79 image lines / 53 apps / 66 distinct; 23 with no patch
|
||||
component). One arguable 24th named rather than rounded away.
|
||||
6. **The catalog history figures** — spot-checked and **all correct**: 153 commits, 53 apps, 0 files
|
||||
with upgrade metadata, and all four multi-major bumps confirmed by commit hash.
|
||||
7. **My own method, corrected in-flight:** the first customer-page search used `grep -o "2.8.6"`, whose
|
||||
unescaped `.` produced two false hits; `grep -F` gives zero. The controls caught it.
|
||||
|
||||
## 4. The two questions
|
||||
---
|
||||
|
||||
`documentation/runbooks/provider-questions-2026-09-01.md` — both drafted ready to paste, linked from
|
||||
R-95 and R-433, **not sent**, and **no provider API was called** (§11-D stands).
|
||||
## 7. Evidence handling
|
||||
|
||||
* **Q1 — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?**
|
||||
*Why:* if it cannot, the only route is a rollback that hits every customer on the box and destroys
|
||||
newer snapshots — the snapshots would then protect almost nobody. **Decides how urgent R-95 is.**
|
||||
* **Q2 — on `rclone serve restic --stdio`, is `--append-only` enforced server-side or chosen by the
|
||||
client?** *Why:* if enforced, the box cannot delete its own backups, with no new machine and no
|
||||
data move. **Could make R-95 disappear.**
|
||||
**Evidence was written directly into `documentation/audits/evidence-spike-app-update-2026-09-01/` on
|
||||
DooPlex as each phase produced it — before every revert, including the intermediate ones.** The
|
||||
catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all
|
||||
happened after their evidence was already off the machine. **Nothing was lost and nothing had to be
|
||||
reproduced.**
|
||||
|
||||
Each carries how to read either answer, including the branch where the lead is worthless and should
|
||||
be recorded as dead rather than left looking promising. **A dated DUE-CHECKS row (R-433, 2026-09-15)
|
||||
tracks the reply** — the 2026-07-27 snapshot check that sat unconfirmed for 36 days is the scar that
|
||||
block exists for, and it was never entered.
|
||||
|
||||
## 5. The register
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| file lines (`wc -l`) | **620** | **688** |
|
||||
| total `R-` rows | **181** | **182** |
|
||||
| open-state rows | **170** | **169** |
|
||||
| closed / decided / answered rows | **11** | **13** |
|
||||
|
||||
*(The brief's "620 lines and 181 open" matches the file length and the row count exactly; the count
|
||||
of rows whose LEADING verdict is actually open was 170, which is the number that matters and is not
|
||||
the same thing.)*
|
||||
|
||||
**A correction to my own commit messages, made here rather than by rewriting history:** the
|
||||
`db38f4c` message says `621 -> 688` and the `1a1b32b` message says `688 -> 700`. Both are wrong. I
|
||||
mixed two measures — `wc -l` and a Python line split that counts the trailing newline as a line — and
|
||||
then carried the error forward. **The table above is `wc -l` throughout and is the number to trust:
|
||||
620 → 688.** It changes nothing about the work and it is exactly the class of sloppiness this project
|
||||
files rows about, so it is named rather than quietly fixed.
|
||||
|
||||
**Rows closed:** **R-434** — with the deletion-not-replacement reasoning, and with the fact that its
|
||||
own "blocked on R-433" verdict was wrong recorded in the closing cell.
|
||||
|
||||
**Rows kept open and marked:**
|
||||
|
||||
* **R-95**, **R-433** — **`BLOCKED-ON-PROVIDER`**, both pointing at Part 3's file. R-95 records that
|
||||
it spent **one day demoted on a clause that did not hold**, and that on today's evidence it belongs
|
||||
back near the top. **That is a proposal. I have not re-ranked it; the order is Viktor's.**
|
||||
* **R-430** — **`LATENT`**, with the trigger stated as a trigger: it becomes live **the moment delete
|
||||
is withdrawn from the box**, which is exactly what R-95's remedy does by either route. So it is a
|
||||
**precondition on the R-95 build, not a follow-up** — settle it in the same change or the
|
||||
crash-lock self-heal ships already broken.
|
||||
* **R-435** — open. The limitation is now in the code's own documentation, but **documenting a blind
|
||||
spot is not covering it**, and `STATUS.md` and R-431 both still say "noticed within a day".
|
||||
|
||||
**Row filed:** **R-437** — the register compression sweep, owed and scoped.
|
||||
|
||||
**Compression was measured and deliberately not run, and the reason is one line as the standing rule
|
||||
requires:** only **12 rows / ~25 KB of 316 KB (about 7 %)** carry a closed leading verdict, so the
|
||||
sweep buys little and touches everything — and it is the exact operation that misfiled seven rows in
|
||||
August (R-378, and the seventh, R-87, sat wrong for nine days — R-405). Running it as the tail end of
|
||||
a session about something else is how that happened the first time. R-437 carries the scope so it can
|
||||
be picked up cold.
|
||||
|
||||
**Also corrected in place:** the ranking paragraph's item 1 said the snapshot mitigation was *"armed
|
||||
(daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong** — seven
|
||||
snapshots exist (R-429) and "armed" was withdrawn the same day. **The ORDER of that list is
|
||||
unchanged.**
|
||||
|
||||
## 6. Hub deploy and its verification
|
||||
|
||||
| step | result |
|
||||
|---|---|
|
||||
| image built + pushed | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` (25 M) |
|
||||
| clean-tree gate before build | `git status --porcelain` empty, `HEAD == origin/main` @ `db38f4c` |
|
||||
| manifest bumped | `manifests/hub.yaml:128` → `:0.111.1`, committed `0f65f7a`, pushed |
|
||||
| ArgoCD hard refresh + **deliberate** sync | `sync=Synced` `health=Healthy` `rev=0f65f7a…` — the revision equals HEAD |
|
||||
| rollout | `deployment "hub" successfully rolled out` |
|
||||
| deployed image | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` |
|
||||
| running binary | `felhom-hub 0.111.1 (built 2026-09-01T16:35:04Z)` |
|
||||
|
||||
No `kubectl set image`, no `kubectl apply`.
|
||||
|
||||
### The live trigger was NOT run, and the reason is not a shortcut
|
||||
|
||||
The brief asked me to *"trigger the alarm once through the hub's own path"*. **I did not, and I am
|
||||
naming it rather than quietly substituting.**
|
||||
|
||||
**This alarm only fires on a real fall in a real customer's snapshot count.** Firing it live therefore
|
||||
means one of two things: deleting a real customer's off-site backups (the destructive act this
|
||||
session explicitly is not), or **POSTing a falsified report claiming `demo-hp` had lost its
|
||||
backups** — which would write a fabricated data point permanently into that customer's report
|
||||
history, move its drop latch and baseline, and send Viktor a **second** alarm mail about a real box,
|
||||
five hours after the first one already confused him (`STATUS.md` item 5). Fabricating customer
|
||||
telemetry to satisfy a verification step is the wrong trade in a project whose whole doctrine is that
|
||||
a measurement must mean what it says.
|
||||
|
||||
**What was done instead covers both halves of what the live trigger was for:**
|
||||
|
||||
1. **Does the deployed artifact carry the sentence?** Proved on the bytes of the running pod's
|
||||
binary — §1, with both controls.
|
||||
2. **Does the production path emit it?** Proved by three tests that drive
|
||||
`saveOffsiteReport → Check() → notify` and read the message the operator would receive, plus the
|
||||
stored-row test that pins mail and audit row together — and all three red-proofed.
|
||||
|
||||
**What is therefore still unproven, stated plainly:** that the *dispatcher* delivers this particular
|
||||
message to a mailbox. That leg was exercised for real on 2026-09-01 at 12:29 UTC by the previous
|
||||
session with the old text, so the routing is known good; only the new wording has not travelled it.
|
||||
|
||||
## 7. Explicitly
|
||||
|
||||
**No controller release. No agent release. No image built for either. No golden baked, and none
|
||||
owed** — the golden-currency gate reads `newest released controller 0.232.0 / newest golden baked
|
||||
0.232.0`. **The floor is unchanged at 0.232.0.** `felhom-controller` and `felhom-agent` working trees
|
||||
were not touched.
|
||||
|
||||
`python3 scripts/unproven.py --summary`: **NOT WALKED 35 of 55 — unchanged.** No claim moved, which
|
||||
is correct: nothing was proven today. All **14** `felhom.eu` gates green before every push.
|
||||
|
||||
**CI, checked by run id against `head_sha` rather than assumed** (the R-417 recipe — the
|
||||
`actions/jobs` endpoint, paged to the end):
|
||||
|
||||
| run id | commit | conclusion |
|
||||
|---|---|---|
|
||||
| **500** | `db38f4c` — hub v0.111.1 + the stopping line | **success** |
|
||||
| **501** | `0f65f7a` — manifest bump to 0.111.1 | **success** |
|
||||
| **502** | `1a1b32b` — REPORT + R-437 | **success** |
|
||||
---
|
||||
|
||||
## 8. Observations
|
||||
|
||||
1. **`strings` stops at the em dash, so the alarm sentence appeared truncated in the deployed
|
||||
binary and briefly looked like a bad build.** `strings` scans ASCII by default and the message's
|
||||
`—` terminates the run; the fix is `LC_ALL=C grep -aoP` on the bytes.
|
||||
**NOT-A-FINDING:** this is the project's already-recorded accented-grep trap appearing on a new
|
||||
surface, so it needs no new row — it is a hazard of my verification method, not a defect in any
|
||||
product code. Recorded here so the next person grepping a binary for Hungarian or em-dashed copy
|
||||
does not read a truncation as a bad build, which is exactly how it read for a minute.
|
||||
1. **The restore path writes the recovery unit's OLD image pin into the live stack dir, and the
|
||||
catalog syncer overwrites it again within 15 minutes.** The overwrite half is measured on demo-hp;
|
||||
that the restore writes to that same path is read at `cmd/controller/main.go:2570`, not measured —
|
||||
both halves are graded as such in the row. **FILED: R-441**
|
||||
|
||||
2. **The dispatcher's cooldowns are in-memory and are lost on every hub restart**, so any deploy
|
||||
re-arms every alarm's 6-hour cooldown. **NOT-A-FINDING:** it is stated in
|
||||
`hub/internal/notify/dispatcher.go:18` as a deliberate, accepted trade. Noted because it is why a
|
||||
live trigger today would definitely have mailed, rather than being absorbed by yesterday's
|
||||
cooldown — it changed the decision in §6.
|
||||
2. **`remove_hdd_data: true` removed nothing.** 128 MB of app data stayed on the drive while the API
|
||||
returned 200 with `hdd_paths_removed:null, hdd_paths_preserved:null`. Root cause established with
|
||||
controls: `Paths.HDDPath` has no default and demo-hp's `controller.yaml` does not set it, so
|
||||
`ParseComposeHDDMounts` returns nil on its first line. **FILED: R-442**
|
||||
|
||||
3. **The brief's baseline for `felhom-agent` was two commits stale** — it names `058b945`
|
||||
(2026-08-23) while `main` is `4586f0f` (2026-09-01). **NOT-A-FINDING:** the agent was untouched
|
||||
either way, and `058b945` is a real commit, so nothing was ambiguous. Flagged only so the number
|
||||
is not copied forward into the next brief.
|
||||
3. **An update can report success over an app it has just broken.** HTTP 200 and "updated
|
||||
successfully" while the container was already crash-looping; the truth reached the customer 5m16s
|
||||
later through the dead-app alarm rather than through the update itself. **FILED: R-443**
|
||||
|
||||
4. **`OPEN-ITEMS.md` has no section called *"Decided — and what would reopen each"*; `STATUS.md`
|
||||
does.** The brief sent the register text to that section by name.
|
||||
**NOT-A-FINDING:** resolved in the session by writing a new register section in that section's
|
||||
voice — the decision, then the condition that reopens it — which is plainly what the instruction
|
||||
meant, so there is nothing left to file. Recorded only so the next session does not hunt
|
||||
`OPEN-ITEMS.md` for a heading that has never existed there.
|
||||
4. **The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed**, and `fstrim` from
|
||||
inside the unprivileged container is refused; `pct fstrim` from the host reclaimed it. Nothing runs
|
||||
it on the fleet. **FILED: R-444**
|
||||
|
||||
## 9. My own mistakes
|
||||
5. **The hub kept app telemetry for an app that no longer exists anywhere**, and it now sets a
|
||||
fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway.
|
||||
Retained rather than cleared, because the reset is irreversible and on the operator's own surface.
|
||||
**FILED: R-445**
|
||||
|
||||
* **I wrote "blocked on R-433" into R-434 yesterday, and it was wrong.** It cost nothing because the
|
||||
block lasted one day, but the reasoning error is the interesting part: I treated *"we do not know
|
||||
what is true"* as a reason not to touch a sentence that was **known to be false**. Removing a false
|
||||
claim never needs the true one. The row and the code comment now say so, and it is the only part of
|
||||
this session I would call a lesson rather than a task.
|
||||
* **I asserted the brief's "six rows" instead of counting, for about ten minutes.** I wrote the
|
||||
register block with "six" in it before running the count that produced eleven, and only caught it
|
||||
because I decided to enumerate the rows rather than describe them — which the brief had insisted
|
||||
on for a different reason. **The instruction that saved it was not the one aimed at this.**
|
||||
* **My first `[BETA-DEFERRED]` claim over-promised.** I wrote in the register that the grep "returns
|
||||
the set as a group and nothing else", then found the grep returns eight lines because §8's own
|
||||
header defines the marker. Corrected in place with the real count. It is a small instance of
|
||||
exactly the class the register spent 2026-09-01 documenting — a claim about an instrument that had
|
||||
not been run.
|
||||
* **I preserved `REPORT.md` twice in one day and should have noticed the pattern the first time.**
|
||||
`REPORT.md` held the only copy of the R-331 report this morning and the only copy of the drill
|
||||
report this afternoon; both are now siblings (`REPORT-r331-backup-card.md`,
|
||||
`REPORT-drill-r95-recovery-2026-09-01.md`). The convention says durable content may not live only
|
||||
in the overwritten file, and it has now been violated twice in a day by two different sessions —
|
||||
which suggests the convention needs a gate, not more diligence. **Not filed:** I am not filing a
|
||||
row for it in a session that already declined a register sweep; it is named here for whoever picks
|
||||
up R-437.
|
||||
6. **The syncer's debug hash line cannot show what it claims to show.** `logFileHashes`
|
||||
(`internal/sync/sync.go:386`) reads the destination *after* the write, so it printed
|
||||
`src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word
|
||||
"changed". **NOT-A-FINDING:** it is DEBUG-only and the `Updated <app>/<file>` INFO line immediately
|
||||
above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the
|
||||
next person reading a sync log does not try to learn from those two hashes, which cannot differ.
|
||||
|
||||
---
|
||||
|
||||
## 9. Teardown — all three layers
|
||||
|
||||
**Layer 1 — the machine.** `bentopdf` restored to its catalog tag `v2.8.6`, digest
|
||||
`sha256:eaeea1e447205a79…`, **byte-identical to the run's baseline**. The throwaway `nextcloud` stack
|
||||
removed via the product's own endpoint; containers, all three volumes and `app.yaml` gone. The 128 MB
|
||||
the product failed to remove (observation 2) deleted by hand along with its backup dirs;
|
||||
`find /mnt -iname "*nextcloud*"` returns nothing. Five images this run pulled removed by **targeted
|
||||
`docker rmi`** — **no `prune` of any kind was run anywhere**. No guest was created; guest 9201 was
|
||||
hard-reset once by design and returned with all nine apps.
|
||||
|
||||
**Layer 2 — the host.** `local-lvm` 68.97% → 70.91% during the run → **26.78%** after `pct fstrim
|
||||
9201`. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to
|
||||
pre-run values (`/` 957 M, `/mnt/sys_drive` 12 G, `hdd_1` 5.5 G).
|
||||
|
||||
**Layer 3 — the hub. This run provisioned NOTHING.** No customer record and no appliance record was
|
||||
created; the customers list is unchanged at five rows, identical to the list read at the start. The
|
||||
existing `demo-hp` customer was used. What the run *did* create is hub **events** (`app_start_failed`,
|
||||
plus deploy/remove for the throwaway app) — **retained deliberately**, because the event log is an
|
||||
append-only record and deleting from it to tidy a test damages the surface this project relies on for
|
||||
history. One residue is **retained rather than cleared** with the reason and the exact one-line command
|
||||
recorded in the findings doc §13 (observation 5).
|
||||
|
||||
**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified. Peti's box was never contacted.**
|
||||
|
||||
---
|
||||
|
||||
## 10. Final verification
|
||||
|
||||
```
|
||||
felhom-controller: git status --porcelain → empty
|
||||
HEAD = origin/main = 960d29b0612c
|
||||
go build ./... → OK
|
||||
go vet ./... → OK
|
||||
go test ./... → rc=0, 28 packages, 0 FAIL
|
||||
```
|
||||
|
||||
**The controller tree was left untouched, and that is proven rather than asserted.**
|
||||
|
||||
---
|
||||
|
||||
## 11. My own mistakes
|
||||
|
||||
1. **My first customer-page search used a regex where I needed a literal.** `grep -o "2.8.6"` treats
|
||||
`.` as a wildcard and reported two hits on a page that contains none. I caught it only because the
|
||||
task requires a control on every such search, and re-ran with `grep -F`. **The rule earned its
|
||||
place in the same session it was applied.**
|
||||
2. **My first poll for "nextcloud is healthy" matched the wrong container.** The break condition
|
||||
matched `healthy` anywhere in the line and fired on `nextcloud-redis`. Corrected to an exact-name
|
||||
filter. No result depended on it.
|
||||
3. **I renumbered `STATUS.md` badly on the first attempt**, leaving the list running 1–6, 8, 9 with no
|
||||
item 7, and left the section's own lead line saying one thing was waiting when there were two.
|
||||
Both fixed. **This is the ranking-paragraph-goes-stale trap (R-405) in miniature**, and it appeared
|
||||
within minutes of my writing about it.
|
||||
|
||||
@@ -18,7 +18,7 @@ not an evening's work.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **One thing is waiting on you — item 4 (send two e-mails).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision, new today).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
the real machines:
|
||||
- the background job that could delete a live restore's lock now waits its turn — and the check
|
||||
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
||||
@@ -70,7 +70,48 @@ nothing.*
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
7. **Where should the safety go before an app updates? This is the one decision from today's
|
||||
measurement, and it is a design choice, not a bug report.**
|
||||
|
||||
**What I measured.** The box downloads new app versions by itself every 15 minutes and writes them
|
||||
into the customer's files, whether the app is running or not. Nothing tells the customer. Then the
|
||||
**Restart** button — not just Update — installs that new version. I watched it download a version
|
||||
that was not on the machine and swap the app onto it in 18 seconds. **And the box does it on its
|
||||
own** when an app fails to come back after a crash: nobody pressed anything.
|
||||
|
||||
**One fear is smaller than we thought, and you should have that too.** A plain power cut does
|
||||
**not** upgrade anything. The apps come back on their old version. It only happens when an app
|
||||
fails to return.
|
||||
|
||||
**One fear is bigger.** I tested whether we can undo an app update. **We cannot.** Once an app has
|
||||
moved its data to the new version, putting the old version back gives an app that will not start at
|
||||
all. So **"rollback" is the wrong word** and I have struck it. The only way back is to restore the
|
||||
customer's data from a copy taken *before* the update — and today no update takes one.
|
||||
|
||||
**The decision, in one sentence: should the safety copy sit under the Update button only, or under
|
||||
everything that can install a new version?**
|
||||
|
||||
- **Under the button only.** Cheap and quick. Covers the case a customer causes. **Leaves the
|
||||
unattended path uncovered** — the one where an app that failed to come back is upgraded with
|
||||
nobody watching.
|
||||
- **Under everything.** Covers all of it. Costs more, and it has a hard limit I measured: for a
|
||||
big app a copy is roughly 30 minutes and about twice the app's size, against a standard box that
|
||||
ships with 20 GB. **A copy of a large app does not fit.** So this option cannot be built without
|
||||
also answering where the copy lives.
|
||||
|
||||
**My pick: under everything — but decide the "where does it live" question first**, because the
|
||||
answer decides whether the rest is even buildable.
|
||||
|
||||
**If you do nothing:** nothing breaks today, and no customer is at risk this week — the fleet is
|
||||
young and its running versions match the catalog. But the exposure is real and dated: **Peti's box
|
||||
has a one-major upgrade of `rallly` queued behind its next boot**, from a catalog change made three
|
||||
days after it went quiet. **Who else is blocked:** nobody can spec the safe-update work until this
|
||||
is answered, because the two options produce different products.
|
||||
|
||||
Full measurement, with the controls and the quoted output:
|
||||
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
|
||||
|
||||
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
|
||||
@@ -99,7 +99,10 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Scenario | Components | Status | Evidence | Gap / roadmap |
|
||||
|---|---|---|---|---|
|
||||
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
|
||||
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3` | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
|
||||
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
|
||||
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller | **PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
|
||||
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
|
||||
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
|
||||
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
|
||||
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
|
||||
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
|
||||
|
||||
@@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap
|
||||
### `sync/`
|
||||
| File | Class | Reason | Risk |
|
||||
|---|---|---|---|
|
||||
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). | clean |
|
||||
| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean |
|
||||
|
||||
### `system/` — split per-function (not per-file)
|
||||
| File | Class | Reason | Risk |
|
||||
@@ -503,3 +503,77 @@ own; every caller that is not the customer must decide for itself whether the ap
|
||||
- **S6:** §5(3) self-restore-test → **status-display only**; the agent owns orchestration.
|
||||
- **Self-update resolved (03 §11):** `updater.go` → **DELETE(→agent)**, `state.go` →
|
||||
DELETE(obsolete), `version.go` KEEP; §6 + §5(2) updated (bulk = `backup=0` mountpoint recipe).
|
||||
|
||||
|
||||
---
|
||||
|
||||
## The app-definition seam — what happens to a deployed app when its compose file is rewritten
|
||||
|
||||
> **STATUS: MEASURED BEHAVIOUR, NOT A RECORDED DESIGN.** This section describes what the system was
|
||||
> observed to do on 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`), with controls. It is
|
||||
> deliberately **not** marked `[DESIGN]`, because the operator has not ruled on whether the behaviour
|
||||
> was intended. **Do not read this as an endorsement, and do not spec against it as though it were
|
||||
> settled.** The open questions are R-438 and R-441.
|
||||
|
||||
Until this was measured, no architecture document said what happens here, and the gap itself is
|
||||
R-438. The three facts below are the ones a reader needs before touching any of it.
|
||||
|
||||
### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle
|
||||
|
||||
`Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies
|
||||
`docker-compose.yml` and `.felhom.yml` into the matching stack folder. **There is no test of whether
|
||||
the app is deployed.** The only guard is a sha256 content compare (`copyIfChanged`) and the only
|
||||
exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus
|
||||
one immediate sync at controller start (`sync.go:98`).
|
||||
|
||||
**It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing
|
||||
else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree,
|
||||
and they stay that way until something else acts.
|
||||
|
||||
### 2. Every lifecycle path resolves that disagreement, silently, by upgrading
|
||||
|
||||
`StartStack`, `RestartStack` and `UpdateStack` (`stacks/manager.go:1029 / 1133 / 1170`) all end in
|
||||
`docker compose up -d`. `up -d` makes the container match the file, **and pulls the image itself if it
|
||||
is absent** — measured at 18.3 s with a pull versus 0.5 s without, against a negative control
|
||||
(unchanged file) that did not even recreate the container.
|
||||
|
||||
`RestartStack` carries an explicit in-source comment saying this is deliberate — *"so that … any
|
||||
template changes (new images, healthchecks) are picked up"*. **That is a fact about the source, not an
|
||||
operator ruling**, and it speaks only for the customer-pressed restart.
|
||||
|
||||
**Thirteen non-API call sites across nine files** reach `up -d` without anyone pressing anything. The
|
||||
full table is §8 of the spike doc; the three that matter most are:
|
||||
|
||||
- `bootrecon.Reconciler.Run` → `StartStack` — `bootrecon/bootrecon.go:269`
|
||||
- `backup.AppStopGuard.Recover` → `StartStack` — `backup/appstop_marker.go:283`
|
||||
- the drive-return gate, `web.Server.restartStacks` → `StartStack` — `web/intermediary.go:222`
|
||||
|
||||
**A plain power cut does NOT trigger this.** Docker's `restart: unless-stopped` restores the existing
|
||||
containers on the old image, the reconciler finds no orphan, and it logs so
|
||||
(`no boot-orphaned apps (nothing to start)`). The unattended upgrade needs the narrower precondition
|
||||
*"and the app did not come back"* — which was measured, and does upgrade.
|
||||
|
||||
### 3. Nothing takes a copy first, and the tag cannot be put back afterwards
|
||||
|
||||
`UpdateStack` runs `compose pull` then `compose up -d --remove-orphans` and does nothing else — no
|
||||
dump, no copy, no hold. The R-361 safety dump (`backup/offbox_reconstitute.go:207`) is **database-only**
|
||||
and is not on this path at all; an app with no database gets nothing from it even on the paths where it
|
||||
does run.
|
||||
|
||||
**And restoring the old tag is not a rollback.** Once an app has migrated its data, the old image
|
||||
refuses to start — measured on Nextcloud: *"the version of the data (32.0.9.2) is higher than the
|
||||
docker image version (31.0.14.1) and downgrading is not supported"*. The data itself survives; only the
|
||||
downgrade is blocked. So the only route back is a **data** restore from a copy taken before the
|
||||
update.
|
||||
|
||||
**And that route fights this seam** (R-441): `stackAdapter.RecreateStackDefinitionFromUnit`
|
||||
(`cmd/controller/main.go:2570`) writes the recovery unit's captured compose — with the OLD pin — into
|
||||
the live stack dir, and `copyIfChanged` overwrites it again on the next tick. The overwrite is
|
||||
measured; that the restore writes to that path is read, not measured.
|
||||
|
||||
### What a change here must not break
|
||||
|
||||
- The syncer's overwrite is also the **repair** path: a locally-corrupted compose is replaced by the
|
||||
catalog's version within 15 minutes (observed). Any "don't touch deployed apps" rule loses that.
|
||||
- `up -d`-on-restart is what injects `app.yaml` env into a running stack. Reverting to
|
||||
`docker compose restart` would silently stop doing that.
|
||||
|
||||
@@ -0,0 +1,626 @@
|
||||
# SPIKE — what an app update actually does, and which other paths do it too (2026-09-01)
|
||||
|
||||
> **THE ANSWER TO PHASE 1, IN ONE SENTENCE: YES — the Restart button upgrades the app, and so does
|
||||
> the box's own repair when an app fails to come back, because every one of these paths ends in
|
||||
> `docker compose up -d`, and `up -d` makes the container match whatever the file now says, pulling
|
||||
> the image itself if it is missing.**
|
||||
>
|
||||
> **AND THE SECOND ANSWER, WHICH THE OPERATOR PAGE GOT WRONG IN THE HELPFUL DIRECTION: a plain power
|
||||
> cut does NOT do this.** Docker's own `restart: unless-stopped` puts the existing containers back on
|
||||
> the OLD image, so the boot reconciler finds nothing to repair and never runs `up -d`. The unattended
|
||||
> upgrade happens only in the narrower case where an app does **not** come back by itself — and there
|
||||
> it happens with nobody pressing anything.
|
||||
>
|
||||
> **AND THE THIRD, WHICH DECIDES THE VOCABULARY OF THIS WHOLE ARC: app data CANNOT be rolled back.**
|
||||
> Measured on Nextcloud: once a migration has actually run, putting the old image tag back produces a
|
||||
> container that refuses to start — *"the version of the data (32.0.9.2) is higher than the docker
|
||||
> image version (31.0.14.1) and downgrading is not supported"*. **So "rollback" is the wrong word and
|
||||
> should be struck before anyone specs against it.** The shape that is actually available is
|
||||
> *pre-update copy plus refuse-and-explain*.
|
||||
|
||||
**Scope.** A spike. **No production code was written, in any repo.** The controller was read only:
|
||||
`felhom-controller` is at `960d29b0612c` before and after, working tree clean, and
|
||||
`go build ./... && go vet ./... && go test ./...` is green (28 packages, 0 FAIL) — run at the end
|
||||
precisely to prove the tree was left untouched.
|
||||
|
||||
**Method.** Endpoint-level throughout, which is the standard method here because there is no browser on
|
||||
DooPlex: every action was the exact endpoint the UI invokes (`POST /api/stacks/{name}/restart`,
|
||||
`/update`, `/start`, `/stop`, `/deploy`, `/remove`), driven over HTTPS with the customer's own session
|
||||
cookie and CSRF token. **No raw agent attach, no hand-set state — the F9 bypass was not used.** The two
|
||||
places where state was staged by hand are named as such at the point of use (Phase 1's compose edits,
|
||||
and Phase 1c-ii's `docker rm -f`), and each says exactly what was staged and why.
|
||||
|
||||
**Target.** All destructive work ran on **`demo-hp` (192.168.0.104, `ssh hp`; guest 9201 =
|
||||
192.168.0.138)**, which `runbooks/target-selection.md` classes **Tier 0 — disposable**. DooPlex,
|
||||
`demo-felhom`, `ep0` and Peti's box were not mutated. Peti's box was not touched at all.
|
||||
|
||||
**Evidence.** `audits/evidence-spike-app-update-2026-09-01/` — written straight onto DooPlex as each
|
||||
phase produced it, before every revert, so R-96 rule 5 is satisfied continuously rather than at the
|
||||
end. Nothing was lost and nothing had to be reproduced.
|
||||
|
||||
---
|
||||
|
||||
## 1. Confirmed baselines — all four matched §1 of the task, none had moved
|
||||
|
||||
| Repo | `main` @ commit at start | at end | note |
|
||||
|---|---|---|---|
|
||||
| felhom-controller | `960d29b0612c` | `960d29b0612c` | **read only — unchanged** |
|
||||
| felhom.eu | `1d59353df437` | moved (this spike's documents) | Phase 0 + Phase 7 |
|
||||
| app-catalog-felhom.eu | `29edad9c5bf4` | moved, then **reverted to the same tree** | `214d448` → `30bd892` → `5d8f25f` |
|
||||
| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched |
|
||||
|
||||
Live controller on demo-hp: **0.232.0**, matching the newest release. Sync interval **15m**
|
||||
(`internal/config/config.go:351`, the default, no override on the box).
|
||||
|
||||
---
|
||||
|
||||
## 2. Phase 1 — THE GATE. Does `compose up -d` upgrade an app whose compose file already moved?
|
||||
|
||||
App: **`bentopdf`** — one container, no database, no volume, no data of any kind, and deployed on
|
||||
`demo-hp` only. Chosen so the gate could be measured with zero data risk anywhere.
|
||||
|
||||
Baseline: container `9473b6e4a18f`, `ghcr.io/alam00000/bentopdf:v2.8.6`, digest
|
||||
`sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3`, started `04:05:50Z`.
|
||||
|
||||
| # | Variant | Result | Discriminator |
|
||||
|---|---|---|---|
|
||||
| **1d** | **negative control** — file NOT edited, then `restart` | **no change** | container id, image id, digest and `StartedAt` all IDENTICAL. `up -d` saw no change and did not even recreate. |
|
||||
| **1a** | `restart`, target image **ABSENT** from the local store | **UPGRADED, and it pulled** | new container `93e9db74e97f`, image `v2.8.5`, digest `sha256:2d867aac…`, and `v2.8.5` **appeared in `docker images`** where it had not been. **18.3 s.** |
|
||||
| **1b** | `restart`, target image **ALREADY PRESENT** | **UPGRADED, no pull** | new container `222f8178dbd0`, image id unchanged from baseline `4baadf01bd68`. **0.5 s.** |
|
||||
| **1c** | **boot recovery** — hard guest reset, app left running | **NO upgrade** | SAME container `222f8178dbd0` restored by Docker at `17:47:45Z` on `v2.8.6` while the file said `v2.8.5`. |
|
||||
| **1c-ii** | boot recovery where the app did **not** come back | **UPGRADED, unattended** | new container `aa443836a4cd` on `v2.8.5` at `17:55:44Z`. Nobody pressed anything. |
|
||||
|
||||
**The 18.3 s vs 0.5 s split is the quantitative discriminator** and it is worth more than the tags:
|
||||
the restart path contains no pull step in the source, yet variant 1a spent 18 seconds and ended with a
|
||||
new image in the local store. `up -d` pulled it.
|
||||
|
||||
Quoted live, variant 1a (`phase1-1a-controller-log.txt`):
|
||||
|
||||
```
|
||||
17:35:45 router.go:566: [INFO] [api] restart requested for stack: bentopdf
|
||||
17:35:45 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
|
||||
17:36:03 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 18.3s)
|
||||
17:36:07 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running
|
||||
```
|
||||
|
||||
Quoted live, variant 1c-ii — **the unattended upgrade, attributed to the exact symbol**
|
||||
(`phase1-1cii-orphan.txt`):
|
||||
|
||||
```
|
||||
17:55:44 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 10s — sweeping
|
||||
17:55:44 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]
|
||||
17:55:44 manager.go:1039: [INFO] [stacks] Starting stack: bentopdf
|
||||
17:55:44 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
|
||||
17:55:48 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running
|
||||
```
|
||||
|
||||
**1c's negative result is a POSITIVE observable, not an absent log line** — the reconciler says so in
|
||||
its own words, which is exactly what it was built to do:
|
||||
|
||||
```
|
||||
17:48:34 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
|
||||
```
|
||||
|
||||
### What 1c-ii staged, stated plainly
|
||||
|
||||
`docker rm -f bentopdf` removed the container so it could not be restored by Docker's restart policy,
|
||||
which is the shape a power cut leaves when an app does not come back. `desired_state: running` was left
|
||||
untouched in `app.yaml`. The controller was then restarted, because that is what runs the boot
|
||||
reconciler. **Nothing else was staged**, and the reconciler selected the app on its own.
|
||||
|
||||
### Why 1c could not use a hand-edited file, and how that was handled
|
||||
|
||||
At controller start the initial catalog sync runs immediately (`sync.go:98`, `17:47:49Z`) while the
|
||||
boot reconciler waits out `bootReconcileSettle` = 5 s plus a settle window (`17:48:34Z`). **The sync
|
||||
wins by ~45 s**, so a hand-edited compose would have been overwritten before the reconciler ever saw
|
||||
it, and 1c would have measured nothing. 1c and 1c-ii were therefore run inside Phase 2's window, where
|
||||
the **catalog itself** carried the new tag — which is also the faithful production shape.
|
||||
|
||||
### The design intent is already stated in the source, and it is not hidden
|
||||
|
||||
`Manager.RestartStack` (`internal/stacks/manager.go:1133`) carries this comment:
|
||||
|
||||
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
|
||||
> changes (new images, healthchecks) are picked up."*
|
||||
|
||||
**So the restart behaviour was chosen, deliberately, and written down.** What is NOT written down
|
||||
anywhere is the consequence once the catalog syncer moves the file underneath a deployed app. That gap
|
||||
is R-438, and whether the choice should extend to the unattended paths is the operator's ruling.
|
||||
|
||||
---
|
||||
|
||||
## 3. Phase 2 — the sync overwrite, confirmed live
|
||||
|
||||
A real tag change (`bentopdf` `v2.8.6` → `v2.8.5`) was pushed to the catalog `main` at **17:38:22Z**
|
||||
(`214d448`) and left to travel the real 15-minute cycle. No hand-edit; the catalog was the only input.
|
||||
|
||||
| time (UTC) | file on demo-hp | running container |
|
||||
|---|---|---|
|
||||
| 17:38:35 (pre-sync) | `v2.8.6` | `v2.8.6` |
|
||||
| 17:45:15 | `v2.8.6` | `v2.8.6` |
|
||||
| **17:45:37 (post-sync)** | **`v2.8.5`** | **`v2.8.6`** |
|
||||
|
||||
The sync's own line, and the file's mtime, agree to the second:
|
||||
|
||||
```
|
||||
17:45:17 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
|
||||
mtime: 2026-09-01 17:45:17.646017366 +0000 /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
container started: 2026-09-01T17:36:35Z (unchanged)
|
||||
```
|
||||
|
||||
`Syncer.copyTemplates` (`internal/sync/sync.go:319`) has **no deployed check of any kind**. Its only
|
||||
guard is a sha256 content compare in `copyIfChanged`, and its only exclusion is `app.yaml`. The
|
||||
post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the sync does not
|
||||
restart anything**, which is why the file and the container can disagree indefinitely.
|
||||
|
||||
### Was the customer told? **No — and the page does not even show a version.**
|
||||
|
||||
Searched with ASCII fragments and both controls, per the standing rule.
|
||||
|
||||
| page | `BentoPDF` (positive control) | `zzz-never-present` (negative control) | `v2.8.5` | `v2.8.6` |
|
||||
|---|---|---|---|---|
|
||||
| `/apps/bentopdf` (40 360 B) | 4 | 0 | **0** | **0** |
|
||||
| `/stacks` (140 883 B) | 1 | 0 | **0** | **0** |
|
||||
|
||||
**A correction to my own method, made here rather than buried:** the first pass used
|
||||
`grep -o "2.8.6"`, where `.` is a regex wildcard, and it reported 2 hits. Re-run with `grep -F` the
|
||||
count is **0**. The controls are what exposed it. The customer's pages carry **no version string at
|
||||
all** — not the running one, not the pending one.
|
||||
|
||||
No event, no notification and no email were emitted by the sync. The only `Friss…` strings on the
|
||||
pages are the catalog-sync toast and the `Frissítés` button label.
|
||||
|
||||
### The revert, verified on the box
|
||||
|
||||
`30bd892` reverted the pin; the sync landed it at **18:10:29Z**
|
||||
(`[INFO] [sync] Updated bentopdf/docker-compose.yml`, `Sablonok frissítve — frissítve: bentopdf`) and
|
||||
the file read `v2.8.6` again. **The catalog tree is byte-identical to `29edad9c5bf4`.**
|
||||
|
||||
**An unlooked-for second result:** that same sync also overwrote a hand-broken compose file (Phase 3's
|
||||
`alpine:3.20` edit) with the catalog's version. **So the syncer is also the repair path for a broken
|
||||
app definition — within 15 minutes, automatically.** The container stays broken until something runs
|
||||
`up -d`, but the definition heals itself.
|
||||
|
||||
---
|
||||
|
||||
## 4. Phase 3 — what a failed update looks like to the customer
|
||||
|
||||
### 3a — the pull fails. **Loud, honest, and harmless.**
|
||||
|
||||
Compose pointed at `ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist`, then `POST .../update`:
|
||||
|
||||
```
|
||||
HTTP 500
|
||||
{"ok":false,"error":"pulling images for bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ... Error manifest unknown\nError response from daemon: manifest unknown"}
|
||||
```
|
||||
|
||||
**The app was untouched:** same container `aa443836a4cd`, `v2.8.5`, `status=running`, `RestartCount=0`.
|
||||
`UpdateStack` returns after the failed `pull` and never reaches `up -d`.
|
||||
|
||||
**3a-ii, the landmine that was not one.** The failed update leaves the file pointing at an image that
|
||||
does not exist, while the page shows a green healthy app with a Restart button beside the Update
|
||||
button. Pressing Restart in that state was measured: **HTTP 500, and the app still ran.** `compose up
|
||||
-d` resolves every image before it touches a container, so a bad reference fails before anything stops.
|
||||
**This is the reassuring half of Phase 3 and it should be said as plainly as the bad half.**
|
||||
|
||||
**What the customer is shown, though, is raw English Docker output** — `exit code 1`, `stderr:`,
|
||||
`manifest unknown`, `Error response from daemon` — on a product whose every other error string is
|
||||
Hungarian.
|
||||
|
||||
### 3b — the pull succeeds and the app does not. **This is the one that matters.**
|
||||
|
||||
Compose pointed at `alpine:3.20` — a real image that pulls cleanly and then exits at once, so
|
||||
`restart: unless-stopped` puts it in a crash loop. This is the shape of a real upstream image whose
|
||||
configuration the template can no longer supply, and the catalog's own history contains exactly that
|
||||
case (`7350cd9`, *"wger: revert 2.6 -> 2.3 (2.6 needs a full DB config the template cannot supply)"*).
|
||||
|
||||
```
|
||||
POST /api/stacks/bentopdf/update
|
||||
HTTP 200
|
||||
{"ok":true,"message":"Stack bentopdf update completed"}
|
||||
|
||||
18:00:47 manager.go:1195: [INFO] [stacks] Stack bentopdf updated successfully (took 3.5s)
|
||||
```
|
||||
|
||||
Reality 37 seconds later: `status=restarting`, `RestartCount=9`, `alpine:3.20`.
|
||||
|
||||
**The controller's own post-start line told the truth** — and it ran *after* the API had already
|
||||
answered `ok:true`:
|
||||
|
||||
```
|
||||
18:00:50 manager.go:1403: [INFO] [stacks] bentopdf alpine:3.20 restarting Restarting (0) Less than a second ago
|
||||
```
|
||||
|
||||
**What the customer's page said** (`/stacks`, live HTML):
|
||||
|
||||
> `BentoPDF` · `pdf.enkisfelhom.hu` · **„URL nem elérhető – útvonal nincs publikálva"** ·
|
||||
> badge **„Újraindítás…"** · `Restarting (0) 15 seconds ago` ·
|
||||
> buttons: **`Frissítés` `Újraindítás` `Leállítás` `Naplók` `Részletek`**
|
||||
|
||||
**„Újraindítás…" reads as transient, not as failure.** Nothing on the page says the update broke the
|
||||
app. `isOperationalState` (`internal/web/funcmap.go:90`) counts `StateRestarting` and `StateDegraded`
|
||||
as operational, so the full button row — including the green `Frissítés` — is rendered over a
|
||||
crash-looping app.
|
||||
|
||||
**Is there a route back? Not from the page.** Every button offered re-runs the same broken definition
|
||||
or stops the app. The compose file is not customer-editable. The routes that exist are (a) the catalog
|
||||
being corrected, which then heals the file within 15 minutes, or (b) a restore, which has its own
|
||||
problem — see §7.
|
||||
|
||||
### The alarm DOES fire — 5 minutes 16 seconds later, and by a different road
|
||||
|
||||
This was measured to a positive observable rather than inferred from silence:
|
||||
|
||||
```
|
||||
18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
```
|
||||
|
||||
Update was at `18:00:43`; the alarm at `18:05:59`. The delay is `crashLoopAfter = 5 * time.Minute`
|
||||
(`internal/stacks/manager.go`), and it is deliberate and well-argued in its own comment — a shorter
|
||||
threshold would alarm on every routine deploy. **The severity word is `warning`, which is inside the
|
||||
hub's exact vocabulary**, so this one really does reach the customer (R-328/R-329 class avoided).
|
||||
|
||||
**So the honest summary of 3b is not "silently broken".** It is: *the button lied at the moment it was
|
||||
pressed, the page then described a failure as a restart, and the truth arrived five minutes later
|
||||
through the dead-app alarm rather than through the update the customer actually performed.*
|
||||
|
||||
---
|
||||
|
||||
## 5. Phase 4 — how far behind is the real fleet
|
||||
|
||||
**Read from the running container, never from the file** — the file is the thing that has already
|
||||
moved. Both tag and digest recorded.
|
||||
|
||||
### demo-hp — **zero drift by tag.** Nine deployed apps, every running tag equal to the catalog pin.
|
||||
|
||||
This is not luck: the box was reinstalled 2026-08-21 and its apps were deployed ~13 h before this run.
|
||||
It is a *young* box, and it is the reason Phase 5 could not be priced on it (§6).
|
||||
|
||||
### demo-felhom — one deployed app, `opengist`, running `ghcr.io/thomiceli/opengist:1.13` = catalog pin.
|
||||
|
||||
### But "up to date by tag" is not up to date. **Two floating pins have already moved upstream.**
|
||||
|
||||
Running digests compared against what the registry serves for the same tag today, with fully-pinned
|
||||
tags as the control:
|
||||
|
||||
| image | kind | verdict |
|
||||
|---|---|---|
|
||||
| `postgres:16-alpine` | floating | SAME |
|
||||
| `redis:7-alpine` | floating | SAME |
|
||||
| `mariadb:11.6` | floating | SAME |
|
||||
| `ghcr.io/thomiceli/opengist:1.13` | floating | SAME |
|
||||
| **`mariadb:11.4`** | **floating** | **MOVED** — running `sha256:4f1d8d20…`, upstream now `sha256:611a2fcc…` |
|
||||
| **`mariadb:12.3`** | **floating** | **MOVED** — running `sha256:a02fe89c…`, upstream now `sha256:dd9b303a…` |
|
||||
| `rommapp/romm:5.0.0` | **pinned (CONTROL)** | SAME |
|
||||
| `privatebin/pdo:2.0.5` | **pinned (CONTROL)** | SAME |
|
||||
|
||||
**Both controls held and both positives are floating tags.** So on a box with no visible drift at all,
|
||||
pressing `Frissítés` today would silently swap the **database engine build** under `romm` (MariaDB
|
||||
11.4) and `bookstack` (MariaDB 12.3) — with no catalog change, no version change on any screen, and no
|
||||
record anywhere of what it was before. That is R-440, no longer as an argument but as a measurement.
|
||||
|
||||
### Peti's box — **NOT measurable, and the task's premise here was wrong**
|
||||
|
||||
The task asks Phase 4 to inspect three boxes. **`runbooks/target-selection.md:161` states plainly:
|
||||
*"Currently DOWN, no enrolled host. No access route from DooPlex, and nothing here needs one."*** The
|
||||
hub agrees: the Hosts page lists exactly **two** enrolled hosts (`demo-felhom-8363b5`,
|
||||
`demo-hp-bb76ea`), and the `peti-felhom` customer shows **status DOWN, last report 48 d ago, controller
|
||||
0.115.0**.
|
||||
|
||||
So there is no running container to read, and **the hub does not record image tags at all** — the
|
||||
report's container payload carries name, state, CPU and memory, and no image field. The row is
|
||||
therefore recorded as **UNKNOWN**, not guessed.
|
||||
|
||||
**What IS knowable read-only, and it matters:** the box last reported ~2026-07-15 running `rallly` +
|
||||
`rallly-postgres`. On **2026-07-18 — three days after it went quiet — the catalog moved
|
||||
`rallly` `3.11.2` → `4.11.1` [MAJOR]** (`e3f3a81`). So a one-major upgrade is queued behind that box's
|
||||
next boot. **On this spike's own measurements that upgrade will NOT fire on the power-on itself**
|
||||
(1c: Docker restores the containers on the old image), **but it will fire the moment any app fails to
|
||||
come back, or anyone presses Restart or Update.** Nothing was touched on that box to establish this;
|
||||
it is the catalog's git history plus the hub's own record.
|
||||
|
||||
---
|
||||
|
||||
## 6. Phase 5 — what a pre-update copy would cost
|
||||
|
||||
### The existing safety machinery is DATABASE-ONLY, and that is the headline
|
||||
|
||||
`Manager.writeSafetyDump` (`internal/backup/offbox_reconstitute.go:207`) discovers the app's databases
|
||||
and dumps each one. **An app with no database gets nothing at all** — `len(mine) == 0` returns an empty
|
||||
set. There is no file-level safety copy on the R-361 path.
|
||||
|
||||
### The database half is nearly free. Measured on demo-hp:
|
||||
|
||||
Every `.sql` dump on the box, including the real `pre-restore-*` undo copies from the August restore
|
||||
work: **48 KB – 395 KB**. Largest is `paperless-ngx-postgres.sql` at 395 065 bytes.
|
||||
|
||||
### The file half is the cost, and demo-hp cannot price it. Stated, not papered over.
|
||||
|
||||
The box holds a few MB of app *content*; the bulk is engine data directories:
|
||||
|
||||
| app | biggest components | total |
|
||||
|---|---|---|
|
||||
| kimai | `kimai_db_data` 166 M + `kimai_var` 56 M | **≈ 222 MB** |
|
||||
| romm | `romm_db_data` 166 M + `romm_redis_data` 25 M + roms 1.1 M | **≈ 192 MB** |
|
||||
| bookstack | `bookstack_db_data` 166 M + `bookstack_config` 6.8 M | **≈ 173 MB** |
|
||||
|
||||
**These are not customer-scale numbers** — 166 MB is a fresh MariaDB's preallocated files, not
|
||||
anyone's data. So rather than over-claim from a young box, the arc's already-measured figures are
|
||||
cited:
|
||||
|
||||
> **`CAMPAIGN-10-two-storage-soak-2026-07-31.md` §6, 66 restores plus an M-band point 327× larger:
|
||||
> backup ≈ 29 s + 17.4 s/GB, and a DB-backed app's recovery unit is 1.90× its data**
|
||||
> (21.1 GB of data produced a 40.2 GB unit).
|
||||
|
||||
Applying that to demo-hp's three heaviest: **≈ 32–33 s each**, unit ≈ 330–420 MB. Trivial.
|
||||
|
||||
**And that is exactly what makes the ceiling the real finding.** For a large catalog app — Immich,
|
||||
Nextcloud, Plex — the same arithmetic gives, for 100 GB of data, **≈ 29 minutes and a ~190 GB copy**,
|
||||
against a default appliance whose `/mnt/sys_drive` ships at **20 GB**. **A pre-update copy of a large
|
||||
app does not fit on a default box, and no amount of tuning the copy changes that.** Whatever is
|
||||
designed here has to answer that before it answers anything else. Where the copy lives, and whether it
|
||||
is a copy at all or a snapshot, is a design decision and is not proposed here.
|
||||
|
||||
---
|
||||
|
||||
## 7. Phase 6 — can app data be rolled back at all? **NO.**
|
||||
|
||||
Run on operator confirmation, on a throwaway Nextcloud on `demo-hp`. Nothing else on any box was
|
||||
involved; the app was created, used and destroyed inside this phase.
|
||||
|
||||
**Seeded on 31.0.14.1** (`occ status`), two independent markers — a database-backed system config value
|
||||
and a real file on disk:
|
||||
|
||||
```
|
||||
System config value spike_marker set to string SPIKE-SEEDED-ON-31.0.14-2026-09-01
|
||||
/var/www/html/data/admin/files/spike-marker.txt = SPIKE-FILE-CONTENT-31.0.14-2026-09-01
|
||||
occ files:scan admin → 5 Folders, 53 Files, 2 Updated, 0 Errors
|
||||
```
|
||||
|
||||
### 6a — the 3-major jump (31.0.14 → 34.0.1). **Refused by the app, reported as success by the product.**
|
||||
|
||||
```
|
||||
POST /api/stacks/nextcloud/update → HTTP 200 {"ok":true,"message":"Stack nextcloud update completed"}
|
||||
```
|
||||
|
||||
The container then crash-looped, saying:
|
||||
|
||||
> **`Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.`**
|
||||
> **`It is only possible to upgrade one major version at a time.`**
|
||||
|
||||
**This is R-40, measured.** The catalog really does carry this jump today: `5e2c1ae`, 2026-07-18,
|
||||
*"nextcloud: 31.0.14-apache -> 34.0.1-apache [MAJOR]"*.
|
||||
|
||||
**It was fully recoverable** — putting `31.0.14` back gave a healthy app in 9 seconds with both markers
|
||||
intact. **But that is only because nothing migrated.** The refusal is Nextcloud's own safety net doing
|
||||
its job, and it is why 6a is the *easy* case.
|
||||
|
||||
### 6b — a migration that actually runs (31.0.14 → 32.0.9, one major). **It ran.**
|
||||
|
||||
```
|
||||
Initializing nextcloud 32.0.9.2 ...
|
||||
Upgrading nextcloud from 31.0.14.1 ...
|
||||
Updated database
|
||||
Updated <dav> to 1.34.2 · <files> to 2.4.0 · <files_sharing> to 1.24.1 · … (20+ apps)
|
||||
```
|
||||
|
||||
Healthy in 31 s on `32.0.9.2`, both markers still readable.
|
||||
|
||||
### 6c — THE ROLLBACK ATTEMPT. **Refused. The old image will not start on migrated data.**
|
||||
|
||||
```
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker
|
||||
image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the
|
||||
newest image version?
|
||||
```
|
||||
|
||||
Crash loop, indefinitely.
|
||||
|
||||
### 6d — positive control: **the data is not destroyed, only the downgrade is blocked**
|
||||
|
||||
Putting `32.0.9` back gave a healthy app in 46 s, `occ status` `32.0.9.2`, and **both markers read back
|
||||
byte-identical to what was seeded on 31**. So the failure in 6c is a refusal, not corruption — which is
|
||||
the distinction that decides the remedy.
|
||||
|
||||
### What this settles
|
||||
|
||||
**Putting the old image tag back is not a rollback and must not be described as one.** The only route
|
||||
back from a migration that has run is **restoring the DATA from a copy taken before the update** —
|
||||
which is precisely the thing the update path does not take (§2 of the task, confirmed by reading
|
||||
`Manager.UpdateStack`, and confirmed live: no dump, no copy, no hold, in any of the six updates run
|
||||
here).
|
||||
|
||||
---
|
||||
|
||||
## 8. The exact symbols — found by reading, as required
|
||||
|
||||
Every path that ends in `compose up -d` on a customer's stack folder, with the enclosing function.
|
||||
|
||||
| what brings the app back | symbol | line |
|
||||
|---|---|---|
|
||||
| the four compose callers | `Manager.StartStack` / `StopStack` / `RestartStack` / `UpdateStack` | `internal/stacks/manager.go:1029 / 1106 / 1133 / 1170` |
|
||||
| partial start (R-47 DB-only window) | `Manager.StartStackServices` | `internal/stacks/manager.go:1082` |
|
||||
| **the boot reconciler** | `Reconciler.Run` → `r.stacks.StartStack(name)` | `internal/bootrecon/bootrecon.go:223` → **`:269`** |
|
||||
| its scheduler | `runBootReconcile` → `bootReconcileFn` | `cmd/controller/main.go:2127` → `:2172`, called at `:450` |
|
||||
| **the app-stop guard** | `AppStopGuard.Recover` → `g.starter.StartStack(name)` | `internal/backup/appstop_marker.go:264` → **`:283`** |
|
||||
| its hold-aware wrapper | `gatedAppStopStarter.StartStack` | `cmd/controller/main.go:2008` |
|
||||
| **the drive-return gate** | `Server.restartStacks` → `s.stackMgr.StartStack(name)` | `internal/web/intermediary.go:220` → **`:222`** |
|
||||
| guest-boot change handler | `Server.processGuestBootChange` | `internal/web/intermediary.go:395` → `:458` |
|
||||
| quiesce restart-after-backup | `Loop.restartAll` | `internal/quiesce/quiesce.go:730` → `:733` |
|
||||
| off-site reconstitution | `Manager.ReconstituteFromOffsite` | `internal/backup/offbox_reconstitute.go:480` → `:692` |
|
||||
| restore from unit / local / tier-2 | `restore_unit.go:381`, `restore.go:76`, `tier2_restore.go:418`, `backup.go:895` | — |
|
||||
| `.fab` export / import | `appexport/export.go:286`, `appexport/restore.go:461` | — |
|
||||
| integrations | `onlyoffice_filebrowser.go:61, :96` (`RestartStack`) | — |
|
||||
|
||||
**CORRECTION TO THE TASK: not five other paths — thirteen.** Excluding the three API actions
|
||||
(`start`/`restart`/`update`) and excluding interface declarations and adapters, there are **13 call
|
||||
sites across 9 files** that call `StartStack` or `RestartStack`, and every one of them ends in
|
||||
`docker compose up -d` against the live compose file.
|
||||
|
||||
---
|
||||
|
||||
## 9. R-439, confirmed by reading and refined
|
||||
|
||||
`Router.actionStack` (`internal/api/router.go:565`) checks the hold under
|
||||
`if action == "start" || action == "restart"` — **`update` is absent** and falls through to
|
||||
`UpdateStack`. The design intent is stated in `RestoreHoldFor`'s own comment
|
||||
(`internal/backup/offbox_reconstitute.go:323`):
|
||||
|
||||
> *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the
|
||||
> boot reconciler — because a hold that only one path honours is not a hold."*
|
||||
|
||||
**The severity argument in the task is correct, but for a narrower reason than it states.** The task
|
||||
says the UI hides `Frissítés` unless the app is operational, and a held app is stopped. That is right —
|
||||
but `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded`
|
||||
as operational too**, which was observed live in Phase 3b: the green `Frissítés` button was rendered
|
||||
over a crash-looping app. So the button is hidden specifically because a held app is `StateStopped`,
|
||||
not because broken apps hide it. **The conclusion (LOW, not customer-reachable) survives; the reason
|
||||
needs stating precisely, and the fix needs a test pinning it or the comment stays a wish.**
|
||||
|
||||
---
|
||||
|
||||
## 10. Every claim in the task that turned out to be wrong, named
|
||||
|
||||
1. **"a power cut on a sleeping customer's box is an unattended three-major-version upgrade"** —
|
||||
**NOT AS STATED.** Measured: a hard guest reset upgraded nothing, because Docker restored the
|
||||
containers itself and the reconciler had no orphan. The unattended upgrade is real but needs the
|
||||
narrower precondition *"and the app did not come back"*. **The exposure is smaller than the operator
|
||||
page claims, and saying so is more useful than leaving the scarier version standing.**
|
||||
2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites across nine
|
||||
files (§8).
|
||||
3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and has **no
|
||||
access route from DooPlex** by the project's own runbook. Two boxes were measured live; Peti's row
|
||||
is UNKNOWN, with what *is* knowable recorded from the hub and the catalog history (§5).
|
||||
4. **"R-440 — 23 catalog image pins float"** — **CONFIRMED EXACTLY** (79 `image:` lines, 53 apps, 66
|
||||
distinct; 23 with no patch component). One arguable 24th is recorded rather than rounded away:
|
||||
`ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly and
|
||||
leaves the PostgreSQL patch floating.
|
||||
5. **The catalog history numbers** — spot-checked and **all correct**: 153 commits touching
|
||||
`templates/`, 53 apps, 0 `.felhom.yml` carrying any upgrade metadata (the only `upgrade` matches are
|
||||
prose about STARTTLS), and all four multi-major bumps confirmed on 2026-07-18 —
|
||||
nextcloud `5e2c1ae`, grafana `b789acc`, calcom `147cee7`, vikunja `3fa63cd`.
|
||||
6. **"The Update button … takes no safety copy, cannot undo itself, does not stop the app if it goes
|
||||
wrong"** — **CONFIRMED**, by reading and across six live updates.
|
||||
7. **A methodological correction of my own, not the task's:** the first customer-page search used
|
||||
`grep -o "2.8.6"` and the unescaped `.` produced two false hits. Re-run with `grep -F`: zero. The
|
||||
controls caught it (§3).
|
||||
|
||||
---
|
||||
|
||||
## 11. What was NOT measured
|
||||
|
||||
- **Whether the drive-return gate and `AppStopGuard.Recover` upgrade in practice.** Both were located
|
||||
by reading (§8) and both call `StartStack`, which is the same function variant 1c-ii measured
|
||||
upgrading an app. **The mechanism is measured; these two specific entry points were not exercised
|
||||
live.** Naming them as read-not-measured is deliberate.
|
||||
- **Whether `RecreateStackDefinitionFromUnit`'s rollback is really undone by the sync in a live
|
||||
restore** — see §12, which states which half is measured and which is read.
|
||||
- **Any behaviour on Peti's box** (§5).
|
||||
|
||||
---
|
||||
|
||||
## 12. Observations — noticed, documented, not acted on
|
||||
|
||||
**O1 — the restore path and the catalog sync disagree about the image, and the sync wins.**
|
||||
`stackAdapter.RecreateStackDefinitionFromUnit` (`cmd/controller/main.go:2570`) writes the recovery
|
||||
unit's captured `docker-compose.yml` — **carrying the OLD image pin** — straight into the live stack
|
||||
dir (`os.WriteFile(filepath.Join(stackDir, fname), data, 0644)`), and `restore_unit.go:317` says so:
|
||||
*"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But
|
||||
`copyIfChanged` overwrites any file whose content differs from the catalog, on the next 15-minute tick.
|
||||
**MEASURED:** a locally-modified compose (Phase 3's `alpine:3.20`) was overwritten by the sync at
|
||||
`18:10:29Z`. **READ, not measured:** that the restore writes to that same path. So a restore's
|
||||
image-level rollback has a **≤15-minute half-life**, and then the next `up -d` from any source
|
||||
re-applies the catalog pin. Filed as **R-441**.
|
||||
|
||||
**O2 — `remove_hdd_data: true` is inert on this box, and the response says neither removed nor
|
||||
preserved.** Removing the Phase 6 Nextcloud with `{"remove_hdd_data":true,"remove_backups":true}`
|
||||
returned `HTTP 200` with `"hdd_paths_removed":null,"hdd_paths_preserved":null`, and left **128 MB** at
|
||||
`/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **Root cause, with controls:** `Paths.HDDPath`
|
||||
(`internal/config/config.go:117`) has **no default** — only an env override at `:403` — and demo-hp's
|
||||
`controller.yaml` `paths:` block contains only `data_dir`, `stacks_dir`, `system_data_path`. The
|
||||
container has no `FELHOM_PATHS_*` variable at all (measured, count 0). So `cfg.Paths.HDDPath == ""` and
|
||||
`ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its first line —
|
||||
`[INFO] found 0 HDD mounts` — for a compose that plainly contains
|
||||
`- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal also
|
||||
no-op'd:** `[WARN] Refusing to remove backup path outside expected directory:
|
||||
/mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. Filed as **R-442**.
|
||||
|
||||
**O3 — an update can report success over a broken app, and the truth arrives by another road 5 minutes
|
||||
later.** §4. Filed as **R-443**.
|
||||
|
||||
**O4 — demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** This run
|
||||
added ~1.05 GiB that `local-lvm` did not reclaim; `fstrim` inside the unprivileged container is refused
|
||||
(`FITRIM ioctl failed: Operation not permitted`), and `pct fstrim 9201` from the host then trimmed
|
||||
30.2 GiB + 57 GiB and took `local-lvm` from **70.91% → 26.78%** — i.e. **23.8 GB below this run's own
|
||||
starting point**. Nothing runs `pct fstrim` on the fleet. Filed as **R-444**.
|
||||
|
||||
**O5 — the hub keeps app telemetry for an app that no longer exists anywhere.** The throwaway Nextcloud
|
||||
now sets a **fleet-wide** `Suggested Limit (P95×1.2) = 352 MB` for Nextcloud, from ~15 minutes of a
|
||||
crash-looping instance, plus three MariaDB `io_uring` "Known Issues" attributed to demo-hp. **Retained
|
||||
deliberately, not cleared** — see §13. Filed as **R-445**.
|
||||
|
||||
**O6 — NOT-A-FINDING: the sync's debug hash line cannot show what changed.** `logFileHashes`
|
||||
(`internal/sync/sync.go:386`) reads the destination *after* the write, so it prints
|
||||
`src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word "changed".
|
||||
Harmless (DEBUG only, and the `Updated <app>/<file>` INFO line above it carries the fact), but it
|
||||
cannot serve the purpose its name implies. Not filed; recorded here so the next person does not trust
|
||||
it.
|
||||
|
||||
---
|
||||
|
||||
## 13. Teardown — all three layers
|
||||
|
||||
**Layer 1 — the machine.**
|
||||
- `bentopdf` **restored to its catalog tag** `v2.8.6`, digest `sha256:eaeea1e447205a79…` — **byte-identical
|
||||
to the run's baseline**. Container `02c80375fcba`, running, healthy.
|
||||
- The throwaway `nextcloud` stack **removed** via `POST /api/stacks/nextcloud/remove` (after the
|
||||
required stop): containers gone, all three named volumes gone, `app.yaml` gone. The stack dir holds
|
||||
only the catalog template (`docker-compose.yml`, `.felhom.yml`), i.e. not deployed.
|
||||
- The **128 MB the product did not remove** (O2) was deleted by hand, together with the nextcloud
|
||||
backup dirs on both drives. `find /mnt -iname "*nextcloud*"` returns nothing.
|
||||
- **Four images this run pulled were removed by targeted `docker rmi`** — `nextcloud:31.0.14-apache`,
|
||||
`32.0.9-apache`, `34.0.1-apache`, `bentopdf:v2.8.5`, plus `alpine:3.20`. **No `prune` of any kind was
|
||||
run, anywhere.**
|
||||
- **No guest was created.** Guest 9201 was hard-reset once, by design (variant 1c), and came back with
|
||||
all nine apps.
|
||||
|
||||
**Layer 2 — the host.** `pvesm status`, before → after:
|
||||
|
||||
| pool | before | after trim | note |
|
||||
|---|---|---|---|
|
||||
| `local` | 44.42% | 44.44% | unchanged in substance |
|
||||
| `local-lvm` | **68.97%** (38 959 729 KiB) | **26.78%** (15 127 469 KiB) | **the run's ~1.05 GiB was returned, and `pct fstrim 9201` reclaimed 23.8 GB more that predated this run** (O4) |
|
||||
|
||||
Guest: `/` 957 M used (baseline 957 M), `/mnt/sys_drive` 12 G / 18%, `/mnt/felhom-drives/hdd_1` 5.5 G —
|
||||
all back to their pre-run values.
|
||||
|
||||
**Layer 3 — the hub. This run provisioned NOTHING.**
|
||||
- **No customer record and no appliance record was created.** The customers list is unchanged at five
|
||||
rows: `demo-felhom`, `demo-hp`, `drill-r50`, `peti-felhom`, `tester-1` — identical to the list read at
|
||||
the start of the run. The existing `demo-hp` customer was used throughout.
|
||||
- **What the run DID create is events** — `app_start_failed` (warning) for BentoPDF at `18:05:59Z`,
|
||||
plus deploy/remove events for the throwaway Nextcloud. **These are retained deliberately:** the event
|
||||
log is an append-only record and deleting from it to tidy up a test would damage the very surface
|
||||
this project relies on for history.
|
||||
- **One piece of residue is retained rather than cleared, and the reason is stated:** the Nextcloud app
|
||||
telemetry row (O5). The hub offers `POST /apps/nextcloud/reset-telemetry`, whose own confirm text is
|
||||
**"Delete all telemetry data for nextcloud? This cannot be undone."** That is an irreversible write on
|
||||
the operator's own surface, and the operator authorised Phase 6, not this. **The exact command is
|
||||
recorded here so it is a one-line decision:**
|
||||
`curl -u ":$HUB_PW" -X POST http://<hub-clusterIP>:8080/apps/nextcloud/reset-telemetry`
|
||||
|
||||
**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified.** Peti's box was never contacted.
|
||||
|
||||
---
|
||||
|
||||
## 14. What goes to the operator
|
||||
|
||||
**One decision, and it is not a bug report.** §2 shows the restart behaviour was *chosen* and is stated
|
||||
in the source. §7 shows the word "rollback" does not describe anything this product can do. The two
|
||||
together mean the safety question is not "fix the Update button" — it is **where the safety belongs**,
|
||||
and that is in `STATUS.md`, phrased as one answerable question with what happens if nothing is done.
|
||||
|
||||
**No design is proposed here, deliberately.** Four production designs in this project were specced
|
||||
against unvalidated mechanisms and all four were wrong; this spike exists so the fifth is not.
|
||||
@@ -0,0 +1,12 @@
|
||||
### date: 2026-09-01T17:35:06Z
|
||||
### container:
|
||||
id=9473b6e4a18f5f6faf0640357c923a347a478029de38b979d0443332a4b9901a image=ghcr.io/alam00000/bentopdf:v2.8.6 imageid=sha256:4baadf01bd68b6fe61fd2c74612e800209907e2e1c16e58b770479af2cc3ae22 started=2026-09-01T04:05:50.905652046Z
|
||||
### repodigest:
|
||||
|
||||
ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
### compose image line:
|
||||
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
### compose sha256:
|
||||
39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1 /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
### local bentopdf images:
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68 2 months ago
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
2026/09/01 17:35:45 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/bentopdf/restart
|
||||
2026/09/01 17:35:45 router.go:80: [DEBUG] [api] POST /api/stacks/bentopdf/restart (path=/stacks/bentopdf/restart)
|
||||
2026/09/01 17:35:45 router.go:566: [INFO] [api] restart requested for stack: bentopdf
|
||||
2026/09/01 17:35:45 router.go:80: [DEBUG] [api] actionStack: action=restart name=bentopdf
|
||||
2026/09/01 17:35:45 manager.go:1140: [DEBUG] [stacks] RestartStack bentopdf: current state=running deployed=true containers=1
|
||||
2026/09/01 17:35:45 manager.go:1143: [INFO] [stacks] Restarting stack: bentopdf
|
||||
2026/09/01 17:35:45 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 17:36:03 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 18.3s)
|
||||
2026/09/01 17:36:04 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=bentopdf, waiting 5s for state refresh
|
||||
2026/09/01 17:36:06 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 17:36:07 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
|
||||
2026/09/01 17:36:07 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running Up 3 seconds (health: starting)
|
||||
2026/09/01 17:36:09 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=bentopdf integrationsFound=0
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
### 1a — hand-edit compose to v2.8.5 (ABSENT from the local docker store), then RESTART
|
||||
file now:
|
||||
11: image: ghcr.io/alam00000/bentopdf:v2.8.5
|
||||
local images (v2.8.5 must be ABSENT):
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
container BEFORE:
|
||||
id=9473b6e4a18f5f6faf0640357c923a347a478029de38b979d0443332a4b9901a image=ghcr.io/alam00000/bentopdf:v2.8.6 started=2026-09-01T04:05:50.905652046Z
|
||||
t0=2026-09-01T17:35:45Z
|
||||
### POST /api/stacks/bentopdf/restart
|
||||
{"ok":true,"message":"Stack bentopdf restart completed"}
|
||||
|
||||
HTTP=200
|
||||
### AFTER 1a, date: 2026-09-01T17:36:16Z
|
||||
id=93e9db74e97f1bb354108cc553b9438f52f7a4a36201bc8c7d74d24838ae1d3d image=ghcr.io/alam00000/bentopdf:v2.8.5 imageid=sha256:ab6eeed172e49347fa0b9d8471772ffd0b3df648ceaa83d3054f3d80edf3fae9 started=2026-09-01T17:36:03.775567517Z status=running
|
||||
digest=ghcr.io/alam00000/bentopdf@sha256:2d867aacb8ab5b196d00ee86944b1899d09d72df355384c5e15cf974737963a0
|
||||
local images now:
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68 2 months ago
|
||||
ghcr.io/alam00000/bentopdf:v2.8.5 ab6eeed172e4 3 months ago
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
### 1b — hand-edit compose to v2.8.6 (image ALREADY PRESENT locally), then RESTART
|
||||
file now:
|
||||
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
local images (v2.8.6 must be PRESENT):
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68
|
||||
ghcr.io/alam00000/bentopdf:v2.8.5 ab6eeed172e4
|
||||
container BEFORE:
|
||||
id=93e9db74e97f1bb354108cc553b9438f52f7a4a36201bc8c7d74d24838ae1d3d image=ghcr.io/alam00000/bentopdf:v2.8.5 started=2026-09-01T17:36:03.775567517Z
|
||||
t0=2026-09-01T17:36:35Z
|
||||
### POST /api/stacks/bentopdf/restart
|
||||
{"ok":true,"message":"Stack bentopdf restart completed"}
|
||||
|
||||
HTTP=200
|
||||
### AFTER 1b, date: 2026-09-01T17:36:42Z
|
||||
id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 image=ghcr.io/alam00000/bentopdf:v2.8.6 imageid=sha256:4baadf01bd68b6fe61fd2c74612e800209907e2e1c16e58b770479af2cc3ae22 started=2026-09-01T17:36:35.916981126Z status=running
|
||||
digest=ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68
|
||||
ghcr.io/alam00000/bentopdf:v2.8.5 ab6eeed172e4
|
||||
@@ -0,0 +1,50 @@
|
||||
### 1c PRE-RESET state (catalog=v2.8.5, file on box=v2.8.5, container=v2.8.6)
|
||||
2026-09-01T17:47:20Z
|
||||
image: ghcr.io/alam00000/bentopdf:v2.8.5
|
||||
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:36:35.916981126Z
|
||||
bentopdf bookstack bookstack-db calibre-web cloudflared docmost docmost-postgres docmost-redis felhom-controller filebrowser kimai kimai-db opengist paperless-postgres paperless-redis paperless-webserver privatebin romm romm-db romm-redis traefik
|
||||
### 1c — HARD RESET of guest 9201 on demo-hp (pct stop = ungraceful, then start)
|
||||
2026-09-01T17:47:27Z
|
||||
stop_rc=0
|
||||
start_rc=0
|
||||
VMID Status Lock Name
|
||||
9201 running demo-hp
|
||||
2026-09-01T17:47:41Z
|
||||
2026-09-01T17:47:48Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:48:05Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:48:22Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:48:39Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:48:56Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:49:12Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:49:28Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:49:45Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:50:01Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:50:18Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:50:34Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:50:51Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:51:07Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:51:23Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:51:40Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:51:56Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:52:13Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:52:29Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:52:46Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:53:02Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:53:18Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:53:35Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:53:51Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:54:08Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
2026-09-01T17:54:24Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:47:45.053040988Z
|
||||
### 1c — the boot reconciler's OWN verdict line (positive observable, not an absent log)
|
||||
2026/09/01 17:47:59 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
|
||||
2026/09/01 17:48:04 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
|
||||
2026/09/01 17:48:09 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
|
||||
2026/09/01 17:48:14 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
|
||||
2026/09/01 17:48:24 main.go:2151: [DEBUG] [bootrecon] boot window: fleet still changing (9 app(s)) — resampling
|
||||
2026/09/01 17:48:34 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 40s (3 identical samples 5s apart) — sweeping
|
||||
2026/09/01 17:48:34 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start)
|
||||
|
||||
### did the initial sync run before it, and did it rewrite bentopdf?
|
||||
2026/09/01 17:47:49 sync.go:98: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
|
||||
2026/09/01 17:47:49 sync.go:189: [INFO] [sync] Starting catalog sync
|
||||
2026/09/01 17:47:50 sync.go:103: [INFO] [sync] Initial sync: Sablonok naprakészek — nincs változás
|
||||
@@ -0,0 +1,46 @@
|
||||
### 1c-ii — SIMULATED boot orphan. Staged explicitly: 'docker rm -f bentopdf' removes the
|
||||
### container so it does NOT come back on its own (the shape a power cut leaves when docker
|
||||
### cannot restore an app). Desired state stays 'running'. Then the CONTROLLER is restarted,
|
||||
### which is what runs the boot reconciler. Nothing else is staged.
|
||||
2026-09-01T17:55:18Z
|
||||
file says:
|
||||
image: ghcr.io/alam00000/bentopdf:v2.8.5
|
||||
container BEFORE:
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6 222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2
|
||||
desired state on disk:
|
||||
desired_state: running
|
||||
REMOVED bentopdf container
|
||||
(empty above = gone)
|
||||
### restarting the controller so the boot reconciler runs
|
||||
2026-09-01T17:55:27Z
|
||||
felhom-controller
|
||||
2026-09-01T17:55:30Z bentopdf: STILL ABSENT
|
||||
2026-09-01T17:55:46Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:56:02Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:56:19Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:56:35Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:56:52Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:57:08Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:57:25Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:57:41Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:57:57Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:58:14Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
2026-09-01T17:58:30Z container=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 started=2026-09-01T17:55:44.913250874Z
|
||||
### which component started bentopdf at 17:55:44?
|
||||
2026/09/01 17:55:29 deploy.go:1051: [DEBUG] [stacks] InjectMissingFields: checking stack bentopdf — 2 deploy fields, 2 existing env vars
|
||||
2026/09/01 17:55:29 backup.go:450: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin]
|
||||
2026/09/01 17:55:29 sync.go:371: [DEBUG] [sync] bentopdf/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:55:29 sync.go:371: [DEBUG] [sync] bentopdf/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:55:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
2026/09/01 17:55:44 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 10s (3 identical samples 5s apart) — sweeping
|
||||
2026/09/01 17:55:44 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf] — up to 2 attempt(s)
|
||||
2026/09/01 17:55:44 manager.go:1036: [DEBUG] [stacks] StartStack bentopdf: current state=stopped deployed=true
|
||||
2026/09/01 17:55:44 manager.go:1039: [INFO] [stacks] Starting stack: bentopdf
|
||||
2026/09/01 17:55:44 manager.go:1046: [DEBUG] [stacks] StartStack bentopdf: prepared 8 env vars for compose
|
||||
2026/09/01 17:55:44 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 17:55:45 manager.go:1054: [INFO] [stacks] Stack bentopdf started successfully (took 0.3s)
|
||||
2026/09/01 17:55:45 bootrecon.go:274: [INFO] [bootrecon] Boot reconciliation attempt 1/2: started "bentopdf" (took 0.4s)
|
||||
2026/09/01 17:55:45 bootrecon.go:303: [INFO] [bootrecon] Boot reconciliation complete: 1 app(s) recovered in 1 attempt(s): [bentopdf]
|
||||
2026/09/01 17:55:48 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 17:55:48 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
|
||||
2026/09/01 17:55:48 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 running Up 3 seconds (health: starting)
|
||||
@@ -0,0 +1,11 @@
|
||||
=== file immediately BEFORE the POST ===
|
||||
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1 /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
=== POST /api/stacks/bentopdf/restart ===
|
||||
{"ok":true,"message":"Stack bentopdf restart completed"}
|
||||
|
||||
HTTP=200
|
||||
### AFTER 1d, date: 2026-09-01T17:35:31Z
|
||||
id=9473b6e4a18f5f6faf0640357c923a347a478029de38b979d0443332a4b9901a image=ghcr.io/alam00000/bentopdf:v2.8.6 imageid=sha256:4baadf01bd68b6fe61fd2c74612e800209907e2e1c16e58b770479af2cc3ae22 started=2026-09-01T04:05:50.905652046Z
|
||||
digest=ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6 4baadf01bd68
|
||||
@@ -0,0 +1,40 @@
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=true composePath=/opt/docker/stacks/privatebin/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=true composePath=/opt/docker/stacks/romm/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=false composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml
|
||||
2026/09/01 17:36:17 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml
|
||||
2026/09/01 17:36:27 healthprobe.go:153: [DEBUG] Health probe bentopdf: HTTP GET :8080/ → 200 (5ms)
|
||||
2026/09/01 17:36:35 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/bentopdf/restart
|
||||
2026/09/01 17:36:35 router.go:80: [DEBUG] [api] POST /api/stacks/bentopdf/restart (path=/stacks/bentopdf/restart)
|
||||
2026/09/01 17:36:35 router.go:566: [INFO] [api] restart requested for stack: bentopdf
|
||||
2026/09/01 17:36:35 router.go:80: [DEBUG] [api] actionStack: action=restart name=bentopdf
|
||||
2026/09/01 17:36:35 manager.go:1140: [DEBUG] [stacks] RestartStack bentopdf: current state=running deployed=true containers=1
|
||||
2026/09/01 17:36:35 manager.go:1143: [INFO] [stacks] Restarting stack: bentopdf
|
||||
2026/09/01 17:36:35 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, SUBDOMAIN, IMPORT_PATH] (8 app + 0 system)
|
||||
2026/09/01 17:36:35 manager.go:1320: [DEBUG] Running: docker compose up -d (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 17:36:36 manager.go:1340: [DEBUG] Command completed: docker compose up -d (took 0.5s)
|
||||
2026/09/01 17:36:36 manager.go:1157: [INFO] [stacks] Stack bentopdf restarted successfully (took 0.5s)
|
||||
2026/09/01 17:36:36 lifecycle.go:70: [DEBUG] [integrations] OnStackStart: stack=bentopdf, waiting 5s for state refresh
|
||||
2026/09/01 17:36:39 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DOMAIN, SUBDOMAIN, IMPORT_PATH] (8 app + 0 system)
|
||||
2026/09/01 17:36:39 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 17:36:39 manager.go:1340: [DEBUG] Command completed: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (took 0.1s)
|
||||
2026/09/01 17:36:39 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
|
||||
2026/09/01 17:36:39 manager.go:1403: [INFO] [stacks] bentopdf ghcr.io/alam00000/bentopdf:v2.8.6 running Up 3 seconds (health: starting)
|
||||
2026/09/01 17:36:41 lifecycle.go:86: [DEBUG] [integrations] OnStackStart: stack=bentopdf integrationsFound=0
|
||||
2026/09/01 17:36:47 healthprobe.go:153: [DEBUG] Health probe bentopdf: HTTP GET :8080/ → 200 (6ms)
|
||||
2026/09/01 17:36:57 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bentopdf — last check 10s ago, effective interval 5m0s, healthy=true
|
||||
@@ -0,0 +1,4 @@
|
||||
pre-sync 2026-09-01T17:38:35Z
|
||||
11: image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
39679e28cdd6ee8f46e359290b4d631cf234a1cfdae374cd84c8fe7588870ed1 /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2
|
||||
@@ -0,0 +1,12 @@
|
||||
2026-09-01T17:42:02Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:42:24Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:42:45Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:43:07Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:43:28Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:43:50Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:44:11Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:44:32Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:44:54Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:45:15Z image: ghcr.io/alam00000/bentopdf:v2.8.6 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
2026-09-01T17:45:37Z image: ghcr.io/alam00000/bentopdf:v2.8.5 container=ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
>>> SYNC LANDED
|
||||
@@ -0,0 +1,46 @@
|
||||
### controller log around the 17:45:17 sync
|
||||
2026/09/01 17:45:07 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting
|
||||
2026/09/01 17:45:07 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting
|
||||
2026/09/01 17:45:07 manager.go:621: [INFO] [stacks] Status refresh: 20 containers across 56 stacks
|
||||
2026/09/01 17:45:17 sync.go:189: [INFO] [sync] Starting catalog sync
|
||||
2026/09/01 17:45:17 sync.go:277: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
|
||||
2026/09/01 17:45:17 sync.go:279: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache
|
||||
2026/09/01 17:45:17 sync.go:449: [DEBUG] [sync] Running: git fetch --depth 1 origin main
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈4456MB avail≈21442MB
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25429852 → total=25898MB avail=21442MB used=4456MB (17.2%)
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.7GB used=11.9GB avail=53.3GB (17.3%)
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.5GB avail=884.6GB (0.6%)
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.23 1.00 0.81 7/963 133176" → 1m=1.23 5m=1.00 15m=0.81
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 54.2°C (hwmon2)
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo done in 74ms — mem=4456MB/25898MB (17.2%), rootDisk=11.9GB/68.7GB (17.3%), load=1.23/1.00/0.81, temp=54.2°C (hwmon2), cpu=2.7%
|
||||
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: backup-cache
|
||||
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: offsite-credential-retry
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job offsite-credential-retry: execution starting
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job backup-cache: execution starting
|
||||
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: hub-report
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job hub-report: execution starting
|
||||
2026/09/01 17:45:17 builder.go:37: [INFO] [report] Building system report
|
||||
2026/09/01 17:45:17 builder.go:40: [DEBUG] [report] BuildReport: starting — version=0.232.0, storagePaths=1
|
||||
2026/09/01 17:45:17 scheduler.go:363: [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/09/01 17:45:17 builder.go:62: [DEBUG] [report] BuildReport: configHash=042b71a3123d... (2989 bytes)
|
||||
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: system-health
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job system-health: execution starting
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting
|
||||
2026/09/01 17:45:17 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
|
||||
2026/09/01 17:45:17 scheduler.go:346: [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/09/01 17:45:17 scheduler.go:67: [DEBUG] [scheduler] job agent-channel-health: execution starting
|
||||
2026/09/01 17:45:17 manager.go:621: [INFO] [stacks] Status refresh: 20 containers across 56 stacks
|
||||
2026/09/01 17:45:17 backup.go:450: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [bentopdf, bookstack, docmost, kimai, opengist, privatebin]
|
||||
2026/09/01 17:45:17 backup.go:450: [DEBUG] groupStacksByDrive: /mnt/felhom-drives/hdd_1 → [calibre-web, paperless-ngx, romm]
|
||||
2026/09/01 17:45:17 backup.go:1083: [INFO] [backup] Found 13 DB dump files across drives
|
||||
|
||||
### file state on the box
|
||||
2ebbbda3765b2b216f41ea6203fc7419cabee2fdf9dbdd60f46e2dd97839b30a /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
2026-09-01 17:45:17.646017366 +0000 /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=222f8178dbd078ba18590e682aafc85974c404f8e92acd121f04d6d95b7fcea2 started=2026-09-01T17:36:35.916981126Z
|
||||
@@ -0,0 +1,26 @@
|
||||
### the sync's own lines (grep: Updated / frissitve / event / notify)
|
||||
2026/09/01 17:45:17 sync.go:189: [INFO] [sync] Starting catalog sync
|
||||
2026/09/01 17:45:17 sync.go:277: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main)
|
||||
2026/09/01 17:45:17 sync.go:279: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache
|
||||
2026/09/01 17:45:17 sync.go:449: [DEBUG] [sync] Running: git fetch --depth 1 origin main
|
||||
2026/09/01 17:45:17 sync.go:449: [DEBUG] [sync] Running: git reset --hard origin/main
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] actualbudget/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] actualbudget/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] adventurelog/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] adventurelog/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] audiobookshelf/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] audiobookshelf/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
|
||||
2026/09/01 17:45:17 sync.go:398: [DEBUG] [sync] bentopdf/docker-compose.yml: src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] bentopdf/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] bookstack/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] bookstack/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calcom/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calcom/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calibre-web/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] calibre-web/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] claper/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] claper/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] code-server/docker-compose.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] code-server/.felhom.yml: hash match, skipped
|
||||
2026/09/01 17:45:17 sync.go:371: [DEBUG] [sync] crafty-controller/docker-compose.yml: hash match, skipped
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
### LITERAL (grep -F) search — my earlier grep used '.' as a wildcard and over-counted. Redone.
|
||||
--- page_apps_bentopdf.html ---
|
||||
BentoPDF hits=4
|
||||
zzz-never-present hits=0
|
||||
v2.8.5 hits=0
|
||||
v2.8.6 hits=0
|
||||
2.8.5 hits=0
|
||||
2.8.6 hits=0
|
||||
Friss hits=1
|
||||
bentopdf/update hits=0
|
||||
bentopdf/restart hits=0
|
||||
--- page_stacks.html ---
|
||||
BentoPDF hits=1
|
||||
zzz-never-present hits=0
|
||||
v2.8.5 hits=0
|
||||
v2.8.6 hits=0
|
||||
2.8.5 hits=0
|
||||
2.8.6 hits=0
|
||||
Friss hits=10
|
||||
bentopdf/update hits=0
|
||||
bentopdf/restart hits=0
|
||||
@@ -0,0 +1,2 @@
|
||||
### context of every 2.8.6 / Friss hit on /apps/bentopdf
|
||||
--- Friss ---
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
### Phase 2 revert verification + Phase 3 teardown
|
||||
2026-09-01T18:12:52Z
|
||||
### did the catalog revert reach the box's file?
|
||||
image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
container=alpine:3.20 status=restarting
|
||||
2026/09/01 18:10:29 sync.go:189: [INFO] [sync] Starting catalog sync
|
||||
2026/09/01 18:10:29 sync.go:366: [INFO] [sync] Updated bentopdf/docker-compose.yml
|
||||
2026/09/01 18:10:29 sync.go:244: [INFO] [sync] Catalog sync complete
|
||||
2026/09/01 18:10:29 sync.go:116: [INFO] [sync] Periodic sync: Sablonok frissítve — frissítve: bentopdf
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
### 3a-ii — the LANDMINE: the update failed and left the file pointing at a tag that
|
||||
### does not exist. The page still shows a green healthy app with a Restart button.
|
||||
### What does that Restart button do?
|
||||
2026-09-01T18:00:23Z
|
||||
### POST /api/stacks/bentopdf/restart:
|
||||
{"ok":false,"error":"restarting stack bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Error manifest unknown\nError response from daemon: manifest unknown"}
|
||||
|
||||
HTTP=500
|
||||
### AFTER:
|
||||
2026-09-01T18:00:29Z
|
||||
image=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running
|
||||
bentopdf ghcr.io/alam00000/bentopdf:v2.8.5 Up 4 minutes (healthy)
|
||||
@@ -0,0 +1,18 @@
|
||||
### 3a — the PULL FAILS. Compose pointed at a tag that does not exist.
|
||||
2026-09-01T17:59:50Z
|
||||
file now:
|
||||
image: ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist
|
||||
container BEFORE:
|
||||
ghcr.io/alam00000/bentopdf:v2.8.5 aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running
|
||||
### POST /api/stacks/bentopdf/update — EXACT status and body:
|
||||
{"ok":false,"error":"pulling images for bentopdf: exit code 1\nstderr: Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Pulling \n Image ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist Error manifest unknown\nError response from daemon: manifest unknown"}
|
||||
|
||||
HTTP=500
|
||||
### 3a AFTER — is the app still running?
|
||||
2026-09-01T18:00:05Z
|
||||
image=ghcr.io/alam00000/bentopdf:v2.8.5 id=aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running restarts=0
|
||||
image: ghcr.io/alam00000/bentopdf:v9.9.9-does-not-exist
|
||||
|
||||
### what the customer's app-list row says now (state badge, buttons)
|
||||
http=200
|
||||
|Pi kompatibilis||||Telepítés||Részletek|||||||||Audiobookshelf||||Nincs telepítve|||Hangoskönyv és podcast kezelő szerver|||~100M||Pi kompatibilis||HDD szükséges||||Telepítés||Részletek|||||||||BentoPDF||pdf.enkisfelhom.hu ↗||||Fut|||Adatvédelmi fókuszú PDF eszköztár|||~100M||Pi kompatibilis|||||bentopdf||Up 4 minutes (healthy)|||||Frissítés||Újraindítás||Leállítás||Naplók||Részletek|||||||<div
|
||||
@@ -0,0 +1,4 @@
|
||||
### does anything ALARM on a crash-looping app after a 'successful' update?
|
||||
2026-09-01T18:01:48Z
|
||||
2026/09/01 18:00:59 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/09/01 18:01:29 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
@@ -0,0 +1,49 @@
|
||||
--- 2026-09-01T18:02:31Z ---
|
||||
--- 2026-09-01T18:03:17Z ---
|
||||
--- 2026-09-01T18:04:04Z ---
|
||||
--- 2026-09-01T18:04:50Z ---
|
||||
--- 2026-09-01T18:05:37Z ---
|
||||
--- 2026-09-01T18:06:23Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
--- 2026-09-01T18:07:09Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:07:56Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:08:42Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:09:29Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:10:15Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:11:02Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:11:48Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
--- 2026-09-01T18:12:35Z ---
|
||||
2026/09/01 18:05:59 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
|
||||
2026/09/01 18:05:59 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/09/01 18:05:59 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BentoPDF
|
||||
2026/09/01 18:06:29 main.go:1785: [INFO] [deadapp] check alive: 20 scans since boot, 9 deployed app(s) evaluated, 1 currently down
|
||||
@@ -0,0 +1,37 @@
|
||||
### 3b — the PULL SUCCEEDS AND THE APP DOES NOT. The compose is pointed at alpine:3.20:
|
||||
### a real image that pulls cleanly, then exits at once, so restart:unless-stopped puts it
|
||||
### into a crash-loop. This is the shape of a real upstream image whose config the template
|
||||
### can no longer supply — the catalog's own 2026-07-18 wger 2.6 revert was exactly that.
|
||||
2026-09-01T18:00:42Z
|
||||
file now:
|
||||
image: alpine:3.20
|
||||
alpine present locally?
|
||||
0
|
||||
container BEFORE:
|
||||
ghcr.io/alam00000/bentopdf:v2.8.5 aa443836a4cd612a6f5000f8d0481f3daf15fcc2a3faf59ccbffb941b5485db9 status=running
|
||||
### POST /api/stacks/bentopdf/update — EXACT status and body:
|
||||
{"ok":true,"message":"Stack bentopdf update completed"}
|
||||
|
||||
HTTP=200
|
||||
### 3b AFTER — what the customer actually has
|
||||
2026-09-01T18:01:20Z
|
||||
image=alpine:3.20 id=436a8ae4fe639581dd05febd183f0476ad4aff91c8180063582ad9162ecacb3e status=restarting restarts=9 exit=0
|
||||
|
||||
bentopdf | alpine:3.20 | Restarting (0) 5 seconds ago
|
||||
|
||||
### the controller's own post-start status line (what it logged about the result)
|
||||
2026/09/01 18:00:43 auth.go:134: [DEBUG] [web] auth: valid session for POST /api/stacks/bentopdf/update
|
||||
2026/09/01 18:00:43 router.go:80: [DEBUG] [api] POST /api/stacks/bentopdf/update (path=/stacks/bentopdf/update)
|
||||
2026/09/01 18:00:43 router.go:566: [INFO] [api] update requested for stack: bentopdf
|
||||
2026/09/01 18:00:43 router.go:80: [DEBUG] [api] actionStack: action=update name=bentopdf
|
||||
2026/09/01 18:00:43 manager.go:1176: [INFO] [stacks] Updating stack: bentopdf
|
||||
2026/09/01 18:00:43 manager.go:1439: [INFO] [stacks] Deploying stack bentopdf — checking 1 images...
|
||||
2026/09/01 18:00:43 manager.go:1320: [DEBUG] Running: docker compose pull (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 18:00:46 manager.go:1320: [DEBUG] Running: docker compose up -d --remove-orphans (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 18:00:47 manager.go:1195: [INFO] [stacks] Stack bentopdf updated successfully (took 3.5s)
|
||||
2026/09/01 18:00:50 manager.go:1320: [DEBUG] Running: docker compose ps -a --format table {{.Name}} {{.Image}} {{.State}} {{.Status}} (in /opt/docker/stacks/bentopdf)
|
||||
2026/09/01 18:00:50 manager.go:1400: [INFO] [stacks] Stack bentopdf post-start status:
|
||||
2026/09/01 18:00:50 manager.go:1403: [INFO] [stacks] bentopdf alpine:3.20 restarting Restarting (0) Less than a second ago
|
||||
### 3b — the EXACT customer-facing row while the app is crash-looping
|
||||
ataset.fallback&&this.dataset.step==='1'){this.dataset.step='2';this.src=this.dataset.fallback;}else{this.onerror=null;this.style.visibility='hidden';}">||BentoPDF||pdf.enkisfelhom.hu ↗||URL nem elérhető – útvonal nincs publikálva||||Újraindítás...|||Adatvédelmi fókuszú PDF eszköztár|||~100M||Pi kompatibilis|||||bentopdf||Restarting (0) 15 seconds ago|||||Frissítés||Újraindítás||Leállítás||Naplók||Részletek||||||<img class="stack-logo-lg" src="/static/assets/bookstack-logo.svg" alt="" data-fallback="/static/app-placeholder.svg"
|
||||
onerror="if(!thi
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.232.0 Up 9 hours (healthy)
|
||||
opengist ghcr.io/thomiceli/opengist:1.13 Up 9 hours (healthy)
|
||||
filebrowser gtstef/filebrowser:1.3.3-stable Up 3 weeks (healthy)
|
||||
cloudflared cloudflare/cloudflared:2026.6.0 Up 3 weeks
|
||||
traefik traefik:v3.6.7 Up 3 weeks
|
||||
+49
@@ -0,0 +1,49 @@
|
||||
=== demo-hp 2026-09-01T17:40:38Z ===
|
||||
bentopdf|ghcr.io/alam00000/bentopdf:v2.8.6|ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
felhom-controller|gitea.dooplex.hu/admin/felhom-controller:0.232.0|gitea.dooplex.hu/admin/felhom-controller@sha256:51b2403520e425e0a7d4bc00cfd395d3b06078d79cf7927bee1035d662814b8c
|
||||
romm|rommapp/romm:5.0.0|rommapp/romm@sha256:91f6611eca5a4dafc4f4a1d72a1ed7dd66a11375d939f28410dc1d1de0b80b1b
|
||||
romm-db|mariadb:11.4|mariadb@sha256:4f1d8d202fcf7bcb3902f63af09f9c1a050c2922a89652f22abaec0d4f015e83
|
||||
romm-redis|redis:7-alpine|redis@sha256:ff02b58f971e7d7d156a1267e283fcbbeee91773b6aa36c49dac28ecfe28eadf
|
||||
privatebin|privatebin/pdo:2.0.5|privatebin/pdo@sha256:8a2cac16eff6caed4dc622e7b3ebd0c0ffcb28dfc0eba0b9c335636552eadac1
|
||||
paperless-webserver|ghcr.io/paperless-ngx/paperless-ngx:2.20.15|ghcr.io/paperless-ngx/paperless-ngx@sha256:6c86cad803970ea782683a8e80e7403444c5bf3cf70de63b4d3c8e87500db92f
|
||||
paperless-redis|redis:7-alpine|redis@sha256:ff02b58f971e7d7d156a1267e283fcbbeee91773b6aa36c49dac28ecfe28eadf
|
||||
paperless-postgres|postgres:16-alpine|postgres@sha256:cf78e76683b9ca8c5733cbbdce6c9262b45b6767934dd0a95e671f9a0fc20685
|
||||
opengist|ghcr.io/thomiceli/opengist:1.13|ghcr.io/thomiceli/opengist@sha256:dddc26031d1320ebb4bc5b913b3c42a9cb84c7528192d387f99ddcbbe57b0085
|
||||
kimai|kimai/kimai2:apache-2.57.0|kimai/kimai2@sha256:efa66c5eadf948b94d8031e9436feae491d90b2ceaa3d5e45e4d7c1a59dd8c38
|
||||
kimai-db|mariadb:11.6|mariadb@sha256:bfb1298c06cd15f446f1c59600b3a856dae861705d1a2bd2a00edbd6c74ba748
|
||||
docmost|docmost/docmost:0.95.0|docmost/docmost@sha256:41c8d777cf23c74e78f94e676aec328b7d7856f48df5e573543dac68d371e37c
|
||||
docmost-postgres|postgres:16-alpine|postgres@sha256:cf78e76683b9ca8c5733cbbdce6c9262b45b6767934dd0a95e671f9a0fc20685
|
||||
docmost-redis|redis:7-alpine|redis@sha256:ff02b58f971e7d7d156a1267e283fcbbeee91773b6aa36c49dac28ecfe28eadf
|
||||
calibre-web|crocodilestick/calibre-web-automated:v4.0.6|crocodilestick/calibre-web-automated@sha256:c31a738b6d5ec6982c050063dd3f063b6943eb1051fc81144789f840d9093a8d
|
||||
bookstack|lscr.io/linuxserver/bookstack:26.05.2|lscr.io/linuxserver/bookstack@sha256:3db259db582808ab498d49ae96b0a63f935d9cf3635c9d5bd8b8815c6ff1f8a1
|
||||
bookstack-db|mariadb:12.3|mariadb@sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
|
||||
filebrowser|gtstef/filebrowser:1.3.3-stable|gtstef/filebrowser@sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
|
||||
cloudflared|cloudflare/cloudflared:2026.6.0|cloudflare/cloudflared@sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
|
||||
traefik|traefik:v3.6.7|traefik@sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
|
||||
=== compose file image lines for DEPLOYED stacks ===
|
||||
--- bentopdf
|
||||
image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
--- bookstack
|
||||
image: lscr.io/linuxserver/bookstack:26.05.2
|
||||
image: mariadb:12.3
|
||||
--- calibre-web
|
||||
image: crocodilestick/calibre-web-automated:v4.0.6
|
||||
--- docmost
|
||||
image: docmost/docmost:0.95.0
|
||||
image: postgres:16-alpine
|
||||
image: redis:7-alpine
|
||||
--- kimai
|
||||
image: kimai/kimai2:apache-2.57.0
|
||||
image: mariadb:11.6
|
||||
--- opengist
|
||||
image: ghcr.io/thomiceli/opengist:1.13
|
||||
--- paperless-ngx
|
||||
image: ghcr.io/paperless-ngx/paperless-ngx:2.20.15
|
||||
image: postgres:16-alpine
|
||||
image: redis:7-alpine
|
||||
--- privatebin
|
||||
image: privatebin/pdo:2.0.5
|
||||
--- romm
|
||||
image: rommapp/romm:5.0.0
|
||||
image: mariadb:11.4
|
||||
image: redis:7-alpine
|
||||
@@ -0,0 +1,14 @@
|
||||
image ref floating? verdict
|
||||
----------------------------------------------------------------------------------------------------
|
||||
docker.io/library/postgres:16-alpine FLOATING SAME
|
||||
docker.io/library/redis:7-alpine FLOATING SAME
|
||||
docker.io/library/mariadb:11.4 FLOATING MOVED
|
||||
running sha256:4f1d8d202fcf7bcb3902f63af09f9c1a050c2922a89652f22abaec0d4f015e83
|
||||
upstream sha256:611a2fcc5fa7c6ceb8644c6f74b25ede004ff6c3a6b38c8f8c23d3bbf6c26430
|
||||
docker.io/library/mariadb:11.6 FLOATING SAME
|
||||
docker.io/library/mariadb:12.3 FLOATING MOVED
|
||||
running sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
|
||||
upstream sha256:dd9b303aed4f4890ed09f766d8ca9ddfd176c0c6f6267feff53b3192ec65a979
|
||||
ghcr.io/thomiceli/opengist:1.13 FLOATING SAME
|
||||
docker.io/rommapp/romm:5.0.0 pinned SAME
|
||||
docker.io/privatebin/pdo:2.0.5 pinned SAME
|
||||
+17
@@ -0,0 +1,17 @@
|
||||
### Phase 4 — Peti's box: NOT measurable live (down, not enrolled, no access route).
|
||||
### What IS knowable read-only: it last reported 48d ago (2026-07-15) running rallly.
|
||||
### So: how far has the CATALOG's rallly pin moved since then?
|
||||
|
||||
current catalog pin:
|
||||
image: lukevella/rallly:4.11.1
|
||||
image: postgres:16-alpine
|
||||
|
||||
commits touching rallly since 2026-07-15:
|
||||
d7ffcbe 2026-07-19 rallly: raise app memory 256M -> 768M (OOM-killed at 256M)
|
||||
e3f3a81 2026-07-18 rallly: 3.11.2 -> 4.11.1 [MAJOR]
|
||||
(empty above = the pin has NOT moved since Peti's box last reported)
|
||||
|
||||
### demo-felhom (opengist, the only deployed app there)
|
||||
catalog:
|
||||
image: ghcr.io/thomiceli/opengist:1.13
|
||||
running: ghcr.io/thomiceli/opengist:1.13 (from phase4-demo-felhom-running.txt)
|
||||
+72
@@ -0,0 +1,72 @@
|
||||
### Phase 5 — what a pre-update copy would cost, measured on demo-hp 2026-09-01T18:13:33Z
|
||||
|
||||
### 1. app data on disk, per deployed app (the thing a FILE copy would have to move)
|
||||
56K /mnt/felhom-drives/hdd_1/appdata/
|
||||
5.9M /mnt/felhom-drives/hdd_1/userdata/
|
||||
742M /mnt/felhom-drives/hdd_1/backups/
|
||||
748M /mnt/sys_drive/felhom-data/
|
||||
|
||||
### 2. the EXISTING safety machinery is DATABASE-ONLY (writeSafetyDump -> DumpOne).
|
||||
### Existing dumps on the box, with size and age — the cost proxy for a DB copy:
|
||||
62270 2026-08-21 21:02 /mnt/felhom-drives/hdd_1/backups/primary/romm/db-dumps/pre-restore-20260821T210246Z-romm-mariadb.sql
|
||||
62270 2026-09-01 02:15 /mnt/felhom-drives/hdd_1/backups/primary/romm/db-dumps/romm-mariadb.sql
|
||||
58775 2026-09-01 00:30 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/bookstack-mariadb.sql
|
||||
58775 2026-08-22 14:24 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/pre-restore-20260822T142418Z-bookstack-mariadb.sql
|
||||
58775 2026-08-22 16:25 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/pre-restore-20260822T162555Z-bookstack-mariadb.sql
|
||||
58775 2026-08-22 21:55 /mnt/felhom-drives/hdd_1/backups/secondary/bookstack/recovery-unit/db-dumps/pre-restore-20260822T215540Z-bookstack-mariadb.sql
|
||||
142277 2026-09-01 00:30 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql
|
||||
141363 2026-08-22 16:23 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||
141363 2026-08-22 16:27 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||
141363 2026-08-22 21:54 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||
48217 2026-09-01 00:30 /mnt/felhom-drives/hdd_1/backups/secondary/kimai/recovery-unit/db-dumps/kimai-mariadb.sql
|
||||
58775 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/bookstack-mariadb.sql
|
||||
58775 2026-08-22 14:24 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/pre-restore-20260822T142418Z-bookstack-mariadb.sql
|
||||
58775 2026-08-22 16:25 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/pre-restore-20260822T162555Z-bookstack-mariadb.sql
|
||||
58775 2026-08-22 21:55 /mnt/sys_drive/felhom-data/backups/primary/bookstack/db-dumps/pre-restore-20260822T215540Z-bookstack-mariadb.sql
|
||||
142277 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/docmost-postgres.sql
|
||||
141363 2026-08-22 16:23 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||
141363 2026-08-22 16:27 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||
141363 2026-08-22 21:54 /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||
48217 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/primary/kimai/db-dumps/kimai-mariadb.sql
|
||||
312381 2026-08-22 07:38 /mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql
|
||||
395065 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx/recovery-unit/db-dumps/paperless-ngx-postgres.sql
|
||||
312957 2026-08-22 07:56 /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx/recovery-unit/db-dumps/pre-restore-20260822T075658Z-paperless-ngx-postgres.sql
|
||||
62270 2026-08-21 21:02 /mnt/sys_drive/felhom-data/backups/secondary/romm/recovery-unit/db-dumps/pre-restore-20260821T210246Z-romm-mariadb.sql
|
||||
62270 2026-09-01 02:15 /mnt/sys_drive/felhom-data/backups/secondary/romm/recovery-unit/db-dumps/romm-mariadb.sql
|
||||
### per-app FILE data on demo-hp (this is what a pre-update copy would have to move)
|
||||
== /mnt/sys_drive/felhom-data/userdata
|
||||
12K /mnt/sys_drive/felhom-data/userdata/import/
|
||||
== /mnt/felhom-drives/hdd_1/appdata
|
||||
12K /mnt/felhom-drives/hdd_1/appdata/romm/
|
||||
40K /mnt/felhom-drives/hdd_1/appdata/paperless/
|
||||
== /mnt/felhom-drives/hdd_1/userdata
|
||||
4.0K /mnt/felhom-drives/hdd_1/userdata/documents/
|
||||
4.0K /mnt/felhom-drives/hdd_1/userdata/downloads/
|
||||
12K /mnt/felhom-drives/hdd_1/userdata/import/
|
||||
1.1M /mnt/felhom-drives/hdd_1/userdata/roms/
|
||||
4.8M /mnt/felhom-drives/hdd_1/userdata/media/
|
||||
### docker named volumes (the other place app data lives)
|
||||
76K /var/lib/docker/volumes/ee40750a8284cf0481f03f3c8c357d2fdc9ab0c68fafd25321b7256c70218bae/
|
||||
76K /var/lib/docker/volumes/f55b60c2fecb068286c83c306f07695557a5524a66542c548316a76535338ebe/
|
||||
76K /var/lib/docker/volumes/f98ca17ff5c961ae51783dee774bd48f7e98ccc21a172d819183e73f187fa62e/
|
||||
76K /var/lib/docker/volumes/fc61c5e41e1aad481c18ec4f5c983e8c095d3ad24c427b1d6cec7044a5be4bbc/
|
||||
92K /var/lib/docker/volumes/filebrowser_filebrowser_data/
|
||||
228K /var/lib/docker/volumes/opengist_opengist_data/
|
||||
960K /var/lib/docker/volumes/calibre-web_calibre_web_config/
|
||||
2.1M /var/lib/docker/volumes/privatebin_privatebin_data/
|
||||
5.0M /var/lib/docker/volumes/paperless-ngx_paperless_data/
|
||||
6.4M /var/lib/docker/volumes/paperless-ngx_paperless_redis_data/
|
||||
6.8M /var/lib/docker/volumes/bookstack_bookstack_config/
|
||||
25M /var/lib/docker/volumes/romm_romm_redis_data/
|
||||
40M /var/lib/docker/volumes/felhom-controller-data/
|
||||
55M /var/lib/docker/volumes/docmost_docmost_redis_data/
|
||||
56M /var/lib/docker/volumes/kimai_kimai_var/
|
||||
67M /var/lib/docker/volumes/docmost_docmost_postgres_data/
|
||||
69M /var/lib/docker/volumes/paperless-ngx_paperless_postgres_data/
|
||||
166M /var/lib/docker/volumes/bookstack_bookstack_db_data/
|
||||
166M /var/lib/docker/volumes/kimai_kimai_db_data/
|
||||
166M /var/lib/docker/volumes/romm_romm_db_data/
|
||||
|
||||
### LIMITATION, stated rather than papered over: demo-hp was reinstalled 2026-08-21 and its
|
||||
### apps were deployed ~13 h before this run, so it holds a few MB of app data in total.
|
||||
### It CANNOT price a pre-update copy at customer scale. Prior measured numbers are cited instead.
|
||||
@@ -0,0 +1,19 @@
|
||||
### Phase 6 — pin the stack to nextcloud:31.0.14-apache BEFORE deploying
|
||||
2026-09-01T18:58:18Z
|
||||
image: nextcloud:31.0.14-apache
|
||||
image: mariadb:11.6
|
||||
image: redis:7-alpine
|
||||
### POST /api/stacks/nextcloud/deploy (file verified at nextcloud:31.0.14-apache above)
|
||||
2026-09-01T18:58:40Z
|
||||
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
|
||||
HTTP=202 time=0.205833s
|
||||
2026-09-01T18:58:40Z
|
||||
2026-09-01T18:58:48Z
|
||||
2026-09-01T18:59:10Z
|
||||
2026-09-01T18:59:31Z nextcloud|nextcloud:31.0.14-apache|Created nextcloud-db|mariadb:11.6|Up 9 seconds (health: starting) nextcloud-redis|redis:7-alpine|Up 9 seconds (healthy)
|
||||
>>> HEALTHY
|
||||
2026-09-01T18:59:39Z nextcloud:31.0.14-apache|Up 6 seconds (health: starting)
|
||||
2026-09-01T19:00:00Z nextcloud:31.0.14-apache|Up 28 seconds (health: starting)
|
||||
2026-09-01T19:00:22Z nextcloud:31.0.14-apache|Up 49 seconds (healthy)
|
||||
>>> nextcloud HEALTHY
|
||||
@@ -0,0 +1,22 @@
|
||||
### Phase 6 — SEED identifiable data on nextcloud 31.0.14-apache
|
||||
2026-09-01T19:00:36Z
|
||||
- installed: true
|
||||
- version: 31.0.14.1
|
||||
- versionstring: 31.0.14
|
||||
- edition:
|
||||
- maintenance: false
|
||||
- needsDbUpgrade: false
|
||||
- productname: Nextcloud
|
||||
- extendedSupport: false
|
||||
--- set a system config marker ---
|
||||
System config value spike_marker set to string SPIKE-SEEDED-ON-31.0.14-2026-09-01
|
||||
--- write a real file into the admin user data dir ---
|
||||
Starting scan for user 1 out of 1 (admin)
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
| Folders | Files | New | Updated | Removed | Errors | Elapsed time |
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
| 5 | 53 | 0 | 2 | 0 | 0 | 00:00:00 |
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
--- read both back ---
|
||||
SPIKE-SEEDED-ON-31.0.14-2026-09-01
|
||||
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
|
||||
@@ -0,0 +1,60 @@
|
||||
### Phase 6 — THE 3-MAJOR JUMP: 31.0.14-apache -> 34.0.1-apache, via the real Update button
|
||||
2026-09-01T19:00:53Z
|
||||
file now:
|
||||
image: nextcloud:34.0.1-apache
|
||||
container BEFORE:
|
||||
nextcloud:31.0.14-apache status=running
|
||||
### POST /api/stacks/nextcloud/update:
|
||||
{"ok":true,"message":"Stack nextcloud update completed"}
|
||||
|
||||
HTTP=200
|
||||
2026-09-01T19:01:40Z
|
||||
### what the app itself says after the 3-major jump
|
||||
2026-09-01T19:01:50Z nextcloud:34.0.1-apache|Restarting (1) Less than a second ago
|
||||
2026-09-01T19:02:06Z nextcloud:34.0.1-apache|Restarting (1) 10 seconds ago
|
||||
2026-09-01T19:02:23Z nextcloud:34.0.1-apache|Restarting (1) 13 seconds ago
|
||||
2026-09-01T19:02:39Z nextcloud:34.0.1-apache|Restarting (1) 4 seconds ago
|
||||
2026-09-01T19:02:56Z nextcloud:34.0.1-apache|Restarting (1) 20 seconds ago
|
||||
2026-09-01T19:03:12Z nextcloud:34.0.1-apache|Restarting (1) 36 seconds ago
|
||||
|
||||
### THE EXACT MESSAGE — container logs since the update
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
=> Configuring PHP session handler...
|
||||
==> Using Redis as PHP session handler...
|
||||
Initializing nextcloud 34.0.1.2 ...
|
||||
Can't start Nextcloud because upgrading from 31.0.14.1 to 34.0.1.2 is not supported.
|
||||
It is only possible to upgrade one major version at a time. For example, if you want to upgrade from version 14 to 16, you will have to upgrade from version 14 to 15, then from 15 to 16.
|
||||
@@ -0,0 +1,36 @@
|
||||
### Phase 6 — CAN IT GO BACK? put the old tag back and press the button
|
||||
2026-09-01T19:03:43Z
|
||||
file now:
|
||||
image: nextcloud:31.0.14-apache
|
||||
### POST /api/stacks/nextcloud/restart:
|
||||
{"ok":true,"message":"Stack nextcloud restart completed"}
|
||||
|
||||
HTTP=200
|
||||
2026-09-01T19:03:46Z
|
||||
2026-09-01T19:03:55Z nextcloud:31.0.14-apache|Up 9 seconds (healthy)
|
||||
2026-09-01T19:04:11Z nextcloud:31.0.14-apache|Up 26 seconds (healthy)
|
||||
2026-09-01T19:04:28Z nextcloud:31.0.14-apache|Up 42 seconds (healthy)
|
||||
2026-09-01T19:04:44Z nextcloud:31.0.14-apache|Up 59 seconds (healthy)
|
||||
2026-09-01T19:05:01Z nextcloud:31.0.14-apache|Up About a minute (healthy)
|
||||
2026-09-01T19:05:17Z nextcloud:31.0.14-apache|Up About a minute (healthy)
|
||||
|
||||
### READ THE SEEDED DATA BACK
|
||||
- installed: true
|
||||
- version: 31.0.14.1
|
||||
- versionstring: 31.0.14
|
||||
- edition:
|
||||
- maintenance: false
|
||||
- needsDbUpgrade: false
|
||||
- productname: Nextcloud
|
||||
- extendedSupport: false
|
||||
--- system config marker ---
|
||||
SPIKE-SEEDED-ON-31.0.14-2026-09-01
|
||||
--- the seeded file ---
|
||||
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
|
||||
--- is the file still known to the DB? ---
|
||||
Starting scan for user 1 out of 1 (admin)
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
| Folders | Files | New | Updated | Removed | Errors | Elapsed time |
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
| 5 | 53 | 0 | 0 | 0 | 0 | 00:00:00 |
|
||||
+---------+-------+-----+---------+---------+--------+--------------+
|
||||
+58
@@ -0,0 +1,58 @@
|
||||
### Phase 6b — THE DECISIVE HALF: a migration that ACTUALLY RUNS. 31.0.14 -> 32.0.9 (one major).
|
||||
### The refused 3-major jump above was recoverable precisely because nothing migrated.
|
||||
2026-09-01T19:05:59Z
|
||||
file now:
|
||||
image: nextcloud:32.0.9-apache
|
||||
### POST /api/stacks/nextcloud/update:
|
||||
{"ok":true,"message":"Stack nextcloud update completed"}
|
||||
|
||||
HTTP=200
|
||||
2026-09-01T19:06:44Z
|
||||
2026-09-01T19:06:54Z nextcloud:32.0.9-apache|Up 10 seconds (health: starting)
|
||||
2026-09-01T19:07:16Z nextcloud:32.0.9-apache|Up 31 seconds (healthy)
|
||||
2026-09-01T19:07:37Z nextcloud:32.0.9-apache|Up 53 seconds (healthy)
|
||||
2026-09-01T19:07:58Z nextcloud:32.0.9-apache|Up About a minute (healthy)
|
||||
2026-09-01T19:08:20Z nextcloud:32.0.9-apache|Up About a minute (healthy)
|
||||
2026-09-01T19:08:41Z nextcloud:32.0.9-apache|Up About a minute (healthy)
|
||||
2026-09-01T19:09:03Z nextcloud:32.0.9-apache|Up 2 minutes (healthy)
|
||||
2026-09-01T19:09:24Z nextcloud:32.0.9-apache|Up 2 minutes (healthy)
|
||||
2026-09-01T19:09:46Z nextcloud:32.0.9-apache|Up 3 minutes (healthy)
|
||||
2026-09-01T19:10:07Z nextcloud:32.0.9-apache|Up 3 minutes (healthy)
|
||||
### did the migration RUN? (occ status + upgrade log)
|
||||
- installed: true
|
||||
- version: 32.0.9.2
|
||||
- versionstring: 32.0.9
|
||||
- edition:
|
||||
- maintenance: false
|
||||
- needsDbUpgrade: false
|
||||
- productname: Nextcloud
|
||||
- extendedSupport: false
|
||||
--- markers ---
|
||||
SPIKE-SEEDED-ON-31.0.14-2026-09-01
|
||||
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
|
||||
### the upgrade evidence in the container log
|
||||
Initializing nextcloud 32.0.9.2 ...
|
||||
Upgrading nextcloud from 31.0.14.1 ...
|
||||
=> Searching for hook scripts (*.sh) to run, located in the folder "/docker-entrypoint-hooks.d/pre-upgrade"
|
||||
==> Skipped: the "pre-upgrade" folder is empty (or does not exist)
|
||||
Updated database
|
||||
Updated <federation> to 1.22.0
|
||||
Updated <lookup_server_connector> to 1.20.0
|
||||
Updated <oauth2> to 1.20.0
|
||||
Updated <password_policy> to 4.0.0
|
||||
Updated <photos> to 5.0.0
|
||||
Updated <activity> to 5.0.0
|
||||
Updated <circles> to 32.0.0
|
||||
Updated <cloud_federation_api> to 1.16.0
|
||||
Updated <dav> to 1.34.2
|
||||
Updated <files> to 2.4.0
|
||||
Updated <files_sharing> to 1.24.1
|
||||
Updated <files_trashbin> to 1.22.0
|
||||
Updated <files_versions> to 1.25.0
|
||||
Updated <sharebymail> to 1.22.0
|
||||
Updated <webhook_listeners> to 1.3.0
|
||||
Updated <workflowengine> to 2.14.0
|
||||
Updated <comments> to 1.22.0
|
||||
Updated <logreader> to 5.0.0
|
||||
Updated <nextcloud_announcements> to 4.0.0
|
||||
Updated <notifications> to 5.0.0
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
### Phase 6c — THE ROLLBACK ATTEMPT: put 31.0.14 back after a migration that RAN
|
||||
2026-09-01T19:10:45Z
|
||||
file BEFORE my edit (the sync may have flipped it):
|
||||
image: nextcloud:34.0.1-apache
|
||||
file now:
|
||||
image: nextcloud:31.0.14-apache
|
||||
### POST /api/stacks/nextcloud/restart:
|
||||
{"ok":true,"message":"Stack nextcloud restart completed"}
|
||||
|
||||
HTTP=200
|
||||
2026-09-01T19:10:49Z
|
||||
2026-09-01T19:10:58Z nextcloud:31.0.14-apache|Restarting (1) Less than a second ago
|
||||
2026-09-01T19:11:15Z nextcloud:31.0.14-apache|Restarting (1) 10 seconds ago
|
||||
2026-09-01T19:11:31Z nextcloud:31.0.14-apache|Restarting (1) 13 seconds ago
|
||||
2026-09-01T19:11:48Z nextcloud:31.0.14-apache|Restarting (1) 4 seconds ago
|
||||
2026-09-01T19:12:04Z nextcloud:31.0.14-apache|Restarting (1) 20 seconds ago
|
||||
2026-09-01T19:12:21Z nextcloud:31.0.14-apache|Restarting (1) 37 seconds ago
|
||||
2026-09-01T19:12:37Z nextcloud:31.0.14-apache|Restarting (1) 2 seconds ago
|
||||
2026-09-01T19:12:53Z nextcloud:31.0.14-apache|Restarting (1) 18 seconds ago
|
||||
|
||||
### THE EXACT MESSAGE from the app when the old version meets migrated data
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
Configuring Redis as session handler
|
||||
Can't start Nextcloud because the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported. Are you sure you have pulled the newest image version?
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
### Phase 6d — POSITIVE CONTROL: the data is NOT destroyed, only the downgrade is refused.
|
||||
### Put 32.0.9 back and read the markers.
|
||||
2026-09-01T19:13:27Z
|
||||
image: nextcloud:32.0.9-apache
|
||||
{"ok":true,"message":"Stack nextcloud restart completed"}
|
||||
|
||||
HTTP=200
|
||||
nextcloud:32.0.9-apache|Up 46 seconds (healthy)
|
||||
- installed: true
|
||||
- version: 32.0.9.2
|
||||
- versionstring: 32.0.9
|
||||
SPIKE-SEEDED-ON-31.0.14-2026-09-01
|
||||
SPIKE-FILE-CONTENT-31.0.14-2026-09-01
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
### OBSERVATION — remove_hdd_data:true, yet the HDD data dir is still there and the response
|
||||
### reported neither removed nor preserved (hdd_paths_removed:null, hdd_paths_preserved:null).
|
||||
total 88592
|
||||
drwxrwx--- 4 www-data www-data 4096 Sep 1 19:07 .
|
||||
drwxr-xr-x 5 root root 4096 Sep 1 18:59 ..
|
||||
-rw-rw-r-- 1 www-data www-data 542 Sep 1 19:06 .htaccess
|
||||
-rw-rw-r-- 1 www-data www-data 52 Sep 1 19:06 .ncdata
|
||||
drwxr-xr-x 3 www-data www-data 4096 Sep 1 18:59 admin
|
||||
drwxr-xr-x 5 www-data www-data 4096 Sep 1 19:07 appdata_ocrus79ehwsx
|
||||
-rw-rw-r-- 1 www-data www-data 0 Sep 1 19:06 index.html
|
||||
-rw-r----- 1 www-data www-data 90692753 Sep 1 19:07 nextcloud.log
|
||||
128M /mnt/felhom-drives/hdd_1/appdata/nextcloud/
|
||||
--- for contrast, the OTHER apps own theirs legitimately ---
|
||||
128M /mnt/felhom-drives/hdd_1/appdata/nextcloud/
|
||||
40K /mnt/felhom-drives/hdd_1/appdata/paperless/
|
||||
12K /mnt/felhom-drives/hdd_1/appdata/romm/
|
||||
2026/09/01 19:14:56 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, SUBDOMAIN, DB_PASSWORD, DOMAIN, HDD_PATH, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_PASSWORD, NEXTCLOUD_ADMIN_USER, USERDATA_PATH, IMPORT_PATH] (14 app + 0 system)
|
||||
2026/09/01 19:15:06 router.go:80: [DEBUG] [api] removeStack: name=nextcloud
|
||||
2026/09/01 19:15:06 router.go:80: [DEBUG] [api] removeStack: name=nextcloud removeHDDData=true removeBackups=true
|
||||
2026/09/01 19:15:06 delete.go:289: [DEBUG] [stacks] RemoveStack called: name="nextcloud", removeHDDData=true, backupPathsToRemove=1
|
||||
2026/09/01 19:15:06 delete.go:303: [DEBUG] [stacks] RemoveStack nextcloud: state=stopped, deployed=true, orphaned=false, deploying=false
|
||||
2026/09/01 19:15:06 delete.go:327: [INFO] Removing deployed stack: nextcloud (removeHDDData=true, backupPaths=1)
|
||||
2026/09/01 19:15:06 delete.go:337: [DEBUG] [stacks] RemoveStack nextcloud: found 0 HDD mounts from compose file
|
||||
2026/09/01 19:15:06 manager.go:1311: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, DB_PASSWORD, DOMAIN, HDD_PATH, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_PASSWORD, NEXTCLOUD_ADMIN_USER, SUBDOMAIN, USERDATA_PATH, IMPORT_PATH] (14 app + 0 system)
|
||||
2026/09/01 19:15:07 delete.go:347: [DEBUG] [stacks] RemoveStack nextcloud: compose down output:
|
||||
2026/09/01 19:15:07 delete.go:402: [DEBUG] [stacks] RemoveStack nextcloud: processing 1 backup paths for removal (base=backups)
|
||||
2026/09/01 19:15:07 delete.go:408: [WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps
|
||||
2026/09/01 19:15:07 delete.go:426: [DEBUG] [stacks] RemoveStack nextcloud: removing app.yaml at /opt/docker/stacks/nextcloud/app.yaml
|
||||
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
|
||||
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.6GB avail=884.5GB (0.6%)
|
||||
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/hdd_1" total=937.8 GB used=5.6 GB avail=884.5GB (0.6%)
|
||||
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
|
||||
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/hdd_1" bsize=4096 total=937.8GB used=5.6GB avail=884.5GB (0.6%)
|
||||
2026/09/01 19:15:09 healthcheck.go:42: [DEBUG] [monitor] Raw values: disk=23.6%, hdd=0.6% (configured=true), mem=15.9% (4106MB/25898MB), cpu=4.1%, temp=52.9°C (hwmon1)
|
||||
2026/09/01 19:15:09 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/hdd_1" total=937.8 GB used=5.6 GB avail=884.5GB (0.6%)
|
||||
2026/09/01 19:15:29 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/hdd_1", hasCPUCollector=true)
|
||||
### ROOT CAUSE of the leftover — positive + negative control
|
||||
### POSITIVE control: the env override IS the only way HDDPath can be set, and it is:
|
||||
0
|
||||
(0 = the env override is ABSENT)
|
||||
(no FELHOM_PATHS_* var at all)
|
||||
|
||||
### and controller.yaml paths: has no hdd_path key (read above).
|
||||
### config.go:117 HDDPath has NO default — only envStr at :403. So cfg.Paths.HDDPath == "".
|
||||
### ParseComposeHDDMounts (delete.go:600-603) returns nil on the FIRST line when hddPath == "".
|
||||
|
||||
### NEGATIVE control — this is not nextcloud-specific. The same 0-mounts line was logged
|
||||
### for bentopdf earlier in this session, which genuinely has no HDD mount:
|
||||
2026/09/01 17:55:29 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/bentopdf/docker-compose.yml
|
||||
### ParseComposeHDDMounts (delete.go:600-603) returns nil on the FIRST line when hddPath == "".
|
||||
@@ -0,0 +1,8 @@
|
||||
### TEARDOWN — restore bentopdf to its catalog tag via the product's own restart path
|
||||
{"ok":true,"message":"Stack bentopdf restart completed"}
|
||||
|
||||
HTTP=200
|
||||
2026-09-01T18:13:18Z
|
||||
image: ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
container=ghcr.io/alam00000/bentopdf:v2.8.6 id=02c80375fcba293c6bd06ee826e6fc851f030bf8f5a2d021394c54964507c69a status=running
|
||||
digest=ghcr.io/alam00000/bentopdf@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
|
||||
@@ -0,0 +1,89 @@
|
||||
### TEARDOWN LAYER 3 — the hub. This run created NO customer and NO appliance record.
|
||||
### It used the EXISTING demo-hp customer. What it DID create is events. Listing them:
|
||||
404 page not found
|
||||
|
||||
### customers list unchanged (5 rows, same as at the start of the run):
|
||||
demo-felhom
|
||||
demo-hp
|
||||
drill-r50
|
||||
peti-felhom
|
||||
tester-1
|
||||
### hub-side residue from the run — the nextcloud telemetry row and the events it pushed
|
||||
nextcloud — Felhom Hub
|
||||
Felhom |Hub
|
||||
Dashboard
|
||||
Customers
|
||||
Apps
|
||||
Hosts
|
||||
Offsite
|
||||
Configuration
|
||||
← Apps
|
||||
24h
|
||||
7d
|
||||
30d
|
||||
Nextcloud
|
||||
Reset Telemetry
|
||||
App Name
|
||||
nextcloud
|
||||
Deployments
|
||||
Catalog Estimate
|
||||
256M
|
||||
Catalog Limit
|
||||
1024M
|
||||
Suggested Limit (P95×1.2)
|
||||
352 MB
|
||||
Avg Memory
|
||||
208 MB
|
||||
P95 Memory
|
||||
280 MB
|
||||
Avg CPU
|
||||
1%
|
||||
Memory Trend
|
||||
Known Issues |— filtered: demo-hp
|
||||
Clear filter (fleet view)
|
||||
Show dismissed
|
||||
Severity
|
||||
Message
|
||||
Occurrences (all customers)
|
||||
Affected Customers
|
||||
First Seen
|
||||
Last Seen
|
||||
warn
|
||||
0 [Warning] InnoDB: liburing disabled: falling back to innodb_use_native_aio=OFF
|
||||
9 min ago
|
||||
9 min ago
|
||||
Details
|
||||
fingerprint: |0 [warning] innodb: liburing disabled: falling back to innodb_use_native_aio=off
|
||||
· severity: warn
|
||||
· first seen: 2026-09-01 19:10:30
|
||||
· last seen: 2026-09-01 19:10:30
|
||||
affected customers:
|
||||
demo-hp
|
||||
Full message:
|
||||
Copy
|
||||
0 [Warning] InnoDB: liburing disabled: falling back to innodb_use_native_aio=OFF
|
||||
No context captured (pre-v0.111 report or warn-severity issue).
|
||||
warn
|
||||
0 [Warning] mariadbd: io_uring_queue_init() failed with errno 0
|
||||
9 min ago
|
||||
9 min ago
|
||||
Details
|
||||
fingerprint: |0 [warning] mariadbd: io_uring_queue_init() failed with errno 0
|
||||
· severity: warn
|
||||
· first seen: 2026-09-01 19:10:30
|
||||
· last seen: 2026-09-01 19:10:30
|
||||
affected customers:
|
||||
demo-hp
|
||||
Full message:
|
||||
Copy
|
||||
0 [Warning] mariadbd: io_uring_queue_init() failed with errno 0
|
||||
No context captured (pre-v0.111 report or warn-severity issue).
|
||||
warn
|
||||
0 [Warning] mariadbd: io_uring_queue_init() failed with errno 2
|
||||
9 min ago
|
||||
9 min ago
|
||||
Details
|
||||
fingerprint: |0 [warning]
|
||||
|
||||
### current containers demo-hp reports (nextcloud must be ABSENT):
|
||||
3
|
||||
@@ -0,0 +1,25 @@
|
||||
### remove the one other image this spike pulled: bentopdf:v2.8.5
|
||||
removed
|
||||
removed alpine:3.20
|
||||
bentopdf images now:
|
||||
ghcr.io/alam00000/bentopdf:v2.8.6
|
||||
|
||||
### TEARDOWN LAYER 2 — the host: space returned
|
||||
Filesystem Size Used Avail Use% Mounted on
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 957M 29G 4% /
|
||||
/dev/mapper/pve-vm--9201--disk--1 69G 12G 54G 18% /mnt/sys_drive
|
||||
/dev/nvme0n1 938G 5.5G 885G 1% /mnt/felhom-drives/hdd_1
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
local dir active 40453376 17977988 20388272 44.44%
|
||||
local-lvm lvmthin active 56487936 40055595 16432340 70.91%
|
||||
### the 1.05 GiB the guest freed did NOT return to the host thin pool. Trying fstrim.
|
||||
before: local-lvm used 40055595 KiB (baseline before this spike: 38959729 KiB)
|
||||
fstrim: /mnt/felhom-drives/hdd_1: FITRIM ioctl failed: Operation not permitted
|
||||
fstrim: /var/lib/felhom: FITRIM ioctl failed: Operation not permitted
|
||||
fstrim: /: FITRIM ioctl failed: Operation not permitted
|
||||
local-lvm lvmthin active 56487936 40055595 16432340 70.91%
|
||||
/var/lib/lxc/9201/rootfs/: 30.2 GiB (32480210944 bytes) trimmed
|
||||
/var/lib/lxc/9201/rootfs/var/lib/felhom: 57 GiB (61255426048 bytes) trimmed
|
||||
rc=0
|
||||
local-lvm lvmthin active 56487936 15127469 41360466 26.78%
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
### TEARDOWN LAYER 1 — the machine: remove the throwaway nextcloud stack this spike created
|
||||
2026-09-01T19:14:36Z
|
||||
{"ok":false,"error":"stack \"nextcloud\" is still running — stop it first before removing"}
|
||||
|
||||
HTTP=409
|
||||
### verify gone
|
||||
nextcloud nextcloud:32.0.9-apache Up About a minute (healthy)
|
||||
nextcloud-db mariadb:11.6 Up 15 minutes (healthy)
|
||||
nextcloud-redis redis:7-alpine Up 15 minutes (healthy)
|
||||
(empty above = containers gone)
|
||||
nextcloud_nextcloud_db_data
|
||||
nextcloud_nextcloud_html
|
||||
nextcloud_nextcloud_redis_data
|
||||
(empty above = volumes gone)
|
||||
app.yaml
|
||||
docker-compose.yml
|
||||
nextcloud
|
||||
paperless
|
||||
romm
|
||||
### stop first, then remove
|
||||
{"ok":true,"message":"Stack nextcloud stop completed"}
|
||||
|
||||
HTTP=200
|
||||
{"ok":true,"data":{"removed":"nextcloud","volumes_removed":null,"hdd_paths_removed":null,"hdd_paths_preserved":null},"message":"Stack nextcloud removed"}
|
||||
|
||||
HTTP=200
|
||||
### verify gone
|
||||
containers:
|
||||
volumes:
|
||||
hdd appdata:
|
||||
nextcloud
|
||||
paperless
|
||||
romm
|
||||
stack dir:
|
||||
docker-compose.yml
|
||||
### removing the spike's own leftover (128 MB the product did not remove)
|
||||
2026-09-01T19:17:28Z
|
||||
128M /mnt/felhom-drives/hdd_1/appdata/nextcloud/
|
||||
after:
|
||||
paperless
|
||||
romm
|
||||
any nextcloud left anywhere under /mnt:
|
||||
stack dir (template only, app.yaml gone = not deployed):
|
||||
docker-compose.yml
|
||||
images:
|
||||
nextcloud:34.0.1-apache
|
||||
nextcloud:32.0.9-apache
|
||||
nextcloud:31.0.14-apache
|
||||
### remove ONLY the three nextcloud images this spike pulled (targeted rmi, never a prune)
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 958M 29G 4% /
|
||||
removed nextcloud:31.0.14-apache
|
||||
removed nextcloud:32.0.9-apache
|
||||
removed nextcloud:34.0.1-apache
|
||||
remaining nextcloud images:
|
||||
(none)
|
||||
/dev/mapper/pve-vm--9201--disk--0 32G 958M 29G 4% /
|
||||
stack dir now:
|
||||
.
|
||||
..
|
||||
.felhom.yml
|
||||
docker-compose.yml
|
||||
@@ -675,9 +675,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
|
||||
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
|
||||
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
|
||||
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
|
||||
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
|
||||
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
|
||||
| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** |
|
||||
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
|
||||
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user