diff --git a/STATUS.md b/STATUS.md index 1cd31774..f3f5c4dd 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,78 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-21 — the one screen that stopped an English speaker is fixed, the floor is raised, and both boxes have it.** +**Updated 2026-09-21 (afternoon) — the update feature: a box can no longer be offered an older version as an "update", and I measured how far behind everything actually is.** + +**A decision I took on my own — you can reverse it.** When we move an app back to an older version in +our catalog, a machine that already took the newer one used to show **"Update available"** — and the +button behind it would have put the **older** version back, on top of data the newer one may already +have changed. We cannot undo that. So from today such a machine says **"Up to date"**, and the button +**refuses**, with a sentence saying the version is newer than ours and that going back needs you. It +only refuses when it is certain; in every unclear case it behaves exactly as before. + +**What I measured, and the numbers are not comfortable.** + +- **Your two demo machines are perfectly up to date** — all ten apps, both machines, nothing behind. +- **But our catalog is far behind the world.** Of the app versions we pin, **46 of 58 have a newer + version out there**. **Seven are a big jump** — Nextcloud, Paperless, Claper, Gokapi, Homepage and + SparkyFitness — and all seven have sat still since **18 July, 65 days**. +- **"Up to date" is not always true, and now I know by how much.** Some of our pins name a *line* + rather than an exact version. **Six of the seven I could check have quietly changed underneath us.** + On the HP machine, four apps say "up to date" over a database engine that has moved. The label is + answering a narrower question than the one a household hears. +- **One app's image is gone from the internet.** Plant-it. It was already marked abandoned, so this is + the expected ending, not a fault. A machine already running it keeps running it; it can never be + installed again. + +**A rule I lifted, and half of one I kept.** Since 13 September our catalog has refused to move a +database engine to a new major version, because the Update button took no backup. **It takes one +now**, so I lifted that ban **for MariaDB** — and only when the engine moves **on its own**, never in +the same change as the app itself, because two changes behind one step give an unreadable failure. +**I did not lift it for PostgreSQL.** That engine does not convert your data by itself and simply +refuses to start on the old data — eleven apps would go down at once. A backup is a way back, not a +conversion. + +**I proved it on a real machine, not only in tests.** On the spare machine I installed a throwaway +app, moved our catalog forward, updated it, **cut the power mid-update**, and then moved the catalog +back. Three results: +- **The power cut is handled honestly.** The machine came back, put the app's version back where it + was, ran it, and told the household in plain words that the update was interrupted and the old + version is running. Nothing was left half-done. +- **The new refusal works.** With the machine running the newer version, the page said "Up to date" + and the Update button was refused, in Hungarian and in English. +- **Nothing moves when it refuses.** I pressed it three times; the app was untouched every time. + +**One new fault found while doing it, and I did not paper over it.** On an English page, the sentence +explaining *why* an update stopped is still **Hungarian** — sitting directly under an English label. +Every sentence the update feature shows in that situation has the same problem, including one that +tells a household **whether their files will come back**. I wrote it down as work to do; I did not fix +it today, because one released version per day is the rule and today's was already out. + +**Two things in the brief I was given were wrong, and I want you to know I checked rather than +assumed.** It said the version label was still Hungarian on English pages — it has been fixed since +yesterday and only our note was out of date. And a number describing our catalog, repeated in four +places, had not matched reality for some time. Both corrected. + +**Rows.** Four closed, three now waiting on you, one corrected, two opened. 299 in total. + +**Needs you — two things.** + +1. **Seven questions about updating by itself.** The machine still cannot update an app on its own, + and 39 of the versions we are behind on are small, safe steps that nobody will press a button for + 39 times. I have written each question down with two or three choices, what each costs, and which + one I would pick. They are in the design notes. + **If you do nothing:** nothing breaks, and nothing updates itself either — the gap to the world + keeps widening, quietly. +2. **The fleet version.** Today's fix is on the two demo machines only. Raising the fleet version + would carry it to the rest — including the tester's real machine — and that is your switch, not + mine. + **If you do nothing:** the other machines keep offering a downgrade as an update. That is a label + and a button, not your data at risk, so waiting is cheap. + + +## Previous note + + +**Updated 2026-09-21 (morning) — the one screen that stopped an English speaker is fixed, the floor is raised, and both boxes have it.** > **Ready for an English-speaking tester: yes — nothing known now stands in their way.** > Ready for a Hungarian volunteer: yes, unchanged. diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index bf70033a..ea166136 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -186,6 +186,174 @@ These are rulings, not proposals. Anything specced against a different assumptio own unit → off-site) stands. Either way the hold sentence ends with what the named copy holds, so a customer is never sent to a copy that cannot bring the data back without being told so. +### 2026-09-21 — decided by CC unattended, operator may reverse + +10. **A box AHEAD of the catalog reads „Naprakész", and the guarded Update refuses to move a pin + backwards** (R-524, controller v0.260.0). *One sentence:* when the catalog is reverted under a box + that already updated, is that a "Frissítés elérhető"? **Options:** (a) leave it — the label + compares for difference, as §5.4 says; (b) show „Naprakész" and let the button still run; + (c) show „Naprakész" and refuse the button. **Costs:** (a) is free and offers a household a + downgrade onto a datadir the newer version may have migrated, which §4 says cannot be undone; + (b) removes the invitation but leaves the loaded gun; (c) costs one comparison and can, wrongly + applied, block a legitimate update. **Why (c):** the direction was already settled — §3 decision 3 + says a version change that cannot be undone needs a human, and this is one. The risk in (c) is + bounded by making the Ahead verdict NARROW: every differing service must be orderable AND newer, + or the answer falls back to today's behaviour. Reversible, no customer-data risk, and it only ever + withholds an act. **Implementation:** `stacks.CatalogOrder`, one verdict read by both the badge + and `UpdatePreflight`. + +--- + +## 3b. OPEN — the seven questions Slices 6 and 7 need answered + +**These are questions, not rulings. CC does not decide them.** Each is one answerable sentence, the +options, what each costs, the recommendation, and what happens if nothing is decided. The measurement +behind them is `audits/UPDATE-ARC-STATE-2026-09-21.md`; the short version is that **46 of the +catalog's 58 exact pins are behind upstream today and 39 of those are within a major** — the +population §3 decision 3 already says may move without a human, and nobody presses 39 buttons. + +### Q1 — When may a box update itself? + +*May the box run the guarded Update by itself between 02:30 and 05:00, nightly?* + +| option | cost | +|---|---| +| **02:30–05:00 nightly, after the backup legs** | the update leans on a copy made hours earlier the same night, which is the freshest the box ever has. The app is down for the health wait in the middle of the night. | +| a weekly window | fewer interruptions; a box sits up to 7 days on a version the catalog already moved past, which widens the support window §3 decision 2 runs on | +| the household picks the window | one more setting on a page that already has several, for a choice almost nobody will change | + +**Recommendation: 02:30–05:00 nightly.** The DB dump runs 02:30 and restic 03:00 on a demo box, so a +window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy of the night without a +new mechanism. **If nothing is decided:** Slice 6 cannot be built at all — every other question below +is downstream of this one. + +### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES? + +*The button's rule and the automatic rule can differ. Should they?* + +**The mechanism, verified at source this session, because an earlier draft had it backwards:** the +guard does **not** refuse these apps. Since v0.239.0/v0.241.0 (§3 decisions 8–9) the Update is refused +only when no copy exists on ANY tier and none can be taken. For an app whose data is bind-mounted +files, `Manager.UpdateTierOrderFor` (`controller/internal/backup/update_guard.go:136-141`) walks +second drive → off-site → **own unit last**, and when the own unit is the copy chosen, +`UpdateCopyHolds` (`:145-165`) ends the hold sentence with *„csak a beállításokat és az adatbázist +tartalmazza, a fájlokat nem"* — it holds the settings and the database and **not the files**. So the +update **proceeds**, and the household is told what the copy holds. + +**With a human pressing, that is an informed choice. With nobody pressing, nobody was informed.** + +| option | cost | +|---|---| +| **automatic requires a fresh copy that HOLDS THE FILES; the button keeps today's rule** | the nine file-leg apps (and any other bind-data app) update automatically only on a box with a second drive or off-site; on a one-drive box they wait for a person. Two rules to hold in one's head. | +| one rule for both — automatic follows the button | simpler; a file-leg app can be updated unattended against a copy that cannot bring its files back, and the sentence saying so is read by nobody | +| automatic skips bind-data apps entirely | simplest; the seven file-leg apps that are behind never move by themselves even when a good copy exists | + +**Recommendation: the first.** It is the smallest rule that keeps the promise the hold sentence makes. +**If nothing is decided:** Slice 6 must be built for the safe subset only, and the file-leg apps stay +manual — which is the third option by default, without anyone choosing it. + +### Q3 — What counts as "within a major" when the tag is not a version number? + +*§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`, +`kimai/kimai2:apache-2.57.0`, a date stamp, a digest?* + +**And the test is per compose SERVICE, with ALL of them having to pass.** An app bump that is minor +while its `mariadb:` sidecar moves a major is **ACROSS** — that sidecar now converts the customer's +datadir by itself (R-459), so the edge carries a migration whatever the app's own number says. + +| option | cost | +|---|---| +| **an unorderable tag on ANY service makes the whole edge ACROSS → human** | the 8 floating pins and every suffix-versioned image stay manual. Conservative, and it is the same rule v0.260.0's `CompareImageRefs` already implements and tests. | +| teach the comparator each shape | every new shape is a new rule, and a wrong rule silently automates a major | +| compare digests instead | needs Q6 first, and a digest carries no order at all — it can say "different", never "newer" | + +**Recommendation: the first**, reusing `stacks.CompareImageRefs` rather than writing a second rule. +*One small extension is needed and is named here so it is not discovered late:* v0.260.0's +`CompareImageRefs` answers *orderable?* and *newer?*, which is all R-524 needed. Slice 6 also needs +*same major?*, so the parsed major has to be exposed from the same normaliser — **an addition to the +one comparator, never a second one.** +**If nothing is decided:** Slice 6 would have to invent a rule under time pressure, which is how a +major gets automated by accident. + +### Q4 — A held app: who is told, when, and does the box try again? + +*An automatic update that ends HELD happened while everyone was asleep.* + +| option | cost | +|---|---| +| **the household on the app page and by mail ONCE; the operator by event; NO retry until the catalog moves again or a person presses** | one mail per held app. The app stays down until someone acts — which is already true of a held update today. | +| retry the next night | a broken edge takes the app down every night and mails every morning; the hold exists precisely because the box cannot fix it | +| tell only the operator | the household finds their app down and has no sentence explaining it | + +**Recommendation: the first.** It is what the manual hold already does (`settings.RestoreHold` with +`reason: update_failed`), plus one mail. **If nothing is decided:** the safe default is no automatic +update at all, because a hold nobody is told about is worse than a version nobody moved. + +### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? + +*Eleven templates, and the image performs no conversion — it refuses to start on an older major's +datadir (R-463).* + +| option | cost | +|---|---| +| **a scripted `pg_upgrade` edge in the harness, proven on all eleven, before the catalog may move** | real work: eleven fixtures, and `pg_upgrade` needs both major's binaries present. The engine-major gate keeps the rule until it exists. | +| move the pin and let the update HOLD honestly | every one of the eleven apps goes down on the same night and comes back only by a restore | +| never move PostgreSQL majors | the fleet sits on an engine that eventually loses upstream support | + +**Recommendation: the first, and the gate stays until it lands.** As of 2026-09-21 the engine-major +rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of it and +`MARIADB_AUTO_UPGRADE=1`); this half is exactly what stays. **If nothing is decided:** nothing breaks +— the gate refuses the move — but the eleven apps drift further from upstream every month. + +### Q6 — Should the catalog record each pin's DIGEST at push time? + +*So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.* + +**This is no longer theoretical. Measured 2026-09-21: six of the seven measurable floating pins have +been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), +`postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, +`postgis:16-3.5-alpine`. On demo-hp today, four apps read „Naprakész" over a database engine image +that has demonstrably moved. + +| option | cost | +|---|---| +| **the catalog records the digest at push time; the box compares digests** | one field per pin. `check-image-resolvable.py` already resolves the digest, so the producer exists. §8.1's rule — the box never queries a registry — is untouched. | +| the box queries registries | breaks §8.1 outright: a page that cannot render without eight upstream registries | +| leave it | the badge stays right about the question it asks and wrong about the one a household hears | + +**Recommendation: yes.** It is the cheapest real improvement on this list and it closes R-446. +**If nothing is decided:** „Naprakész" keeps meaning "the reference matches", which is measurably not +what it sounds like. + +### Q7 — What does the hub's report need to carry for a fleet view? + +*Slice 7 lets the operator SEE and MOVE how far behind every box is.* + +**Verified both sides this session:** the controller's report payload carries name, state, CPU and +memory and no image (`controller/internal/report/types.go` L98–103), and the hub's +`Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts**. **But +the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent — +what is missing is the denormalisation and the page, not the transport. + +| option | cost | +|---|---| +| **per app: installed reference + catalog reference + badge state; the hub lists boxes behind, with a "move" that is the same guarded Update, operator-triggered** | additive on both sides; the report grows by a few fields per app | +| badge state only | smaller payload; the operator cannot see WHAT is behind, only that something is | +| leave it to per-box pages | free today at two boxes; unusable at twenty | + +**Recommendation: the first, and it stays P3-LOW until the fleet grows.** **If nothing is decided:** +the only way to answer "is the fleet current?" is what this session did — read both boxes' files by +hand. + +--- + +### Not a question — already ruled + +**R-462's scope was decided on 2026-09-13.** §3 decision 6: the upgrade test goes to **all** apps +through the nightly rotation, explicitly *not* "database apps first". The register row R-462 still +says *"VIKTOR rules on scope"* — **that row is stale and is corrected to cite decision 6.** The update +night below proposes an ORDER *inside* that ruling; it does not reopen it. + ## 4. The vocabulary ruling — "rollback" is struck **App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually @@ -329,7 +497,7 @@ earlier feature is the failure mode to look for whenever a file changes meaning. |---|---|---| | **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** | | **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** | -| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** | +| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02); English since v0.258.0 (R-589); a FOURTH verdict — AHEAD — and the downgrade refusal in v0.260.0 (R-524, §3 decision 10)** | | **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 | | **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | **SHIPPED + PROVEN LIVE, controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (2026-09-13); any backup tier since v0.239.0 (§3 decision 8)** — §6.1 | | **5** | **An upgrade test that runs again** — a harness that upgrades a real app with real data in it and asks the app for the data back. | **SHIPPED, `app-catalog/scripts/upgrade-test.py` (2026-09-06)** — 7 edges, 3 apps; see §4.1 and §10 | @@ -484,6 +652,82 @@ breaks. --- +### 6.2 Slice 6, as it would be built (OPEN — R-450; needs Q1–Q4) + +**Not a design yet; the shape the seven questions bound.** Written down so the answers have somewhere +to land. + +**Nothing new happens to the app.** Inside the window, for an app that qualifies, the box runs +**exactly the guarded Update of §6.1** — same precondition, same safety dump, same pin journal, same +health wait, same hold. Slice 6 adds a *caller*, not a *path*. That is the whole reason it is +affordable: every failure mode was measured in slice 4 and every one of them already ends in a hold +the household can read. + +**An app qualifies when ALL of these hold** (each clause is a question above, not a decision taken): + +1. the window is open (Q1); +2. `stacks.CatalogOrder` says **Behind** — never Unknown, never Ahead (v0.260.0 gives all four); +3. the edge is **within a major for EVERY compose service**, the engine sidecar included (Q3), + judged by `stacks.CompareImageRefs` — one unorderable service makes the whole edge *across*; +4. a fresh copy exists on a tier that holds what this app's data actually is (Q2); +5. the app's own switch is on (Q2's default: on). + +**What the household sees.** An event and a line on the app page's timeline, before and after, in both +languages: *„Automatikus frissítés 03:12-kor — sikeres"* / *„— megállítva, a másolat 2026-09-20-i"*. +A held app is not retried until the catalog moves again or a person presses (Q4). + +**Where it would live.** A scheduler beside the existing nightly legs, reading `settings` for the +window and the per-app switch, and calling `Manager.StartGuardedUpdate`. **It must respect the same +`isHeld`/`SetUpdatingCheck` interlocks v0.238.1 added** — the nightly capture running *inside* an +update's health wait is the defect that release fixed, and a second unattended caller is exactly the +shape that finds it again. + +**Ships behind `auto_update: off` with no UI until the operator answers Q1.** + +### 6.3 Slice 7, as it would be built (OPEN — R-451; needs Q7) + +Three additive pieces, and the transport already exists (§3b Q7): + +1. **controller** — the report's per-app object gains installed reference, catalog reference and badge + state. Additive; an older hub ignores it. +2. **hub** — denormalise those out of the raw report it already stores whole, and list boxes by how + far behind they are. +3. **hub → box** — a "move" button that is the same guarded Update, operator-triggered, through the + existing command path. **Not a second update mechanism**, and not automatic. + +Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the state audit is what +that looks like today: the only way to answer *"is the fleet current?"* was to read both boxes' files +by hand. + +### 6.4 The update night — a drill brief outline, costed from R-462's real numbers + +**The ruling is decision 6: all 53 apps, through the nightly rotation.** This is an ORDER inside that +ruling, not a scope change. The database apps go first because they are the ones where a wrong answer +costs data rather than uptime. + +**The real numbers this rests on** (R-462, measured 2026-09-06): a successful edge takes +**6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly 8×, because a negative is +only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**. **Machine +time is not the cost — fixtures are.** Two of the three apps needed a bespoke non-browser seed route, +one needed two attempts and a discarded approach, and one (bookstack) can only ever be half-proven +headlessly (R-460). + +| leg | what | cost | +|---|---|---| +| A | the **15 database services** — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus `adventurelog`'s postgis — one edge each, fixture per app | **15–25 CC-hours**, dominated by seed routes; ~30 min machine time at the median; ~25 GB | +| B | one **power cut mid-update** on a real version change, in `pulling` and again in `starting` | 1–2 CC-hours (R-520 — the first half is measured in this session) | +| C | one **PostgreSQL `pg_upgrade` rehearsal**, the Q5 edge, on one app before any of the eleven | 3–4 CC-hours | +| D | one **downgrade refusal** | **already done** — v0.260.0, proven live 2026-09-21 | +| E | the **automatic night** on a throwaway: one app, one real catalog step, inside a simulated window, with the guard and the hold; then the same edge made to fail → HOLD, the event, no retry loop | 2–3 CC-hours | +| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised | + +**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C +and E are the ones that unblock a decision; leg A is the one that takes the time. + +**Venue:** a scratch host, never a customer box — `demo-hp`'s guest 9202 for the box-side legs, the +harness on DooPlex for the image-side ones. + + ## 7. What slices 1 and 2 actually built ### 7.1 The record (slice 1) @@ -547,12 +791,25 @@ Version strings stay in the logs, the API and the hub. ## 8. Known limitations, stated plainly -1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and - queries no registry — a customer's box must not depend on reaching eight upstream registries to - render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be - identical while the image behind it has moved. **Measured, not theorised:** spike §5 found - `mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls - holding. Digest-level comparison needs a registry query and is deferred — **R-446**. +1. **„Naprakész" can be FALSE for the floating pins, and 2026-09-21 measured HOW false.** The + comparison is reference-to-reference and queries no registry — a customer's box must not depend on + reaching eight upstream registries to render a page. For `postgres:16-alpine`, `mariadb:11.6` and + the others the reference can be identical while the image behind it has moved. + **NUMBERS, 2026-09-21** (`audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3). **The count this document + carried — "23 of 66" — is STALE and matched no definition the catalog supports today.** Recounted + at catalog `18a6d2d8`, with the definition stated so it can be rechecked: a pin FLOATS when its + tag names a version LINE rather than an exact release. Of 66 unique pins, 48 are full `X.Y.Z`, + **6 are two-part lines** (`mariadb:11.4`/`11.6`/`12.3`, `claper:2.5`, `opengist:1.13`, + `wger/server:2.6`) and **4 are major lines** (`postgres:15-alpine`, `postgres:16-alpine`, + `redis:7-alpine`, `postgis:16-3.5-alpine`) — **10 float**. The remaining 8 are exact versions + wearing a variant suffix (`ghost:6.53.0-alpine`, `nextcloud:34.0.1-apache`, …), which do not float + by this definition. Of the 8 database and cache engine pins the sweep measured, 7 were measurable + and **6 have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` + (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`. Only `mariadb:11.6` has not. + The 8th, immich's own ghcr build, is UNMEASURED — ghcr exposes no anonymous last-modified. So on + demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved. + Digest-level comparison needs the catalog to record the digest at push time — **R-446**, put to the + operator as §3b **Q6**, recommended YES. 2. **~~Nothing enforces `catalog_since`.~~ Enforced by the pre-push hook since 2026-09-13 (R-452, `app-catalog-felhom.eu/scripts/check-catalog-since.py`); CI's shallow clone still skips it out loud.** A commit that moves an `image:` line and forgets the date under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent commit to diff against, so the gate needs a deeper fetch — **R-452**. diff --git a/documentation/audits/UPDATE-ARC-STATE-2026-09-21.md b/documentation/audits/UPDATE-ARC-STATE-2026-09-21.md new file mode 100644 index 00000000..a2005e79 --- /dev/null +++ b/documentation/audits/UPDATE-ARC-STATE-2026-09-21.md @@ -0,0 +1,346 @@ +# UPDATE ARC — the state, measured (2026-09-21) + +Phase 0 of the resumed update arc. **Nothing here is an estimate.** Every number was read from a +live box, a live registry or live source on 2026-09-21. Raw captures: +`audits/update-arc-2026-09-21/`. + +Baselines at the start: controller `19ef0329ab66` v0.259.0 · agent `d9864a94bf62` v0.132.0 · +felhom.eu `bcdd5b205875` hub v0.119.0 · catalog `18a6d2d8243e`. All four trees clean and level with +`origin/main`. + +--- + +## 1. Three claims in the brief that turned out wrong — named first + +**1.1 R-589 is NOT open. It shipped in v0.258.0.** The brief said, reviewer-verified, that +`updatebadge.go` L82–L99 builds the badge from four raw Hungarian literals and that R-589 is +therefore open. **The literals are real and the conclusion does not follow.** Those literals are the +Hungarian form, and they are deliberately frozen — that IS the localisation parity guarantee. The +ENGLISH form has been rebuilt from the bundle in `web.localeFuncs` (`internal/web/i18n_web.go`, +the `"updateBadge"` entry) since **v0.258.0**, with `badge.update.*` in both `hu.json` and `en.json`, +pinned by `TestUpdateBadgeFollowsTheLanguage`, and **proven live on a fresh box the same morning** — +`DRILL-first-hour-en-0258-2026-09-20.md` item 9 reads *"PASS — 'Up to date' in English (R-589, fixed +this morning, proven on a fresh box)"*, and R-561's own row records R-589 among the three defects +Part 0 of v0.258.0 fixed. + +**The register row is stale, not the code.** It is closed in this session with that citation. The +lesson is the general one: *a reviewer who reads one producer cannot see a second producer that +overrides it.* Reading `updatebadge.go` alone gives exactly the wrong answer, and the file now says +so in its own comment. + +**1.2 The chaos-night canary is NOT a defect to fix.** The brief asked why +`check-image-resolvable` and `check-volume-persistence` returned INCONCLUSIVE on 2026-09-17 and +whether the fix is bounded. **They behaved exactly as designed.** Both scripts' own headers state the +rule: a detector that cannot prove itself must refuse to report rather than guess — +`check-image-resolvable.py` cites the 2026-07-21 incident where treating a Docker Hub throttle as +failure swept 24 of 65 pins as falsely dead, and the phrase in the drill doc, *"its own canary +failed"*, is `check-volume-persistence.py`'s own literal error text (its `self_test`, which needs to +`docker build` a canary image and therefore needs registry access). The most-supported reading is +that registry access was constrained at that moment; **that is inferred from the code plus the +documented throttle precedent, not observed** — no raw stdout of the gate run survives in either +evidence directory. Round 7's response — substitute `use` for `update`, log the deviation, leave the +tree clean — is what the scripts' own documentation asks for. + +**No register row named it, and there is one real gap**, which is diagnostic rather than behavioural: +`catalog_gates.py` collapses *"the harness refused to run at all"* and *"a per-app result is +genuinely undetermined"* into one `INCONCLUSIVE` label, so a reader cannot tell which happened +without re-running with output captured. Filed as **R-605**. + +**1.3 "The hub report carries no image field" — CONFIRMED, on both sides, with one useful nuance.** +The brief confirmed the controller half (`internal/report/types.go` L98–103: name, state, CPU, +memory). The hub half was this session's to check. `Store.SaveReport` +(`hub/internal/store/store.go:965`) **stores the raw report JSON whole** and denormalises only a +fixed list out of it — controller version, URL, language, CPU/memory percent, container **total and +running counts**, last snapshot, health status. `Containers` is two integers; there is no per-app +structure anywhere, and no hub code reads an image tag. + +**The nuance matters for costing Slice 7:** because the raw payload is kept whole, a new +controller-side field would already be *stored* the day the controller sends it — what is missing is +the denormalisation and the page, not the transport. That makes Slice 7 additive on both sides, as +`09` §8.9 says, and slightly cheaper than "a hub-side change" suggests. + +--- + +## 2. How far behind are the two demo boxes? Not at all. + +Read on-disk on both boxes (`app.yaml` `installed_images` against the syncer's own catalog clone at +`/catalog-cache/templates//docker-compose.yml` — the exact two inputs +`stacks.CatalogOrder` reads), controller v0.259.0 on both, catalog cache at `18a6d2d8` on both. + +| box | deployed apps | behind | unknown (no record) | oldest `catalog_since` | +|---|---|---|---|---| +| N100 `demo-felhom-8363b5` | 1 (opengist) | **0** | **0** | 2026-07-18 | +| HP t740 `demo-hp-bb76ea` | 9 | **0** | **0** | 2026-07-12 | + +**Every deployed app on the fleet reads „Naprakész" today, and every one of them has a record.** The +v0.234.0 backfill has done its job: the "unknown on a quiet box" residue of §8.3 is zero here. + +**This is the honest caveat, and it is the whole of §8.1:** „Naprakész" on these boxes means *the +reference matches*, not *the image has not moved*. Section 3 measures exactly how false that can be. + +--- + +## 3. Upstream drift — the number that decides how urgent Slice 6 is + +53 apps · 79 `image:` lines · **66 unique pins**. Measured from DooPlex +against the public registries, never from a box (§8.1's rule applies to the product, not to an +audit). The dead-image sweep used the catalog's own `check-image-resolvable.py --all`, unmodified. + +### 3.1 The exact pins + +| | count | +|---|---| +| exact pins evaluated | 58 | +| **behind upstream today** | **46** | +| …of those, a MAJOR step | **7** | +| …of those, a minor/patch step | **39** | +| already newest | 10 | +| Felhom's own internal image (not upstream) | 1 | +| inconclusive | 1 | + +**39 of 46 are within a major — that is the population Slice 6's automatic update would serve.** +Seven cross a major and are human work by the 2026-09-02 ruling: nextcloud 34→35, paperless-ngx +2.20→3.2, claper 2.5→3.0, gokapi 1.9→2.2, homepage 1.13→2.4, sparkyfitness 0.17→1.7 (two images). +**All seven were pinned 2026-07-18 — 65 days ago.** + +### 3.2 By class + +| class | apps | pins | floating | behind (of exact) | major | minor/patch | +|---|---|---|---|---|---|---| +| database | 14 | 23 | 7 | 14 of 16 | 5 | 9 | +| file-leg | 9 | 9 | 0 | 7 of 9 | 0 | 7 | +| other | 30 | 34 | 1 | 25 of 33 | 2 | 23 | + +*Classification caveat, stated rather than folded in:* `adventurelog` runs a `postgis/postgis` +sidecar, which is PostgreSQL, but the literal substring rule puts it in `other`. It is the one app +that moves if postgis counts as a database (`other` 30→29, `database` 14→15). **v0.240.0's R-484 +already treats PostGIS/pgvector/TimescaleDB as Postgres for the logical dump**, so the product is +right and only this table's rule is literal. + +### 3.3 The floating pins — §8.1 is now MEASURED, and its COUNT was stale + +**First, the count itself.** `09` §8.1, `REUSE.md`, the badge's own comment and R-446 all said +*"23 of the catalog's 66 distinct pins float"*. **That number matches no definition the catalog +supports today**, and it has been corrected in all four places. Recounted at catalog `18a6d2d8`, +with the definition written down so it can be rechecked: **a pin FLOATS when its tag names a version +LINE rather than an exact release.** + +| shape | count | examples | +|---|---|---| +| full `X.Y.Z` | 48 | `privatebin/pdo:2.0.5` | +| **two-part line** | **6** | `mariadb:11.4`/`11.6`/`12.3`, `claper:2.5`, `opengist:1.13`, `wger/server:2.6` | +| **major line** | **4** | `postgres:15-alpine`, `postgres:16-alpine`, `redis:7-alpine`, `postgis:16-3.5-alpine` | +| exact version + variant suffix | 8 | `ghost:6.53.0-alpine`, `nextcloud:34.0.1-apache`, `kimai/kimai2:apache-2.57.0` | + +**10 of 66 float.** The last row does not, by this definition — those name an exact release and only +wear a flavour. + +**Then, the movement.** The sweep measured the 8 database and cache engine pins, of which 7 were +measurable: + +For each floating tag, the registry's own last-push date against the date the catalog set the pin: + +| pin | apps | pinned | last repushed upstream | moved? | +|---|---|---|---|---| +| `postgres:16-alpine` | 8 apps | 2026-07-18 | 2026-09-21 | **yes** | +| `postgres:15-alpine` | sparkyfitness | 2026-07-18 | 2026-09-21 | **yes** | +| `redis:7-alpine` | 6 apps | 2026-07-18 | 2026-09-18 | **yes** | +| `mariadb:11.4` | romm | 2026-07-18 | 2026-09-18 | **yes** | +| `mariadb:12.3` | bookstack | 2026-07-18 | 2026-09-18 | **yes** | +| `postgis/postgis:16-3.5-alpine` | adventurelog | 2026-07-18 | 2026-08-31 | **yes** | +| `mariadb:11.6` | kimai, nextcloud | 2026-07-18 | 2025-02-04 | no | +| `ghcr.io/immich-app/postgres:16-vectorchord…` | immich | 2026-07-18 | — | **UNMEASURED** | + +**Six of the seven measurable floating tags have been repushed since the catalog pinned them.** So on +demo-hp today, `adventurelog`, `bookstack`, `docmost` and `romm` read „Naprakész" over a database +engine image that has demonstrably moved. The badge is not lying — it answers the question it was +built to answer — but the question a household hears is the other one. **R-446 is no longer a +theoretical blind spot: it is six pins out of seven.** The eighth (immich's own ghcr build) could not +be measured: ghcr's anonymous API exposes no last-modified timestamp and no authenticated token was +used. Stated unmeasured rather than guessed. + +### 3.4 One image is gone upstream + +`msdeluise/plant-it:0.10.0` — the gate says INCONCLUSIVE (a `denied`/`unauthorized` stderr, which it +refuses to read as "dead"), and two independent signals say it really is gone: Docker Hub's catalog +API answers `object not found` for the whole repository, and the upstream GitHub repo 404s. The app +is already `lifecycle: abandoned`, so this is the expected end state, not an incident. **A deployed +plant-it would survive — nothing re-pulls a running image — but it can never be redeployed.** + +--- + +## 4. What this session then did + +- **controller v0.260.0** — R-524 closed: a box AHEAD of the catalog reads „Naprakész", and the + guarded Update refuses to move a pin backwards (`downgrade`, 409). The comparison moved to + `stacks.CatalogOrder` so the badge and the refusal cannot drift. Three red-proofs, each seen to + fail. See the repo's CHANGELOG. +- **catalog** — R-469: the MariaDB half of the engine-major rule lifted, with R-450's own-edge clause + enforced in its place; PostgreSQL and MySQL stay refused, now citing R-463 rather than the shipped + R-448. Two new decoys, two red-proofs. +- **R-589** closed as already-shipped (§1.1). **R-605** filed (§1.2). + +Live proof of v0.260.0 and the R-520 power cut: `audits/update-arc-2026-09-21/`. + +--- + +## 4a. R-524 — PROVEN LIVE, both languages, both surfaces + +The state was built the honest way, not by editing a file on the box: a throwaway `uptime-kuma` on +the scratch guest, a real catalog move **2.4.0 → 2.5.0**, the guarded Update pressed and run to +completion, then the catalog **reverted to 2.4.0** — exactly the BIGNIGHT Phase 6 shape that R-524 +recorded. The box then runs 2.5.0 against a 2.4.0 catalog. + +**The badge — app page AND apps list, both languages:** + +``` +[hu] Naprakész +[en] Up to date +``` + +`tag-warn` is **absent** on all four captures (negative control), and the AHEAD title is a *different +sentence* from the ordinary up-to-date title captured at baseline — which is the v0.260.0 behaviour, +not a coincidence of wording. + +**The refusal — the endpoint the button invokes:** + +| request | code | body | +|---|---|---| +| `POST /api/stacks/uptime-kuma/update` | **409** | „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére." | +| the same with `?lang=en` | **409** | "This version is newer than the one in the catalog — moving back needs the operator." | +| the same with `?lang=hu` (control) | **409** | the Hungarian sentence | + +**The English one is the half that was nearly missed.** The new refusal was born as a bundle key, but +`api.Router` rendered update refusals from `ref.Message` — the Hungarian fallback — so without the +one-line change to `errText` in the same release the key would have been a seam built and never +wired. The `?lang=en` row above is the proof that it is wired. + +**The log names both maps, so the verdict can be audited without guessing:** + +``` +[ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the +catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 …}] +catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) +``` + +**And NOTHING MOVED — which is the property that makes the refusal safe to add.** After three refused +POSTs the app is still up and healthy on 2.5.0, the live compose line still reads 2.5.0, and **no +update journal exists** (with a positive control proving the directory searched was the real one). +That is §6.1's phase-0 contract intact: a refusal taken *before* the intent is recorded leaves no +trace, so a household that presses a refused button has changed nothing. + + +--- + +## 4b. R-520 — a power cut during a REAL version change. CLOSED, and it behaved. + +**What R-520 asked for**, and why the 2026-09-14 attempt could not answer it: that night the only +Update the catalog allowed was a **same-version** one, so nothing could have gone wrong and nothing +did. The row asked for the same cut on a **real bump**. + +**The measurement**, on the scratch guest 9202 (demo-hp), controller v0.260.0, with a throwaway +`uptime-kuma` and a real one-step catalog move 2.4.0 → 2.5.0 (reverted in the same session): + +| | | +|---|---| +| 11:04:54.436Z | poller sees `updating: true`, `update_phase: pulling` | +| 11:04:54.439Z | `pct stop 9202` — the cut | +| 11:04:58.207Z | stopped | + +*Honest limit on the instrument, recorded rather than smoothed:* `pct stop` took 3.8 s to return, so +the phase **at the moment of the decision** is observed and the phase **at the moment the kernel +froze** is inferred. + +**After `pct start`, the box said so itself** — the POSITIVE observable, not an absence: + +``` +update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) + — nothing had run; putting the pin back +pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0 +update uptime-kuma: pin and definition PUT BACK to the pre-update version +``` + +- `pinned_images` **and** `installed_images` both 2.4.0; the live compose line back to 2.4.0. +- No hold (correctly — nothing ran), `update_phase: failed`. +- The app is **running and healthy on 2.4.0**, 31 s after boot. +- The household is told, and the sentence is true: *„A frissítés megszakadt, mert a vezérlő + újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* + +**R-520 closes.** A power cut in `pulling` during a real version change is recovered honestly: the +pin goes back, the old version runs, and the page says what happened. + +### The instrument trap this run found, and it is worth more than the result + +The first post-crash read, taken from the host against the **stopped** guest, reported the update +journal **ABSENT** — which would have made this a defect finding. It was a **false negative**. + +**`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts.** +Guest 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume; from the host that +path exists and is an **empty stub**. The journal was at +`/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the +box's own recovery read it at the next boot — which is the proof that it was there all along. + +**It was caught by running `find` over the whole rootfs instead of trusting one constructed path**, +and by demanding a positive control for the directory being searched. **An empty directory is not +evidence of an absent file** — it is equally consistent with "you are looking at an unmounted stub". +That is R-96 rule 3 in a new place. + +--- + +## 4c. A defect the live run found — filed, not fixed today + +**On the English app page, the badge is English and the sentence under it is Hungarian.** Captured +verbatim at step 5, directly above one another: + +``` +[en] Update available — today + A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. + Az alkalmazás a korábbi verzióval fut tovább. +``` + +`Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it from +eight raw literals in `update.go`, and **both** `app_info.html` and `stacks.html` render it verbatim. +The same path carries `UpdatePhaseLabel` and the HOLD sentence — including +`backup.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"*, +**which is a promise about whether the customer's files come back**. + +**v0.260.0 did not close this.** That release routed the 409 REFUSAL through `errText`, and a refusal +is the path where nothing happened; these are the sentences for when something did. **Filed as +R-606 (P2), not fixed** — one release per repo per session, and today's had already shipped. + +--- + +## 4d. Teardown — three layers, stated + +| layer | state | +|---|---| +| **machine** (guest 9202) | throwaway `uptime-kuma` removed **through the product** — the API refused the remove while it ran (409, *"stop it first"*), so it was stopped through the product and then removed, taking its volume and its 24K backup. Only the three protected infra containers remain, as at baseline. No update journal. | +| **host** (demo-hp) | guest 9202 left running, as it was found. **Guest 9201 untouched** — 24 containers before and after. `felhom-pve` not involved. | +| **hub** | nothing provisioned, nothing enrolled, nothing discarded. Guest 9202 reports to no hub by design. | + +**The catalog is back where it started:** `louislam/uptime-kuma:2.4.0`, tree clean and level with +`origin/main`. `catalog_since` reads 2026-09-21 rather than 2026-07-18 — the gate requires an image +move to carry the day's date, and the bump and its revert are two moves that net to zero. No box is +behind on that app, so no age count is affected. + +--- + +## 5. What the numbers say about the two open slices + +**Slice 6 (automatic within a major) is worth more than its P2 suggests, and the reason is §3.1:** +39 of 58 exact pins are one minor step behind, every one of them a change the 2026-09-02 ruling +already says may happen without a human. Nobody presses 39 buttons. The support window (§3 decision +2) runs on how far behind the catalog a box is, so today every box is inside it only because the +catalog has not moved since 2026-09-15 — the moment the catalog catches up with upstream, every box +on the fleet is 46 pins behind at once. + +**Slice 7 (the fleet view) is still correctly P3-LOW** at two enrolled boxes — but §2 shows why it +will not stay cheap: the only reason a person could answer "is the fleet current?" today is that a +session read both boxes' files by hand. + +**R-446 (digests) is the cheapest real improvement on this list.** §3.3 measured the blind spot at +six pins of seven, and the catalog's `check-image-resolvable.py` already resolves a digest for every +pin at push time. Recording it costs one field and closes the gap without the box ever reaching a +registry. diff --git a/documentation/audits/update-arc-2026-09-21/00-drift.py b/documentation/audits/update-arc-2026-09-21/00-drift.py new file mode 100644 index 00000000..05ce539f --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/00-drift.py @@ -0,0 +1,225 @@ +#!/usr/bin/env python3 +"""Measure upstream drift for app-catalog-felhom.eu image pins. Read-only, DooPlex-only. + +Reuses the auth-handling approach of scripts/check-image-resolvable.py (per-registry anonymous +bearer-token flow against the Docker Registry v2 API), but instead of `docker manifest inspect` +per-ref, lists the full upstream tag catalogue so we can compare the pinned tag against the newest +one available, and classify the step as same-major (minor/patch) or major. +""" +import json +import re +import sys +import time +from pathlib import Path + +import requests + +SCRATCH = Path("/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/5f83c0fb-b1bf-404e-b99e-2e3b28a31615/scratchpad") + +SESSION = requests.Session() +SESSION.headers.update({"User-Agent": "felhom-catalog-drift-audit/1.0 (read-only measurement)"}) + + +def get_bearer_token(www_auth: str) -> str: + assert www_auth.startswith("Bearer ") + parts = www_auth[len("Bearer "):].split(",") + d = {} + for p in parts: + k, v = p.split("=", 1) + d[k] = v.strip('"') + r = SESSION.get(d["realm"], params={"service": d.get("service"), "scope": d.get("scope")}, timeout=20) + r.raise_for_status() + return r.json()["token"] + + +def tags_list_all(host: str, repo: str, max_pages=1000): + """Full tag list via Docker Registry v2 API, following RFC5988 Link pagination.""" + url = f"https://{host}/v2/{repo}/tags/list" + params = {"n": 1000} + r = SESSION.get(url, params=params, timeout=30) + headers = {} + if r.status_code == 401: + tok = get_bearer_token(r.headers["WWW-Authenticate"]) + headers = {"Authorization": f"Bearer {tok}"} + r = SESSION.get(url, params=params, headers=headers, timeout=30) + r.raise_for_status() + j = r.json() + tags = list(j.get("tags") or []) + link = r.headers.get("Link") + pages = 1 + while link and pages < max_pages: + m = re.search(r'<([^>]+)>;\s*rel="next"', link) + if not m: + break + nexturl = m.group(1) + if nexturl.startswith("/"): + nexturl = f"https://{host}{nexturl}" + r = SESSION.get(nexturl, headers=headers, timeout=30) + r.raise_for_status() + j = r.json() + tags.extend(j.get("tags") or []) + link = r.headers.get("Link") + pages += 1 + return tags + + +def manifest_digest(host: str, repo: str, ref: str): + """Resolve one ref's manifest digest (for floating-tag drift: 'has the tag moved'). Read-only.""" + url = f"https://{host}/v2/{repo}/manifests/{ref}" + accept = ("application/vnd.docker.distribution.manifest.v2+json," + "application/vnd.docker.distribution.manifest.list.v2+json," + "application/vnd.oci.image.manifest.v1+json," + "application/vnd.oci.image.index.v1+json") + r = SESSION.head(url, headers={"Accept": accept}, timeout=30) + if r.status_code == 401: + tok = get_bearer_token(r.headers["WWW-Authenticate"]) + r = SESSION.head(url, headers={"Accept": accept, "Authorization": f"Bearer {tok}"}, timeout=30) + if r.status_code != 200: + return None, f"HTTP {r.status_code}" + return r.headers.get("Docker-Content-Digest"), None + + +# host resolution per repo prefix, mirroring how these refs are actually pulled +def resolve_host_repo(ref_repo: str): + if ref_repo.startswith("ghcr.io/"): + return "ghcr.io", ref_repo[len("ghcr.io/"):] + if ref_repo.startswith("lscr.io/"): + return "lscr.io", ref_repo[len("lscr.io/"):] + if ref_repo.startswith("registry.gitlab.com/"): + return "registry.gitlab.com", ref_repo[len("registry.gitlab.com/"):] + if ref_repo.startswith("gitea.dooplex.hu/"): + return "gitea.dooplex.hu", ref_repo[len("gitea.dooplex.hu/"):] + if ref_repo.startswith("quay.io/"): + return "quay.io", ref_repo[len("quay.io/"):] + # Docker Hub + if "/" not in ref_repo: + return "registry-1.docker.io", f"library/{ref_repo}" + return "registry-1.docker.io", ref_repo + + +VERSION_CORE_RE = re.compile(r'(\d+(?:\.\d+){1,3})') +UNSTABLE_MARKERS = ("rc", "beta", "alpha", "dev", "nightly", "canary", "edge", "preview", "snapshot", "-pr", "test") + + +def split_tag(tag: str): + """(prefix, core_version_tuple, suffix) — first maximal dotted-numeric run of len>=2 is the core.""" + m = VERSION_CORE_RE.search(tag) + if not m: + return None + prefix = tag[:m.start()] + core = tuple(int(x) for x in m.group(1).split(".")) + suffix = tag[m.end():] + return prefix, core, suffix + + +def is_unstable(tag: str) -> bool: + low = tag.lower() + return any(mk in low for mk in UNSTABLE_MARKERS) + + +def newest_matching(all_tags, current_tag): + """Find the newest tag sharing the current tag's prefix/suffix 'shape', by core version tuple.""" + cur = split_tag(current_tag) + if not cur: + return None, "current tag has no parseable version core" + cur_prefix, cur_core, cur_suffix = cur + + # suffix "shape": if suffix contains a long hex-looking token, treat it as a wildcard hash + def suffix_shape(suf): + if re.search(r'[0-9a-f]{7,40}', suf.lower()) and re.search(r'[a-f]', suf.lower()): + return re.sub(r'[0-9a-f]{7,40}', '', suf.lower()) + return suf + + cur_shape = suffix_shape(cur_suffix) + + best = None + best_core = None + for t in all_tags: + if t == current_tag: + continue + if is_unstable(t): + continue + parsed = split_tag(t) + if not parsed: + continue + p, core, suf = parsed + if p != cur_prefix: + continue + if suffix_shape(suf) != cur_shape: + continue + if len(core) != len(cur_core): + continue + # Some registries (linuxserver.io/lscr.io in particular) publish EXTRA tags for a + # same-app-version base-image rebuild, shaped like the release tag plus a YYYYMMDD-ish + # trailing component (e.g. bookstack 26.05.2 coexists with 26.05.20260608). That is not a + # newer application release — it is a rebuild of an existing one — so any component that + # looks like a calendar date (>= 20000, comfortably above any real release counter we saw + # across all 58 pinned-semver images, all of which stayed under 10000) is excluded as a + # candidate "newest" rather than misread as a huge version jump. + if any(c >= 20000 for c in core): + continue + if best_core is None or core > best_core: + best_core = core + best = t + if best is None: + return None, "no matching newer tag found (same prefix/suffix shape)" + if best_core <= cur_core: + return None, "up to date (no tag newer than current within the same shape)" + return (best, best_core, cur_core), None + + +def main(): + pins = json.load(open(SCRATCH / "pins.json")) + results = {} + for i, (ref, p) in enumerate(sorted(pins.items())): + repo = p["repo"] + tag = p["tag"] + host, hrepo = resolve_host_repo(repo) + entry = {"ref": ref, "repo": repo, "tag": tag, "kind": p["kind"], "host": host, "hrepo": hrepo} + if repo.startswith("gitea.dooplex.hu"): + entry["status"] = "internal-not-upstream" + results[ref] = entry + print(f"[{i+1}/{len(pins)}] SKIP internal: {ref}") + continue + try: + if p["kind"] == "floating": + dig, err = manifest_digest(host, hrepo, tag) + if err: + entry["status"] = "error" + entry["error"] = err + else: + entry["status"] = "ok" + entry["current_digest"] = dig + print(f"[{i+1}/{len(pins)}] FLOAT {ref} -> {entry.get('current_digest') or entry.get('error')}") + else: + tags = tags_list_all(host, hrepo) + entry["n_upstream_tags"] = len(tags) + res, err = newest_matching(tags, tag) + if err: + entry["status"] = "no-newer-or-unparsed" + entry["detail"] = err + else: + best, best_core, cur_core = res + entry["status"] = "behind" + entry["newest_tag"] = best + entry["newest_core"] = list(best_core) + entry["current_core"] = list(cur_core) + entry["major_step"] = best_core[0] != cur_core[0] + print(f"[{i+1}/{len(pins)}] SEMVER {ref} ({len(tags)} tags) -> {entry.get('newest_tag') or entry.get('detail')}") + except requests.HTTPError as e: + entry["status"] = "error" + entry["error"] = f"HTTP error: {e}" + print(f"[{i+1}/{len(pins)}] ERROR {ref}: {e}") + except Exception as e: + entry["status"] = "error" + entry["error"] = f"{type(e).__name__}: {e}" + print(f"[{i+1}/{len(pins)}] ERROR {ref}: {e}") + results[ref] = entry + time.sleep(0.15) + + json.dump(results, open(SCRATCH / "results.json", "w"), indent=2) + print("\nwrote", SCRATCH / "results.json") + + +if __name__ == "__main__": + main() diff --git a/documentation/audits/update-arc-2026-09-21/00-upstream-drift-raw.txt b/documentation/audits/update-arc-2026-09-21/00-upstream-drift-raw.txt new file mode 100644 index 00000000..d3a7279e --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/00-upstream-drift-raw.txt @@ -0,0 +1,192 @@ +upstream image drift audit — app-catalog-felhom.eu — 2026-09-21 +====================================================================== + +Read-only measurement, run from DooPlex against public registries only +(Docker Hub, ghcr.io, lscr.io, registry.gitlab.com). Nothing was queried +from a customer/demo box. The catalog repo was not modified and nothing +was committed. + +PRE-EXISTING REPO STATE NOTE: at the start of this session `git status +--porcelain` in app-catalog-felhom.eu already showed 3 modified files +(CLAUDE.md, scripts/check-engine-major.py, scripts/test_gate_decoys.py) — +an in-progress, uncommitted R-469 change (MariaDB major-upgrade gate) +that predates this session. This audit did not touch those files and +made no edits/commits anywhere in the repo; flagging only because the +task said 'do not modify' and the tree was already not clean. + +1. INVENTORY +---------------------------------------------------------------------- +Apps (templates/*/docker-compose.yml): 53 +Total `image:` lines across all composes: 79 +Unique image refs (pins): 66 + + floating (DB-engine family tag, e.g. postgres:16-alpine, mariadb:11.6, + redis:7-alpine, postgis:16-3.5-alpine): 8 + pinned-semver (full app release tag): 58 + +App classification (from compose + .felhom.yml, not memory): + database = has a postgres/mariadb/mysql/mongo sidecar image + file-leg = binds host media/library paths (HDD_PATH / USERDATA_PATH), + no DB sidecar + other = neither + database: 14 apps -> bookstack, calcom, claper, docmost, immich, kimai, nextcloud, outline, paperless-ngx, rallly, romm, sparkyfitness, tandoor, zipline + file-leg: 9 apps -> audiobookshelf, calibre-web, emby, jellyfin, komga, navidrome, plex, radarr, sonarr + other: 30 apps -> actualbudget, adventurelog, bentopdf, code-server, crafty-controller, ghost, gitea, glance, gokapi, grafana, gramps-web, home-assistant, homebox, homepage, mealie, n8n, onlyoffice, opengist, papra, plant-it, privatebin, recipe-importer, seerr, termix, uptime-kuma, vaultwarden, vikunja, wanderer, wger, wishlist + +EDGE CASE flagged, not silently folded in: `adventurelog` runs a +postgis/postgis (postgres-derivative) DB sidecar, but the string +'postgis' does not contain 'postgres'/'mariadb'/'mysql'/'mongo', so by +the literal substring rule it lands in 'other', not 'database'. If you +want postgis apps folded into 'database', adventurelog is the one app +that moves (other: 30->29, database: 14->15). + +2. DEAD-IMAGE CHECK (does the pin still resolve at all) +---------------------------------------------------------------------- +Ran the repo's own gate, unmodified, exactly as documented in CLAUDE.md: + $ python3 scripts/check-image-resolvable.py --all +Result: 65 of 66 pins verified via `docker manifest inspect`, ALL resolve. +NO dead images among those checked. + +The 1 unchecked: msdeluise/plant-it:0.10.0 (app `plant-it`, +lifecycle: abandoned). The gate calls it INCONCLUSIVE (rc=0 but +stderr contains 'denied'/'unauthorized', which its own trap-1 logic +refuses to read as success). Independently: hub.docker.com's own +catalog API returns {"message":"object not found"} for the whole +msdeluise/plant-it repository (not just the tag), and the upstream +GitHub repo msdeluise/plant-it now 404s too — two positive signals the +image is gone, stronger than the registry gate alone captures, but I'm +stating both rather than asserting 'dead' from a registry API that +itself only said INCONCLUSIVE. This matches its already-documented +`lifecycle: abandoned` state — expected, not a new incident. + +3. PINNED-SEMVER: IS THERE A NEWER TAG UPSTREAM +---------------------------------------------------------------------- +pinned-semver pins evaluated: 58 + behind: 46 + no-newer-or-unparsed: 10 + error: 1 + internal-not-upstream: 1 + +Of the 46 pins with a newer upstream tag: 7 cross a MAJOR version, +39 are MINOR/PATCH steps (same leading version number). + +NOTE ON THE MAJOR/MINOR RULE: 'major' = the first dot-separated numeric +component changed. This is literal, not semantic — it correctly reads a +real major bump for apps like nextcloud/claper/paperless-ngx, but for +calendar-versioned images (home-assistant: YYYY.MM.PATCH, +actualbudget: YY.MM.PATCH) the 'first number' is a year, so a same-year +update is correctly minor/patch and would only misread as 'major' at a +year rollover — noted, not a defect found in this sweep's data. + +4. APPS BEHIND ON A PINNED-SEMVER IMAGE, sorted furthest-behind first +---------------------------------------------------------------------- +STEP IMAGE REPO CURRENT NEWEST APP(S) PINNED SINCE +MAJOR nextcloud 34.0.1-apache 35.0.0-apache nextcloud 2026-07-18 +MAJOR f0rc3/gokapi v1.9.6 v2.2.4 gokapi 2026-07-18 +MAJOR ghcr.io/gethomepage/homepage v1.13.2 v2.4.0 homepage 2026-07-18 +MAJOR codewithcj/sparkyfitness_server v0.17.3 v1.7.2 sparkyfitness 2026-07-18 +MAJOR codewithcj/sparkyfitness v0.17.3 v1.7.2 sparkyfitness 2026-07-18 +MAJOR ghcr.io/paperless-ngx/paperless-ngx 2.20.15 3.2.1 paperless-ngx 2026-07-18 +MAJOR ghcr.io/claperco/claper 2.5 3.0 claper 2026-07-18 +minor/patch plexinc/pms-docker 1.41.4.9463-630c9f557 -> 1.43.4.10903-e5521bd8c plex 2026-02-15 +minor/patch emby/embyserver 4.10.0.20 4.11.0.1 emby 2026-07-18 +minor/patch getmeili/meilisearch v1.36.0 v1.54.0 wanderer 2026-07-21 +minor/patch ghost 6.53.0-alpine 6.64.0-alpine ghost 2026-07-18 +minor/patch n8nio/n8n 2.31.3 2.40.4 n8n 2026-07-18 +minor/patch lscr.io/linuxserver/code-server 4.129.0 4.138.0 code-server 2026-07-18 +minor/patch ghcr.io/mealie-recipes/mealie v3.20.1 v3.27.0 mealie 2026-07-18 +minor/patch lukevella/rallly 4.11.1 4.15.1 rallly 2026-07-18 +minor/patch ghcr.io/lukegus/termix 2.5.0 2.8.0 termix 2026-07-12 +minor/patch vikunja/vikunja 2.3.0 2.6.0 vikunja 2026-07-18 +minor/patch ghcr.io/home-assistant/home-assistant 2026.7.2 2026.9.3 home-assistant 2026-07-18 +minor/patch actualbudget/actual-server 26.7.0 26.9.0 actualbudget 2026-07-18 +minor/patch gotson/komga 1.25.0 1.27.0 komga 2026-07-18 +minor/patch rommapp/romm 5.0.0 5.2.0 romm 2026-07-18 +minor/patch ghcr.io/immich-app/immich-server v3.0.3 v3.2.2 immich 2026-07-18 +minor/patch ghcr.io/immich-app/immich-machine-learning v3.0.3 v3.2.2 immich 2026-07-18 +minor/patch louislam/uptime-kuma 2.4.0 2.5.5 uptime-kuma 2026-07-18 +minor/patch lscr.io/linuxserver/radarr 6.3.0 6.4.4 radarr 2026-07-18 +minor/patch vaultwarden/server 1.36.0-alpine 1.37.3-alpine vaultwarden 2026-07-18 +minor/patch grafana/grafana 13.1.0 13.2.2 grafana 2026-07-18 +minor/patch ghcr.io/advplyr/audiobookshelf 2.35.1 2.36.1 audiobookshelf 2026-07-18 +minor/patch docmost/docmost 0.95.0 0.96.0 docmost 2026-07-18 +minor/patch outlinewiki/outline 1.9.1 1.10.1 outline 2026-07-18 +minor/patch ghcr.io/cmintey/wishlist v0.66.0 v0.67.0 wishlist 2026-07-19 +minor/patch ghcr.io/seanmorley15/adventurelog-backend v0.12.1 v0.13.0 adventurelog 2026-07-18 +minor/patch ghcr.io/seanmorley15/adventurelog-frontend v0.12.1 v0.13.0 adventurelog 2026-07-18 +minor/patch ghcr.io/diced/zipline 4.6.1 4.7.0 zipline 2026-07-18 +minor/patch deluan/navidrome 0.63.2 0.64.0 navidrome 2026-07-18 +minor/patch registry.gitlab.com/crafty-controller/crafty-4 4.10.7 4.11.0 crafty-controller 2026-06-26 +minor/patch lscr.io/linuxserver/bookstack 26.05.2 26.05.5 bookstack 2026-07-18 +minor/patch gitea/gitea 1.27.0 1.27.3 gitea 2026-07-18 +minor/patch ghcr.io/alam00000/bentopdf v2.8.6 v2.8.8 bentopdf 2026-07-12 +minor/patch ghcr.io/thomiceli/opengist 1.13 1.15 opengist 2026-07-18 +minor/patch ghcr.io/tandoorrecipes/recipes 2.6.13 2.6.15 tandoor 2026-07-18 +minor/patch glanceapp/glance v0.8.5 v0.8.6 glance 2026-07-18 +minor/patch ghcr.io/papra-hq/papra 26.6.1-rootless 26.6.2-rootless papra 2026-07-12 +minor/patch privatebin/pdo 2.0.5 2.0.6 privatebin 2026-09-14 +minor/patch lscr.io/linuxserver/sonarr 4.0.19 4.0.20 sonarr 2026-07-18 +minor/patch wger/server 2.6 2.7 wger 2026-07-19 + +5. BY APP CLASS +---------------------------------------------------------------------- +database (14 apps, 23 unique pins): + floating: 7 pinned-semver: 16 + of pinned-semver -> behind: 14 (major: 5, minor/patch: 9), up to date: 2, other: 0 + +file-leg (9 apps, 9 unique pins): + floating: 0 pinned-semver: 9 + of pinned-semver -> behind: 7 (major: 0, minor/patch: 7), up to date: 2, other: 0 + +other (30 apps, 34 unique pins): + floating: 1 pinned-semver: 33 + of pinned-semver -> behind: 25 (major: 2, minor/patch: 23), up to date: 6, other: 2 + +6. FLOATING PINS — has the tag's digest moved since the catalog set it +---------------------------------------------------------------------- +Method: compare each floating tag's registry-reported last-push date +(Docker Hub API tag metadata; ghcr.io has no equivalent public +timestamp API reachable anonymously) against the app's own +`catalog_since` (the date THIS repo last moved that pin — CLAUDE.md's +own definition of 'when we set it'). + + mariadb:11.4 apps=romm catalog_since=2026-07-18 registry last push=2026-09-18T02:36:29.445608Z + mariadb:11.6 apps=kimai, nextcloud catalog_since=2026-07-18 registry last push=2025-02-04T13:35:57.325706Z + mariadb:12.3 apps=bookstack catalog_since=2026-07-18 registry last push=2026-09-18T02:37:26.402424Z + postgres:15-alpine apps=sparkyfitness catalog_since=2026-07-18 registry last push=2026-09-21T07:07:40.622579Z + postgres:16-alpine apps=8 apps (see below) catalog_since=2026-07-18 registry last push=2026-09-21T04:07:46.010722Z + redis:7-alpine apps=6 apps (see below) catalog_since=2026-07-18 registry last push=2026-09-18T03:04:58.155237Z + postgis/postgis:16-3.5-alpine apps=adventurelog catalog_since=2026-07-18 registry last push=2026-08-31T11:37:24.0266Z + +Verdict (comparing dates above): 6 of 7 Docker-Hub-hosted floating pins +have had their tag repushed (new digest) at least once since the catalog +set them (mariadb:11.4, mariadb:12.3, postgres:15-alpine, +postgres:16-alpine, redis:7-alpine, postgis:16-3.5-alpine — pushes on or +after 2026-08-31, vs. catalog_since 2026-07-18 for all of them). Only +mariadb:11.6 (kimai, nextcloud) has NOT been repushed since pinning — +its last Docker Hub push was 2025-02-04, well before catalog_since. + +ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 (immich): +COULD NOT TELL. ghcr.io's anonymous v2 API exposes no last-modified +timestamp, and I did not have an authenticated GitHub Packages token to +query the versions API. Digest resolved today: + sha256:1a078b237c1d9b420b0ee59147386b4aa60d3a07a8e6a402fc84a57e41b043a4 +but I have no earlier digest recorded anywhere to compare it to, so +'has it moved since catalog_since' is unmeasured for this one pin — +stated as unmeasured rather than guessed. + +7. REPRODUCIBLE COMMANDS +---------------------------------------------------------------------- +Dead-image check (the repo's own gate, unmodified): + cd app-catalog-felhom.eu && python3 scripts/check-image-resolvable.py --all + +Newer-tag-upstream check (this audit's script, not part of the repo, +kept in the scratchpad — reuses the registry-v2 anonymous-bearer-token +auth flow from scripts/check-image-resolvable.py, but lists the full +upstream tag set per repo instead of testing one ref): + python3 /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/5f83c0fb-b1bf-404e-b99e-2e3b28a31615/scratchpad/drift.py + (writes results.json next to it; apps.json/pins.json are the inventory + extracted from the catalog's compose/.felhom.yml files) + +Floating-tag last-push check (Docker Hub API only, no docker CLI): + curl -s https://hub.docker.com/v2/repositories///tags/ diff --git a/documentation/audits/update-arc-2026-09-21/01-baseline.txt b/documentation/audits/update-arc-2026-09-21/01-baseline.txt new file mode 100644 index 00000000..97a7d864 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/01-baseline.txt @@ -0,0 +1,20 @@ +=== 9202 stacks: deployed==true, before the drill (2026-09-21) === +total templates: 55 + traefik deployed= False running= None protected= True orphaned= False + +=== catalog cache pin on the box for uptime-kuma === +11: image: louislam/uptime-kuma:2.4.0 +11:# catalog_since: the date THIS repo last changed this app's pinned images. Any commit that +13:catalog_since: "2026-07-18" + +=== DooPlex catalog repo HEAD === +5ff36d098cbcca8a025bb7852e20c10f9ecec691 +(clean) +repo pin: image: louislam/uptime-kuma:2.4.0 +repo catalog_since: # catalog_since: the date THIS repo last changed this app's pinned images. Any commit that +catalog_since: "2026-07-18" + +=== controller version === +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 Up 4 minutes (healthy) +filebrowser gtstef/filebrowser:1.3.3-stable Up 7 days (healthy) +traefik traefik:v3.6.7 Up 7 days diff --git a/documentation/audits/update-arc-2026-09-21/02-deploy-post.txt b/documentation/audits/update-arc-2026-09-21/02-deploy-post.txt new file mode 100644 index 00000000..5d047407 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/02-deploy-post.txt @@ -0,0 +1,3 @@ +{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} + +HTTPCODE:202 diff --git a/documentation/audits/update-arc-2026-09-21/03-app-state-after-deploy.txt b/documentation/audits/update-arc-2026-09-21/03-app-state-after-deploy.txt new file mode 100644 index 00000000..61ae3371 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/03-app-state-after-deploy.txt @@ -0,0 +1,136 @@ +{ + "ok": true, + "data": { + "name": "uptime-kuma", + "meta": { + "display_name": "Uptime Kuma", + "description": "Szolg\u00e1ltat\u00e1s \u00e9s weboldal monitoring", + "category": "dashboard", + "subdomain": "status", + "slug": "uptime-kuma", + "catalog_since": "2026-07-18", + "resources": { + "mem_request": "50M", + "mem_limit": "256M", + "pi_compatible": true, + "needs_hdd": false, + "hungarian_ui": false + }, + "deploy_fields": [ + { + "env_var": "DOMAIN", + "label": "Domain", + "type": "domain", + "generate": "", + "default": "", + "required": true, + "placeholder": "", + "description": "A szerver domain neve", + "locked_after_deploy": true + }, + { + "env_var": "SUBDOMAIN", + "label": "Aldomain", + "type": "subdomain", + "generate": "", + "default": "status", + "required": true, + "placeholder": "", + "description": "Az alkalmaz\u00e1s aldomainje", + "locked_after_deploy": true + } + ], + "app_info": { + "tagline": "Szolg\u00e1ltat\u00e1s monitoring - \u00e9rtes\u00edt\u00e9s ha valami nem m\u0171k\u00f6dik", + "use_cases": [ + "Weboldalak \u00e9s szolg\u00e1ltat\u00e1sok el\u00e9rhet\u0151s\u00e9g\u00e9nek figyel\u00e9se", + "HTTP, TCP, DNS, ping \u00e9s egy\u00e9b protokollok t\u00e1mogat\u00e1sa", + "\u00c9rtes\u00edt\u00e9sek emailben, Telegramon, Discord-on, stb.", + "Nyilv\u00e1nos st\u00e1tusz oldal a felhaszn\u00e1l\u00f3knak", + "Sz\u00e9p grafikonok a rendelkez\u00e9sre \u00e1ll\u00e1sr\u00f3l" + ], + "first_steps": [ + "Nyisd meg a status.DOMAIN c\u00edmet a b\u00f6ng\u00e9sz\u0151ben", + "Hozd l\u00e9tre az admin fi\u00f3kot az els\u0151 megnyit\u00e1skor", + "Add hozz\u00e1 az els\u0151 monitort (pl. a saj\u00e1t weboldalad)", + "\u00c1ll\u00edts be \u00e9rtes\u00edt\u00e9seket (email, Telegram, stb.)", + "Opcion\u00e1lisan: hozz l\u00e9tre egy nyilv\u00e1nos st\u00e1tusz oldalt" + ], + "prerequisites": null, + "default_creds": "", + "docs_url": "https://github.com/louislam/uptime-kuma/wiki" + }, + "optional_config": null, + "healthcheck": { + "interval": "5m", + "checks": [ + { + "type": "http", + "port": 3001, + "path": "/", + "method": "" + } + ] + } + }, + "compose_path": "/opt/docker/stacks/uptime-kuma/docker-compose.yml", + "state": "running", + "deployed": true, + "protected": false, + "orphaned": false, + "containers": [ + { + "name": "uptime-kuma", + "image": "louislam/uptime-kuma:2.4.0", + "state": "running", + "status": "Up About a minute (healthy)" + } + ], + "app_config": { + "deployed": true, + "deployed_at": "2026-09-21T10:57:12Z", + "env": { + "DOMAIN": "enkisfelhom.hu", + "SUBDOMAIN": "status" + }, + "locked_fields": [ + "DOMAIN", + "SUBDOMAIN" + ], + "desired_state": "running", + "installed_images": { + "uptime-kuma": { + "ref": "louislam/uptime-kuma:2.4.0", + "digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985", + "at": "2026-09-21T10:57:45Z" + } + }, + "pinned_images": { + "uptime-kuma": "louislam/uptime-kuma:2.4.0" + } + }, + "deploying": false, + "updating": false, + "health_probe": { + "healthy": true, + "last_check": "2026-09-21T10:58:06.444077946Z", + "details": [ + { + "type": "http", + "target": ":3001/", + "healthy": true, + "status": 302, + "latency": "3ms" + } + ] + }, + "last_updated": "2026-09-21T10:59:16.464548662Z", + "restarting_since": "0001-01-01T00:00:00Z", + "template_images": { + "uptime-kuma": "louislam/uptime-kuma:2.4.0" + }, + "catalog_images": { + "uptime-kuma": "louislam/uptime-kuma:2.4.0" + } + } +} diff --git a/documentation/audits/update-arc-2026-09-21/04-baseline-badge.txt b/documentation/audits/update-arc-2026-09-21/04-baseline-badge.txt new file mode 100644 index 00000000..0f637736 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/04-baseline-badge.txt @@ -0,0 +1,44 @@ +### STEP 2 — BASELINE BADGE (installed 2.4.0 == catalog 2.4.0) + +--- app page /apps/uptime-kuma --- +Naprakész + [raw badge markup, grep 'tag-ok\|tag-warn' context] +ugrep: error: error at position 84 +xbf][\x80-\xbf]*){0,200} + \___exceeds complexity limits + + + +--- apps list /stacks (uptime-kuma row) --- + +--- app page /apps/uptime-kuma?lang=en --- +Up to date + [raw badge markup, grep 'tag-ok\|tag-warn' context] +ugrep: error: error at position 84 +xbf][\x80-\xbf]*){0,200} + \___exceeds complexity limits + + + +--- apps list /stacks?lang=en (uptime-kuma row) --- + +### STEP 2b — apps list page /stacks, both languages +--- stacks-hu.html (133504 bytes) --- + badges present: + 51 class="tag tag-off"> + 1 class="tag tag-ok" title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja.">Naprakész + 3 class="tag tag-run"> + uptime-kuma mentioned: 9 times +--- stacks-en.html (132005 bytes) --- + badges present: + 51 class="tag tag-off"> + 1 class="tag tag-ok" title="This app is running the newest version available.">Up to date + 3 class="tag tag-run"> + uptime-kuma mentioned: 9 times + +### SEARCH CONTROLS (ASCII-only fragments, R-rule) +POSITIVE control 'Naprak' in stacks-hu.html : 1 (expect >0) +POSITIVE control 'Up to date' in stacks-en.html: 1 (expect >0) +NEGATIVE control 'Frissites elerheto' (deaccented, never in output): 0 (expect 0) +NEGATIVE control 'tag-warn' in stacks-hu.html at baseline: 0 (expect 0 for uptime-kuma) +NEGATIVE control 'ZZZ-not-present': 0 (expect 0) diff --git a/documentation/audits/update-arc-2026-09-21/05-catalog-bump-push-gates.txt b/documentation/audits/update-arc-2026-09-21/05-catalog-bump-push-gates.txt new file mode 100644 index 00000000..b5731d9c --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/05-catalog-bump-push-gates.txt @@ -0,0 +1,72 @@ +============================================================================== +== gate: image-pins (check-image-pins.py) +============================================================================== +image-pin gate OK — 53 templates, 0 unpinned images + +============================================================================== +== gate: engine-major (check-engine-major.py --range=bd22749e501540cfe2526408613832fef29f8de7..89304ab89617e4467726f055901d0d0ee83531b7) +============================================================================== +engine-major gate — range bd22749e501540cfe2526408613832fef29f8de7..89304ab89617e4467726f055901d0d0ee83531b7: 1 compose file(s) changed, 0 engine pin(s) compared +engine-major gate OK — no forbidden engine major, and no engine major bundled with another image move (rule: CLAUDE.md, amended by R-469 2026-09-21) + +============================================================================== +== gate: catalog-since (check-catalog-since.py --range=bd22749e501540cfe2526408613832fef29f8de7..89304ab89617e4467726f055901d0d0ee83531b7) +============================================================================== +catalog-since gate — range bd22749e501540cfe2526408613832fef29f8de7..89304ab89617e4467726f055901d0d0ee83531b7: 1 compose file(s) changed, 1 image move(s) dated +catalog-since gate OK — every image move in the range carries a catalog_since on or after its commit day + +============================================================================== +== gate: copy-i18n (check-copy-i18n.py) +============================================================================== +copy-i18n: matcher controls OK (positive „Jelentkezz be…" convicts, negative „Encrypted notes…" does not) +copy-i18n: retrieval-promise vocabulary matches the sibling clone +copy-i18n: 53 apps · 1032 Hungarian copy strings frozen · 1031 with English +copy-i18n: translated so far — actualbudget 15/15, adventurelog 17/17, audiobookshelf 21/21, bentopdf 14/14, bookstack 19/19, calcom 21/21, calibre-web 24/24, claper 17/17, code-server 20/20, crafty-controller 24/24, docmost 19/19, emby 22/22, ghost 16/16, gitea 16/16, glance 15/15, gokapi 19/19, grafana 19/19, gramps-web 16/16, home-assistant 18/18, homebox 16/16, homepage 15/15, immich 24/24, jellyfin 22/22, kimai 20/20, komga 21/21, mealie 17/17, n8n 17/17, navidrome 21/21, nextcloud 28/28, onlyoffice 24/24, opengist 15/15, outline 21/21, paperless-ngx 33/33, papra 13/14, plant-it 16/16, plex 25/25, privatebin 14/14, radarr 22/22, rallly 17/17, recipe-importer 19/19, romm 42/42, seerr 18/18, sonarr 22/22, sparkyfitness 20/20, tandoor 18/18, termix 11/11, uptime-kuma 16/16, vaultwarden 23/23, vikunja 17/17, wanderer 19/19, wger 18/18, wishlist 15/15, zipline 20/20 +copy-i18n: OK + +============================================================================== +== summary +============================================================================== + image-pins OK (exit 0) + engine-major OK (exit 0) + catalog-since OK (exit 0) + copy-i18n OK (exit 0) + +all catalog gates OK +pre-push [app-catalog-felhom.eu]: gates OK - push proceeding. +remote: . Processing 1 references +remote: Processed 1 references in total +To https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + bd22749..89304ab main -> main +PUSH_RC=0 + +[exited with code 0] + +=== commit === +89304ab DRILL: uptime-kuma 2.4.0 -> 2.5.0 for the update-arc measurement (reverted in this session) +bd22749 REPORT for the R-469 rule lift +=== the bump diff === +commit 89304ab89617e4467726f055901d0d0ee83531b7 +Author: kisfenyo +Date: Mon Sep 21 13:03:42 2026 +0200 + + DRILL: uptime-kuma 2.4.0 -> 2.5.0 for the update-arc measurement (reverted in this session) + + A temporary image move so a live box can be walked through the update arc on demo-hp LXC 9202 + (audit update-arc-2026-09-21): the badge going to "Frissites elerheto - ma", a power cut mid-pull + (R-520), then the revert that leaves the box AHEAD of the catalog (R-524). + + 2.5.0 is the next REAL released upstream tag: 2.4.1 and 2.4.2 do not exist on Docker Hub + (docker manifest inspect, all three checked). catalog_since moves to today as the catalog-since + gate requires of any commit that moves an image: line. + + Co-Authored-By: Claude Opus 5 + Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS + + templates/uptime-kuma/.felhom.yml | 2 +- + templates/uptime-kuma/docker-compose.yml | 2 +- + 2 files changed, 2 insertions(+), 2 deletions(-) +-catalog_since: "2026-07-18" ++catalog_since: "2026-09-21" +- image: louislam/uptime-kuma:2.4.0 ++ image: louislam/uptime-kuma:2.5.0 diff --git a/documentation/audits/update-arc-2026-09-21/06-sync-after-bump.txt b/documentation/audits/update-arc-2026-09-21/06-sync-after-bump.txt new file mode 100644 index 00000000..119918fb --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/06-sync-after-bump.txt @@ -0,0 +1,16 @@ +### STEP 3 — force catalog sync on 9202 +-- POST /api/sync -- +{"ok":true,"data":{"ok":true,"updated":["uptime-kuma"],"message":"Sablonok frissítve — frissítve: uptime-kuma"},"message":"Sablonok frissítve — frissítve: uptime-kuma"} + +HTTPCODE:200 + +-- box catalog cache AFTER sync -- +11: image: louislam/uptime-kuma:2.5.0 +13:catalog_since: "2026-09-21" + +-- what the controller now reports as catalog_images vs installed_images -- +catalog_images : {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} +template_images : {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'} +installed_images: {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'} +pinned_images : {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'} +meta.catalog_since: 2026-09-21 diff --git a/documentation/audits/update-arc-2026-09-21/07-badge-after-bump.txt b/documentation/audits/update-arc-2026-09-21/07-badge-after-bump.txt new file mode 100644 index 00000000..df76581b --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/07-badge-after-bump.txt @@ -0,0 +1,22 @@ +### STEP 4 — BADGE AFTER THE BUMP (installed 2.4.0, catalog 2.5.0, catalog_since = today) + +--- app page /apps/uptime-kuma [hu] --- +Frissítés elérhető — ma +--- apps list /stacks [hu] --- + 51 class="tag tag-off"> + 3 class="tag tag-run"> + 1 class="tag tag-warn" title="Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.">Frissítés elérhető — ma + +--- app page /apps/uptime-kuma [en] --- +Update available — today +--- apps list /stacks [en] --- + 51 class="tag tag-off"> + 3 class="tag tag-run"> + 1 class="tag tag-warn" title="A newer version of this app is available. Select the Update button to start it.">Update available — today + +### SEARCH CONTROLS +POSITIVE 'tag-warn' in app-behind-hu.html : 1 (expect >0) +POSITIVE 'tag-warn' in app-behind-en.html : 1 (expect >0) +NEGATIVE 'Naprak' in app-behind-hu.html : 0 (expect 0 — it moved off 'up to date') +NEGATIVE 'Up to date' in app-behind-en.html: 0 (expect 0) +NEGATIVE 'ZZZ-not-present' : 0 (expect 0) diff --git a/documentation/audits/update-arc-2026-09-21/08-precut-state.txt b/documentation/audits/update-arc-2026-09-21/08-precut-state.txt new file mode 100644 index 00000000..dd247de9 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/08-precut-state.txt @@ -0,0 +1,23 @@ +### STEP 5 PRE-CUT STATE (host clock: Mon Sep 21 11:04:42 UTC 2026) +-- update-journal.json present before the update? -- +ABSENT +-- app.yaml before the update -- +# Auto-generated by felhom-controller — do not edit locked fields manually +deployed: true +deployed_at: "2026-09-21T10:57:12Z" +env: + DOMAIN: enkisfelhom.hu + SUBDOMAIN: status +locked_fields: + - DOMAIN + - SUBDOMAIN +desired_state: running +installed_images: + uptime-kuma: + ref: louislam/uptime-kuma:2.4.0 + digest: sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985 + at: "2026-09-21T10:57:45Z" +pinned_images: + uptime-kuma: louislam/uptime-kuma:2.4.0 +-- live compose image line before the update -- + image: louislam/uptime-kuma:2.4.0 diff --git a/documentation/audits/update-arc-2026-09-21/09-the-cut.txt b/documentation/audits/update-arc-2026-09-21/09-the-cut.txt new file mode 100644 index 00000000..e766c94e --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/09-the-cut.txt @@ -0,0 +1,79 @@ +### STEP 5 — THE CUT (poller observations, host UTC) +11:04:49.596Z poller armed, waiting for update_phase=pulling +11:04:50.561Z "updating":false +11:04:52.501Z "updating":false +11:04:54.436Z "updating":true "update_phase":"pulling" +11:04:54.439Z >>> CUTTING POWER: pct stop 9202 +11:04:58.207Z >>> pct stop returned rc=0 ; status: status: stopped + +NOTE: pct stop is not instantaneous — it returned 3.8 s after the decision, so the guest + had up to ~3.8 s more of life inside the pulling phase. The phase at the moment of + the decision is OBSERVED; the phase at the moment the kernel actually froze is INFERRED. + +### POST-CRASH ON-DISK STATE, read from the STOPPED guest's rootfs (pct mount, no boot yet) +-- update-journal.json -- +ABSENT + +-- app.yaml as the crash left it -- +# Auto-generated by felhom-controller — do not edit locked fields manually +deployed: true +deployed_at: "2026-09-21T10:57:12Z" +env: + DOMAIN: enkisfelhom.hu + SUBDOMAIN: status +locked_fields: + - DOMAIN + - SUBDOMAIN +desired_state: running +installed_images: + uptime-kuma: + ref: louislam/uptime-kuma:2.4.0 + digest: sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985 + at: "2026-09-21T10:57:45Z" +pinned_images: + uptime-kuma: louislam/uptime-kuma:2.5.0 + +-- live docker-compose.yml image line as the crash left it -- + image: louislam/uptime-kuma:2.5.0 + +-- pre-update compose copy, if one was saved -- +total 32 +drwxr-xr-x 2 100000 100000 4096 Sep 21 13:04 . +drwxr-xr-x 57 100000 100000 4096 Sep 13 22:22 .. +-rw-r--r-- 1 100000 100000 3318 Sep 21 13:04 .felhom.yml +-rw------- 1 100000 100000 506 Sep 21 13:04 app.yaml +-rw-r--r-- 1 100000 100000 1597 Sep 21 13:04 applied-compose.yml +-rw-r--r-- 1 100000 100000 1597 Sep 21 13:04 docker-compose.yml +-rw-r--r-- 1 100000 100000 1597 Sep 21 13:04 pre-update-applied.yml +-rw-r--r-- 1 100000 100000 1597 Sep 21 13:04 pre-update-compose.yml +(unmounted) + +### CORRECTION — the journal read above was WRONG. + /var/lib/docker inside the guest is a separate MOUNT (not a symlink — my first wording was + wrong too): mp0 mounts a 70G disk at /var/lib/felhom and its /docker subdirectory is bound + onto /var/lib/docker (findmnt: /dev/loop1[/docker] /var/lib/docker ext4). + 'pct mount 9202' maps only the ROOTFS and does not apply the guest's internal mounts, so + /var/lib/docker is an EMPTY MOUNTPOINT STUB from the host. 'test -f' therefore + answered ABSENT for a file that exists — a FALSE NEGATIVE, not evidence. + From the host the bytes are at /var/lib/felhom/docker/volumes/.../_data/data/. + Found by running 'find' over the whole rootfs instead of trusting the one path. + Path provenance and positive controls: 12-journal-path-provenance.txt. + +### POST-CRASH update-journal.json, read from the STOPPED guest (correct path, still no boot) +-rw------- 1 100000 100000 439 Sep 21 13:04 /var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/update-journal.json +--- contents --- +{ + "updates": { + "uptime-kuma": { + "phase": "pulling", + "started_at": "2026-09-21T11:04:52.924206274Z", + "prev_pin": { + "uptime-kuma": "louislam/uptime-kuma:2.4.0" + }, + "prev_compose": "/opt/docker/stacks/uptime-kuma/pre-update-compose.yml", + "prev_applied": "/opt/docker/stacks/uptime-kuma/pre-update-applied.yml", + "proven_copy_at": "2026-09-21T11:01:46Z", + "proven_tier": 1 + } + } +} diff --git a/documentation/audits/update-arc-2026-09-21/10-controller-log-full-after-cut.txt b/documentation/audits/update-arc-2026-09-21/10-controller-log-full-after-cut.txt new file mode 100644 index 00000000..3030edd4 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/10-controller-log-full-after-cut.txt @@ -0,0 +1,244 @@ +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026/09/21 11:06:12 main.go:334: [INFO] felhom-controller 0.260.0 starting (customer: demo-hp, domain: enkisfelhom.hu) +2026/09/21 11:06:12 settings.go:693: [INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json +2026/09/21 11:06:12 settings.go:1611: [DEBUG] [settings] AutoDiscoverStoragePaths discovered=[] fallback="" existing=1 +2026/09/21 11:06:12 main.go:360: [INFO] Encryption key loaded from /opt/docker/felhom-controller/data/encryption.key +2026/09/21 11:06:12 manager.go:283: [INFO] [stacks] Using compose command: docker compose +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=false composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=false composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=false composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=false composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=false composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=false composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=true composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/09/21 11:06:12 manager.go:1620: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/09/21 11:06:12 manager.go:630: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/09/21 11:06:12 manager.go:661: [INFO] [stacks] ScanStacks complete: 55 stacks found (1 deployed, 54 available) +2026/09/21 11:06:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:12 deploy.go:1086: [DEBUG] [stacks] InjectMissingFields: checking 1 stacks +2026/09/21 11:06:12 deploy.go:1108: [DEBUG] [stacks] InjectMissingFields: checking stack uptime-kuma — 2 deploy fields, 2 existing env vars +2026/09/21 11:06:12 deploy.go:1171: [INFO] [stacks] InjectMissingFields: processed 1 stacks +2026/09/21 11:06:12 manager.go:422: [DEBUG] [stacks] MigrateEncryption: checking 1 deployed stacks for plaintext sensitive values +2026/09/21 11:06:12 manager.go:466: [INFO] [stacks] Encryption migration: no stacks needed migration +2026/09/21 11:06:12 desiredstate.go:137: [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded (state ambiguous — legacy boot behaviour retained) +2026/09/21 11:06:12 installed.go:525: [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 1 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing) +2026/09/21 11:06:12 pin.go:289: [INFO] [stacks] pin adoption: 0 pinned, 1 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers) +2026/09/21 11:06:12 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s) +2026/09/21 11:06:12 sync.go:208: [INFO] [sync] Starting catalog sync +2026/09/21 11:06:12 sync.go:296: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main) +2026/09/21 11:06:12 sync.go:298: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache +2026/09/21 11:06:12 update.go:904: [WARN] [stacks] update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back +2026/09/21 11:06:12 sync.go:592: [DEBUG] [sync] Running: git fetch --depth 1 origin main +2026/09/21 11:06:12 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/uptime-kuma — 2 env vars, 0 encrypted, 0 sensitive fields +2026/09/21 11:06:12 [INFO] [stacks] SaveAppConfig: saved config for uptime-kuma +2026/09/21 11:06:12 pin.go:93: [INFO] [stacks] pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0 +2026/09/21 11:06:12 update.go:705: [INFO] [stacks] update uptime-kuma: pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] CPUCollector.Start: sampleRate=5s +2026/09/21 11:06:12 main.go:511: [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db +2026/09/21 11:06:12 main.go:523: [INFO] Metrics collector started (60s interval) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:06:12 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) +2026/09/21 11:06:12 notifier.go:75: [INFO] Notifier disabled (hub not configured) +2026/09/21 11:06:12 main.go:642: [INFO] [mailrelay] app-email shim unavailable (no hub configured or disabled in config) +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: status-refresh (every 10s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="status-refresh" interval=10s totalJobs=1 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="stack-scan" interval=2m0s totalJobs=2 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: health-probes (every 10s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="health-probes" interval=10s totalJobs=3 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: system-health (every 5m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="system-health" interval=5m0s totalJobs=4 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="deadapp-check" interval=30s totalJobs=5 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: ring-spill (every 30s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="ring-spill" interval=30s totalJobs=6 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-22 02:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="db-dump" schedule="02:30" nextRun=2026-09-22T02:30:00+02:00 totalJobs=7 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="offsite-credential-retry" interval=5m0s totalJobs=8 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="backup-cache" interval=5m0s totalJobs=9 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-22 03:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="tier2-backup" schedule="03:30" nextRun=2026-09-22T03:30:00+02:00 totalJobs=10 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-22 04:15 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offbox-backup" schedule="04:15" nextRun=2026-09-22T04:15:00+02:00 totalJobs=11 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-22 05:10 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-abandon-sweep" schedule="05:10" nextRun=2026-09-22T05:10:00+02:00 totalJobs=12 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-22 06:00 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-integrity" schedule="06:00" nextRun=2026-09-22T06:00:00+02:00 totalJobs=13 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-22 05:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-proof" schedule="05:30" nextRun=2026-09-22T05:30:00+02:00 totalJobs=14 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-22 04:00 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="metrics-prune" schedule="04:00" nextRun=2026-09-22T04:00:00+02:00 totalJobs=15 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-22 03:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="fill-watch" schedule="03:30" nextRun=2026-09-22T03:30:00+02:00 totalJobs=16 +2026/09/21 11:06:12 selftest.go:50: [INFO] ========== Startup Self-Test ========== +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3844MB avail≈22054MB +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26154728 → total=25898MB avail=22054MB used=3844MB (14.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.7GB avail=58.2GB (9.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.2GB (1.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.32 1.05 0.94 11/1204 124" → 1m=1.32 5m=1.05 15m=0.94 +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 60.6°C (hwmon2) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo done in 66ms — mem=3844MB/25898MB (14.8%), rootDisk=6.7GB/68.4GB (9.8%), load=1.32/1.05/0.94, temp=60.6°C (hwmon2), cpu=0.0% +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Docker socket: reachable (v29.8.0) +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Stacks directory: /opt/docker/stacks +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Data directory: /opt/docker/felhom-controller/data (writable) +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] System data path: /mnt/sys_drive +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Storage paths: 1 connected, 0 disconnected +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Git catalog: 53 app definitions found +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Hub connectivity: hub disabled, skipped +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Metrics DB: 4.5 MB +2026/09/21 11:06:12 selftest.go:70: [INFO] ======================================== +2026/09/21 11:06:12 selftest.go:71: [INFO] Self-test complete: 8 passed, 0 warnings, 0 failed +2026/09/21 11:06:12 scheduler.go:206: [INFO] [scheduler] Starting scheduler with 16 jobs +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] scheduler started: periodic=8 daily=8 +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-abandon-sweep: next run at 2026-09-22 05:10:00 CEST (waiting 16h3m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: next run at 2026-09-22 02:30:00 CEST (waiting 13h23m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-proof: next run at 2026-09-22 05:30:00 CEST (waiting 16h23m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-integrity: next run at 2026-09-22 06:00:00 CEST (waiting 16h53m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: next run at 2026-09-22 03:30:00 CEST (waiting 14h23m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: next run at 2026-09-22 04:15:00 CEST (waiting 15h8m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job metrics-prune: next run at 2026-09-22 04:00:00 CEST (waiting 14h53m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job fill-watch: next run at 2026-09-22 03:30:00 CEST (waiting 14h23m48s) +2026/09/21 11:06:12 server.go:250: [DEBUG] [web] NewServer: initializing web server v0.260.0 +2026/09/21 11:06:12 server.go:251: [DEBUG] [web] NewServer: backup=true scheduler=true alertMgr=true notifier=true updater=false +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:06:12 backup.go:465: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [uptime-kuma] +2026/09/21 11:06:12 backup.go:1107: [INFO] [backup] Found 0 DB dump files across drives +2026/09/21 11:06:12 dbdump.go:128: [DEBUG] DiscoverDatabases: running docker ps to find database containers +2026/09/21 11:06:12 sync.go:212: [ERROR] [sync] Catalog sync failed: git fetch: git fetch --depth 1 origin main: exit status 128 +stderr: fatal: unable to access 'https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git/': Could not resolve host: gitea.dooplex.hu +2026/09/21 11:06:12 sync.go:122: [INFO] [sync] Initial sync: Git hiba: git fetch: git fetch --depth 1 origin main: exit status 128 +stderr: fatal: unable to access 'https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git/': Could not resolve host: gitea.dooplex.hu +2026/09/21 11:06:12 dbdump.go:147: [DEBUG] DiscoverDatabases: docker ps output: 7c794623b894 felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 +e87c990f2026 uptime-kuma uptime-kuma louislam/uptime-kuma:2.4.0 +6d206504c99e filebrowser filebrowser gtstef/filebrowser:1.3.3-stable +2515331bc19e traefik traefik traefik:v3.6.7 +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.260.0, not a database) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container uptime-kuma (image=louislam/uptime-kuma:2.4.0, not a database) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/09/21 11:06:12 dbdump.go:203: [DEBUG] DiscoverDatabases: found 0 database(s), skipped 4 non-DB container(s) +2026/09/21 11:06:12 dbdump.go:206: [INFO] [backup] Discovered 0 databases +2026/09/21 11:06:12 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:06:12 [INFO] [backup] Discovered app data: 1 apps +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/sys_drive" total=68.4 GB used=6.7 GB avail=58.2GB (9.8%) +2026/09/21 11:06:12 recovery_unit.go:224: [INFO] [backup] Recovery unit captured for uptime-kuma → /mnt/sys_drive/felhom-data/backups/primary/uptime-kuma (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) +2026/09/21 11:06:12 backup.go:1177: [INFO] [backup] Backup status cache refreshed +2026/09/21 11:06:12 server.go:324: [DEBUG] [web] loadTemplates: lang=hu loaded 81 templates, 1877 markers expanded in 28.728841ms +2026/09/21 11:06:12 server.go:324: [DEBUG] [web] loadTemplates: lang=en loaded 81 templates, 1877 markers expanded in 19.600884ms +2026/09/21 11:06:12 server.go:270: [INFO] [web] Auth: using password from settings.json +2026/09/21 11:06:12 server.go:811: [DEBUG] [web] CatchAllMiddleware: controller host=felhom.enkisfelhom.hu +2026/09/21 11:06:12 main.go:1856: [INFO] Web UI listening on :8080 +2026/09/21 11:06:12 manager.go:258: [DEBUG] [integrations] ReapplyConfigForTarget: target=filebrowser integrations=0 +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3874MB avail≈22024MB +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26119944 → total=25898MB avail=22024MB used=3874MB (15.0%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.7GB avail=58.2GB (9.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.2GB (1.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.32 1.05 0.94 6/1204 179" → 1m=1.32 5m=1.05 15m=0.94 +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 60.6°C (hwmon2) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo done in 79ms — mem=3874MB/25898MB (15.0%), rootDisk=6.7GB/68.4GB (9.8%), load=1.32/1.05/0.94, temp=60.6°C (hwmon2), cpu=0.0% +2026/09/21 11:06:12 healthcheck.go:86: [DEBUG] [monitor] Raw values: disk=9.8%, hdd=1.8% (configured=true), mem=15.0% (3874MB/25898MB), cpu=0.0%, temp=60.6°C (hwmon2) +2026/09/21 11:06:12 healthcheck.go:118: [DEBUG] [monitor] SSD disk: OK (10%) +2026/09/21 11:06:12 healthcheck.go:151: [DEBUG] [monitor] Memory: OK (15%) +2026/09/21 11:06:12 healthcheck.go:187: [DEBUG] [monitor] Temperature: OK (61°C) +2026/09/21 11:06:12 infra.go:228: [INFO] [infra] connected felhom-controller to traefik-public +2026/09/21 11:06:12 infra.go:67: [INFO] [infra] cloudflared skipped — no cf_tunnel_token configured (LAN-only node) +2026/09/21 11:06:12 healthcheck.go:201: [DEBUG] [monitor] Docker daemon: OK +2026/09/21 11:06:12 healthcheck.go:209: [DEBUG] [monitor] Checking 1 protected containers: [traefik] +2026/09/21 11:06:12 healthcheck.go:219: [DEBUG] [monitor] All protected containers running +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/scratch_hdd" total=937.8 GB used=17.0 GB avail=873.2GB (1.8%) +2026/09/21 11:06:12 healthcheck.go:238: [INFO] [monitor] Health check: status=ok +2026/09/21 11:06:12 healthcheck.go:242: [DEBUG] [monitor] Final status: ok (issues=0, warnings=0, info=4) +2026/09/21 11:06:12 handlers.go:3399: [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s)) +2026/09/21 11:06:17 info.go:13: [DEBUG] [system] CPUCollector: first sample — cpu=16.8% (idle=3282 total=3943) +2026/09/21 11:06:17 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:06:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:06:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 1 targets (0 skipped not due, 0 skipped no container) +2026/09/21 11:06:22 healthprobe.go:153: [DEBUG] Health probe uptime-kuma: HTTP GET :3001/ → 302 (4ms) +2026/09/21 11:06:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:22 healthprobe.go:133: [INFO] Health probes: 1 ok (of 1 probed) +2026/09/21 11:06:27 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:27 main.go:2273: [INFO] [bootrecon] boot window: fleet settled after 10s (3 identical samples 5s apart) — sweeping +2026/09/21 11:06:27 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start) +2026/09/21 11:06:32 auth.go:154: [DEBUG] [web] login attempt from 172.18.0.3:56160 (X-Forwarded-For: 172.18.0.1) +2026/09/21 11:06:32 auth.go:202: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/09/21 11:06:32 auth.go:268: [DEBUG] [web] session created, expires=2026-09-28T11:06:32Z, active_sessions=1 +2026/09/21 11:06:32 auth.go:222: [INFO] [web] Login from 172.18.0.3:56160 +2026/09/21 11:06:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:06:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:06:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:06:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:06:32 auth.go:142: [DEBUG] [web] auth: valid session for GET /stacks +2026/09/21 11:06:32 server.go:542: [DEBUG] [web] ServeHTTP: GET /stacks from 172.18.0.3:56160 +2026/09/21 11:06:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks diff --git a/documentation/audits/update-arc-2026-09-21/11-r520-after-restart.txt b/documentation/audits/update-arc-2026-09-21/11-r520-after-restart.txt new file mode 100644 index 00000000..e4d76b17 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/11-r520-after-restart.txt @@ -0,0 +1,78 @@ +### STEP 5 RESULT — after pct start 9202 + +== 1. Controller log lines about the interrupted update (verbatim) == +2026/09/21 11:06:12 installed.go:525: [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 1 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing) +2026/09/21 11:06:12 pin.go:289: [INFO] [stacks] pin adoption: 0 pinned, 1 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers) +2026/09/21 11:06:12 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s) +2026/09/21 11:06:12 sync.go:208: [INFO] [sync] Starting catalog sync +2026/09/21 11:06:12 sync.go:296: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main) +2026/09/21 11:06:12 sync.go:298: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache +2026/09/21 11:06:12 update.go:904: [WARN] [stacks] update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back +2026/09/21 11:06:12 sync.go:592: [DEBUG] [sync] Running: git fetch --depth 1 origin main +2026/09/21 11:06:12 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/uptime-kuma — 2 env vars, 0 encrypted, 0 sensitive fields +2026/09/21 11:06:12 [INFO] [stacks] SaveAppConfig: saved config for uptime-kuma +2026/09/21 11:06:12 pin.go:93: [INFO] [stacks] pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0 +2026/09/21 11:06:12 update.go:705: [INFO] [stacks] update uptime-kuma: pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] CPUCollector.Start: sampleRate=5s +2026/09/21 11:06:12 main.go:511: [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db +2026/09/21 11:06:12 main.go:523: [INFO] Metrics collector started (60s interval) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) + +== 2. app.yaml — pinned_images vs installed_images == +# Auto-generated by felhom-controller — do not edit locked fields manually +deployed: true +deployed_at: "2026-09-21T10:57:12Z" +env: + DOMAIN: enkisfelhom.hu + SUBDOMAIN: status +locked_fields: + - DOMAIN + - SUBDOMAIN +desired_state: running +installed_images: + uptime-kuma: + ref: louislam/uptime-kuma:2.4.0 + digest: sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985 + at: "2026-09-21T10:57:45Z" +pinned_images: + uptime-kuma: louislam/uptime-kuma:2.4.0 + +== live docker-compose.yml image line == + image: louislam/uptime-kuma:2.4.0 + +== 3. update-journal.json after recovery == + +== 4. Was a HOLD set? (API view) == + state 'running' + deployed True + updating False + deploying False + update_phase 'failed' + update_phase_label 'A frissítés nem sikerült' + update_error 'A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább.' + health_probe.healthy True + containers [('louislam/uptime-kuma:2.4.0', 'Up 31 seconds (healthy)')] + installed_images {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'} + pinned_images {'uptime-kuma': 'louislam/uptime-kuma:2.4.0'} + catalog_images {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} + template_images {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} + +== hold-related log lines (positive+negative control) == +grep -ci 'hold' in the log: 0 +(negative control) grep -c 'ZZZ-not-present': 0 +== 5. What the app page says to the household after the cut == +--- [hu] badge + update-state block --- +Frissítés elérhető — ma + sentences shown: +A frissítés indításához nyomd meg a Frissítés gombot.">Frissítés elérhető — ma +A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább. +--- [en] badge + update-state block --- +Update available — today + sentences shown: +A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább. + +== 6. Is the app running AND serving? == +uptime-kuma | louislam/uptime-kuma:2.4.0 | Up 2 minutes (healthy) + HTTP from inside the guest to the app container :3001/ -> 302 + controller health probe: + {'healthy': True, 'last_check': '2026-09-21T11:06:22.247118521Z', 'details': [{'type': 'http', 'target': ':3001/', 'healthy': True, 'status': 302, 'latency': '4ms'}]} diff --git a/documentation/audits/update-arc-2026-09-21/12-journal-path-provenance.txt b/documentation/audits/update-arc-2026-09-21/12-journal-path-provenance.txt new file mode 100644 index 00000000..b13bfcd6 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/12-journal-path-provenance.txt @@ -0,0 +1,66 @@ +### PATH PROVENANCE — exactly where I looked, with positive controls + +== A. What /var/lib/docker actually IS inside the running guest == +drwx--x--- 12 root root 4096 Sep 21 11:06 /var/lib/docker +drwx--x--- 12 root root 4096 Sep 21 11:06 /var/lib/felhom/docker +/var/lib/docker + +== B. Configured data_dir (controller.yaml, path INSIDE the container) == + +== C. POSITIVE CONTROL — the directory I searched really is the controller data dir. + Listing it via the RUNNING guest, where the symlink resolves correctly: +total 9540 +drwxr-xr-x 3 root root 4096 Sep 21 11:06 . +drwxr-xr-x 3 root root 4096 Sep 13 20:20 .. +drwxr-xr-x 9 root root 4096 Sep 21 11:04 catalog-cache +-rw-r--r-- 1 root root 870169 Sep 21 11:06 debug-ring.log +-rw------- 1 root root 32 Sep 13 20:22 encryption.key +-rw-r--r-- 1 root root 4673536 Sep 21 09:22 metrics.db +-rw-r--r-- 1 root root 32768 Sep 21 11:06 metrics.db-shm +-rw-r--r-- 1 root root 4157112 Sep 21 11:06 metrics.db-wal +-rw-r--r-- 1 root root 572 Sep 16 19:18 settings.json +-rw-r--r-- 1 root root 784 Sep 16 19:17 settings.json.bak + + ...and the SAME directory by its physical path (no symlink): +total 9540 +drwxr-xr-x 3 root root 4096 Sep 21 11:06 . +drwxr-xr-x 3 root root 4096 Sep 13 20:20 .. +drwxr-xr-x 9 root root 4096 Sep 21 11:04 catalog-cache +-rw-r--r-- 1 root root 870169 Sep 21 11:06 debug-ring.log +-rw------- 1 root root 32 Sep 13 20:22 encryption.key +-rw-r--r-- 1 root root 4673536 Sep 21 09:22 metrics.db +-rw-r--r-- 1 root root 32768 Sep 21 11:06 metrics.db-shm +-rw-r--r-- 1 root root 4157112 Sep 21 11:06 metrics.db-wal +-rw-r--r-- 1 root root 572 Sep 16 19:18 settings.json +-rw-r--r-- 1 root root 784 Sep 16 19:17 settings.json.bak + +== D. WHY the host-side read missed it: it is a BIND MOUNT, not a symlink == +-- guest 9202 LXC config mountpoints (from the host) -- +mp0: nvme-scratch:9202/vm-9202-disk-1.raw,mp=/var/lib/felhom,backup=1,size=70G +mp8: /mnt/hdd_1/scratch-drives/scratch_hdd,mp=/mnt/felhom-drives/scratch_hdd +mp9: /var/lib/felhom-agent/guests/9202/bootstrap,mp=/etc/felhom-bootstrap,ro=1 +rootfs: nvme-scratch:9202/vm-9202-disk-0.raw,size=32G +-- inside the guest: is /var/lib/docker its own mount? -- +/dev/loop1[/docker] /var/lib/docker ext4 +-- readlink -f says it is NOT a symlink -- +/var/lib/docker + +CONCLUSION ON PATHS: + * INSIDE the running guest the coordinator's path is correct and is what I used for every + live read: /var/lib/docker/volumes/felhom-controller-data/_data/data/ + * FROM THE HOST, reading the STOPPED guest via 'pct mount 9202', that same path is an EMPTY + mountpoint stub, because pct mount maps only the rootfs and does not apply the guest's + internal bind mounts. The bytes live at + /var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/ + * My FIRST host-side read used the stub path and reported ABSENT. That was a FALSE NEGATIVE, + found by running 'find' over the whole rootfs rather than trusting the one path. + +== E. controller.yaml location (the /etc/felhom guess was wrong too) == +/mnt -> /mnt +/etc/felhom-bootstrap -> /etc/felhom-bootstrap +/var/lib/docker/volumes/felhom-controller-data/_data -> /opt/docker/felhom-controller +/opt/docker/stacks -> /opt/docker/stacks +/var/run/docker.sock -> /var/run/docker.sock + +/var/lib/docker/volumes/felhom-controller-data/_data/controller.yaml +/var/lib/felhom/docker/volumes/felhom-controller-data/_data/controller.yaml diff --git a/documentation/audits/update-arc-2026-09-21/13-finding-interrupted-sentence-not-localised.txt b/documentation/audits/update-arc-2026-09-21/13-finding-interrupted-sentence-not-localised.txt new file mode 100644 index 00000000..ee867cab --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/13-finding-interrupted-sentence-not-localised.txt @@ -0,0 +1,29 @@ + +### FINDING — the interrupted-update sentence is NOT localised (English page shows Hungarian) + +-- EN page: does it carry the HUNGARIAN sentence? (ASCII fragments) -- + 'megszakadt' in app-postcut-en.html : 1 (expect >0 = the bug) + 'vez'+'rl' stem in en : 1 (expect >0 = the bug) +-- EN page: is there any ENGLISH equivalent? -- + 'interrupted' in en : 0 (expect 0) + 'restarted' in en : 0 (expect 0) +-- POSITIVE CONTROL: the page IS in English otherwise -- + 'Update available' in en : 1 (expect >0) + 'Naprak' in en : 0 (expect 0) +-- NEGATIVE CONTROL -- + 'ZZZ-not-present' in en : 0 (expect 0) + +-- source: the sentence is a bare Hungarian literal stored in Stack.UpdateError, + not a bundle key, so errText never sees it (it is stored state, rendered directly): -- +95: MsgUpdateInterrupted = "A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." +902: m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted) +907: m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted) +936: m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted) +943: m.finishUpdate(name, UpdatePhaseFailed, MsgUpdateInterrupted) +-- does hu.json/en.json carry a key for it? -- +/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/i18n/locales/en.json:115: "app_info.megszakadt": "Stopped:", +/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/i18n/locales/en.json:387: "app_import.a_feltoltes_megszakadt_ellenorizze_a": "The upload stopped — check the connection, then try again.", +/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/i18n/locales/en.json:844: "backups_restore.megszakadt_visszaallitas": "Interrupted restore", +/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/i18n/locales/en.json:1149: "storage.megszakadt": "Stopped:", +/mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/i18n/locales/hu.json:115: "app_info.megszakadt": "Megszakadt:", + (grep count in i18n bundle dir: 2 files) diff --git a/documentation/audits/update-arc-2026-09-21/14-update-complete.txt b/documentation/audits/update-arc-2026-09-21/14-update-complete.txt new file mode 100644 index 00000000..64d945ef --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/14-update-complete.txt @@ -0,0 +1,64 @@ +### STEP 6 — POST the update again, let it run to completion +-- POST /api/stacks/uptime-kuma/update -- +{"ok":true,"data":{"accepted":true,"completed":false},"message":"Frissítés elindult – az állapot a kártyán követhető"} + +HTTPCODE:202 + +-- phase timeline observed (poller) -- +11:08:47Z "update_phase":"pulling" +11:08:58Z "update_phase":"verifying" +11:09:09Z "update_phase":"done" + +[exited with code 0] + +-- state after the update completed -- + update_phase : done | Frissítve + update_error : None + updating : False + state : running + containers : [('louislam/uptime-kuma:2.5.0', 'Up 25 seconds (healthy)')] + installed_images: {'uptime-kuma': ('louislam/uptime-kuma:2.5.0', 'sha256:a8610b3b4c38077922b...')} + pinned_images : {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} + catalog_images : {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} + health_probe : {'healthy': True, 'last_check': '2026-09-21T11:09:07.392077664Z', 'details': [{'type': 'http', 'target': ':3001/', 'healthy': True, 'status': 302, 'latency': '3ms'}]} + +-- app.yaml on disk -- +installed_images: + uptime-kuma: + ref: louislam/uptime-kuma:2.5.0 + digest: sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 + at: "2026-09-21T11:09:07Z" +pinned_images: + uptime-kuma: louislam/uptime-kuma:2.5.0 + +-- live compose image line -- + image: louislam/uptime-kuma:2.5.0 + +-- controller log, the completed update -- +2026/09/21 11:06:12 update.go:904: [WARN] [stacks] update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back +2026/09/21 11:06:12 pin.go:93: [INFO] [stacks] pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0 +2026/09/21 11:06:12 update.go:705: [INFO] [stacks] update uptime-kuma: pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container uptime-kuma (image=louislam/uptime-kuma:2.4.0, not a database) +2026/09/21 11:06:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:06:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 20s ago, effective interval 5m0s, healthy=true +2026/09/21 11:06:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 30s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 40s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 50s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m0s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m20s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m30s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m40s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:42 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/uptime-kuma/update +2026/09/21 11:08:42 router.go:81: [DEBUG] [api] POST /api/stacks/uptime-kuma/update (path=/stacks/uptime-kuma/update) +2026/09/21 11:08:42 router.go:582: [INFO] [api] update requested for stack: uptime-kuma + +-- badge now that installed == catalog (2.5.0) -- +Naprakész +Up to date + +-- container now serving -- +uptime-kuma | louislam/uptime-kuma:2.5.0 | Up 40 seconds (healthy) diff --git a/documentation/audits/update-arc-2026-09-21/15-catalog-revert-push.txt b/documentation/audits/update-arc-2026-09-21/15-catalog-revert-push.txt new file mode 100644 index 00000000..ee17d31a --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/15-catalog-revert-push.txt @@ -0,0 +1,19 @@ +ff9717d REVERT the drill bump: uptime-kuma back to 2.4.0 (update-arc measurement finished) + +============================================================================== +== summary +============================================================================== + image-pins OK (exit 0) + engine-major OK (exit 0) + catalog-since OK (exit 0) + copy-i18n OK (exit 0) + +all catalog gates OK +pre-push [app-catalog-felhom.eu]: gates OK - push proceeding. +remote: . Processing 1 references +remote: Processed 1 references in total +To https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + 89304ab..ff9717d main -> main +PUSH_RC=0 + +[exited with code 0] diff --git a/documentation/audits/update-arc-2026-09-21/16-r524-sync-box-ahead.txt b/documentation/audits/update-arc-2026-09-21/16-r524-sync-box-ahead.txt new file mode 100644 index 00000000..c8c52713 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/16-r524-sync-box-ahead.txt @@ -0,0 +1,14 @@ +### STEP 7 — R-524: the box now runs something NEWER than the catalog + +-- force sync -- +{"ok":true,"data":{"ok":true,"message":"Sablonok naprakészek — nincs változás"},"message":"Sablonok naprakészek — nincs változás"} + +HTTPCODE:200 + +-- box catalog cache after the revert -- + image: louislam/uptime-kuma:2.4.0 + +-- the asymmetry the whole row is about -- + installed_images (what the box RUNS) : {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} + catalog_images (what the catalog has): {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} + pinned_images : {'uptime-kuma': 'louislam/uptime-kuma:2.5.0'} diff --git a/documentation/audits/update-arc-2026-09-21/17-r524-badge.txt b/documentation/audits/update-arc-2026-09-21/17-r524-badge.txt new file mode 100644 index 00000000..cc383007 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/17-r524-badge.txt @@ -0,0 +1,24 @@ +### R-524 CAPTURE — box AHEAD (installed 2.5.0 vs catalog 2.4.0) + +== 1+2. THE BADGE AND ITS TITLE, both languages (app page) == +--- [hu] --- +Naprakész +--- [en] --- +Up to date + +== the same badge on the apps LIST page == + [hu] class="tag tag-ok" title="Ez az alkalmazás a katalógusnál újabb változatot futtat, ezért nincs teendőd.">Naprakész + [en] class="tag tag-ok" title="This app is running a version newer than the catalog offers, so there is nothing for you to do.">Up to date + +== CONTRAST: the ORDINARY up-to-date title, captured at step 2 when installed == catalog == + [hu] title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja." + [en] title="This app is running the newest version available." + -> the AHEAD title above is a DIFFERENT sentence, which is the v0.260.0 behaviour. + +== SEARCH CONTROLS == + POSITIVE 'tag-ok' in app-ahead-hu : 1 (expect >0) + POSITIVE 'Naprak' in app-ahead-hu : 1 (expect >0) + POSITIVE 'Up to date' in app-ahead-en: 1 (expect >0) + NEGATIVE 'tag-warn' in app-ahead-hu : 0 (expect 0 — NOT 'update available') + NEGATIVE 'tag-warn' in app-ahead-en : 0 (expect 0) + NEGATIVE 'ZZZ-not-present' : 0 (expect 0) diff --git a/documentation/audits/update-arc-2026-09-21/18-r524-409-refusal.txt b/documentation/audits/update-arc-2026-09-21/18-r524-409-refusal.txt new file mode 100644 index 00000000..1a1cc8cf --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/18-r524-409-refusal.txt @@ -0,0 +1,33 @@ +### R-524 — THE GUARDED UPDATE MUST REFUSE (409), both languages + +== 3a. POST /api/stacks/uptime-kuma/update (Hungarian, the box's saved language) == +{"ok":false,"error":"Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére."} + +HTTPCODE:409 + +== 3b. POST /api/stacks/uptime-kuma/update?lang=en (English) == +{"ok":false,"error":"This version is newer than the one in the catalog — moving back needs the operator."} + +HTTPCODE:409 + +== 3c. control — ?lang=hu explicitly == +{"ok":false,"error":"Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére."} + +HTTPCODE:409 + +== 4. The controller log line for the refusal (verbatim, names BOTH image maps) == +2026/09/21 11:11:06 update.go:292: [ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 2026-09-21T11:09:07Z}] catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) +2026/09/21 11:11:07 update.go:292: [ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 2026-09-21T11:09:07Z}] catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) +2026/09/21 11:11:08 update.go:292: [ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 2026-09-21T11:09:07Z}] catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) + +-- controls -- + POSITIVE 'downgrade' in the log : 3 (expect >0) + POSITIVE 'REFUSED' in the log : 3 (expect 3 = the three POSTs) + NEGATIVE 'ZZZ-not-present' : 0 (expect 0) + +-- proof nothing moved: the app is untouched by the three refused POSTs -- +uptime-kuma | louislam/uptime-kuma:2.5.0 | Up 2 minutes (healthy) + image: louislam/uptime-kuma:2.5.0 + update-journal.json present (should be ABSENT — a refused update records nothing): ABSENT + (positive control that this path is real — catalog-cache is beside it:) +catalog-cache debug-ring.log encryption.key metrics.db metrics.db-shm metrics.db-wal settings.json settings.json.bak \ No newline at end of file diff --git a/documentation/audits/update-arc-2026-09-21/19-controller-log-r524-refusal.txt b/documentation/audits/update-arc-2026-09-21/19-controller-log-r524-refusal.txt new file mode 100644 index 00000000..787e50f7 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/19-controller-log-r524-refusal.txt @@ -0,0 +1,912 @@ +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026/09/21 11:06:12 main.go:334: [INFO] felhom-controller 0.260.0 starting (customer: demo-hp, domain: enkisfelhom.hu) +2026/09/21 11:06:12 settings.go:693: [INFO] [settings] Loaded settings from /opt/docker/felhom-controller/data/settings.json +2026/09/21 11:06:12 settings.go:1611: [DEBUG] [settings] AutoDiscoverStoragePaths discovered=[] fallback="" existing=1 +2026/09/21 11:06:12 main.go:360: [INFO] Encryption key loaded from /opt/docker/felhom-controller/data/encryption.key +2026/09/21 11:06:12 manager.go:283: [INFO] [stacks] Using compose command: docker compose +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=false composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=false composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=false composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=false composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=false composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=false composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=true composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/09/21 11:06:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/09/21 11:06:12 manager.go:1620: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/09/21 11:06:12 manager.go:630: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/09/21 11:06:12 manager.go:661: [INFO] [stacks] ScanStacks complete: 55 stacks found (1 deployed, 54 available) +2026/09/21 11:06:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:12 deploy.go:1086: [DEBUG] [stacks] InjectMissingFields: checking 1 stacks +2026/09/21 11:06:12 deploy.go:1108: [DEBUG] [stacks] InjectMissingFields: checking stack uptime-kuma — 2 deploy fields, 2 existing env vars +2026/09/21 11:06:12 deploy.go:1171: [INFO] [stacks] InjectMissingFields: processed 1 stacks +2026/09/21 11:06:12 manager.go:422: [DEBUG] [stacks] MigrateEncryption: checking 1 deployed stacks for plaintext sensitive values +2026/09/21 11:06:12 manager.go:466: [INFO] [stacks] Encryption migration: no stacks needed migration +2026/09/21 11:06:12 desiredstate.go:137: [INFO] [stacks] desired-state backfill: 0 app(s) recorded as running, 0 left unrecorded (state ambiguous — legacy boot behaviour retained) +2026/09/21 11:06:12 installed.go:525: [INFO] [stacks] installed-images backfill: 0 app(s) recorded, 1 already had a record, 0 left unrecorded (could not be observed completely — unknown, which renders nothing) +2026/09/21 11:06:12 pin.go:289: [INFO] [stacks] pin adoption: 0 pinned, 1 already pinned, 0 left unpinned (0 not completely observed, 0 running something the template no longer offers) +2026/09/21 11:06:12 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s) +2026/09/21 11:06:12 sync.go:208: [INFO] [sync] Starting catalog sync +2026/09/21 11:06:12 sync.go:296: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main) +2026/09/21 11:06:12 sync.go:298: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache +2026/09/21 11:06:12 update.go:904: [WARN] [stacks] update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back +2026/09/21 11:06:12 sync.go:592: [DEBUG] [sync] Running: git fetch --depth 1 origin main +2026/09/21 11:06:12 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/uptime-kuma — 2 env vars, 0 encrypted, 0 sensitive fields +2026/09/21 11:06:12 [INFO] [stacks] SaveAppConfig: saved config for uptime-kuma +2026/09/21 11:06:12 pin.go:93: [INFO] [stacks] pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0 +2026/09/21 11:06:12 update.go:705: [INFO] [stacks] update uptime-kuma: pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] CPUCollector.Start: sampleRate=5s +2026/09/21 11:06:12 main.go:511: [INFO] Metrics store opened at /opt/docker/felhom-controller/data/metrics.db +2026/09/21 11:06:12 main.go:523: [INFO] Metrics collector started (60s interval) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:06:12 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) +2026/09/21 11:06:12 notifier.go:75: [INFO] Notifier disabled (hub not configured) +2026/09/21 11:06:12 main.go:642: [INFO] [mailrelay] app-email shim unavailable (no hub configured or disabled in config) +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: status-refresh (every 10s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="status-refresh" interval=10s totalJobs=1 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: stack-scan (every 2m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="stack-scan" interval=2m0s totalJobs=2 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: health-probes (every 10s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="health-probes" interval=10s totalJobs=3 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: system-health (every 5m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="system-health" interval=5m0s totalJobs=4 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: deadapp-check (every 30s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="deadapp-check" interval=30s totalJobs=5 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: ring-spill (every 30s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="ring-spill" interval=30s totalJobs=6 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-22 02:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="db-dump" schedule="02:30" nextRun=2026-09-22T02:30:00+02:00 totalJobs=7 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="offsite-credential-retry" interval=5m0s totalJobs=8 +2026/09/21 11:06:12 scheduler.go:102: [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="backup-cache" interval=5m0s totalJobs=9 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-22 03:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="tier2-backup" schedule="03:30" nextRun=2026-09-22T03:30:00+02:00 totalJobs=10 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-22 04:15 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offbox-backup" schedule="04:15" nextRun=2026-09-22T04:15:00+02:00 totalJobs=11 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-22 05:10 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-abandon-sweep" schedule="05:10" nextRun=2026-09-22T05:10:00+02:00 totalJobs=12 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-22 06:00 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-integrity" schedule="06:00" nextRun=2026-09-22T06:00:00+02:00 totalJobs=13 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-22 05:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="offsite-proof" schedule="05:30" nextRun=2026-09-22T05:30:00+02:00 totalJobs=14 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-22 04:00 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="metrics-prune" schedule="04:00" nextRun=2026-09-22T04:00:00+02:00 totalJobs=15 +2026/09/21 11:06:12 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-22 03:30 CEST +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job registered: name="fill-watch" schedule="03:30" nextRun=2026-09-22T03:30:00+02:00 totalJobs=16 +2026/09/21 11:06:12 selftest.go:50: [INFO] ========== Startup Self-Test ========== +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3844MB avail≈22054MB +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26154728 → total=25898MB avail=22054MB used=3844MB (14.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.7GB avail=58.2GB (9.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.2GB (1.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.32 1.05 0.94 11/1204 124" → 1m=1.32 5m=1.05 15m=0.94 +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 60.6°C (hwmon2) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo done in 66ms — mem=3844MB/25898MB (14.8%), rootDisk=6.7GB/68.4GB (9.8%), load=1.32/1.05/0.94, temp=60.6°C (hwmon2), cpu=0.0% +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Docker socket: reachable (v29.8.0) +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Stacks directory: /opt/docker/stacks +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Data directory: /opt/docker/felhom-controller/data (writable) +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] System data path: /mnt/sys_drive +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Storage paths: 1 connected, 0 disconnected +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Git catalog: 53 app definitions found +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Hub connectivity: hub disabled, skipped +2026/09/21 11:06:12 selftest.go:68: [INFO] [PASS] Metrics DB: 4.5 MB +2026/09/21 11:06:12 selftest.go:70: [INFO] ======================================== +2026/09/21 11:06:12 selftest.go:71: [INFO] Self-test complete: 8 passed, 0 warnings, 0 failed +2026/09/21 11:06:12 scheduler.go:206: [INFO] [scheduler] Starting scheduler with 16 jobs +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] scheduler started: periodic=8 daily=8 +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-abandon-sweep: next run at 2026-09-22 05:10:00 CEST (waiting 16h3m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: next run at 2026-09-22 02:30:00 CEST (waiting 13h23m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-proof: next run at 2026-09-22 05:30:00 CEST (waiting 16h23m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offsite-integrity: next run at 2026-09-22 06:00:00 CEST (waiting 16h53m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: next run at 2026-09-22 03:30:00 CEST (waiting 14h23m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: next run at 2026-09-22 04:15:00 CEST (waiting 15h8m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job metrics-prune: next run at 2026-09-22 04:00:00 CEST (waiting 14h53m48s) +2026/09/21 11:06:12 scheduler.go:67: [DEBUG] [scheduler] daily job fill-watch: next run at 2026-09-22 03:30:00 CEST (waiting 14h23m48s) +2026/09/21 11:06:12 server.go:250: [DEBUG] [web] NewServer: initializing web server v0.260.0 +2026/09/21 11:06:12 server.go:251: [DEBUG] [web] NewServer: backup=true scheduler=true alertMgr=true notifier=true updater=false +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:06:12 backup.go:465: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [uptime-kuma] +2026/09/21 11:06:12 backup.go:1107: [INFO] [backup] Found 0 DB dump files across drives +2026/09/21 11:06:12 dbdump.go:128: [DEBUG] DiscoverDatabases: running docker ps to find database containers +2026/09/21 11:06:12 sync.go:212: [ERROR] [sync] Catalog sync failed: git fetch: git fetch --depth 1 origin main: exit status 128 +stderr: fatal: unable to access 'https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git/': Could not resolve host: gitea.dooplex.hu +2026/09/21 11:06:12 sync.go:122: [INFO] [sync] Initial sync: Git hiba: git fetch: git fetch --depth 1 origin main: exit status 128 +stderr: fatal: unable to access 'https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git/': Could not resolve host: gitea.dooplex.hu +2026/09/21 11:06:12 dbdump.go:147: [DEBUG] DiscoverDatabases: docker ps output: 7c794623b894 felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 +e87c990f2026 uptime-kuma uptime-kuma louislam/uptime-kuma:2.4.0 +6d206504c99e filebrowser filebrowser gtstef/filebrowser:1.3.3-stable +2515331bc19e traefik traefik traefik:v3.6.7 +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.260.0, not a database) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container uptime-kuma (image=louislam/uptime-kuma:2.4.0, not a database) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/09/21 11:06:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/09/21 11:06:12 dbdump.go:203: [DEBUG] DiscoverDatabases: found 0 database(s), skipped 4 non-DB container(s) +2026/09/21 11:06:12 dbdump.go:206: [INFO] [backup] Discovered 0 databases +2026/09/21 11:06:12 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:06:12 [INFO] [backup] Discovered app data: 1 apps +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/sys_drive" total=68.4 GB used=6.7 GB avail=58.2GB (9.8%) +2026/09/21 11:06:12 recovery_unit.go:224: [INFO] [backup] Recovery unit captured for uptime-kuma → /mnt/sys_drive/felhom-data/backups/primary/uptime-kuma (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) +2026/09/21 11:06:12 backup.go:1177: [INFO] [backup] Backup status cache refreshed +2026/09/21 11:06:12 server.go:324: [DEBUG] [web] loadTemplates: lang=hu loaded 81 templates, 1877 markers expanded in 28.728841ms +2026/09/21 11:06:12 server.go:324: [DEBUG] [web] loadTemplates: lang=en loaded 81 templates, 1877 markers expanded in 19.600884ms +2026/09/21 11:06:12 server.go:270: [INFO] [web] Auth: using password from settings.json +2026/09/21 11:06:12 server.go:811: [DEBUG] [web] CatchAllMiddleware: controller host=felhom.enkisfelhom.hu +2026/09/21 11:06:12 main.go:1856: [INFO] Web UI listening on :8080 +2026/09/21 11:06:12 manager.go:258: [DEBUG] [integrations] ReapplyConfigForTarget: target=filebrowser integrations=0 +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3874MB avail≈22024MB +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26119944 → total=25898MB avail=22024MB used=3874MB (15.0%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.7GB avail=58.2GB (9.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.2GB (1.8%) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.32 1.05 0.94 6/1204 179" → 1m=1.32 5m=1.05 15m=0.94 +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 60.6°C (hwmon2) +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetInfo done in 79ms — mem=3874MB/25898MB (15.0%), rootDisk=6.7GB/68.4GB (9.8%), load=1.32/1.05/0.94, temp=60.6°C (hwmon2), cpu=0.0% +2026/09/21 11:06:12 healthcheck.go:86: [DEBUG] [monitor] Raw values: disk=9.8%, hdd=1.8% (configured=true), mem=15.0% (3874MB/25898MB), cpu=0.0%, temp=60.6°C (hwmon2) +2026/09/21 11:06:12 healthcheck.go:118: [DEBUG] [monitor] SSD disk: OK (10%) +2026/09/21 11:06:12 healthcheck.go:151: [DEBUG] [monitor] Memory: OK (15%) +2026/09/21 11:06:12 healthcheck.go:187: [DEBUG] [monitor] Temperature: OK (61°C) +2026/09/21 11:06:12 infra.go:228: [INFO] [infra] connected felhom-controller to traefik-public +2026/09/21 11:06:12 infra.go:67: [INFO] [infra] cloudflared skipped — no cf_tunnel_token configured (LAN-only node) +2026/09/21 11:06:12 healthcheck.go:201: [DEBUG] [monitor] Docker daemon: OK +2026/09/21 11:06:12 healthcheck.go:209: [DEBUG] [monitor] Checking 1 protected containers: [traefik] +2026/09/21 11:06:12 healthcheck.go:219: [DEBUG] [monitor] All protected containers running +2026/09/21 11:06:12 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/scratch_hdd" total=937.8 GB used=17.0 GB avail=873.2GB (1.8%) +2026/09/21 11:06:12 healthcheck.go:238: [INFO] [monitor] Health check: status=ok +2026/09/21 11:06:12 healthcheck.go:242: [DEBUG] [monitor] Final status: ok (issues=0, warnings=0, info=4) +2026/09/21 11:06:12 handlers.go:3399: [INFO] [web] FileBrowser sync — no config/compose change, ensured running without recreate (1 storage path(s)) +2026/09/21 11:06:17 info.go:13: [DEBUG] [system] CPUCollector: first sample — cpu=16.8% (idle=3282 total=3943) +2026/09/21 11:06:17 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:06:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:06:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 1 targets (0 skipped not due, 0 skipped no container) +2026/09/21 11:06:22 healthprobe.go:153: [DEBUG] Health probe uptime-kuma: HTTP GET :3001/ → 302 (4ms) +2026/09/21 11:06:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:22 healthprobe.go:133: [INFO] Health probes: 1 ok (of 1 probed) +2026/09/21 11:06:27 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:27 main.go:2273: [INFO] [bootrecon] boot window: fleet settled after 10s (3 identical samples 5s apart) — sweeping +2026/09/21 11:06:27 bootrecon.go:255: [INFO] [bootrecon] Boot reconciliation: no boot-orphaned apps (nothing to start) +2026/09/21 11:06:32 auth.go:154: [DEBUG] [web] login attempt from 172.18.0.3:56160 (X-Forwarded-For: 172.18.0.1) +2026/09/21 11:06:32 auth.go:202: [DEBUG] [web] login successful from 172.18.0.1, creating session +2026/09/21 11:06:32 auth.go:268: [DEBUG] [web] session created, expires=2026-09-28T11:06:32Z, active_sessions=1 +2026/09/21 11:06:32 auth.go:222: [INFO] [web] Login from 172.18.0.3:56160 +2026/09/21 11:06:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:06:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:06:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:06:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:06:32 auth.go:142: [DEBUG] [web] auth: valid session for GET /stacks +2026/09/21 11:06:32 server.go:542: [DEBUG] [web] ServeHTTP: GET /stacks from 172.18.0.3:56160 +2026/09/21 11:06:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:06:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:06:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:06:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:06:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 20s ago, effective interval 5m0s, healthy=true +2026/09/21 11:06:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:06:42 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:52 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:06:52 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:06:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:06:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:06:52 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:06:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 30s ago, effective interval 5m0s, healthy=true +2026/09/21 11:06:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:07:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 40s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:07:02 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈4073MB avail≈21825MB +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25883456 → total=25898MB avail=21825MB used=4073MB (15.7%) +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.7GB avail=58.2GB (9.8%) +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.2GB (1.8%) +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="0.94 0.98 0.92 2/1190 371" → 1m=0.94 5m=0.98 15m=0.92 +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 53.8°C (hwmon2) +2026/09/21 11:07:12 info.go:13: [DEBUG] [system] GetInfo done in 62ms — mem=4073MB/25898MB (15.7%), rootDisk=6.7GB/68.4GB (9.8%), load=0.94/0.98/0.92, temp=53.8°C (hwmon2), cpu=15.9% +2026/09/21 11:07:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:07:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:07:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:07:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 50s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:07:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:07:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:07:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:07:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:07:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m0s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:07:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:07:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:07:42 info.go:13: [DEBUG] [system] GetDiskUsage: path="/" total=68.4 GB used=6.7 GB avail=58.2GB (9.8%) +2026/09/21 11:07:42 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/scratch_hdd" total=937.8 GB used=17.0 GB avail=873.2GB (1.8%) +2026/09/21 11:07:42 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/sys_drive" total=68.4 GB used=6.7 GB avail=58.2GB (9.8%) +2026/09/21 11:07:42 fillwatch.go:256: [INFO] [fillwatch] checked 3 filesystem(s), 0 unreadable/skipped, 0 notification(s); bands: all ok +2026/09/21 11:07:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:07:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:07:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:07:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m20s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:07:42 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:07:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:07:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:07:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m30s ago, effective interval 5m0s, healthy=true +2026/09/21 11:07:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:07:52 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:08:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:08:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m40s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:08:02 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:11 auth.go:142: [DEBUG] [web] auth: valid session for GET /apps/uptime-kuma +2026/09/21 11:08:11 server.go:542: [DEBUG] [web] ServeHTTP: GET /apps/uptime-kuma from 172.18.0.3:56160 +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3984MB avail≈21914MB +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25988332 → total=25898MB avail=21914MB used=3984MB (15.4%) +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.7GB avail=58.2GB (9.8%) +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.2GB (1.8%) +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="0.75 0.92 0.91 7/1193 488" → 1m=0.75 5m=0.92 15m=0.91 +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 54.6°C (hwmon2) +2026/09/21 11:08:12 info.go:13: [DEBUG] [system] GetInfo done in 72ms — mem=3984MB/25898MB (15.4%), rootDisk=6.7GB/68.4GB (9.8%), load=0.75/0.92/0.91, temp=54.6°C (hwmon2), cpu=8.7% +2026/09/21 11:08:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:08:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:08:12 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan +2026/09/21 11:08:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:08:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m50s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:12 scheduler.go:67: [DEBUG] [scheduler] job stack-scan: execution starting +2026/09/21 11:08:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:08:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=false composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=false composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=false composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=false composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=false composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=false composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=true composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/09/21 11:08:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/09/21 11:08:12 manager.go:1620: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/09/21 11:08:12 manager.go:630: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/09/21 11:08:12 manager.go:661: [INFO] [stacks] ScanStacks complete: 55 stacks found (1 deployed, 54 available) +2026/09/21 11:08:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:12 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 68ms) +2026/09/21 11:08:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:13 auth.go:142: [DEBUG] [web] auth: valid session for GET /apps/uptime-kuma +2026/09/21 11:08:13 server.go:542: [DEBUG] [web] ServeHTTP: GET /apps/uptime-kuma from 172.18.0.3:56160 +2026/09/21 11:08:16 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:08:16 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:08:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:08:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m0s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:08:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:08:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:08:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m10s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:08:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:08:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:42 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/uptime-kuma/update +2026/09/21 11:08:42 router.go:81: [DEBUG] [api] POST /api/stacks/uptime-kuma/update (path=/stacks/uptime-kuma/update) +2026/09/21 11:08:42 router.go:582: [INFO] [api] update requested for stack: uptime-kuma +2026/09/21 11:08:42 router.go:81: [DEBUG] [api] actionStack: action=update name=uptime-kuma +2026/09/21 11:08:42 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈4062MB avail≈21836MB +2026/09/21 11:08:42 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25896520 → total=25898MB avail=21836MB used=4062MB (15.7%) +2026/09/21 11:08:42 deploy.go:1232: [INFO] [stacks] Memory check: total=25898MB, reserved=384MB, usable=25514MB, committed_used=0MB, new_req=50MB, remaining=25464MB +2026/09/21 11:08:42 info.go:13: [DEBUG] [system] GetDiskUsage: path="/" total=68.4 GB used=6.7 GB avail=58.2GB (9.8%) +2026/09/21 11:08:42 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈4040MB avail≈21858MB +2026/09/21 11:08:42 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25922964 → total=25898MB avail=21858MB used=4040MB (15.6%) +2026/09/21 11:08:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:08:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:08:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:08:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:08:42 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:42 deploy.go:1232: [INFO] [stacks] Memory check: total=25898MB, reserved=384MB, usable=25514MB, committed_used=0MB, new_req=50MB, remaining=25464MB +2026/09/21 11:08:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m20s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:08:42 info.go:13: [DEBUG] [system] GetDiskUsage: path="/" total=68.4 GB used=6.7 GB avail=58.2GB (9.8%) +2026/09/21 11:08:42 update.go:429: [INFO] [stacks] update uptime-kuma: accepted — guarded update started +2026/09/21 11:08:42 update.go:858: [INFO] [stacks] update uptime-kuma: phase checking +2026/09/21 11:08:42 update_guard.go:224: [DEBUG] [backup] update precondition for uptime-kuma: no Tier-2 copy (nincs másodlagos fájlmásolat ehhez az alkalmazáshoz) +2026/09/21 11:08:42 update.go:544: [INFO] [stacks] update uptime-kuma: precondition met — Tier 1 (own recovery unit) copy from 2026-09-21T11:06:12Z (3m0s old, limit 24h0m0s) +2026/09/21 11:08:42 update.go:858: [INFO] [stacks] update uptime-kuma: phase safety-dump +2026/09/21 11:08:42 dbdump.go:128: [DEBUG] DiscoverDatabases: running docker ps to find database containers +2026/09/21 11:08:42 dbdump.go:147: [DEBUG] DiscoverDatabases: docker ps output: 7c794623b894 felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 +e87c990f2026 uptime-kuma uptime-kuma louislam/uptime-kuma:2.4.0 +6d206504c99e filebrowser filebrowser gtstef/filebrowser:1.3.3-stable +2515331bc19e traefik traefik traefik:v3.6.7 +2026/09/21 11:08:42 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.260.0, not a database) +2026/09/21 11:08:42 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container uptime-kuma (image=louislam/uptime-kuma:2.4.0, not a database) +2026/09/21 11:08:42 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/09/21 11:08:42 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/09/21 11:08:42 dbdump.go:203: [DEBUG] DiscoverDatabases: found 0 database(s), skipped 4 non-DB container(s) +2026/09/21 11:08:42 dbdump.go:206: [INFO] [backup] Discovered 0 databases +2026/09/21 11:08:42 update_guard.go:452: [INFO] [backup] update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op) +2026/09/21 11:08:42 update.go:576: [INFO] [stacks] update uptime-kuma: safety dump done (0 file(s)) [] +2026/09/21 11:08:42 update.go:858: [INFO] [stacks] update uptime-kuma: phase pinning +2026/09/21 11:08:42 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/uptime-kuma — 2 env vars, 0 encrypted, 0 sensitive fields +2026/09/21 11:08:42 [INFO] [stacks] SaveAppConfig: saved config for uptime-kuma +2026/09/21 11:08:42 pin.go:362: [INFO] [stacks] update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.5.0) +2026/09/21 11:08:42 update.go:858: [INFO] [stacks] update uptime-kuma: phase pulling +2026/09/21 11:08:42 manager.go:1384: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, SUBDOMAIN, DOMAIN, IMPORT_PATH] (8 app + 0 system) +2026/09/21 11:08:42 manager.go:1393: [DEBUG] Running: docker compose pull (in /opt/docker/stacks/uptime-kuma) +2026/09/21 11:08:47 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:08:47 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:08:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:08:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m30s ago, effective interval 5m0s, healthy=true +2026/09/21 11:08:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:08:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:08:52 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:56 manager.go:1413: [DEBUG] Command completed: docker compose pull (took 13.9s) +2026/09/21 11:08:56 update.go:858: [INFO] [stacks] update uptime-kuma: phase starting +2026/09/21 11:08:56 manager.go:1384: [DEBUG] Env vars for compose: [PATH, HOSTNAME, FELHOM_BOOTSTRAP_PATH, HOME, DOMAIN, SUBDOMAIN, DOMAIN, IMPORT_PATH] (8 app + 0 system) +2026/09/21 11:08:56 manager.go:1393: [DEBUG] Running: docker compose up -d --remove-orphans (in /opt/docker/stacks/uptime-kuma) +2026/09/21 11:08:56 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:08:56 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:08:57 manager.go:1413: [DEBUG] Command completed: docker compose up -d --remove-orphans (took 1.1s) +2026/09/21 11:08:57 update.go:858: [INFO] [stacks] update uptime-kuma: phase verifying +2026/09/21 11:08:57 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:08:58 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:08:58 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:09:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:09:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:09:02 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (0 skipped not due, 0 skipped no container) +2026/09/21 11:09:02 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:07 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:07 healthprobe.go:153: [DEBUG] Health probe uptime-kuma: HTTP GET :3001/ → 302 (3ms) +2026/09/21 11:09:07 update.go:652: [INFO] [stacks] update uptime-kuma: healthy after 10s (the app's health check passed) +2026/09/21 11:09:07 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:09:07 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:09:07 [DEBUG] [stacks] SaveAppConfig: saving /opt/docker/stacks/uptime-kuma — 2 env vars, 0 encrypted, 0 sensitive fields +2026/09/21 11:09:07 [INFO] [stacks] SaveAppConfig: saved config for uptime-kuma +2026/09/21 11:09:07 installed.go:408: [INFO] [stacks] installed-images uptime-kuma: recorded 1 service(s) (uptime-kuma=louislam/uptime-kuma:2.5.0 (sha256:a8610b3b4c38…)) +2026/09/21 11:09:07 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:07 update.go:658: [INFO] [stacks] update uptime-kuma: DONE in 25s +2026/09/21 11:09:09 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:09:09 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:09:11 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:09:11 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:09:12 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:09:12 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈4093MB avail≈21805MB +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=25860040 → total=25898MB avail=21805MB used=4093MB (15.8%) +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.9GB avail=57.9GB (10.2%) +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.1GB (1.8%) +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.07 0.97 0.92 3/1195 889" → 1m=1.07 5m=0.97 15m=0.92 +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 53.0°C (hwmon2) +2026/09/21 11:09:12 info.go:13: [DEBUG] [system] GetInfo done in 63ms — mem=4093MB/25898MB (15.8%), rootDisk=6.9GB/68.4GB (10.2%), load=1.07/0.97/0.92, temp=53.0°C (hwmon2), cpu=6.5% +2026/09/21 11:09:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:09:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:09:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:09:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:09:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 5s ago, effective interval 5m0s, healthy=true +2026/09/21 11:09:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:09:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:09:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:09:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 15s ago, effective interval 5m0s, healthy=true +2026/09/21 11:09:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:09:24 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:09:24 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:09:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:09:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 25s ago, effective interval 5m0s, healthy=true +2026/09/21 11:09:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:09:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:09:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:34 auth.go:142: [DEBUG] [web] auth: valid session for GET /apps/uptime-kuma +2026/09/21 11:09:34 server.go:542: [DEBUG] [web] ServeHTTP: GET /apps/uptime-kuma from 172.18.0.3:56160 +2026/09/21 11:09:36 auth.go:142: [DEBUG] [web] auth: valid session for GET /apps/uptime-kuma +2026/09/21 11:09:36 server.go:542: [DEBUG] [web] ServeHTTP: GET /apps/uptime-kuma from 172.18.0.3:56160 +2026/09/21 11:09:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:09:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:09:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 35s ago, effective interval 5m0s, healthy=true +2026/09/21 11:09:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:09:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:09:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:09:42 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:09:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:09:52 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:09:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 45s ago, effective interval 5m0s, healthy=true +2026/09/21 11:09:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:10:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:10:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 55s ago, effective interval 5m0s, healthy=true +2026/09/21 11:10:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:02 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3939MB avail≈21959MB +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26042720 → total=25898MB avail=21959MB used=3939MB (15.2%) +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.9GB avail=57.9GB (10.2%) +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.1GB (1.8%) +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.17 1.01 0.94 4/1190 1029" → 1m=1.17 5m=1.01 15m=0.94 +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 55.4°C (hwmon2) +2026/09/21 11:10:12 info.go:13: [DEBUG] [system] GetInfo done in 66ms — mem=3939MB/25898MB (15.2%), rootDisk=6.9GB/68.4GB (10.2%), load=1.17/1.01/0.94, temp=55.4°C (hwmon2), cpu=3.4% +2026/09/21 11:10:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:10:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:10:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:10:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m5s ago, effective interval 5m0s, healthy=true +2026/09/21 11:10:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:10:12 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan +2026/09/21 11:10:12 scheduler.go:67: [DEBUG] [scheduler] job stack-scan: execution starting +2026/09/21 11:10:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=false composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=false composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=false composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=false composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=false composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=false composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=true composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/09/21 11:10:12 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/09/21 11:10:12 manager.go:1620: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/09/21 11:10:12 manager.go:630: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/09/21 11:10:12 manager.go:661: [INFO] [stacks] ScanStacks complete: 55 stacks found (1 deployed, 54 available) +2026/09/21 11:10:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:12 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 82ms) +2026/09/21 11:10:12 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/sync +2026/09/21 11:10:12 router.go:81: [DEBUG] [api] POST /api/sync (path=/sync) +2026/09/21 11:10:12 router.go:1002: [INFO] [api] Manual catalog sync requested +2026/09/21 11:10:12 sync.go:208: [INFO] [sync] Starting catalog sync +2026/09/21 11:10:12 sync.go:296: [INFO] [sync] Pulling latest from https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (branch: main) +2026/09/21 11:10:12 sync.go:298: [DEBUG] [sync] git fetch --depth 1 origin main in /opt/docker/felhom-controller/data/catalog-cache +2026/09/21 11:10:12 sync.go:592: [DEBUG] [sync] Running: git fetch --depth 1 origin main +2026/09/21 11:10:12 sync.go:592: [DEBUG] [sync] Running: git reset --hard origin/main +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] actualbudget/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] actualbudget/.felhom.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] adventurelog/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] adventurelog/.felhom.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] audiobookshelf/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] audiobookshelf/.felhom.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] bentopdf/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] bentopdf/.felhom.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] bookstack/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] bookstack/.felhom.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] calcom/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] calcom/.felhom.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] calibre-web/docker-compose.yml: hash match, skipped +2026/09/21 11:10:12 sync.go:416: [DEBUG] [sync] calibre-web/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] claper/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] claper/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] code-server/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] code-server/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] crafty-controller/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] crafty-controller/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] docmost/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] docmost/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] emby/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] emby/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] ghost/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] ghost/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] gitea/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] gitea/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] glance/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] glance/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] gokapi/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] gokapi/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] grafana/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] grafana/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] gramps-web/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] gramps-web/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] home-assistant/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] home-assistant/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] homebox/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] homebox/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] homepage/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] homepage/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] immich/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] immich/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] jellyfin/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] jellyfin/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] kimai/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] kimai/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] komga/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] komga/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] mealie/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] mealie/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] n8n/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] n8n/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] navidrome/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] navidrome/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] nextcloud/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] nextcloud/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] onlyoffice/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] onlyoffice/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] opengist/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] opengist/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] outline/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] outline/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] paperless-ngx/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] paperless-ngx/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] papra/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] papra/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] plant-it/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] plant-it/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] plex/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] plex/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] privatebin/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] privatebin/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] radarr/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] radarr/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] rallly/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] rallly/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] recipe-importer/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] recipe-importer/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] romm/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] romm/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] seerr/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] seerr/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] sonarr/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] sonarr/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] sparkyfitness/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] sparkyfitness/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] tandoor/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] tandoor/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] termix/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] termix/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:509: [DEBUG] [sync] uptime-kuma: catalog has moved past the pin — rendering the stored applied definition +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] uptime-kuma/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] uptime-kuma/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] vaultwarden/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] vaultwarden/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] vikunja/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] vikunja/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] wanderer/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] wanderer/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] wger/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] wger/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] wishlist/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] wishlist/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] zipline/docker-compose.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:416: [DEBUG] [sync] zipline/.felhom.yml: hash match, skipped +2026/09/21 11:10:13 sync.go:263: [INFO] [sync] Catalog sync complete +2026/09/21 11:10:13 router.go:1008: [INFO] [api] Catalog sync completed: Sablonok naprakészek — nincs változás +2026/09/21 11:10:15 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:10:15 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:10:22 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:10:22 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:10:22 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m15s ago, effective interval 5m0s, healthy=true +2026/09/21 11:10:22 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:22 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:32 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:10:32 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:10:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:32 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m25s ago, effective interval 5m0s, healthy=true +2026/09/21 11:10:32 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:32 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/rescan +2026/09/21 11:10:32 router.go:81: [DEBUG] [api] POST /api/stacks/rescan (path=/stacks/rescan) +2026/09/21 11:10:32 router.go:362: [INFO] [api] Manual stack rescan requested +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "actualbudget" deployed=false composePath=/opt/docker/stacks/actualbudget/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "adventurelog" deployed=false composePath=/opt/docker/stacks/adventurelog/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "audiobookshelf" deployed=false composePath=/opt/docker/stacks/audiobookshelf/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bentopdf" deployed=false composePath=/opt/docker/stacks/bentopdf/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=false composePath=/opt/docker/stacks/bookstack/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calcom" deployed=false composePath=/opt/docker/stacks/calcom/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "calibre-web" deployed=false composePath=/opt/docker/stacks/calibre-web/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "claper" deployed=false composePath=/opt/docker/stacks/claper/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "code-server" deployed=false composePath=/opt/docker/stacks/code-server/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "crafty-controller" deployed=false composePath=/opt/docker/stacks/crafty-controller/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "docmost" deployed=false composePath=/opt/docker/stacks/docmost/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "emby" deployed=false composePath=/opt/docker/stacks/emby/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "ghost" deployed=false composePath=/opt/docker/stacks/ghost/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gitea" deployed=false composePath=/opt/docker/stacks/gitea/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "glance" deployed=false composePath=/opt/docker/stacks/glance/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gokapi" deployed=false composePath=/opt/docker/stacks/gokapi/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "grafana" deployed=false composePath=/opt/docker/stacks/grafana/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "gramps-web" deployed=false composePath=/opt/docker/stacks/gramps-web/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "home-assistant" deployed=false composePath=/opt/docker/stacks/home-assistant/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homebox" deployed=false composePath=/opt/docker/stacks/homebox/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "homepage" deployed=false composePath=/opt/docker/stacks/homepage/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "immich" deployed=false composePath=/opt/docker/stacks/immich/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "jellyfin" deployed=false composePath=/opt/docker/stacks/jellyfin/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "kimai" deployed=false composePath=/opt/docker/stacks/kimai/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "komga" deployed=false composePath=/opt/docker/stacks/komga/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "mealie" deployed=false composePath=/opt/docker/stacks/mealie/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "n8n" deployed=false composePath=/opt/docker/stacks/n8n/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "navidrome" deployed=false composePath=/opt/docker/stacks/navidrome/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "nextcloud" deployed=false composePath=/opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "onlyoffice" deployed=false composePath=/opt/docker/stacks/onlyoffice/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "opengist" deployed=false composePath=/opt/docker/stacks/opengist/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "outline" deployed=false composePath=/opt/docker/stacks/outline/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "paperless-ngx" deployed=false composePath=/opt/docker/stacks/paperless-ngx/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "papra" deployed=false composePath=/opt/docker/stacks/papra/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plant-it" deployed=false composePath=/opt/docker/stacks/plant-it/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "plex" deployed=false composePath=/opt/docker/stacks/plex/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "privatebin" deployed=false composePath=/opt/docker/stacks/privatebin/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "radarr" deployed=false composePath=/opt/docker/stacks/radarr/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "rallly" deployed=false composePath=/opt/docker/stacks/rallly/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "recipe-importer" deployed=false composePath=/opt/docker/stacks/recipe-importer/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "romm" deployed=false composePath=/opt/docker/stacks/romm/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "seerr" deployed=false composePath=/opt/docker/stacks/seerr/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sonarr" deployed=false composePath=/opt/docker/stacks/sonarr/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "sparkyfitness" deployed=false composePath=/opt/docker/stacks/sparkyfitness/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "tandoor" deployed=false composePath=/opt/docker/stacks/tandoor/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "termix" deployed=false composePath=/opt/docker/stacks/termix/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "traefik" deployed=false composePath=/opt/docker/stacks/traefik/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "uptime-kuma" deployed=true composePath=/opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vaultwarden" deployed=false composePath=/opt/docker/stacks/vaultwarden/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "vikunja" deployed=false composePath=/opt/docker/stacks/vikunja/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wanderer" deployed=false composePath=/opt/docker/stacks/wanderer/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wger" deployed=false composePath=/opt/docker/stacks/wger/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "wishlist" deployed=false composePath=/opt/docker/stacks/wishlist/docker-compose.yml +2026/09/21 11:10:32 manager.go:563: [DEBUG] [stacks] ScanStacks: found stack "zipline" deployed=false composePath=/opt/docker/stacks/zipline/docker-compose.yml +2026/09/21 11:10:32 manager.go:1620: [DEBUG] [stacks] getCatalogTemplateSlugs: found 53 template slugs in /opt/docker/felhom-controller/data/catalog-cache/templates +2026/09/21 11:10:32 manager.go:630: [DEBUG] [stacks] ScanStacks: catalog has 53 template slugs for orphan detection +2026/09/21 11:10:32 manager.go:661: [INFO] [stacks] ScanStacks complete: 55 stacks found (1 deployed, 54 available) +2026/09/21 11:10:32 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:32 router.go:369: [INFO] [api] Stack rescan completed: 55 stacks found +2026/09/21 11:10:33 auth.go:142: [DEBUG] [web] auth: valid session for GET /api/stacks/uptime-kuma +2026/09/21 11:10:33 router.go:81: [DEBUG] [api] GET /api/stacks/uptime-kuma (path=/stacks/uptime-kuma) +2026/09/21 11:10:42 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:10:42 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:10:42 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:10:42 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:10:42 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:42 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m35s ago, effective interval 5m0s, healthy=true +2026/09/21 11:10:42 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:52 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:10:52 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:10:52 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:10:52 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m45s ago, effective interval 5m0s, healthy=true +2026/09/21 11:10:52 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:10:55 auth.go:142: [DEBUG] [web] auth: valid session for GET /apps/uptime-kuma +2026/09/21 11:10:55 server.go:542: [DEBUG] [web] ServeHTTP: GET /apps/uptime-kuma from 172.18.0.3:56160 +2026/09/21 11:10:56 auth.go:142: [DEBUG] [web] auth: valid session for GET /stacks +2026/09/21 11:10:56 server.go:542: [DEBUG] [web] ServeHTTP: GET /stacks from 172.18.0.3:56160 +2026/09/21 11:10:57 auth.go:142: [DEBUG] [web] auth: valid session for GET /apps/uptime-kuma +2026/09/21 11:10:57 server.go:542: [DEBUG] [web] ServeHTTP: GET /apps/uptime-kuma from 172.18.0.3:56160 +2026/09/21 11:10:58 auth.go:142: [DEBUG] [web] auth: valid session for GET /stacks +2026/09/21 11:10:58 server.go:542: [DEBUG] [web] ServeHTTP: GET /stacks from 172.18.0.3:56160 +2026/09/21 11:11:02 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:11:02 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 1m55s ago, effective interval 5m0s, healthy=true +2026/09/21 11:11:02 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:11:02 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:11:02 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:11:06 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/uptime-kuma/update +2026/09/21 11:11:06 router.go:81: [DEBUG] [api] POST /api/stacks/uptime-kuma/update (path=/stacks/uptime-kuma/update) +2026/09/21 11:11:06 router.go:582: [INFO] [api] update requested for stack: uptime-kuma +2026/09/21 11:11:06 router.go:81: [DEBUG] [api] actionStack: action=update name=uptime-kuma +2026/09/21 11:11:06 update.go:292: [ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 2026-09-21T11:09:07Z}] catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) +2026/09/21 11:11:07 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/uptime-kuma/update +2026/09/21 11:11:07 router.go:81: [DEBUG] [api] POST /api/stacks/uptime-kuma/update (path=/stacks/uptime-kuma/update) +2026/09/21 11:11:07 router.go:582: [INFO] [api] update requested for stack: uptime-kuma +2026/09/21 11:11:07 router.go:81: [DEBUG] [api] actionStack: action=update name=uptime-kuma +2026/09/21 11:11:07 update.go:292: [ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 2026-09-21T11:09:07Z}] catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) +2026/09/21 11:11:08 auth.go:142: [DEBUG] [web] auth: valid session for POST /api/stacks/uptime-kuma/update +2026/09/21 11:11:08 router.go:81: [DEBUG] [api] POST /api/stacks/uptime-kuma/update (path=/stacks/uptime-kuma/update) +2026/09/21 11:11:08 router.go:582: [INFO] [api] update requested for stack: uptime-kuma +2026/09/21 11:11:08 router.go:81: [DEBUG] [api] actionStack: action=update name=uptime-kuma +2026/09/21 11:11:08 update.go:292: [ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 sha256:a8610b3b4c38077922ba51b036691e06887d7cefd91fe620fd3d6d23d03dc240 2026-09-21T11:09:07Z}] catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0]) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3967MB avail≈21931MB +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26008684 → total=25898MB avail=21931MB used=3967MB (15.3%) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.9GB avail=57.9GB (10.2%) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.1GB (1.8%) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.08 1.02 0.95 3/1183 1190" → 1m=1.08 5m=1.02 15m=0.95 +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 52.9°C (hwmon1) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] GetInfo done in 59ms — mem=3967MB/25898MB (15.3%), rootDisk=6.9GB/68.4GB (10.2%), load=1.08/1.02/0.95, temp=52.9°C (hwmon1), cpu=7.2% +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job status-refresh: execution starting +2026/09/21 11:11:12 scheduler.go:346: [INFO] [scheduler] Running job: offsite-credential-retry +2026/09/21 11:11:12 scheduler.go:346: [INFO] [scheduler] Running job: system-health +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job system-health: execution starting +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] GetInfo starting (hddPath="/mnt/felhom-drives/scratch_hdd", hasCPUCollector=true) +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job health-probes: execution starting +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job ring-spill: execution starting +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job offsite-credential-retry: execution starting +2026/09/21 11:11:12 scheduler.go:363: [INFO] [scheduler] Job offsite-credential-retry completed (took 0s) +2026/09/21 11:11:12 scheduler.go:346: [INFO] [scheduler] Running job: backup-cache +2026/09/21 11:11:12 scheduler.go:67: [DEBUG] [scheduler] job backup-cache: execution starting +2026/09/21 11:11:12 manager.go:712: [INFO] [stacks] Status refresh: 3 containers across 55 stacks +2026/09/21 11:11:12 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping uptime-kuma — last check 2m5s ago, effective interval 5m0s, healthy=true +2026/09/21 11:11:12 healthprobe.go:76: [DEBUG] [stacks] RunHealthProbes: collected 0 targets (1 skipped not due, 0 skipped no container) +2026/09/21 11:11:12 backup.go:465: [DEBUG] groupStacksByDrive: /mnt/sys_drive → [uptime-kuma] +2026/09/21 11:11:12 backup.go:1107: [INFO] [backup] Found 0 DB dump files across drives +2026/09/21 11:11:12 dbdump.go:128: [DEBUG] DiscoverDatabases: running docker ps to find database containers +2026/09/21 11:11:12 dbdump.go:147: [DEBUG] DiscoverDatabases: docker ps output: c8741bfa31f7 uptime-kuma uptime-kuma louislam/uptime-kuma:2.5.0 +7c794623b894 felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 +6d206504c99e filebrowser filebrowser gtstef/filebrowser:1.3.3-stable +2515331bc19e traefik traefik traefik:v3.6.7 +2026/09/21 11:11:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container uptime-kuma (image=louislam/uptime-kuma:2.5.0, not a database) +2026/09/21 11:11:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container felhom-controller (image=gitea.dooplex.hu/admin/felhom-controller:0.260.0, not a database) +2026/09/21 11:11:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database) +2026/09/21 11:11:12 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container traefik (image=traefik:v3.6.7, not a database) +2026/09/21 11:11:12 dbdump.go:203: [DEBUG] DiscoverDatabases: found 0 database(s), skipped 4 non-DB container(s) +2026/09/21 11:11:12 dbdump.go:206: [INFO] [backup] Discovered 0 databases +2026/09/21 11:11:12 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/uptime-kuma/docker-compose.yml +2026/09/21 11:11:12 [INFO] [backup] Discovered app data: 1 apps +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/sys_drive" total=68.4 GB used=6.9 GB avail=57.9GB (10.2%) +2026/09/21 11:11:12 recovery_unit.go:224: [INFO] [backup] Recovery unit captured for uptime-kuma → /mnt/sys_drive/felhom-data/backups/primary/uptime-kuma (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) +2026/09/21 11:11:12 backup.go:1177: [INFO] [backup] Backup status cache refreshed +2026/09/21 11:11:12 scheduler.go:363: [INFO] [scheduler] Job backup-cache completed (took 53ms) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readMemInfo: guest cap=25898MB (host total was 30714460KB) → used≈3951MB avail≈21947MB +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readMemInfo: totalKB=30714460 availKB=26027912 → total=25898MB avail=21947MB used=3951MB (15.3%) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/" bsize=4096 total=68.4GB used=6.9GB avail=57.9GB (10.2%) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readDiskUsage: path="/mnt/felhom-drives/scratch_hdd" bsize=4096 total=937.8GB used=17.0GB avail=873.1GB (1.8%) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readLoadAvg: raw="1.08 1.02 0.95 4/1195 1254" → 1m=1.08 5m=1.02 15m=0.95 +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readThermalZones: /sys — found 1 zones +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] readTemperature: found via hwmon at /sys — 52.9°C (hwmon1) +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] GetInfo done in 65ms — mem=3951MB/25898MB (15.3%), rootDisk=6.9GB/68.4GB (10.2%), load=1.08/1.02/0.95, temp=52.9°C (hwmon1), cpu=8.9% +2026/09/21 11:11:12 healthcheck.go:86: [DEBUG] [monitor] Raw values: disk=10.2%, hdd=1.8% (configured=true), mem=15.3% (3951MB/25898MB), cpu=8.9%, temp=52.9°C (hwmon1) +2026/09/21 11:11:12 healthcheck.go:118: [DEBUG] [monitor] SSD disk: OK (10%) +2026/09/21 11:11:12 healthcheck.go:151: [DEBUG] [monitor] Memory: OK (15%) +2026/09/21 11:11:12 healthcheck.go:169: [DEBUG] [monitor] CPU: OK (9%) +2026/09/21 11:11:12 healthcheck.go:187: [DEBUG] [monitor] Temperature: OK (53°C) +2026/09/21 11:11:12 healthcheck.go:201: [DEBUG] [monitor] Docker daemon: OK +2026/09/21 11:11:12 healthcheck.go:209: [DEBUG] [monitor] Checking 1 protected containers: [traefik] +2026/09/21 11:11:12 healthcheck.go:219: [DEBUG] [monitor] All protected containers running +2026/09/21 11:11:12 info.go:13: [DEBUG] [system] GetDiskUsage: path="/mnt/felhom-drives/scratch_hdd" total=937.8 GB used=17.0 GB avail=873.1GB (1.8%) +2026/09/21 11:11:12 healthcheck.go:238: [INFO] [monitor] Health check: status=ok +2026/09/21 11:11:12 healthcheck.go:242: [DEBUG] [monitor] Final status: ok (issues=0, warnings=0, info=5) +2026/09/21 11:11:12 scheduler.go:363: [INFO] [scheduler] Job system-health completed (took 140ms) +2026/09/21 11:11:12 infra.go:67: [INFO] [infra] cloudflared skipped — no cf_tunnel_token configured (LAN-only node) diff --git a/documentation/audits/update-arc-2026-09-21/20-teardown.txt b/documentation/audits/update-arc-2026-09-21/20-teardown.txt new file mode 100644 index 00000000..c5f66012 --- /dev/null +++ b/documentation/audits/update-arc-2026-09-21/20-teardown.txt @@ -0,0 +1,117 @@ +### STEP 8 — TEARDOWN + +== LAYER 1 — the guest (9202): remove the throwaway app through the product == +-- POST /api/stacks/uptime-kuma/remove {remove_hdd_data:true, remove_backups:true} -- +{"ok":false,"error":"stack \"uptime-kuma\" is still running — stop it first before removing"} + +HTTPCODE:409 + -> 409: the product requires a stop first. Doing that through the product too. +-- POST /api/stacks/uptime-kuma/stop -- +{"ok":true,"message":"Stack uptime-kuma stop completed"} + +HTTPCODE:200 + +-- POST /api/stacks/uptime-kuma/remove -- +{"ok":true,"data":{"removed":"uptime-kuma","volumes_removed":["uptime-kuma_uptime_kuma_data"],"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem tárolt saját adatot külső meghajtón, így ott nem volt mit törölni.","backup_paths_removed":["/mnt/sys_drive/felhom-data/backups/primary/uptime-kuma (24K)"]},"message":"Stack uptime-kuma removed"} + +HTTPCODE:200 + +-- verify the guest is back to how I found it -- +felhom-controller | gitea.dooplex.hu/admin/felhom-controller:0.260.0 | Up 5 minutes (healthy) +filebrowser | gtstef/filebrowser:1.3.3-stable | Up 5 minutes (healthy) +traefik | traefik:v3.6.7 | Up 5 minutes + stack dir /opt/docker/stacks/uptime-kuma : STILL-PRESENT + docker volume uptime-kuma_uptime_kuma_data : 0 + deployed stacks now (expect only protected traefik, as at baseline): + total templates: 55 + traefik deployed= False protected= True orphaned= False + + update-journal.json: ABSENT + +== LAYER 2 — the host (demo-hp): guest 9202 left RUNNING, 9201 untouched == +status: running + 9201 -> status: running + (9201 was never a target: no pct exec, no stop, no start was issued against it this session) + helper files I left on the host/guest, now removed: + cleaned + +== LAYER 3 — the hub == + Nothing provisioned. Guest 9202 has hub reporting OFF and no tunnel, so no host, no escrow, + no report and no customer record was created at any point. Nothing to tear down. + +-- NOTE: /opt/docker/stacks/uptime-kuma survived the remove. Contents: -- +total 20 +drwxr-xr-x 2 root root 4096 Sep 21 11:11 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 3318 Sep 21 11:04 .felhom.yml +-rw-r--r-- 1 root root 1597 Sep 21 11:08 applied-compose.yml +-rw-r--r-- 1 root root 1597 Sep 21 11:08 docker-compose.yml +-- is this normal for this box? compare with another never-deployed app's dir -- +actualbudget +adventurelog +audiobookshelf +bentopdf +bookstack +calcom +calibre-web +claper +code-server +crafty-controller +docmost +emby +filebrowser +ghost +gitea +glance +gokapi +grafana +gramps-web +home-assistant +-- does app.yaml still claim deployed? -- +(no app.yaml — the deploy record is gone) +-- POSITIVE CONTROL: a never-deployed app has the SAME shape, so this is not a leftover -- +total 16 +drwxr-xr-x 2 root root 4096 Sep 13 20:22 . +drwxr-xr-x 57 root root 4096 Sep 13 20:22 .. +-rw-r--r-- 1 root root 3517 Sep 20 14:47 .felhom.yml +-rw-r--r-- 1 root root 1412 Sep 13 20:22 docker-compose.yml + -> uptime-kuma now matches: .felhom.yml + docker-compose.yml + applied-compose.yml, NO app.yaml. + -> the pre-update-compose.yml / pre-update-applied.yml copies the update left are GONE. + CORRECTION to the line above: applied-compose.yml is NOT present on both. actualbudget has only + .felhom.yml + docker-compose.yml; uptime-kuma additionally retains applied-compose.yml after the + remove. So the remove leaves ONE residual file that a never-deployed app does not have. + Stated as observed, not diagnosed: I did not establish whether that is intended (a record of what + was last applied) or a small gap in RemoveStack. It carries no customer data. Not called a defect + on one observation. + +== CATALOG REPO — back to the pre-bump image, clean, pushed == +-- git log -- +ff9717d REVERT the drill bump: uptime-kuma back to 2.4.0 (update-arc measurement finished) +89304ab DRILL: uptime-kuma 2.4.0 -> 2.5.0 for the update-arc measurement (reverted in this session) +bd22749 REPORT for the R-469 rule lift +-- working tree -- + (empty above = clean) +-- HEAD vs origin/main -- + HEAD=ff9717d3794974724e09f8fd58abf058d4cdc2d0 + origin/main=ff9717d3794974724e09f8fd58abf058d4cdc2d0 +-- the actual image: line -- +11: image: louislam/uptime-kuma:2.4.0 +-- catalog_since (deliberately today, not 2026-07-18 — the gate's rule) -- +13:catalog_since: "2026-09-21" + +-- diff of the whole drill against the pre-drill commit 5ff36d0 -- +-catalog_since: "2026-07-18" ++catalog_since: "2026-09-21" + (image line identical to pre-drill; only catalog_since differs, as the gate requires) +-- and the box's cached copy agrees -- + image: louislam/uptime-kuma:2.4.0 + +== FINAL STATE CHECK == +9202 status: status: running +felhom-controller | gitea.dooplex.hu/admin/felhom-controller:0.260.0 | Up 6 minutes (healthy) +filebrowser | gtstef/filebrowser:1.3.3-stable | Up 6 minutes (healthy) +traefik | traefik:v3.6.7 | Up 6 minutes +credential file /tmp/.ctlpw on DooPlex: removed + +NOTE: 00-drift.py and 00-upstream-drift-raw.txt in this directory are NOT mine — they predate + my first write (12:53 vs 12:56) and belong to another session. Files 01-20 are mine. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 2cfb4fff..36704a52 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -678,19 +678,19 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | -| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** | -| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | -| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** | +| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. | **WAITING-ON-OPERATOR — `09` §3b Q6; owner: CC once answered** | +| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. | **WAITING-ON-OPERATOR — `09` §3b Q1–Q4 (own-edge half SHIPPED); owner: CC once answered** | +| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. | **WAITING-ON-OPERATOR — `09` §3b Q7; owner: CC once answered** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | | **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** | | **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** | -| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** | +| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | | **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 | **READY — rank P2-MEDIUM; owner: CC** | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** | | **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | -| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** | +| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** | | **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** | @@ -714,10 +714,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** | | **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | -| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** | +| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. | **CLOSED 2026-09-21 — measured on a real version change** | | **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** | | **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** | -| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. **CLOSED 2026-09-21 — controller v0.260.0.** `stacks.CatalogOrder` (`internal/stacks/updateorder.go`) replaces the three-way comparison with FOUR verdicts — Unknown / Current / Behind / **Ahead** — and **moves out of `web` so the badge and the refusal read ONE verdict**; `web.compareInstalledToTemplate` is now a thin wrapper. An app AHEAD reads „Naprakész" / "Up to date" with `tag-ok` (the same word and class as level — there is nothing for the household to do) and a title saying why (`badge.update.ahead.title`, both bundles). `Manager.UpdatePreflight` refuses with reason `downgrade`, HTTP 409, „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.", logged with both image maps. **The API now renders update refusals through `errText`**, so the new key is not a seam built and never wired. **Ahead is NARROW on purpose:** every differing service must be orderable AND newer, or the verdict falls back to Behind — this gate can BLOCK an update, so it errs towards letting one run. Ordering is `util.Version.Compare` (the house rule: one comparator) behind a tag normaliser — `X.Y`/`X.Y.Z`, optional leading `v`, two-part padded with `.0`, and a trailing suffix that must be IDENTICAL on both sides, so `nextcloud:31.0.14-apache → 31.0.15-apache` orders while `postgres:16-alpine`, `26.05.2-ls310 → -ls311`, `kimai/kimai2:apache-2.57.0`, a date stamp and a digest pin do not. **The suffix rule was found by the fixture, not by design** — the first implementation called every real catalog tag unorderable. **Three red-proofs, each SEEN to fail.** Recorded as `09` §3 decision 10 (decided by CC unattended — operator may reverse). | **CLOSED 2026-09-21 — controller v0.260.0** | | **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. **CLOSED 2026-09-16 — the grant works, proven END TO END on a fresh box.** After the narrow grant (DatastoreAdmin for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only — `DatastorePowerUser` was measured to carry Backup+Prune and PBS has no custom roles), a newly installed box for the same rebuilt customer hit the very refusal this row describes, by itself: „pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit Re-issue PBS credentials action" (19:01 CEST). Pressing that action then SUCCEEDED: „tenantsync: reissue ok … token_id=felhom@pbs!tester-1", „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, gen 2; fresh consume-once secret stored)", „pbs token secret consumed by host … (single-use)" — no permission error. This morning the identical action returned „missing Datastore.Modify … status 255" → 502. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt` and `phaseC-reissue.txt`. | **CLOSED 2026-09-16 — grant given and proven end to end** | | **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom. címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. **CLOSED 2026-09-16 — shipped in ISO 1.28.0, published the same day on the operator's explicit yes.** `felhom-bootstrap.sh` prints `print_bound_banner` the moment the bind delivery lands: „a doboz össze van kötve", „a beállítás magától folytatódik", „ezen a gépen nincs több teendőd" — replacing the pairing code on the console. The payload in the published image is byte-identical to repo HEAD and the string is present in it; the boot menu and the install were walked on that exact file. **What it deliberately does NOT do, recorded rather than implied away:** it does not name the dashboard URL (the one-shot bind delivery carries the customer id, passphrase and mode — not the domain), and it does not reflect the later CLAIM, because this unit has exited by then. **And the new banner was never SEEN on a screen** — the box bound itself while the walk was driving it headlessly, so the proof is the shipped payload plus the gate, not a photograph. | **CLOSED 2026-09-16 — shipped in ISO 1.28.0 (published); on-screen effect not photographed** | | **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks//app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0.** `app_deploy_started` is emitted beside the 202; `app_deployed` now fires from the async path's own end, and `app_deploy_failed` (warning) replaces the silence an interrupted install used to get. Both new types are registered in `allowedEventTypes` AND `customerMessages`. The accept-time `app.yaml` is deliberately NOT deleted on failure — it is the crash-safe record with `Deployed:false` and it holds the settings the customer typed; the state every surface reads is `not_deployed`. Red-proofs: the accept-time call back → `TestDeployAcceptance_DoesNotClaimTheAppIsInstalled` fails; the success hook removed → `TestDeployDoneHook_...` fails at „the deploy ended and nothing was told about it". | **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0** | @@ -764,7 +764,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-586** | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** FOUND 2026-09-18 starting slice 4 (R-559): running `scripts/iso/test/bootstrap-modes.sh` unchanged at `183727db9c44` reported `FAIL: R-496: banner painted to the console seam` and `FAIL: R-496: banner names the Tulajdonosi jelmondat`. **Cause:** the script paints each banner with `> "$CONSOLE_DEV"`. On a real console that is a character device and truncation is a no-op, so every paint appears; pointed at a plain file, as the harness did, each paint TRUNCATES. Commit `c033b3b` (ISO 1.28.0, R-535, 2026-09-16) added `print_bound_banner`, which paints immediately after the pairing banner in the same invocation — from that commit the pairing banner was wiped before the check read it. `c033b3b` did not touch the harness. **Why it survived: the harness is in NO gate and NO CI run** — not in `repo_gates.py`, not in `.gitea/`; it runs only when a person runs it, and between 09-16 and 09-18 nobody did. FIXED in the same session: the harness points `FELHOM_CONSOLE_DEV` at a FIFO with a background reader, which restores device semantics (opening a FIFO with `>` truncates nothing) and lets a test see EVERY paint — which slice 4's golden checks then needed anyway. Production code untouched. **What is still open: the harness remains outside every gate.** It needs a container, so it cannot join `repo_gates.py --fast`, which is what both the pre-push hook and CI run — meaning a non-fast entry would still never execute. **Fix shape:** either give CI a container-capable job that runs it, or make the ISO release gate's G16 the place it is required (done for G16 this session — so it now runs at least once per ISO release, which is better than never but later than a push). | **READY - rank P2-MED; owner: CC** | | **R-587** | **[P3-LOW] Two root-password files sit in the directory the public ISO is published FROM.** FOUND 2026-09-18 running the ISO release gate's credential criteria for slice 4: `/mnt/5_hdd/felhom.eu/felhom-iso/out/` holds `felhom-pve-9.2-1-v1.24.0-nested-probe-generic.iso.rootpw.txt` and `felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso.rootpw.txt` from 2026-07-22/23, mode 0600, left by two appliance-mode builds. **They are NOT on the bucket** — both `https://iso.felhom.eu/` return **404**, against a control (`felhom-installer-1.28.0-pve9.2-1.iso.sha256`) that returns 200, so the check is real and not a dead probe. **The only thing keeping them off a PUBLIC bucket is the `--include "felhom-installer-*"` pattern in the publish command**, and they are named `felhom-pve-*`, so the pattern misses them by an accident of naming rather than by design. A publish typed without the include, or an include widened to `felhom-*`, uploads root passwords to a world-readable bucket. **Fix shape:** shred the two files (they are three months old and their VMs are long gone), and make the release build refuse to run — or the publish step refuse to start — while any `*.rootpw.txt` exists in the out directory. A pattern that protects by coincidence is not a control. **The two files were SHREDDED 2026-09-18** from the publish source directory, immediately before the 1.29.0 upload ran from it. The directory now holds no `*.rootpw.txt`. **The mechanism half is still open**: nothing stops the next appliance build leaving one there, and nothing refuses a publish while one exists — the `--include` pattern still protects by coincidence. | **PARTLY DONE 2026-09-18 (files gone; the guard is not built) - rank P3-LOW; owner: CC** | | **R-588** | **[P3-LOW] ISO release records live in two different places, so "was the gate run for this image?" cannot be answered by looking.** FOUND 2026-09-18: the release records for 1.27.0 and 1.27.1 are directories under `documentation/tests/iso-release--/`, and there is none for **1.28.0** — which led me to conclude its gate had not been run. It had: the record is `documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt`, inside an audit about something else. **The gate's own rule is "a criterion with no recorded observation is a criterion that was not run"**, and that rule is unenforceable while the records have no single home — the question it answers has to be settled by a full-text search for a checksum, which is what it took here. **Fix shape:** one home, `documentation/tests/iso-release--/`, and a line in `iso-release-gate.md`'s Result-recording section naming it; optionally a check that every published ISO version has a record directory. | **READY - rank P3-LOW; owner: CC** | -| **R-589** | **[P3-LOW] The update badge is Hungarian on an English app page — and it is the badge, not a corner case.** FOUND 2026-09-20 live on demo-hp (0.257.0) while proving localisation slice 5's pilot: on `/apps/privatebin?lang=en` the only Felhom-authored Hungarian left, apart from the language picker naming itself, was the tag „Naprakész" with the title „Ez az alkalmazás a legfrissebb elérhető változatot futtatja." `controller/internal/web/updatebadge.go` `updateBadgeAt` builds a `MetaBadge` from four RAW Hungarian literals (L82, L84, L87 + the „ — %d napja" suffix, L98-99) with no key, so `executeTemplate`'s language set never sees them. `hu.json` already carries `backups.naprakesz` → 'Up to date' for a DIFFERENT surface, so the English word exists and this producer simply does not use it. Slice 1 listed „Naprakész" among what stays Hungarian (10 §2.3) and slice 2 was the slice that was to take it; slice 2 is CLOSED and it is still there. **Fix shape:** four keys, `MetaBadge.LabelKey`/`TitleKey` (or `Msgf` at the handler, as R-566's page titles did — the „N napja" suffix has a count and needs the same `%d` treatment), with a render test per branch in both languages. Small. | **READY - rank P3-LOW; owner: CC (controller)** | +| **R-589** | **[P3-LOW] The update badge is Hungarian on an English app page — and it is the badge, not a corner case.** FOUND 2026-09-20 live on demo-hp (0.257.0) while proving localisation slice 5's pilot: on `/apps/privatebin?lang=en` the only Felhom-authored Hungarian left, apart from the language picker naming itself, was the tag „Naprakész" with the title „Ez az alkalmazás a legfrissebb elérhető változatot futtatja." `controller/internal/web/updatebadge.go` `updateBadgeAt` builds a `MetaBadge` from four RAW Hungarian literals (L82, L84, L87 + the „ — %d napja" suffix, L98-99) with no key, so `executeTemplate`'s language set never sees them. `hu.json` already carries `backups.naprakesz` → 'Up to date' for a DIFFERENT surface, so the English word exists and this producer simply does not use it. Slice 1 listed „Naprakész" among what stays Hungarian (10 §2.3) and slice 2 was the slice that was to take it; slice 2 is CLOSED and it is still there. **Fix shape:** four keys, `MetaBadge.LabelKey`/`TitleKey` (or `Msgf` at the handler, as R-566's page titles did — the „N napja" suffix has a count and needs the same `%d` treatment), with a render test per branch in both languages. Small. **NOT A DEFECT WHEN THE ROW WAS READ — ALREADY FIXED, AND THE ROW WAS THE STALE PART. VERIFIED 2026-09-21 at controller `19ef0329ab66`.** The four Hungarian literals in `updatebadge.go` are REAL and DELIBERATE: they are the localisation parity guarantee (templateFuncMap's Hungarian output stays byte-identical). The ENGLISH form has been rebuilt from the bundle in `web.localeFuncs`'s `"updateBadge"` entry since **v0.258.0**, with `badge.update.current/.title/.behind/.today/.days/.behind.title` in BOTH `hu.json` and `en.json`, pinned by `TestUpdateBadgeFollowsTheLanguage`, and **proven live on a fresh box the same morning** — `audits/DRILL-first-hour-en-0258-2026-09-20.md` item 9: *"PASS — 'Up to date' in English (R-589, fixed this morning, proven on a fresh box)"*. R-561's own row lists R-589 among the three defects v0.258.0 Part 0 fixed. **The lesson, because it cost a reviewer pass and half a task brief: a reviewer who reads ONE producer cannot see a SECOND producer that overrides it** — reading `updatebadge.go` alone gives exactly the wrong answer. `updatebadge.go` now says so in its own comment, and v0.260.0's new arm was written into BOTH producers with a test that fails if either is missing. | **CLOSED 2026-09-21 — shipped in v0.258.0; the row was stale** | | **R-590** | **[P3-LOW] The data-folder card tells an English household, in Hungarian, whether its files are backed up.** FOUND 2026-09-20 live on demo-hp (0.257.0), same pass as R-589: `/apps/paperless-ngx?lang=en` renders „Ide másold a feldolgozandó fájlokat. Az alkalmazás beolvassa, majd törli innen — ez a mappa átmeneti, és nem készül róla biztonsági mentés." and `/apps/romm?lang=en` renders „Itt tárolódnak a fájljaid. Biztonsági mentés készül róla." `controller/internal/web/datapath_card.go` `consequenceFor` (L33-48) returns four raw Hungarian literals chosen by BACKUP CLASS. The R-75 `label` beside them IS translated now (it is catalog copy), so the English page reads „Documents to read in" followed by a Hungarian sentence — the mixed line is worse than either language whole. **This one carries a promise about the customer's files**, which is why it ranks with R-589 rather than below it: the sentence exists to say a folder is temporary and unbacked, and a household that cannot read it may leave originals in a drop-zone the backup filter drops at every tier. **Fix shape:** four keys, class-keyed exactly as now (the comment's whole point is that the promise is class-driven and never per-app), plus a render test per class in both languages. | **READY - rank P3-LOW; owner: CC (controller)** | | **R-591** | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** FOUND 2026-09-20 while adding the catalog's English overlay: `controller/internal/stacks/manager.go` `Copy()` deep-copies `Meta.DeployFields` (with nested `Options`), `Meta.OptionalConfig` (with nested `Fields`), `Meta.Integrations`, `Meta.HealthCheck` and `Meta.InitialCreds` — and does NOT copy the new `Meta.I18n` map, which the struct assignment leaves shared between the original and the "copy". **It is safe TODAY and that is exactly the shape worth filing:** `Metadata.For` reads the overlay and never writes to it (pinned by `TestForDoesNotMutateTheReceiver`), so nothing can observe the sharing yet. The hole is in the CONTRACT — a function whose whole purpose is "a snapshot the caller may mutate" now has a field that is not one, and the next person to write through an overlay will find a bug with no failing test in front of it. **Fix shape:** deep-copy `I18n` in `Copy()` and pin it with a test that mutates the copy's overlay and asserts the original is unchanged. Alternatively state in `Copy()`'s comment that `I18n` is deliberately shared and immutable, and pin THAT with a test. Either is fine; silence is not. | **READY - rank P3-LOW; owner: CC (controller)** | | **R-592** | **[P3-LOW] Two defects in the new catalog copy gate, each found by its own decoy rather than by reading it.** FOUND AND FIXED 2026-09-20 building `app-catalog-felhom.eu/scripts/check-copy-i18n.py` (R-560 slice 5). (1) **Coverage was counted only for the apps NAMED on the command line**, so `check-copy-i18n.py privatebin` reported 47 more untranslated strings than the same tree unscoped — a ratchet anybody could loosen by naming one app, and either number could have been made to "pass". (2) **The credential check searched for its token as a bare substring**, so a `default_creds` of „admin / adminadmin" translated to "administrator / hunter2" PASSED: "admin" is inside "administrator". (3) **The ASCII-Hungarian stems matched as bare substrings too**, so „ird be" convicted "the third best" and „angol" convicted "Angola". All three now have their own case in `scripts/test_gate_decoys.py` (33 cases, every one seen to convict or pass as intended). **Recorded rather than left in a commit message** because the general form is worth the row: every one of the three was a check that matched a LABEL where the FACT was a word, a login or a whole catalog — the R-421 shapes, inside the gate written to enforce R-421. | **CLOSED 2026-09-20 — fixed in the same session, decoys added** | @@ -778,6 +778,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. | **READY - rank P2-MEDIUM; owner: CC (hub)** | | **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** | | **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | +| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** | +| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** | | **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |