diff --git a/REPORT.md b/REPORT.md index 9ff93ca5..77e11bb5 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,214 +1,63 @@ -# REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01) +# REPORT — update arc slices 1 & 2: documents, register, roadmap, capability map (2026-09-02) -**Task class: Spike.** No production code was written in any repo. Output is a findings doc, register -rows, two architecture updates and one operator decision. +*Overwritten each session. Nothing durable lives only here.* -**Findings doc:** `documentation/audits/SPIKE-app-update-2026-09-01.md` -**Evidence:** `documentation/audits/evidence-spike-app-update-2026-09-01/` (19 files) +## What this session changed in THIS repo ---- - -## 1. The Phase 1 answer, stated unambiguously - -**YES — `docker compose up -d` upgrades an app whose compose file has already moved, and the Restart -button does it.** - -Produced by **variants 1a and 1b, and independently by 1c-ii**: - -- **1a** (target image ABSENT locally): the restart took **18.3 s**, ended with the container on the - new tag, and **left the new image in the local store where it had not been** — so `up -d` performed - a network pull, in a code path that contains no pull step. -- **1b** (target image PRESENT): **0.5 s**, container recreated onto the new tag, image id unchanged. -- **1d, the negative control** (file NOT edited): container id, image id, digest and `StartedAt` all - identical — `up -d` did not even recreate. **The method can show "no change" when nothing changed.** -- **1c-ii** (unattended): the boot reconciler brought a failed app back on the **new** version with - nobody pressing anything. - -**And the answer to 1c is NO, which narrows the exposure:** a hard guest reset upgraded nothing. -Docker's `restart: unless-stopped` restored the existing containers on the old image and the -reconciler logged its own verdict — `Boot reconciliation: no boot-orphaned apps (nothing to start)`. - -**Per the task's own rule, no fix is proposed.** The behaviour may have been chosen — `RestartStack` -says so in a comment — so it goes to the operator as a decision. - ---- - -## 2. Confirmed baselines actually used - -| Repo | at §1 of the task | measured at start | end | moved? | -|---|---|---|---|---| -| felhom-controller | `960d29b0612c` | `960d29b0612c` | `960d29b0612c` | **no — read only** | -| felhom.eu | `1d59353df437` | `1d59353df437` | this session's documents | as planned | -| app-catalog-felhom.eu | `29edad9c5bf4` | `29edad9c5bf4` | `5d8f25f`, **tree identical to `29edad9c5bf4`** | reverted | -| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched | - -**None had moved since the task was written.** Live controller on demo-hp: 0.232.0. - ---- - -## 3. Per-phase results, each with its control - -| Phase | Result | Control | -|---|---|---| -| **0** | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) | -| **1** | **THE GATE: YES.** 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | **1d** — unchanged file, no change at all | -| **2** | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (`BentoPDF`, 4 hits) + negative (`zzz-never-present`, 0) on both customer pages | -| **3a** | Pull failure: HTTP 500, **app untouched and still running** | re-run as a Restart: also 500, app still ran | -| **3b** | **HTTP 200 "update completed" over a crash-looping app**; alarm fired 5m16s later | the alarm is a positive observable (`Event pushed: app_start_failed`), plus the detector heartbeat `1 currently down` | -| **4** | demo-hp and demo-felhom: **zero drift by tag**. But **2 of 6 floating pins have already MOVED upstream** | both fully-pinned images (`romm:5.0.0`, `pdo:2.0.5`) were SAME | -| **5** | Existing safety dump is **database-only**; DB dumps are 48 KB–395 KB; the **file** half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead | -| **6** | **App data CANNOT be rolled back.** Old image refuses to start on migrated data | returning to 32.0.9 restored **both** seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked | - ---- - -## 4. The exact symbols that bring an app back — found by reading - -| path | `file:symbol` | +| file | change | |---|---| -| boot reconciler | `felhom-controller/controller/internal/bootrecon/bootrecon.go:269` — `Reconciler.Run` → `StartStack` | -| its scheduler | `controller/cmd/controller/main.go:2127` — `runBootReconcile`, called at `:450` | -| app-stop guard | `controller/internal/backup/appstop_marker.go:283` — `AppStopGuard.Recover` → `StartStack` | -| its hold-aware wrapper | `controller/cmd/controller/main.go:2008` — `gatedAppStopStarter.StartStack` | -| drive-return gate | `controller/internal/web/intermediary.go:222` — `Server.restartStacks` → `StartStack` | -| guest-boot change | `controller/internal/web/intermediary.go:458` — `Server.processGuestBootChange` | -| quiesce restart | `controller/internal/quiesce/quiesce.go:733` — `Loop.restartAll` | -| off-site reconstitution | `controller/internal/backup/offbox_reconstitute.go:692` — `Manager.ReconstituteFromOffsite` | +| **`documentation/architecture/09-update-architecture.md`** | **CREATED — the deliverable.** Its absence was R-438. A LIVING document: every slice of this arc updates it in the same session. | +| `documentation/tests/VALIDATION-update-slice12-2026-09-02.md` | CREATED — the live evidence, copied off demo-hp at the end of the phase that produced it. | +| `documentation/backlog/OPEN-ITEMS.md` | R-438 and R-440 amended (both stay OPEN); **8 new rows: R-446..R-453**. | +| `documentation/backlog/ROADMAP.md` | the update arc added as one item, naming the capability-map rows it flips. | +| `documentation/architecture/00-capability-map.md` | one new row: what version a box runs, and whether it is behind. | +| `STATUS.md` | one new operator item (9) and a new lead paragraph. | -Full table of all 13 non-API call sites: findings doc §8. +## The architecture document — what it settles ---- +1. **How an update works today, as measured** — `UpdateStack` is `pull` then `up -d --remove-orphans`; + the syncer overwrites a deployed app's compose on a 15-minute cycle with no deployed check; + **thirteen** non-API call sites end in `compose up -d`. +2. **What was chosen, and by whom** — `RestartStack`'s own comment, quoted. **A design decision is not + a defect.** What was never decided is what the syncer does underneath a deployed app. +3. **The three operator rulings of 2026-09-02** — verified backup as a precondition; the support window + runs on how far behind the CATALOG a box is; automatic within a major, never across one. +4. **The vocabulary ruling** — "rollback" is struck. The available shapes are ABORT and RESTORE. +5. **The target shape** — the live compose file becomes derived from a pin in `app.yaml`. +6. **The seven slices**, each with a status line. 1 and 2 are shipped. +7. **Known limitations**, including the floating-tag one. -## 5. Register rows opened, updated or re-ranked +## Register -| row | what | owner | +**194 rows before, 202 after.** Nothing closed, and that is stated rather than implied: R-438 and +R-440 are **amended and stay OPEN** — the mechanism is now documented, not changed — so nothing moved +to `CLOSED-ITEMS.md` and that file is untouched. + +| row | what | state | |---|---|---| -| **R-438** | opened P1-HIGH, then **updated with the live evidence** (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | **VIKTOR rules, CC measures** | -| **R-439** | opened P3-LOW, then **updated — the severity survives but its stated reason was imprecise** (`isOperationalState` counts `restarting`/`degraded` as operational) | CC | -| **R-440** | opened P2-MEDIUM, then **updated: two floating pins have ALREADY moved, with a passing control** | CC | -| **R-441** | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules | -| **R-442** | NEW — **P1-HIGH**: `remove_hdd_data:true` is inert with no `paths.hdd_path`; 128 MB left, API said neither removed nor preserved | CC | -| **R-443** | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules | -| **R-444** | NEW — nothing runs `pct fstrim`; demo-hp's thin pool held ~23.8 GB of freed blocks | CC | -| **R-445** | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements | +| R-446 | „Naprakész" can be FALSE for the 23 floating pins | OPEN, P2-MEDIUM, CC | +| R-447 | slice 3 — make the live compose DERIVED | **BLOCKED** on an operator ruling, P1-HIGH | +| R-448 | slice 4 — a guarded update (subsumes R-443) | READY, P2-MEDIUM | +| R-449 | slice 5 — an upgrade test that runs again | READY, P2-MEDIUM | +| R-450 | slice 6 — version sequence; an engine change gets its own edge | READY, P2-MEDIUM | +| R-451 | slice 7 — a fleet sweep (needs a hub change: no image field is reported) | READY, P3-LOW | +| R-452 | no gate enforces `catalog_since` (`--depth 1` has no parent to diff) | READY, P3-LOW | +| R-453 | **the vaulted dashboard password is stale on BOTH demo boxes** | **WAITING-ON-OPERATOR**, P2-MEDIUM | -**Nothing was closed and nothing was re-ranked.** The ranking of R-438 relative to existing rows is -Viktor's, and I have not moved anything. +## Live validation ---- +Full evidence: `documentation/tests/VALIDATION-update-slice12-2026-09-02.md`. -## 6. Claims in the task that turned out to be wrong, named +**PROVEN LIVE on demo-hp at controller 0.233.0**, through a real production caller (the boot +reconciler — no hand-set state): `bentopdf` recorded 1 service, `bookstack` recorded **2**, keyed by +compose service name, **all three digests matching ground truth read independently beforehand**. +Encrypted secrets byte-identical across the write. Both apps up and healthy; nothing provisioned. -1. **"a power cut … is an unattended three-major-version upgrade"** — not as stated. A plain power cut - upgraded nothing (measured). The unattended upgrade needs *"and the app did not come back"*. - **The exposure is smaller than the operator page claims.** -2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites, nine files. -3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and - `runbooks/target-selection.md:161` says *"No access route from DooPlex"*. Two boxes measured live; - Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history. -4. **R-439's severity reason** — right conclusion, imprecise reason (see §5). -5. **R-440's "23 pins"** — **exactly right** (79 image lines / 53 apps / 66 distinct; 23 with no patch - component). One arguable 24th named rather than rounded away. -6. **The catalog history figures** — spot-checked and **all correct**: 153 commits, 53 apps, 0 files - with upgrade metadata, and all four multi-major bumps confirmed by commit hash. -7. **My own method, corrected in-flight:** the first customer-page search used `grep -o "2.8.6"`, whose - unescaped `.` produced two false hits; `grep -F` gives zero. The controls caught it. +**NOT live-validated: the rendered badge.** The vaulted dashboard password is stale on both demo +controllers (R-453) and there is no operator route to a customer's password. Five attempts are listed +in §4 of the validation file. The render is covered by tests that render the PRODUCTION templates. ---- +## Sibling repos -## 7. Evidence handling - -**Evidence was written directly into `documentation/audits/evidence-spike-app-update-2026-09-01/` on -DooPlex as each phase produced it — before every revert, including the intermediate ones.** The -catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all -happened after their evidence was already off the machine. **Nothing was lost and nothing had to be -reproduced.** - ---- - -## 8. Observations - -1. **The restore path writes the recovery unit's OLD image pin into the live stack dir, and the - catalog syncer overwrites it again within 15 minutes.** The overwrite half is measured on demo-hp; - that the restore writes to that same path is read at `cmd/controller/main.go:2570`, not measured — - both halves are graded as such in the row. **FILED: R-441** - -2. **`remove_hdd_data: true` removed nothing.** 128 MB of app data stayed on the drive while the API - returned 200 with `hdd_paths_removed:null, hdd_paths_preserved:null`. Root cause established with - controls: `Paths.HDDPath` has no default and demo-hp's `controller.yaml` does not set it, so - `ParseComposeHDDMounts` returns nil on its first line. **FILED: R-442** - -3. **An update can report success over an app it has just broken.** HTTP 200 and "updated - successfully" while the container was already crash-looping; the truth reached the customer 5m16s - later through the dead-app alarm rather than through the update itself. **FILED: R-443** - -4. **The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed**, and `fstrim` from - inside the unprivileged container is refused; `pct fstrim` from the host reclaimed it. Nothing runs - it on the fleet. **FILED: R-444** - -5. **The hub kept app telemetry for an app that no longer exists anywhere**, and it now sets a - fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway. - Retained rather than cleared, because the reset is irreversible and on the operator's own surface. - **FILED: R-445** - -6. **The syncer's debug hash line cannot show what it claims to show.** `logFileHashes` - (`internal/sync/sync.go:386`) reads the destination *after* the write, so it printed - `src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word - "changed". **NOT-A-FINDING:** it is DEBUG-only and the `Updated /` INFO line immediately - above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the - next person reading a sync log does not try to learn from those two hashes, which cannot differ. - ---- - -## 9. Teardown — all three layers - -**Layer 1 — the machine.** `bentopdf` restored to its catalog tag `v2.8.6`, digest -`sha256:eaeea1e447205a79…`, **byte-identical to the run's baseline**. The throwaway `nextcloud` stack -removed via the product's own endpoint; containers, all three volumes and `app.yaml` gone. The 128 MB -the product failed to remove (observation 2) deleted by hand along with its backup dirs; -`find /mnt -iname "*nextcloud*"` returns nothing. Five images this run pulled removed by **targeted -`docker rmi`** — **no `prune` of any kind was run anywhere**. No guest was created; guest 9201 was -hard-reset once by design and returned with all nine apps. - -**Layer 2 — the host.** `local-lvm` 68.97% → 70.91% during the run → **26.78%** after `pct fstrim -9201`. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to -pre-run values (`/` 957 M, `/mnt/sys_drive` 12 G, `hdd_1` 5.5 G). - -**Layer 3 — the hub. This run provisioned NOTHING.** No customer record and no appliance record was -created; the customers list is unchanged at five rows, identical to the list read at the start. The -existing `demo-hp` customer was used. What the run *did* create is hub **events** (`app_start_failed`, -plus deploy/remove for the throwaway app) — **retained deliberately**, because the event log is an -append-only record and deleting from it to tidy a test damages the surface this project relies on for -history. One residue is **retained rather than cleared** with the reason and the exact one-line command -recorded in the findings doc §13 (observation 5). - -**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified. Peti's box was never contacted.** - ---- - -## 10. Final verification - -``` -felhom-controller: git status --porcelain → empty - HEAD = origin/main = 960d29b0612c - go build ./... → OK - go vet ./... → OK - go test ./... → rc=0, 28 packages, 0 FAIL -``` - -**The controller tree was left untouched, and that is proven rather than asserted.** - ---- - -## 11. My own mistakes - -1. **My first customer-page search used a regex where I needed a literal.** `grep -o "2.8.6"` treats - `.` as a wildcard and reported two hits on a page that contains none. I caught it only because the - task requires a control on every such search, and re-ran with `grep -F`. **The rule earned its - place in the same session it was applied.** -2. **My first poll for "nextcloud is healthy" matched the wrong container.** The break condition - matched `healthy` anywhere in the line and fired on `nextcloud-redis`. Corrected to an exact-name - filter. No result depended on it. -3. **I renumbered `STATUS.md` badly on the first attempt**, leaving the list running 1–6, 8, 9 with no - item 7, and left the section's own lead line saying one thing was waiting when there were two. - Both fixed. **This is the ranking-paragraph-goes-stale trap (R-405) in miniature**, and it appeared - within minutes of my writing about it. +- `felhom-controller` **v0.233.0** — `8025304acc0a`, deployed to demo-hp and verified healthy. +- `app-catalog-felhom.eu` — `69761cf91bfc` (backfill) + `8220f8d` (REPORT). diff --git a/STATUS.md b/STATUS.md index b7400b8a..511b54bd 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,11 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down +**Updated 2026-09-02 — the box now writes down which version of each app it is running, and shows one +small label saying whether it is up to date. No version numbers, and nothing about updating changed. +The writing-down half is proven on the real machine; for the label I need one password from you +(item 9). Nothing is broken while it waits.** + +**Earlier 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1). ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.** @@ -18,7 +23,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision, new today).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on +1. **Three things are waiting on you — item 4 (send two e-mails), item 7 (one design decision) and item 9 (one password, new today, two minutes).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -111,6 +116,36 @@ nothing.* Full measurement, with the controls and the quoted output: `felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`. +9. **I need the dashboard password for `demo-hp`, or your go-ahead to reset it. Two minutes, and + nothing is broken while it waits.** + + Today the box learned to write down which version of each app it is really running, and to show + the customer one small label: **„Naprakész"** or **„Frissítés elérhető — 45 napja"**. No version + numbers — a household cannot act on `26.05.2`. + + **The writing-down half is proven on the real machine.** I made two apps fail to come back, the box + repaired them by itself, and it wrote down exactly what it installed — one line per container, with + the fingerprint that cannot lie. I checked those fingerprints against the machine independently and + they match. Both apps are up and healthy, and no data was touched. + + **The label half I could not look at.** To open a customer page I have to log in as the customer, + and the password we keep on file no longer works — on `demo-hp` **or** on `demo-felhom`. I tried + both, and the box's own log says "wrong password", not "wrong address". There is no operator route + to a customer's password: it is only ever e-mailed to them. + + **Two ways forward. Pick one:** + + - **Send me the current `demo-hp` dashboard password.** I use it for two page loads and nothing + else. Simplest, and it changes nothing on the box. + - **Let me reset it back to the one on file.** Same thing we did on 9 August. It also repairs the + stored password, so the next session does not lose this half hour again. + + **My pick: let me reset it** — otherwise the same wall is there next week, on both machines. + + **If you do nothing:** nothing breaks and no customer is affected. The label is covered by tests + that render the real pages, so this is a confirmation, not a discovery. The box goes on recording + versions either way. + 8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, as it misled one by an hour. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 0bfa2aab..979b14dc 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -104,6 +104,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | | Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) | +| **What VERSION a box is running, and whether it is behind the catalog** | controller v0.233.0 + catalog `69761cf` | **PROVEN-LIVE for the RECORD; IMPLEMENTED for the BADGE** — and the split is the point, not a hedge. | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: the rendered badge has never been seen on a live page.** The vaulted dashboard password is stale on BOTH demo controllers (`Hibás jelszó`, confirmed against the controller's own log, and the same on demo-felhom), and there is no operator-side route to a customer's dashboard password (R-119). Every INPUT the badge reads is verified live; the render is covered only by tests that render the PRODUCTION templates. **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 | | Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | | | Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) | | App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md new file mode 100644 index 00000000..5b1b3998 --- /dev/null +++ b/documentation/architecture/09-update-architecture.md @@ -0,0 +1,230 @@ +# 09 — How an app update works, and what it is becoming + +> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.** +> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen +> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who +> did not know it was one. + +**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The +source is the truth.** Nothing here is invented: every mechanism claim is cited either to +`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`. + +--- + +## 1. How an update works today, as measured + +### 1.1 The button + +`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose +commands and nothing else: + +``` +compose pull → compose up -d --remove-orphans +``` + +**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live +updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the +failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike +§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443. + +### 1.2 The catalog syncer moves the file underneath a deployed app + +`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies +`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle +(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only +guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is +`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the +sync does not restart anything.** + +That is why a deployed app's compose file and its running containers can disagree **indefinitely**. +Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the +deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3). + +### 1.3 Thirteen other paths end in `compose up -d` + +Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`, +and every one ends in `compose up -d` against the live compose file (spike §8 — the task that +commissioned the spike said five; the count is thirteen). They include the **boot reconciler** +(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return +gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every +restore path. + +**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where +a boot reconciliation started an app on a newer image at 17:55:44Z. + +### 1.4 One fear is measured SMALLER than it was stated + +A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the +existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs +`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant +1c, a positive observable and not an absent log line). + +**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying +so is more useful than leaving the scarier version standing. + +--- + +## 2. What was chosen, and by whom + +`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the +whole arc: + +> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template +> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends +> SIGTERM+start to existing containers without re-reading the compose file or env."* + +**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a +defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer +moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen +unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism; +it does not change it. + +--- + +## 3. The three operator decisions (2026-09-02) + +These are rulings, not proposals. Anything specced against a different assumption is wrong. + +1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the + update path. The guest-snapshot alternative is to be **spiked before anything is designed around + it**. Context: the existing safety machinery (`Manager.writeSafetyDump`, + `internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike + §6 — the file half was never priced, and demo-hp is too young a box to price it. + +2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.** + A customer on the newest version is supported however old that version is. This is why + `catalog_since` exists and why no version string is shown. + +3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human, + because §4 says it cannot be undone. + +--- + +## 4. The vocabulary ruling — "rollback" is struck + +**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually +run, putting the old image tag back produces a container that refuses to start — + +> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and +> downgrading is not supported"* + +with a positive control proving the data is intact, only unreachable by the old version (§7 6d). + +**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are: + +| shape | when it applies | what it does | +|---|---|---| +| **ABORT** | before anything migrated | stop, put the old image back, the app runs again | +| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy | + +There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute +half-life because the syncer overwrites it (**R-441**). + +--- + +## 5. The target shape + +**The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the +syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the +rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths +stop being able to change a version by accident. + +**Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the +prerequisite for judging how urgent it is. + +--- + +## 6. The seven slices + +| # | slice | status | +|---|---|---| +| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** | +| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** | +| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 | +| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 | +| **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 | +| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 | +| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 | + +**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge, +never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in +one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it +breaks. + +--- + +## 7. What slices 1 and 2 actually built + +### 7.1 The record (slice 1) + +`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs +after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`, +and writes `app.yaml`: + +```yaml +installed_images: + web: + ref: lscr.io/linuxserver/bookstack:26.05.2 + digest: sha256:… # "" if the image was never pulled from a registry + at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service +``` + +Three rules, each with its reason: + +- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has + already moved. A record built from it would answer "what will happen next time something runs + `up -d`", which is a different question. +- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`. + Intent refused, observation logged. Refusing to start a customer's app because we could not write + down which version it is trades a real outage for a bookkeeping gap. +- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a + complete record with a partial one. + +**Nothing reads it to take a decision.** Slice 2 reads it to render a label. + +### 7.2 The label (slice 2) + +`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against +what the current template pins, and renders through the existing `meta_badge` partial — no new markup, +no new CSS. + +**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no +record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were +current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by +a test with a companion red-proof. + +**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`). +Version strings stay in the logs, the API and the hub. + +--- + +## 8. Known limitations, stated plainly + +1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and + queries no registry — a customer's box must not depend on reaching eight upstream registries to + render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be + identical while the image behind it has moved. **Measured, not theorised:** spike §5 found + `mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls + holding. Digest-level comparison needs a registry query and is deferred — **R-446**. +2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date + under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent + commit to diff against, so the gate needs a deeper fetch — **R-452**. +3. **The record only appears after the next lifecycle action.** An app that is running and untouched + keeps a legacy `app.yaml` and therefore no badge, until someone restarts, updates or redeploys it. + That is correct — the alternative is inventing a record from the file §1.2 says has already moved — + but it means the fleet view fills in gradually rather than at upgrade. +4. **The hub does not record image tags at all.** Its report's container payload carries name, state, + CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side + change; it is not derivable from what is already reported. + +--- + +## 9. Where the rest lives + +| what | where | +|---|---| +| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` | +| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 | +| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` | +| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` | +| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 901892ae..070ba6e2 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -675,14 +675,22 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** | | **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** | | **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** | -| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** | +| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` **THE DOCUMENT HALF IS NOW DISCHARGED, 2026-09-02: `documentation/architecture/09-update-architecture.md` exists and is a LIVING document, updated by every slice of this arc.** It records the mechanism as measured (§1), quotes the `RestartStack` comment that proves the restart half was CHOSEN (§2), carries the three operator rulings of 2026-09-02 (§3), strikes the word "rollback" (§4), states the target shape (§5) and lists the seven slices with a status each (§6). **THE ROW STAYS OPEN AND THE REASON IS THE POINT: the mechanism is now DOCUMENTED, not CHANGED.** Whether the syncer should go on overwriting a deployed app's compose file is still the operator's ruling, and acting on it is slice 3 (R-447). | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** | | **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | -| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | +| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | | **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** | | **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** | | **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | +| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** | +| **R-447** | **[P1-HIGH] UPDATE ARC SLICE 3 — make the live compose file DERIVED, so the syncer stops changing a deployed app's version.** The target shape (`architecture/09-update-architecture.md` §5): the pin lives in `app.yaml`, the one file `Syncer.copyTemplates` never touches, and the live `docker-compose.yml` becomes an OUTPUT rather than an input. The catalog then proposes and `app.yaml` decides, and the thirteen unattended `compose up -d` call sites (spike §8) stop being able to change a version by accident. **BLOCKED ON AN OPERATOR RULING, and that is the whole reason this is a separate slice:** R-438 established that the restart half of this behaviour was CHOSEN and written down in `Manager.RestartStack`'s own comment. Changing it is not a bug fix; it is reversing a decision, and the decision-maker is Viktor. **Slices 1 and 2 shipped first ON PURPOSE** — the fleet's real state has to be visible before anyone can judge how urgent this is. Do NOT add a deployed check to any of the thirteen paths ahead of the ruling. `architecture/09-update-architecture.md` §5 | **BLOCKED — rank P1-HIGH; owner: VIKTOR rules, CC implements** | +| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** | +| **R-449** | **[P2-MEDIUM] UPDATE ARC SLICE 5 — an upgrade test that proves a real one-major upgrade end to end, INCLUDING the abort path.** Spike §7 performed the pieces by hand on Nextcloud (31.0.14 → 32.0.9 ran the migration; 31.0.14 → 34.0.1 was refused by the app and reported as SUCCESS by the product; the downgrade attempt was refused with the data intact). **None of it is a test that runs again.** A one-off measurement that nothing repeats decays into a claim — this project's most-repeated defect class. The test must assert the CONSEQUENCE (does the app serve after the upgrade? does the abort put it back?), not the mechanism. Venue: a Tier-0 box; `runbooks/target-selection.md` names which. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: CC** | +| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | +| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** | +| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** | +| **R-453** | **[P2-MEDIUM] The vaulted customer dashboard password is STALE ON BOTH DEMO BOXES, and there is no operator-side route to the real one — so no session can drive a customer page.** MEASURED 2026-09-02 while trying to live-validate the update badge: `PASSWORD` from DooPlex `~/.config/credentials` returns **HTTP 200 with the login page and the body string `Hibás jelszó`** against demo-hp guest 9201 (`https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`) **and** against demo-felhom guest 9201 (`https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu`). **The discriminator is the controller's own log, not the status code** — `auth.go:176: [WARN] [web] Failed login` proves wrong PASSWORD rather than wrong Host header, which is the trap this class always presents (a rejected login renders no flash and looks exactly like a routing problem). Every other key in the credentials file was checked and none is a dashboard password (`HUB_PW`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*`, and the `R_*` keys are escrow recovery codes). **THIS IS THE SECOND TIME:** it drifted on demo-hp on 2026-08-09 and was put back on operator instruction by writing a fresh bcrypt hash into the guest's `data/settings.json`; demo-hp was then reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted as well. **Why it is not merely inconvenient: it silently converts "endpoint-level validation" — this project's STANDARD method, because there is no browser on DooPlex — into "unit tests only" for anything that renders a customer page.** The cost is paid per session and rediscovered each time. **The fix is a decision, not a command:** the claim code is bcrypt-hashed hub-side and only e-mailed (R-119), so either the operator records the current demo passwords out-of-band, or a re-set becomes a standing authorisation for the two Tier-0 demo boxes. **NOT TAKEN UNILATERALLY:** re-setting a dashboard password is a decision about a customer account, and it was done under operator instruction last time. Raised in `STATUS.md` item 9. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4, which lists all five attempts. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** |