09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s

R-438's document half. It records how an update works AS MEASURED, quotes the
RestartStack comment that proves the restart half was CHOSEN (a design decision
is not a defect), carries the three operator rulings of 2026-09-02, strikes the
word 'rollback' (once a migration has run the old image will not start), states
the target shape, and lists the seven slices with a status each.

R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not
changed. Nothing closed, so CLOSED-ITEMS.md is untouched.

Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23
floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner),
R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and
R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what
stopped the badge render from being validated live).

Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN
LIVE through the boot reconciler on demo-hp - one entry per compose service,
digests matching ground truth read independently. The badge RENDER is not, and
the five attempts are listed rather than summarised.
This commit is contained in:
2026-09-02 20:32:30 +02:00
parent 56c7e373a3
commit 6035dfcc3a
7 changed files with 497 additions and 203 deletions
+48 -199
View File
@@ -1,214 +1,63 @@
# REPORT — SPIKE: what an app update actually does, and which other paths do it too (2026-09-01)
# REPORT — update arc slices 1 & 2: documents, register, roadmap, capability map (2026-09-02)
**Task class: Spike.** No production code was written in any repo. Output is a findings doc, register
rows, two architecture updates and one operator decision.
*Overwritten each session. Nothing durable lives only here.*
**Findings doc:** `documentation/audits/SPIKE-app-update-2026-09-01.md`
**Evidence:** `documentation/audits/evidence-spike-app-update-2026-09-01/` (19 files)
## What this session changed in THIS repo
---
## 1. The Phase 1 answer, stated unambiguously
**YES — `docker compose up -d` upgrades an app whose compose file has already moved, and the Restart
button does it.**
Produced by **variants 1a and 1b, and independently by 1c-ii**:
- **1a** (target image ABSENT locally): the restart took **18.3 s**, ended with the container on the
new tag, and **left the new image in the local store where it had not been** — so `up -d` performed
a network pull, in a code path that contains no pull step.
- **1b** (target image PRESENT): **0.5 s**, container recreated onto the new tag, image id unchanged.
- **1d, the negative control** (file NOT edited): container id, image id, digest and `StartedAt` all
identical — `up -d` did not even recreate. **The method can show "no change" when nothing changed.**
- **1c-ii** (unattended): the boot reconciler brought a failed app back on the **new** version with
nobody pressing anything.
**And the answer to 1c is NO, which narrows the exposure:** a hard guest reset upgraded nothing.
Docker's `restart: unless-stopped` restored the existing containers on the old image and the
reconciler logged its own verdict — `Boot reconciliation: no boot-orphaned apps (nothing to start)`.
**Per the task's own rule, no fix is proposed.** The behaviour may have been chosen — `RestartStack`
says so in a comment — so it goes to the operator as a decision.
---
## 2. Confirmed baselines actually used
| Repo | at §1 of the task | measured at start | end | moved? |
|---|---|---|---|---|
| felhom-controller | `960d29b0612c` | `960d29b0612c` | `960d29b0612c` | **no — read only** |
| felhom.eu | `1d59353df437` | `1d59353df437` | this session's documents | as planned |
| app-catalog-felhom.eu | `29edad9c5bf4` | `29edad9c5bf4` | `5d8f25f`, **tree identical to `29edad9c5bf4`** | reverted |
| felhom-agent | `4586f0f7f6d1` | `4586f0f7f6d1` | `4586f0f7f6d1` | not touched |
**None had moved since the task was written.** Live controller on demo-hp: 0.232.0.
---
## 3. Per-phase results, each with its control
| Phase | Result | Control |
|---|---|---|
| **0** | R-438, R-439, R-440 filed, each with rank and owner | ids confirmed free by grepping all three register files (R-437 was highest) |
| **1** | **THE GATE: YES.** 1a/1b upgrade; 1c does not; 1c-ii upgrades unattended | **1d** — unchanged file, no change at all |
| **2** | Sync rewrote a DEPLOYED app's file at 17:45:17Z; container unchanged; nothing told the customer | positive (`BentoPDF`, 4 hits) + negative (`zzz-never-present`, 0) on both customer pages |
| **3a** | Pull failure: HTTP 500, **app untouched and still running** | re-run as a Restart: also 500, app still ran |
| **3b** | **HTTP 200 "update completed" over a crash-looping app**; alarm fired 5m16s later | the alarm is a positive observable (`Event pushed: app_start_failed`), plus the detector heartbeat `1 currently down` |
| **4** | demo-hp and demo-felhom: **zero drift by tag**. But **2 of 6 floating pins have already MOVED upstream** | both fully-pinned images (`romm:5.0.0`, `pdo:2.0.5`) were SAME |
| **5** | Existing safety dump is **database-only**; DB dumps are 48 KB–395 KB; the **file** half is the cost, and it does not fit for a large app | demo-hp is too young to price it — stated, and Campaign 10's measured figures cited instead |
| **6** | **App data CANNOT be rolled back.** Old image refuses to start on migrated data | returning to 32.0.9 restored **both** seeded markers byte-identical — the data is not destroyed, only the downgrade is blocked |
---
## 4. The exact symbols that bring an app back — found by reading
| path | `file:symbol` |
| file | change |
|---|---|
| boot reconciler | `felhom-controller/controller/internal/bootrecon/bootrecon.go:269` — `Reconciler.Run` → `StartStack` |
| its scheduler | `controller/cmd/controller/main.go:2127` — `runBootReconcile`, called at `:450` |
| app-stop guard | `controller/internal/backup/appstop_marker.go:283` — `AppStopGuard.Recover` → `StartStack` |
| its hold-aware wrapper | `controller/cmd/controller/main.go:2008` — `gatedAppStopStarter.StartStack` |
| drive-return gate | `controller/internal/web/intermediary.go:222` — `Server.restartStacks` → `StartStack` |
| guest-boot change | `controller/internal/web/intermediary.go:458` — `Server.processGuestBootChange` |
| quiesce restart | `controller/internal/quiesce/quiesce.go:733` — `Loop.restartAll` |
| off-site reconstitution | `controller/internal/backup/offbox_reconstitute.go:692` — `Manager.ReconstituteFromOffsite` |
| **`documentation/architecture/09-update-architecture.md`** | **CREATED — the deliverable.** Its absence was R-438. A LIVING document: every slice of this arc updates it in the same session. |
| `documentation/tests/VALIDATION-update-slice12-2026-09-02.md` | CREATED — the live evidence, copied off demo-hp at the end of the phase that produced it. |
| `documentation/backlog/OPEN-ITEMS.md` | R-438 and R-440 amended (both stay OPEN); **8 new rows: R-446..R-453**. |
| `documentation/backlog/ROADMAP.md` | the update arc added as one item, naming the capability-map rows it flips. |
| `documentation/architecture/00-capability-map.md` | one new row: what version a box runs, and whether it is behind. |
| `STATUS.md` | one new operator item (9) and a new lead paragraph. |
Full table of all 13 non-API call sites: findings doc §8.
## The architecture document — what it settles
---
1. **How an update works today, as measured** — `UpdateStack` is `pull` then `up -d --remove-orphans`;
the syncer overwrites a deployed app's compose on a 15-minute cycle with no deployed check;
**thirteen** non-API call sites end in `compose up -d`.
2. **What was chosen, and by whom** — `RestartStack`'s own comment, quoted. **A design decision is not
a defect.** What was never decided is what the syncer does underneath a deployed app.
3. **The three operator rulings of 2026-09-02** — verified backup as a precondition; the support window
runs on how far behind the CATALOG a box is; automatic within a major, never across one.
4. **The vocabulary ruling** — "rollback" is struck. The available shapes are ABORT and RESTORE.
5. **The target shape** — the live compose file becomes derived from a pin in `app.yaml`.
6. **The seven slices**, each with a status line. 1 and 2 are shipped.
7. **Known limitations**, including the floating-tag one.
## 5. Register rows opened, updated or re-ranked
## Register
| row | what | owner |
**194 rows before, 202 after.** Nothing closed, and that is stated rather than implied: R-438 and
R-440 are **amended and stay OPEN** — the mechanism is now documented, not changed — so nothing moved
to `CLOSED-ITEMS.md` and that file is untouched.
| row | what | state |
|---|---|---|
| **R-438** | opened P1-HIGH, then **updated with the live evidence** (sync overwrite measured; consequence measured; the in-source design intent found and recorded, which narrows it) | **VIKTOR rules, CC measures** |
| **R-439** | opened P3-LOW, then **updated — the severity survives but its stated reason was imprecise** (`isOperationalState` counts `restarting`/`degraded` as operational) | CC |
| **R-440** | opened P2-MEDIUM, then **updated: two floating pins have ALREADY moved, with a passing control** | CC |
| **R-441** | NEW — the restore path and the sync disagree about the image, and the sync wins within 15 minutes | CC measures, VIKTOR rules |
| **R-442** | NEW — **P1-HIGH**: `remove_hdd_data:true` is inert with no `paths.hdd_path`; 128 MB left, API said neither removed nor preserved | CC |
| **R-443** | NEW — the Update button reports success over an app it has broken | CC proposes, VIKTOR rules |
| **R-444** | NEW — nothing runs `pct fstrim`; demo-hp's thin pool held ~23.8 GB of freed blocks | CC |
| **R-445** | NEW — hub app telemetry outlives the app and sets a fleet-wide recommendation | VIKTOR rules, CC implements |
| R-446 | „Naprakész" can be FALSE for the 23 floating pins | OPEN, P2-MEDIUM, CC |
| R-447 | slice 3 — make the live compose DERIVED | **BLOCKED** on an operator ruling, P1-HIGH |
| R-448 | slice 4 — a guarded update (subsumes R-443) | READY, P2-MEDIUM |
| R-449 | slice 5 — an upgrade test that runs again | READY, P2-MEDIUM |
| R-450 | slice 6 — version sequence; an engine change gets its own edge | READY, P2-MEDIUM |
| R-451 | slice 7 — a fleet sweep (needs a hub change: no image field is reported) | READY, P3-LOW |
| R-452 | no gate enforces `catalog_since` (`--depth 1` has no parent to diff) | READY, P3-LOW |
| R-453 | **the vaulted dashboard password is stale on BOTH demo boxes** | **WAITING-ON-OPERATOR**, P2-MEDIUM |
**Nothing was closed and nothing was re-ranked.** The ranking of R-438 relative to existing rows is
Viktor's, and I have not moved anything.
## Live validation
---
Full evidence: `documentation/tests/VALIDATION-update-slice12-2026-09-02.md`.
## 6. Claims in the task that turned out to be wrong, named
**PROVEN LIVE on demo-hp at controller 0.233.0**, through a real production caller (the boot
reconciler — no hand-set state): `bentopdf` recorded 1 service, `bookstack` recorded **2**, keyed by
compose service name, **all three digests matching ground truth read independently beforehand**.
Encrypted secrets byte-identical across the write. Both apps up and healthy; nothing provisioned.
1. **"a power cut … is an unattended three-major-version upgrade"** — not as stated. A plain power cut
upgraded nothing (measured). The unattended upgrade needs *"and the app did not come back"*.
**The exposure is smaller than the operator page claims.**
2. **"Five other code paths end in `compose up -d`"** — **thirteen** non-API call sites, nine files.
3. **"Phase 4 — demo-hp, demo-felhom and Peti's box"** — Peti's box is DOWN, not enrolled, and
`runbooks/target-selection.md:161` says *"No access route from DooPlex"*. Two boxes measured live;
Peti's row is UNKNOWN, with what is knowable taken read-only from the hub and the catalog history.
4. **R-439's severity reason** — right conclusion, imprecise reason (see §5).
5. **R-440's "23 pins"** — **exactly right** (79 image lines / 53 apps / 66 distinct; 23 with no patch
component). One arguable 24th named rather than rounded away.
6. **The catalog history figures** — spot-checked and **all correct**: 153 commits, 53 apps, 0 files
with upgrade metadata, and all four multi-major bumps confirmed by commit hash.
7. **My own method, corrected in-flight:** the first customer-page search used `grep -o "2.8.6"`, whose
unescaped `.` produced two false hits; `grep -F` gives zero. The controls caught it.
**NOT live-validated: the rendered badge.** The vaulted dashboard password is stale on both demo
controllers (R-453) and there is no operator route to a customer's password. Five attempts are listed
in §4 of the validation file. The render is covered by tests that render the PRODUCTION templates.
---
## Sibling repos
## 7. Evidence handling
**Evidence was written directly into `documentation/audits/evidence-spike-app-update-2026-09-01/` on
DooPlex as each phase produced it — before every revert, including the intermediate ones.** The
catalog revert (18:10:29Z), the compose reverts, the guest reset and the Nextcloud teardown all
happened after their evidence was already off the machine. **Nothing was lost and nothing had to be
reproduced.**
---
## 8. Observations
1. **The restore path writes the recovery unit's OLD image pin into the live stack dir, and the
catalog syncer overwrites it again within 15 minutes.** The overwrite half is measured on demo-hp;
that the restore writes to that same path is read at `cmd/controller/main.go:2570`, not measured —
both halves are graded as such in the row. **FILED: R-441**
2. **`remove_hdd_data: true` removed nothing.** 128 MB of app data stayed on the drive while the API
returned 200 with `hdd_paths_removed:null, hdd_paths_preserved:null`. Root cause established with
controls: `Paths.HDDPath` has no default and demo-hp's `controller.yaml` does not set it, so
`ParseComposeHDDMounts` returns nil on its first line. **FILED: R-442**
3. **An update can report success over an app it has just broken.** HTTP 200 and "updated
successfully" while the container was already crash-looping; the truth reached the customer 5m16s
later through the dead-app alarm rather than through the update itself. **FILED: R-443**
4. **The PVE thin pool was holding ~23.8 GB of blocks the guest had already freed**, and `fstrim` from
inside the unprivileged container is refused; `pct fstrim` from the host reclaimed it. Nothing runs
it on the fleet. **FILED: R-444**
5. **The hub kept app telemetry for an app that no longer exists anywhere**, and it now sets a
fleet-wide suggested memory limit for Nextcloud derived from a 15-minute crash-looping throwaway.
Retained rather than cleared, because the reset is irreversible and on the operator's own surface.
**FILED: R-445**
6. **The syncer's debug hash line cannot show what it claims to show.** `logFileHashes`
(`internal/sync/sync.go:386`) reads the destination *after* the write, so it printed
`src=2ebbbda3765b2b21, dst=2ebbbda3765b2b21 (changed)` — the same hash twice, with the word
"changed". **NOT-A-FINDING:** it is DEBUG-only and the `Updated <app>/<file>` INFO line immediately
above it carries the fact correctly, so nothing is lost and no behaviour is wrong. Recorded so the
next person reading a sync log does not try to learn from those two hashes, which cannot differ.
---
## 9. Teardown — all three layers
**Layer 1 — the machine.** `bentopdf` restored to its catalog tag `v2.8.6`, digest
`sha256:eaeea1e447205a79…`, **byte-identical to the run's baseline**. The throwaway `nextcloud` stack
removed via the product's own endpoint; containers, all three volumes and `app.yaml` gone. The 128 MB
the product failed to remove (observation 2) deleted by hand along with its backup dirs;
`find /mnt -iname "*nextcloud*"` returns nothing. Five images this run pulled removed by **targeted
`docker rmi`** — **no `prune` of any kind was run anywhere**. No guest was created; guest 9201 was
hard-reset once by design and returned with all nine apps.
**Layer 2 — the host.** `local-lvm` 68.97% → 70.91% during the run → **26.78%** after `pct fstrim
9201`. The run's ~1.05 GiB was returned and 23.8 GB more that predated it. Guest filesystems back to
pre-run values (`/` 957 M, `/mnt/sys_drive` 12 G, `hdd_1` 5.5 G).
**Layer 3 — the hub. This run provisioned NOTHING.** No customer record and no appliance record was
created; the customers list is unchanged at five rows, identical to the list read at the start. The
existing `demo-hp` customer was used. What the run *did* create is hub **events** (`app_start_failed`,
plus deploy/remove for the throwaway app) — **retained deliberately**, because the event log is an
append-only record and deleting from it to tidy a test damages the surface this project relies on for
history. One residue is **retained rather than cleared** with the reason and the exact one-line command
recorded in the findings doc §13 (observation 5).
**Nothing on `demo-felhom`, `ep0`, DooPlex or Peti's box was modified. Peti's box was never contacted.**
---
## 10. Final verification
```
felhom-controller: git status --porcelain → empty
HEAD = origin/main = 960d29b0612c
go build ./... → OK
go vet ./... → OK
go test ./... → rc=0, 28 packages, 0 FAIL
```
**The controller tree was left untouched, and that is proven rather than asserted.**
---
## 11. My own mistakes
1. **My first customer-page search used a regex where I needed a literal.** `grep -o "2.8.6"` treats
`.` as a wildcard and reported two hits on a page that contains none. I caught it only because the
task requires a control on every such search, and re-ran with `grep -F`. **The rule earned its
place in the same session it was applied.**
2. **My first poll for "nextcloud is healthy" matched the wrong container.** The break condition
matched `healthy` anywhere in the line and fired on `nextcloud-redis`. Corrected to an exact-name
filter. No result depended on it.
3. **I renumbered `STATUS.md` badly on the first attempt**, leaving the list running 1–6, 8, 9 with no
item 7, and left the section's own lead line saying one thing was waiting when there were two.
Both fixed. **This is the ranking-paragraph-goes-stale trap (R-405) in miniature**, and it appeared
within minutes of my writing about it.
- `felhom-controller` **v0.233.0** — `8025304acc0a`, deployed to demo-hp and verified healthy.
- `app-catalog-felhom.eu` — `69761cf91bfc` (backfill) + `8220f8d` (REPORT).
+37 -2
View File
@@ -1,6 +1,11 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
**Updated 2026-09-02 — the box now writes down which version of each app it is running, and shows one
small label saying whether it is up to date. No version numbers, and nothing about updating changed.
The writing-down half is proven on the real machine; for the label I need one password from you
(item 9). Nothing is broken while it waits.**
**Earlier 2026-09-01 (third pass) — the backup work is FINISHED for beta, and I have written down
where it stops. One alarm that was telling you something untrue is fixed and live (hub 0.111.1).
ONE THING is waiting on you: two short e-mails to Hetzner, drafted and ready to paste.**
@@ -18,7 +23,7 @@ not an evening's work.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision, new today).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
1. **Three things are waiting on you — item 4 (send two e-mails), item 7 (one design decision) and item 9 (one password, new today, two minutes).** Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on
the real machines:
- the background job that could delete a live restore's lock now waits its turn — and the check
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
@@ -111,6 +116,36 @@ nothing.*
Full measurement, with the controls and the quoted output:
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`.
9. **I need the dashboard password for `demo-hp`, or your go-ahead to reset it. Two minutes, and
nothing is broken while it waits.**
Today the box learned to write down which version of each app it is really running, and to show
the customer one small label: **„Naprakész"** or **„Frissítés elérhető — 45 napja"**. No version
numbers — a household cannot act on `26.05.2`.
**The writing-down half is proven on the real machine.** I made two apps fail to come back, the box
repaired them by itself, and it wrote down exactly what it installed — one line per container, with
the fingerprint that cannot lie. I checked those fingerprints against the machine independently and
they match. Both apps are up and healthy, and no data was touched.
**The label half I could not look at.** To open a customer page I have to log in as the customer,
and the password we keep on file no longer works — on `demo-hp` **or** on `demo-felhom`. I tried
both, and the box's own log says "wrong password", not "wrong address". There is no operator route
to a customer's password: it is only ever e-mailed to them.
**Two ways forward. Pick one:**
- **Send me the current `demo-hp` dashboard password.** I use it for two page loads and nothing
else. Simplest, and it changes nothing on the box.
- **Let me reset it back to the one on file.** Same thing we did on 9 August. It also repairs the
stored password, so the next session does not lose this half hour again.
**My pick: let me reset it** — otherwise the same wall is there next week, on both machines.
**If you do nothing:** nothing breaks and no customer is affected. The label is covered by tests
that render the real pages, so this is a confirmation, not a discovery. The box goes on recording
versions either way.
8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.
@@ -104,6 +104,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
| Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) |
| **What VERSION a box is running, and whether it is behind the catalog** | controller v0.233.0 + catalog `69761cf` | **PROVEN-LIVE for the RECORD; IMPLEMENTED for the BADGE** — and the split is the point, not a hedge. | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: the rendered badge has never been seen on a live page.** The vaulted dashboard password is stale on BOTH demo controllers (`Hibás jelszó`, confirmed against the controller's own log, and the same on demo-felhom), and there is no operator-side route to a customer's dashboard password (R-119). Every INPUT the badge reads is verified live; the render is covered only by tests that render the PRODUCTION templates. **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 |
| Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | |
| Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) |
| App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven |
@@ -0,0 +1,230 @@
# 09 — How an app update works, and what it is becoming
> **LIVING DOCUMENT. Every slice of the update arc updates this file in the same session.**
> Opened 2026-09-02 with slices 1 and 2. Its absence was **R-438**: the update mechanism was chosen
> deliberately and written down nowhere, which is how a deliberate design gets "fixed" by someone who
> did not know it was one.
**This file carries the REASONING. The register (`backlog/OPEN-ITEMS.md`) carries the work. The
source is the truth.** Nothing here is invented: every mechanism claim is cited either to
`audits/SPIKE-app-update-2026-09-01.md`, which measured it live, or to live source at `file:symbol`.
---
## 1. How an update works today, as measured
### 1.1 The button
`Manager.UpdateStack` (`felhom-controller/controller/internal/stacks/manager.go:1199`) is two compose
commands and nothing else:
```
compose pull → compose up -d --remove-orphans
```
**No safety copy. No rollback. No hold. No verification.** Confirmed by reading and across six live
updates (spike §10 item 6). A pull FAILURE is handled correctly — `UpdateStack` returns after the
failed pull and never reaches `up -d`, so the running app survives untouched (measured twice, spike
§4 3a). A pull that succeeds over an image that then fails to RUN is the bad case, and it is R-443.
### 1.2 The catalog syncer moves the file underneath a deployed app
`Syncer.copyTemplates` (`felhom-controller/controller/internal/sync/sync.go:319`) copies
`docker-compose.yml` and `.felhom.yml` into **every** stack folder on a 15-minute cycle
(`internal/config/config.go:351`, default `15m`). **It has no deployed check of any kind.** Its only
guard is a sha256 content compare in `copyIfChanged` (`sync.go:403`) and its only exclusion is
`app.yaml`. The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else — **the
sync does not restart anything.**
That is why a deployed app's compose file and its running containers can disagree **indefinitely**.
Measured live 2026-09-01: a real catalog pin change travelled the real cycle, the sync rewrote the
deployed app's file at 17:45:17Z, and the container went on running the old image (spike §3).
### 1.3 Thirteen other paths end in `compose up -d`
Excluding the three API actions, **13 call sites across 9 files** call `StartStack` or `RestartStack`,
and every one ends in `compose up -d` against the live compose file (spike §8 — the task that
commissioned the spike said five; the count is thirteen). They include the **boot reconciler**
(`bootrecon.go:269`), the **app-stop guard's** recovery (`appstop_marker.go:283`), the **drive-return
gate** (`intermediary.go:222`), the quiesce restart-after-backup, off-site reconstitution, and every
restore path.
**So an upgrade can happen with nobody pressing anything** — measured, spike §2 variant 1c-ii, where
a boot reconciliation started an app on a newer image at 17:55:44Z.
### 1.4 One fear is measured SMALLER than it was stated
A plain power cut does **not** upgrade anything. Docker's own `restart: unless-stopped` puts the
existing containers back on the OLD image, so the boot reconciler finds no orphan and never runs
`up -d` — it says so in its own words: `no boot-orphaned apps (nothing to start)` (spike §2, variant
1c, a positive observable and not an absent log line).
**The unattended upgrade needs the narrower precondition: *"and the app did not come back."*** Saying
so is more useful than leaving the scarier version standing.
---
## 2. What was chosen, and by whom
`Manager.RestartStack` (`internal/stacks/manager.go:1161`) carries this comment, and it predates the
whole arc:
> *"Use `up -d` instead of bare `restart` so that env vars from app.yaml are injected and any template
> changes (new images, healthchecks) are picked up. Plain `docker compose restart` only sends
> SIGTERM+start to existing containers without re-reading the compose file or env."*
**So the restart behaviour was chosen, deliberately, and written down. A design decision is not a
defect.** What was never decided — and is recorded nowhere — is what happens once the catalog syncer
moves the file underneath a *deployed* app, and whether the choice was meant to extend to the thirteen
unattended call sites. **That gap is R-438, and it stays open**: this document records the mechanism;
it does not change it.
---
## 3. The three operator decisions (2026-09-02)
These are rulings, not proposals. Anything specced against a different assumption is wrong.
1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the
update path. The guest-snapshot alternative is to be **spiked before anything is designed around
it**. Context: the existing safety machinery (`Manager.writeSafetyDump`,
`internal/backup/offbox_reconstitute.go:207`) is **database-only**, which is the headline of spike
§6 — the file half was never priced, and demo-hp is too young a box to price it.
2. **The support window runs on HOW FAR BEHIND THE CATALOG a box is, not on how old its version is.**
A customer on the newest version is supported however old that version is. This is why
`catalog_since` exists and why no version string is shown.
3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human,
because §4 says it cannot be undone.
---
## 4. The vocabulary ruling — "rollback" is struck
**App data CANNOT be rolled back.** Measured on Nextcloud (spike §7): once a migration has actually
run, putting the old image tag back produces a container that refuses to start —
> *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and
> downgrading is not supported"*
with a positive control proving the data is intact, only unreachable by the old version (§7 6d).
**So "rollback" must not appear in any spec for this arc.** The two shapes actually available are:
| shape | when it applies | what it does |
|---|---|---|
| **ABORT** | before anything migrated | stop, put the old image back, the app runs again |
| **RESTORE FROM A COPY** | after a migration ran | the data restore is the whole remedy |
There is no third. And per §5 below, a restore's image-level undo currently has a ≤15-minute
half-life because the syncer overwrites it (**R-441**).
---
## 5. The target shape
**The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the
syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the
rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths
stop being able to change a version by accident.
**Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the
prerequisite for judging how urgent it is.
---
## 6. The seven slices
| # | slice | status |
|---|---|---|
| **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** |
| **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** |
| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 |
| **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 |
| **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 |
| **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 |
| **7** | **A fleet sweep pipeline** — the operator can see, and move, how far behind every box is. | OPEN — R-451 |
**The rule slice 6 inherits, recorded now while it is cheap:** an engine change gets its own edge,
never bundled with an app version bump. `bookstack` moved the application *and* MariaDB 11.6 → 12.3 in
one commit (`0b73e5e`); that is two migrations behind one edge, and an unreadable failure when it
breaks.
---
## 7. What slices 1 and 2 actually built
### 7.1 The record (slice 1)
`Manager.recordInstalledImages` (`felhom-controller/controller/internal/stacks/installed.go`) runs
after a successful compose up from `StartStack`, `RestartStack`, `UpdateStack` and `runComposeDeploy`,
and writes `app.yaml`:
```yaml
installed_images:
web:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:… # "" if the image was never pulled from a registry
at: "2026-09-02T18:41:03Z" # when this ref+digest was FIRST seen for this service
```
Three rules, each with its reason:
- **It reads the CONTAINER, never `docker-compose.yml`.** §1.2 is why: that file is the value that has
already moved. A record built from it would answer "what will happen next time something runs
`up -d`", which is a different question.
- **A failed write NEVER refuses the action** — deliberately the opposite of `SetDesiredState`.
Intent refused, observation logged. Refusing to start a customer's app because we could not write
down which version it is trades a real outage for a bookkeeping gap.
- **It is NOT called from `StartStackServices`** — the R-47 DB-only restore window would overwrite a
complete record with a partial one.
**Nothing reads it to take a decision.** Slice 2 reads it to render a label.
### 7.2 The label (slice 2)
`web.updateBadge` (`internal/web/updatebadge.go`) compares the recorded reference per service against
what the current template pins, and renders through the existing `meta_badge` partial — no new markup,
no new CSS.
**Absent means UNKNOWN and never means current.** Every `app.yaml` written before v0.233.0 has no
record, so a fall-through to „Naprakész" would have told the whole fleet their months-old apps were
current. This is the R-166 lesson applied to an observation instead of an intent, and it is pinned by
a test with a companion red-proof.
**No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
Version strings stay in the logs, the API and the hub.
---
## 8. Known limitations, stated plainly
1. **„Naprakész" can be FALSE for the 23 floating pins.** The comparison is reference-to-reference and
queries no registry — a customer's box must not depend on reaching eight upstream registries to
render a page. For `postgres:16-alpine`, `mariadb:11.6` and 21 others the reference can be
identical while the image behind it has moved. **Measured, not theorised:** spike §5 found
`mariadb:11.4` and `mariadb:12.3` had both already moved upstream, with two fully-pinned controls
holding. Digest-level comparison needs a registry query and is deferred — **R-446**.
2. **Nothing enforces `catalog_since`.** A commit that moves an `image:` line and forgets the date
under-reports how far behind a box is. The gates runner fetches at `--depth 1` and has no parent
commit to diff against, so the gate needs a deeper fetch — **R-452**.
3. **The record only appears after the next lifecycle action.** An app that is running and untouched
keeps a legacy `app.yaml` and therefore no badge, until someone restarts, updates or redeploys it.
That is correct — the alternative is inventing a record from the file §1.2 says has already moved —
but it means the fleet view fills in gradually rather than at upgrade.
4. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
---
## 9. Where the rest lives
| what | where |
|---|---|
| the measurements this document rests on | `audits/SPIKE-app-update-2026-09-01.md` |
| the work | `backlog/OPEN-ITEMS.md` — R-438..R-445, R-446..R-452 |
| the syncer, described accurately but without the consequence | `architecture/02-controller-module-map.md` |
| what the lifecycle actions are proven to do | `architecture/00-capability-map.md` |
| the implementation | `felhom-controller/controller/README.md` §"What is installed, and is it current?" |
+10 -2
View File
@@ -675,14 +675,22 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` **THE DOCUMENT HALF IS NOW DISCHARGED, 2026-09-02: `documentation/architecture/09-update-architecture.md` exists and is a LIVING document, updated by every slice of this arc.** It records the mechanism as measured (§1), quotes the `RestartStack` comment that proves the restart half was CHOSEN (§2), carries the three operator rulings of 2026-09-02 (§3), strikes the word "rollback" (§4), states the target shape (§5) and lists the seven slices with a status each (§6). **THE ROW STAYS OPEN AND THE REASON IS THE POINT: the mechanism is now DOCUMENTED, not CHANGED.** Whether the syncer should go on overwriting a deployed app's compose file is still the operator's ruling, and acting on it is slice 3 (R-447). | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** |
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-447** | **[P1-HIGH] UPDATE ARC SLICE 3 — make the live compose file DERIVED, so the syncer stops changing a deployed app's version.** The target shape (`architecture/09-update-architecture.md` §5): the pin lives in `app.yaml`, the one file `Syncer.copyTemplates` never touches, and the live `docker-compose.yml` becomes an OUTPUT rather than an input. The catalog then proposes and `app.yaml` decides, and the thirteen unattended `compose up -d` call sites (spike §8) stop being able to change a version by accident. **BLOCKED ON AN OPERATOR RULING, and that is the whole reason this is a separate slice:** R-438 established that the restart half of this behaviour was CHOSEN and written down in `Manager.RestartStack`'s own comment. Changing it is not a bug fix; it is reversing a decision, and the decision-maker is Viktor. **Slices 1 and 2 shipped first ON PURPOSE** — the fleet's real state has to be visible before anyone can judge how urgent this is. Do NOT add a deployed check to any of the thirteen paths ahead of the ruling. `architecture/09-update-architecture.md` §5 | **BLOCKED — rank P1-HIGH; owner: VIKTOR rules, CC implements** |
| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** |
| **R-449** | **[P2-MEDIUM] UPDATE ARC SLICE 5 — an upgrade test that proves a real one-major upgrade end to end, INCLUDING the abort path.** Spike §7 performed the pieces by hand on Nextcloud (31.0.14 → 32.0.9 ran the migration; 31.0.14 → 34.0.1 was refused by the app and reported as SUCCESS by the product; the downgrade attempt was refused with the data intact). **None of it is a test that runs again.** A one-off measurement that nothing repeats decays into a claim — this project's most-repeated defect class. The test must assert the CONSEQUENCE (does the app serve after the upgrade? does the abort put it back?), not the mechanism. Venue: a Tier-0 box; `runbooks/target-selection.md` names which. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: CC** |
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
| **R-453** | **[P2-MEDIUM] The vaulted customer dashboard password is STALE ON BOTH DEMO BOXES, and there is no operator-side route to the real one — so no session can drive a customer page.** MEASURED 2026-09-02 while trying to live-validate the update badge: `PASSWORD` from DooPlex `~/.config/credentials` returns **HTTP 200 with the login page and the body string `Hibás jelszó`** against demo-hp guest 9201 (`https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`) **and** against demo-felhom guest 9201 (`https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu`). **The discriminator is the controller's own log, not the status code** — `auth.go:176: [WARN] [web] Failed login` proves wrong PASSWORD rather than wrong Host header, which is the trap this class always presents (a rejected login renders no flash and looks exactly like a routing problem). Every other key in the credentials file was checked and none is a dashboard password (`HUB_PW`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*`, and the `R_*` keys are escrow recovery codes). **THIS IS THE SECOND TIME:** it drifted on demo-hp on 2026-08-09 and was put back on operator instruction by writing a fresh bcrypt hash into the guest's `data/settings.json`; demo-hp was then reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted as well. **Why it is not merely inconvenient: it silently converts "endpoint-level validation" — this project's STANDARD method, because there is no browser on DooPlex — into "unit tests only" for anything that renders a customer page.** The cost is paid per session and rediscovered each time. **The fix is a decision, not a command:** the claim code is bcrypt-hashed hub-side and only e-mailed (R-119), so either the operator records the current demo passwords out-of-band, or a re-set becomes a standing authorisation for the two Tier-0 demo boxes. **NOT TAKEN UNILATERALLY:** re-setting a dashboard password is a decision about a customer account, and it was done under operator instruction last time. Raised in `STATUS.md` item 9. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4, which lists all five attempts. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
+1
View File
@@ -52,6 +52,7 @@
| ID | Item | Size | Status | Notes |
|----|------|------|--------|-------|
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1 & 2 SHIPPED 2026-09-02 (controller v0.233.0 + catalog `69761cf`); slices 3–7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 (R-447) is BLOCKED ON AN OPERATOR RULING and that is deliberate:** changing what the syncer does to a deployed app reverses a decision, not a bug, and slices 1 and 2 shipped first so the fleet's real state is visible before anyone judges how urgent it is. |
| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. |
| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/<share>`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. |
| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** |
@@ -0,0 +1,170 @@
# VALIDATION — update arc slices 1 & 2, live on demo-hp (2026-09-02)
**Controller `0.233.0`, guest 9201 on `demo-hp` (Tier 0 — disposable).** Evidence copied off the box
at the end of the phase that produced it, not at the end of the session.
**Method: endpoint/production-path level.** `claude-in-chrome` is not available on DooPlex.
**Which path was used, stated exactly: the BOOT RECONCILER**, `bootrecon.Run → StackProvider.StartStack
→ compose up -d → recordInstalledImages` — a REAL production caller, the same one
`SPIKE-app-update-2026-09-01` §2 variant 1c-ii used, and no hand-set state anywhere.
**Why not the customer's Restart button: see §4 — the vaulted dashboard password no longer opens
either demo controller.**
---
## 1. Deployed version
```
$ pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'
gitea.dooplex.hu/admin/felhom-controller:0.233.0 Up 19 seconds (healthy)
```
Previous: `0.232.0`. Image digest `sha256:df5940ccf5a548ceee9065a2cc55a467f8941c636138441e0d77e320dc2023ab`.
## 2. The record — PROVEN LIVE, on a single-service AND a multi-service app
### 2.0 Before — nine deployed apps, ZERO with a record
```
bentopdf deployed=1 installed_images=0 paperless-ngx deployed=1 installed_images=0
bookstack deployed=1 installed_images=0 privatebin deployed=1 installed_images=0
calibre-web deployed=1 installed_images=0 romm deployed=1 installed_images=0
docmost deployed=1 installed_images=0
kimai deployed=1 installed_images=0
opengist deployed=1 installed_images=0
```
**That is the legacy case, and it is the state the badge must render NOTHING for.**
### 2.0b Ground truth, read from the containers BEFORE anything was touched
```
bentopdf ref=ghcr.io/alam00000/bentopdf:v2.8.6 rd=…@sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
bookstack ref=lscr.io/linuxserver/bookstack:26.05.2 rd=…@sha256:3db259db582808ab498d49ae96b0a63f935d9cf3635c9d5bd8b8815c6ff1f8a1
bookstack-db ref=mariadb:12.3 rd=…@sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
```
**This is the discriminator.** Every digest below is compared against these, taken independently.
### 2.1 Single service — `bentopdf`
Staged exactly as the spike stages a boot orphan: `docker rm -f bentopdf` (chosen because it has **no
database, no volume and no data of any kind**), then the controller restarted so the reconciler runs.
`desired_state: running` was left untouched. **Nothing else was staged; the reconciler selected the app
on its own.**
```
18:27:32 main.go:2162: [INFO] [bootrecon] boot window: fleet settled after 30s — sweeping
18:27:32 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]
18:27:32 manager.go:1066: [INFO] [stacks] Starting stack: bentopdf
18:27:32 manager.go:1081: [INFO] [stacks] Stack bentopdf started successfully (took 0.3s)
18:27:32 installed.go:408: [INFO] [stacks] installed-images bentopdf: recorded 1 service(s) (bentopdf=ghcr.io/alam00000/bentopdf:v2.8.6 (sha256:eaeea1e44720…))
```
`/opt/docker/stacks/bentopdf/app.yaml` after:
```yaml
installed_images:
bentopdf:
ref: ghcr.io/alam00000/bentopdf:v2.8.6
digest: sha256:eaeea1e447205a79cb61d7efdc6966f37311dc1bc9c36a3a5c897bf79107c2c3
at: "2026-09-02T18:27:32Z"
```
**Digest MATCHES §2.0b exactly.**
### 2.2 Multi-service — `bookstack`. **The case that matters.**
**One entry for the whole stack is the wrong outcome, and only a multi-container app can tell.**
Same staging (`docker rm -f bookstack bookstack-db`; both volumes are NAMED —
`bookstack_bookstack_config`, `bookstack_bookstack_db_data` — and were verified present, so no data
was at risk), then the controller restarted.
```
18:28:48 bootrecon.go:259: [INFO] [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bookstack]
18:28:48 manager.go:1066: [INFO] [stacks] Starting stack: bookstack
18:28:54 installed.go:408: [INFO] [stacks] installed-images bookstack: recorded 2 service(s) (bookstack=lscr.io/linuxserver/bookstack:26.05.2 (sha256:3db259db5828…), bookstack-db=mariadb:12.3 (sha256:a02fe89cb597…))
18:28:54 bootrecon.go:303: [INFO] [bootrecon] Boot reconciliation complete: 1 app(s) recovered in 1 attempt(s)
```
```yaml
installed_images:
bookstack:
ref: lscr.io/linuxserver/bookstack:26.05.2
digest: sha256:3db259db582808ab498d49ae96b0a63f935d9cf3635c9d5bd8b8815c6ff1f8a1
at: "2026-09-02T18:28:54Z"
bookstack-db:
ref: mariadb:12.3
digest: sha256:a02fe89cb597d4375812b2eac90cf9d0775d4686daa7f7cc750ebbcad7525bbc
at: "2026-09-02T18:28:54Z"
```
**TWO entries, keyed by COMPOSE SERVICE NAME, both digests matching §2.0b exactly.**
**A NOTE ON WHY IT TOOK TWO PASSES, because it is a real fact about the reconciler and not a mishap.**
The first pass removed only the `bookstack` app container and left `bookstack-db` running; the
reconciler did **not** select bookstack — it found `[bentopdf]` alone. A stack with one live member is
not a boot orphan to it. Removing the DB container as well made the whole stack orphaned and it was
selected on the next pass. Both apps are running and healthy at the end (§5).
### 2.3 The encrypted secrets survived the write — checked, not assumed
`bookstack`'s `app.yaml` holds two encrypted values. Before and after the record was written, both
`ENC:` strings are **byte-identical**, and `deployed`, `deployed_at`, `locked_fields` and
`desired_state` are unchanged. This is the copy-and-overlay property (R-100) holding under a new field.
## 3. `catalog_since` reached the box by the NORMAL 15-minute sync
No force, no hand-edit:
```
$ grep -n catalog_since /opt/docker/stacks/bookstack/.felhom.yml
13:catalog_since: "2026-07-18"
```
## 4. The badge render — **NOT LIVE-VALIDATED. What was tried, in full.**
**A "no access" claim must list its attempts.** These are the attempts:
| # | attempt | result |
|---|---|---|
| 1 | `POST /login` to demo-hp guest `https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`, `-k`, password from DooPlex `~/.config/credentials` `PASSWORD` (extracted with `sed`, never `cut` — the values are quoted) | **HTTP 200 with the login page and the body string `Hibás jelszó`** |
| 2 | the controller's own log, as the discriminator between "wrong host header" and "wrong password" | `auth.go:176: [WARN] [web] Failed login from 172.18.0.3` — **wrong password, not a routing problem** |
| 3 | the same password against the OTHER demo box, demo-felhom guest `https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu` | **HTTP 200 + `Hibás jelszó` as well** |
| 4 | every other key in `~/.config/credentials` — `HUB_PW`, `R_DEMO-HP`, `R_DEMO-FELHOM`, `R_C11_REWALK`, `R_PART4`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*` | none is a dashboard password; the `R_*` keys are escrow recovery codes |
| 5 | reading the source for an unauthenticated route to an app page | only `/claim`, `/claim/request-new-code`, `/api/health` and `/static/` are exempt (`internal/web/auth.go`) |
**So the vaulted `PASSWORD` is stale on BOTH demo controllers.** It was already recorded as having
drifted once (memory `demo-hp-guest-controller-access`, 2026-08-09, put back on operator instruction);
demo-hp was reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted too.
**There is no operator-side route to a customer's dashboard password** — the claim code is
bcrypt-hashed hub-side and only emailed (R-119). Re-setting it means writing a new bcrypt hash into the
guest's `data/settings.json`, which is a decision about a customer account and was done under operator
instruction last time. **It is therefore a HUMAN step, by this project's own rule**, and it is not
taken here.
**What IS established about the badge, and it stops short of the render:** every INPUT the badge reads
is verified live and consistent on this box — `installed_images` present with refs matching the
template's pins exactly (§2.2 vs the compose file's `lscr.io/linuxserver/bookstack:26.05.2` and
`mariadb:12.3`), and `catalog_since: "2026-07-18"` present (§3). So bookstack's inputs are the
`Naprakész` case and the seven undisturbed apps are the no-record case. **That is an inference from
verified inputs, not an observation of the rendered page, and it is not counted as evidence.**
The render itself is covered by unit tests that render the PRODUCTION templates
(`TestGroupD_BadgeRendersOnBothSurfaces`, both surfaces, all states, with a companion red-proof), which
is the strongest statement available without the password.
## 5. End state — nothing left broken, nothing provisioned
```
$ pct exec 9201 -- docker ps --format '{{.Names}}|{{.Image}}|{{.Status}}' | grep -E 'bookstack|bentopdf'
bookstack|lscr.io/linuxserver/bookstack:26.05.2|Up 10 seconds (health: starting)
bookstack-db|mariadb:12.3|Up 16 seconds (healthy)
bentopdf|ghcr.io/alam00000/bentopdf:v2.8.6|Up About a minute (healthy)
```
All three containers are back on the SAME images they ran before, both named volumes untouched, and
both apps carry a complete record. **This run provisioned nothing** — no guest, no storage, no hub
record — so there is no teardown to report on any of the three layers.