09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the RestartStack comment that proves the restart half was CHOSEN (a design decision is not a defect), carries the three operator rulings of 2026-09-02, strikes the word 'rollback' (once a migration has run the old image will not start), states the target shape, and lists the seven slices with a status each. R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not changed. Nothing closed, so CLOSED-ITEMS.md is untouched. Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23 floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner), R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what stopped the badge render from being validated live). Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN LIVE through the boot reconciler on demo-hp - one entry per compose service, digests matching ground truth read independently. The badge RENDER is not, and the five attempts are listed rather than summarised.
This commit is contained in:
@@ -675,14 +675,22 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
|
||||
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
|
||||
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
|
||||
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
|
||||
| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` **THE DOCUMENT HALF IS NOW DISCHARGED, 2026-09-02: `documentation/architecture/09-update-architecture.md` exists and is a LIVING document, updated by every slice of this arc.** It records the mechanism as measured (§1), quotes the `RestartStack` comment that proves the restart half was CHOSEN (§2), carries the three operator rulings of 2026-09-02 (§3), strikes the word "rollback" (§4), states the target shape (§5) and lists the seven slices with a status each (§6). **THE ROW STAYS OPEN AND THE REASON IS THE POINT: the mechanism is now DOCUMENTED, not CHANGED.** Whether the syncer should go on overwriting a deployed app's compose file is still the operator's ruling, and acting on it is slice 3 (R-447). | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** |
|
||||
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
|
||||
| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** |
|
||||
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
|
||||
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
|
||||
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-447** | **[P1-HIGH] UPDATE ARC SLICE 3 — make the live compose file DERIVED, so the syncer stops changing a deployed app's version.** The target shape (`architecture/09-update-architecture.md` §5): the pin lives in `app.yaml`, the one file `Syncer.copyTemplates` never touches, and the live `docker-compose.yml` becomes an OUTPUT rather than an input. The catalog then proposes and `app.yaml` decides, and the thirteen unattended `compose up -d` call sites (spike §8) stop being able to change a version by accident. **BLOCKED ON AN OPERATOR RULING, and that is the whole reason this is a separate slice:** R-438 established that the restart half of this behaviour was CHOSEN and written down in `Manager.RestartStack`'s own comment. Changing it is not a bug fix; it is reversing a decision, and the decision-maker is Viktor. **Slices 1 and 2 shipped first ON PURPOSE** — the fleet's real state has to be visible before anyone can judge how urgent this is. Do NOT add a deployed check to any of the thirteen paths ahead of the ruling. `architecture/09-update-architecture.md` §5 | **BLOCKED — rank P1-HIGH; owner: VIKTOR rules, CC implements** |
|
||||
| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** |
|
||||
| **R-449** | **[P2-MEDIUM] UPDATE ARC SLICE 5 — an upgrade test that proves a real one-major upgrade end to end, INCLUDING the abort path.** Spike §7 performed the pieces by hand on Nextcloud (31.0.14 → 32.0.9 ran the migration; 31.0.14 → 34.0.1 was refused by the app and reported as SUCCESS by the product; the downgrade attempt was refused with the data intact). **None of it is a test that runs again.** A one-off measurement that nothing repeats decays into a claim — this project's most-repeated defect class. The test must assert the CONSEQUENCE (does the app serve after the upgrade? does the abort put it back?), not the mechanism. Venue: a Tier-0 box; `runbooks/target-selection.md` names which. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-453** | **[P2-MEDIUM] The vaulted customer dashboard password is STALE ON BOTH DEMO BOXES, and there is no operator-side route to the real one — so no session can drive a customer page.** MEASURED 2026-09-02 while trying to live-validate the update badge: `PASSWORD` from DooPlex `~/.config/credentials` returns **HTTP 200 with the login page and the body string `Hibás jelszó`** against demo-hp guest 9201 (`https://192.168.0.138:443`, `Host: felhom.enkisfelhom.hu`) **and** against demo-felhom guest 9201 (`https://192.168.0.149:443`, `Host: felhom.demo-felhom.eu`). **The discriminator is the controller's own log, not the status code** — `auth.go:176: [WARN] [web] Failed login` proves wrong PASSWORD rather than wrong Host header, which is the trap this class always presents (a rejected login renders no flash and looks exactly like a routing problem). Every other key in the credentials file was checked and none is a dashboard password (`HUB_PW`, `TS_KEY`, `HETZNER_API`, `ISO_S3_*`, and the `R_*` keys are escrow recovery codes). **THIS IS THE SECOND TIME:** it drifted on demo-hp on 2026-08-09 and was put back on operator instruction by writing a fresh bcrypt hash into the guest's `data/settings.json`; demo-hp was then reinstalled and re-claimed on 2026-08-21, and demo-felhom has now drifted as well. **Why it is not merely inconvenient: it silently converts "endpoint-level validation" — this project's STANDARD method, because there is no browser on DooPlex — into "unit tests only" for anything that renders a customer page.** The cost is paid per session and rediscovered each time. **The fix is a decision, not a command:** the claim code is bcrypt-hashed hub-side and only e-mailed (R-119), so either the operator records the current demo passwords out-of-band, or a re-set becomes a standing authorisation for the two Tier-0 demo boxes. **NOT TAKEN UNILATERALLY:** re-setting a dashboard password is a decision about a customer account, and it was done under operator instruction last time. Raised in `STATUS.md` item 9. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4, which lists all five attempts. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
@@ -52,6 +52,7 @@
|
||||
|
||||
| ID | Item | Size | Status | Notes |
|
||||
|----|------|------|--------|-------|
|
||||
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1 & 2 SHIPPED 2026-09-02 (controller v0.233.0 + catalog `69761cf`); slices 3–7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 (R-447) is BLOCKED ON AN OPERATOR RULING and that is deliberate:** changing what the syncer does to a deployed app reverses a decision, not a bug, and slices 1 and 2 shipped first so the fleet's real state is visible before anyone judges how urgent it is. |
|
||||
| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. |
|
||||
| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/<share>`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. |
|
||||
| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** |
|
||||
|
||||
Reference in New Issue
Block a user