From 417df06f352921745b968e92b9248bb18156c2dc Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 6 Sep 2026 10:37:33 +0200 Subject: [PATCH] slice 3 docs: the ruling, the shipped mechanism, and four rows closed 09-update-architecture.md gains the fourth dated operator ruling (2026-09-06, Option 1) and its section 5 is rewritten from a proposed shape into the shipped one: the pin, the stored definition, the render table, the four writers, the startup ordering, and the trap this slice set for slice 2 - the live compose file is now the frozen one, so a badge comparing against it would answer Naprakesz on exactly the apps that are behind. 02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being true today, so it is corrected, and the two sections describing the old seam now carry a banner saying they describe v0.234.0 and below - kept because every box under v0.235.0 still behaves that way and because they are the measured account of why it changed. R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458 opened for the .felhom.yml asymmetry, with what would settle it by measurement. Live evidence: two real catalog pushes travelling the real 15-minute cycle, both reverted, the tree byte-identical afterwards. The restart that used to take 18.3 seconds and pull a new image now takes 0.1 seconds and pulls nothing. --- REPORT.md | 97 ++++------ STATUS.md | 44 ++++- .../architecture/00-capability-map.md | 3 +- .../architecture/02-controller-module-map.md | 15 +- .../architecture/09-update-architecture.md | 138 +++++++++++++- documentation/backlog/CLOSED-ITEMS.md | 4 + documentation/backlog/OPEN-ITEMS.md | 5 +- documentation/backlog/ROADMAP.md | 2 +- .../VALIDATION-update-slice3-2026-09-06.md | 179 ++++++++++++++++++ 9 files changed, 411 insertions(+), 76 deletions(-) create mode 100644 documentation/tests/VALIDATION-update-slice3-2026-09-06.md diff --git a/REPORT.md b/REPORT.md index f8307e0d..5b93e8b7 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,4 +1,4 @@ -# REPORT — update arc slices 1, 1b & 2: documents, register, roadmap, capability map (2026-09-03) +# REPORT — update arc slice 3: the version freeze (2026-09-06) *Overwritten each session. Nothing durable lives only here.* @@ -6,70 +6,57 @@ | file | change | |---|---| -| **`documentation/architecture/09-update-architecture.md`** | **CREATED — the deliverable.** Its absence was R-438. A LIVING document: every slice of this arc updates it in the same session. | -| `documentation/tests/VALIDATION-update-slice12-2026-09-02.md` | CREATED — the live evidence, copied off demo-hp at the end of the phase that produced it. | -| `documentation/backlog/OPEN-ITEMS.md` | R-438 and R-440 amended (both stay OPEN); **8 new rows: R-446..R-453**. | -| `documentation/backlog/ROADMAP.md` | the update arc added as one item, naming the capability-map rows it flips. | -| `documentation/architecture/00-capability-map.md` | one new row: what version a box runs, and whether it is behind. | -| `STATUS.md` | one new operator item (9) and a new lead paragraph. | +| `documentation/architecture/09-update-architecture.md` | **§3 gains the fourth dated operator ruling** (2026-09-06, Option 1); **§5 rewritten** from a proposed shape into the shipped mechanism — the pin, the stored definition, the render table, the four writers, the startup ordering, and §5.6 the trap it set for slice 2; slice 3 moved to shipped; two limitations added (§8.4 frozen-whole, §8.5 the `.felhom.yml` asymmetry). | +| `documentation/architecture/02-controller-module-map.md` | its *"copy compose + `.felhom.yml`"* line stopped being true and is corrected; §1/§2 of the app-definition seam now carry a banner saying they describe ≤ v0.234.0 and why they are kept. | +| `documentation/tests/VALIDATION-update-slice3-2026-09-06.md` | CREATED — the live evidence. | +| `documentation/backlog/OPEN-ITEMS.md` | **R-447, R-441, R-438 and R-455 closed and moved out**; **R-458 opened**. | +| `documentation/backlog/CLOSED-ITEMS.md` | the four closed rows, compressed, each naming `bc47dd4ef997` as the commit whose `git show` returns the original. | +| `documentation/backlog/ROADMAP.md` | the arc item collapsed to slices 1/1b/2/3 shipped; **R-448 named as the new head of the arc**. | +| `documentation/architecture/00-capability-map.md` | one new row, **PROVEN-LIVE**; the pre-v0.235.0 row relabelled rather than deleted, because every box below v0.235.0 still behaves that way. | +| `STATUS.md` | new lead + item 10; **items 4 and R-455 closed**. | -## The architecture document — what it settles +## The ruling, recorded -1. **How an update works today, as measured** — `UpdateStack` is `pull` then `up -d --remove-orphans`; - the syncer overwrites a deployed app's compose on a 15-minute cycle with no deployed check; - **thirteen** non-API call sites end in `compose up -d`. -2. **What was chosen, and by whom** — `RestartStack`'s own comment, quoted. **A design decision is not - a defect.** What was never decided is what the syncer does underneath a deployed app. -3. **The three operator rulings of 2026-09-02** — verified backup as a precondition; the support window - runs on how far behind the CATALOG a box is; automatic within a major, never across one. -4. **The vocabulary ruling** — "rollback" is struck. The available shapes are ABORT and RESTORE. -5. **The target shape** — the live compose file becomes derived from a pin in `app.yaml`. -6. **The seven slices**, each with a status line. 1 and 2 are shipped. -7. **Known limitations**, including the floating-tag one. +**2026-09-06, Option 1: freeze the version, keep the fixes flowing.** The document records *what the +ruling looked at*, not just its outcome: the old behaviour had two halves — a restart silently +changing an app's VERSION (unwanted), and template corrections plus self-healing reaching a deployed +app (worth keeping) — and the ruling keeps the second while removing the first. -## Register +## Register — 203 open rows before, 203 after; closed 163 → 167 -**194 rows before, 206 after** (205 on 2026-09-02, plus R-457 on 2026-09-03). Nothing closed, and that is stated rather than implied: R-438 and -R-440 are **amended and stay OPEN** — the mechanism is now documented, not changed — so nothing moved -to `CLOSED-ITEMS.md` and that file is untouched. +Four closed and moved, one opened. The count is level by coincidence, not by inaction. -| row | what | state | -|---|---|---| -| R-446 | „Naprakész" can be FALSE for the 23 floating pins | OPEN, P2-MEDIUM, CC | -| R-447 | slice 3 — make the live compose DERIVED | **BLOCKED** on an operator ruling, P1-HIGH | -| R-448 | slice 4 — a guarded update (subsumes R-443) | READY, P2-MEDIUM | -| R-449 | slice 5 — an upgrade test that runs again | READY, P2-MEDIUM | -| R-450 | slice 6 — version sequence; an engine change gets its own edge | READY, P2-MEDIUM | -| R-451 | slice 7 — a fleet sweep (needs a hub change: no image field is reported) | READY, P3-LOW | -| R-452 | no gate enforces `catalog_since` (`--depth 1` has no parent to diff) | READY, P3-LOW | -| R-453 | **CORRECTED** — the password was fine; `~/.config/credentials` uses SINGLE quotes and the strip was half-applied. The finding is the instrumentation lesson, not the typo | READY, P3-LOW | -| R-454 | five `gofmt`-unclean `internal/web` test files, and no gate notices | READY, P3-LOW | -| R-455 | **DooPlex has no Docker Hub login — the ceiling now blocks BUILDS, not just gates.** The mirror workaround is **MEASURED byte-identical** to Hub, so it is a sound fallback, not a hope | **WAITING-ON-OPERATOR**, P2-MEDIUM | -| R-456 | a partly-dead stack is not a boot orphan, and that rule is written down nowhere | READY, P3-LOW | -| R-457 | **a test that hardcodes a date and asserts an age is green only on the day it is written** — one proven, six candidate files named | READY, P3-LOW | +| row | disposition | +|---|---| +| **R-447** | **CLOSED** — slice 3 shipped, controller v0.235.0 | +| **R-441** | **CLOSED BY MEASUREMENT** — the restore now pins to what the unit captured; the render obeys it. The live half is §3/§6b of the validation file, not a code reading | +| **R-438** | **CLOSED** — both halves discharged: documented 2026-09-02, behaviour changed 2026-09-06 | +| **R-455** | **CLOSED** — the operator added a Docker Hub PAT | +| **R-458** | **OPENED** (P3-LOW, CC) — `.felhom.yml` keeps flowing to a frozen app, so it can receive a health check written for a newer version. **False alarm, never data loss.** The row states what would settle it by measurement rather than by code | ## Live validation -Full evidence: `documentation/tests/VALIDATION-update-slice12-2026-09-02.md`. +Full evidence: `documentation/tests/VALIDATION-update-slice3-2026-09-06.md`. **Two REAL catalog pushes +travelling the REAL 15-minute cycle**, both reverted in the same session; the catalog tree is +byte-identical to `8220f8d` afterwards. -**PROVEN LIVE on demo-hp at controller 0.233.0**, through a real production caller (the boot -reconciler — no hand-set state): `bentopdf` recorded 1 service, `bookstack` recorded **2**, keyed by -compose service name, **all three digests matching ground truth read independently beforehand**. -Encrypted secrets byte-identical across the write. Both apps up and healthy; nothing provisioned. +- **Scenario A** — a non-image change reached the pinned app (08:01:51Z), container untouched. +- **Scenario B** — an image change did not (08:20:29Z), and the restart afterwards took **0.1 s, did + not recreate the container, and never pulled the new image** — against **18.3 s with a pull** for the + identical sequence measured before the change. +- **Scenario D** — the Update button still moved the version, with the pin advancing **17 s before** + the pull completed. +- **Scenario G** — the frozen app read „Frissítés elérhető — 56 napja" while the other eight read + „Naprakész". +- **Teardown** — the container is back on the baseline digest `sha256:eaeea1e4…`, byte for byte. -**The badge is ALSO proven live**, on both surfaces and in three of its four states — including the -load-bearing one: a deployed app with **no record renders no badge at all**. „Frissítés elérhető — -**52 napja**" was produced by editing bentopdf's compose tag only (no restart, no `up -d`) and reverted -byte-identically. **One correction of my own is stated in §4.0 of the validation file rather than -buried: I first reported the vaulted password as stale on both boxes. It was not — I stripped only -double quotes from a single-quoted value.** The operator caught it in one line. +**A real gap was found by the live run and fixed in it:** the stored definition did not follow the +fixes delivered after the pin, so the first freeze would have reverted them — silently undoing the +half of the ruling that says fixes keep flowing. Section 6 of the validation file keeps the +observation that proved it. ## Sibling repos -- `felhom-controller` **v0.233.0** — `8025304acc0a`, deployed to demo-hp and verified healthy. -- `felhom-controller` **v0.234.0** — `38d28b5b624c`, the startup backfill. Deployed to **both** demo - boxes and verified: 2 apps seeded on demo-hp with digests matching ground truth, 7 left untouched; - all 9 then badged. **Its cause is worth keeping: `09-update-architecture.md` §8.3 was written as an - accepted limitation on 2026-09-02 and was a defect by the next morning** — the operator found an app - that simply ran and therefore showed nothing. -- `app-catalog-felhom.eu` — `69761cf91bfc` (backfill) + `8220f8d` (REPORT). +- `felhom-controller` **v0.235.0** — `8a0e0a59adc7` + `2a56f557d048`, deployed to demo-hp, healthy. +- `app-catalog-felhom.eu` — `dc7e548`, `09b4ff5` (live-test) and `1798ce6`, `17cc784` (reverts), plus + `7b9b9b3` recording them as a measurement rather than a release. diff --git a/STATUS.md b/STATUS.md index 77ea7127..c45a9ac6 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,8 +1,13 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-03 — you spotted that OpenGist had no label. You were right, and it was a real gap: +**Updated 2026-09-06 — a restart no longer changes which version an app runs. Fixes still arrive +every 15 minutes, and a broken app definition still repairs itself. Only the Update button moves a +version now. Live on the HP (0.235.0). NOTHING IS WAITING ON YOU — and two old items are now closed: +the Hetzner e-mails are answered, and the Docker Hub login is in place.** + +**Earlier 2026-09-03 — you spotted that OpenGist had no label. You were right, and it was a real gap: the label only appeared on apps something had restarted. Fixed and live (0.234.0). Every app on both -machines now carries one. NOTHING IS WAITING ON YOU.** +machines now carries one.** **Earlier 2026-09-02 — the box now writes down which version of each app it is running, and shows one small label saying whether it is up to date: „Naprakész" or „Frissítés elérhető — 52 napja". No version @@ -26,7 +31,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Two things are waiting on you — item 4 (send two e-mails) and item 7 (one design decision).** Item 9 is new today and needs nothing from you. Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on +1. **Nothing is waiting on you.** Item 4 (the Hetzner e-mails) is answered and is being handled in a separate session. Item 7 — the safety-copy decision — is the one open question, and it is **not urgent any more**: the thing that made it urgent was that a restart could upgrade an app behind your back, and as of today it cannot. Item 10 is new and needs nothing from you. Item 5's alarm mail can now be ignored for good. Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -45,7 +50,7 @@ nothing.* two register lines in the hub (already live). No customer action, no data migration, no credential change. -4. **Please send two short e-mails to Hetzner. They are written for you.** +4. **~~Please send two short e-mails to Hetzner.~~ DONE — you sent them and Hetzner replied (2026-09-06). The reply is being worked in a separate session; nothing about it belongs to the update work.** The background below is kept because it is why the questions were asked. `felhom.eu/documentation/runbooks/provider-questions-2026-09-01.md` — open it, copy, send. No password or key is in that file, and none should be added. @@ -155,6 +160,37 @@ nothing.* **If you do nothing:** nothing. This one is closed. +10. **Nothing here needs you. The version of an app is now frozen, and only the Update button moves + it.** + + **What used to happen.** The box downloaded the catalog every 15 minutes and wrote each app's new + definition straight over yours — including a new version. From then on, **thirteen** different + things could install that new version: the Update button, the Restart button, or one of eleven + repairs the box performs on its own. Nobody had to press anything. + + **What happens now.** The version is pinned to what you have. Only a deliberate Update moves it. + + **What deliberately did NOT change, because it was worth keeping.** Corrections to an app's + definition — a fixed health check, a memory limit, a new setting — still arrive on the same + 15-minute cycle, and a broken definition still repairs itself. That was a real benefit of the old + behaviour and it is intact. In one sentence: **while the catalog offers the same version you are + running, its fixes reach you; the moment it moves to a newer version, you stay where you are until + you choose to update.** + + **Proven on the real machine**, twice over: I pushed a genuine catalog change with no version in + it and watched it arrive on the normal cycle; then I pushed a version change and watched the app + refuse it. + + **Two things this did NOT do, so they are not read as done.** The Update button is exactly as safe + as it was yesterday — no backup, no undo. That is the next piece of work, and it is item 7. And an + app the box could not confidently pin would have been left behaving exactly as before, loudly; on + the HP that was none of the nine. + + **One rough edge I chose to write down rather than fix:** a frozen app still receives the small + metadata file that carries its health check, so it can be given a check written for a newer + version and look unwell when it is fine. **It cannot lose data — the worst case is a false + alarm.** Freezing that file too would break the „Frissítés elérhető" label, which is a worse trade. + 8. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, as it misled one by an hour. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 51934edc..6c888811 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -100,11 +100,12 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis |---|---|---|---|---| | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. What they do to app DATA is now measured too, and it is a separate row-worth of facts (below).** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove live in `CAMPAIGN-3`; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | -| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller | **PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | +| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | | Protected infra stacks can't be stopped/removed from UI | controller | **PROVEN-LIVE** | `CAMPAIGN-nomercy` + `RERUN-p1p3` T-SEC-PROTECTED (refuse stop/remove, stay Up) | (Cited `CAMPAIGN-2` T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal) | | **What VERSION a box is running, and whether it is behind the catalog** | controller **v0.234.0** + catalog `69761cf` | **PROVEN-LIVE — both the record and the rendered badge.** | **`tests/VALIDATION-update-slice12-2026-09-02.md`** — on demo-hp 0.233.0, through a REAL production caller (`bootrecon → StartStack → compose up -d → recordInstalledImages`, no hand-set state): `bentopdf` recorded **1** service and `bookstack` recorded **2**, keyed by compose SERVICE name, and **all three digests match the ground truth read independently from the containers before anything was touched**. `catalog_since` reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. **v0.234.0 startup backfill, PROVEN LIVE 2026-09-03 on demo-hp:** the record was stripped from `privatebin` (1 service) and `romm` (3 services) to recreate the pre-0.233.0 shape, the controller restarted, and the backfill re-seeded **exactly** those two — every digest matching the ground truth read from the containers beforehand — while logging `2 app(s) recorded, 7 already had a record, 0 left unrecorded`. **All 9 deployed apps then carried „Naprakész" on `/stacks`** (ASCII fragments with a negative control at 0). On demo-felhom the operator's own case, OpenGist, now renders the badge. **Why the backfill exists at all: without it the label never reached an app that simply runs**, which the operator found the morning after v0.233.0. Unit side: `installed_test.go` + `updatebadge_test.go`, incl. a wiring test through a real `RestartStack`, an AST walk of all four call sites, and three companion red-proofs. | **THE UNEXERCISED LEG, NAMED: one badge STATE of four.** „Frissítés elérhető" WITHOUT an age needs an app whose `catalog_since` is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (`TestGroupF`). The other three are live: „Naprakész" ×2 on `/stacks` and on `/apps/bookstack`; **NO badge at all on `/apps/docmost`**, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — **52 napja**" on both surfaces, the age being real arithmetic on bentopdf's `catalog_since` 2026-07-12. **The behind state was staged by editing bentopdf's compose tag ONLY** — no container restarted, no `up -d` — then reverted byte-identically (`sha256` equal, `diff` empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with `grep -oF` ASCII fragments plus positive AND negative controls. **The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state.** **Absent means UNKNOWN, never current** — a legacy `app.yaml` renders NOTHING, red-proved. **No version number is shown to the customer** and **no registry is queried**, so „Naprakész" CAN BE FALSE for the 23 floating pins (**R-446**). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: `architecture/09-update-architecture.md`; remaining slices R-447..R-452 | +| **An app's VERSION is frozen to what the customer has; only a deliberate Update moves it, while template CORRECTIONS still arrive** | controller v0.235.0 | **PROVEN-LIVE (2026-09-06)** — by two REAL catalog pushes travelling the REAL 15-minute cycle, not a hand-edited file | **`tests/VALIDATION-update-slice3-2026-09-06.md`** — a non-image catalog change REACHED the pinned app (08:01:51Z) with the container untouched; an image change did NOT (08:20:29Z); and the restart afterwards took **0.1 s, did not recreate the container, and never pulled the new image**, against the spike's **18.3 s with a pull** for the identical sequence before. The Update button still moved the version (pin advanced 17 s BEFORE the pull completed) and the teardown update returned the container to the baseline digest byte for byte. The freeze holds in BOTH directions. Unit side: `internal/sync/render_test.go` (the whole render table, incl. self-healing in BOTH branches), `internal/stacks/pin_test.go` (adoption never guesses; the update advances the pin BEFORE the pull), `TestGroupG` (the badge reads the catalog, not the frozen file), `TestGroupH` (an AST walk of `cmd/controller/main.go` asserting the seam, adoption, and their ORDER against `syncer.Start()`). **Three companion red-proofs**, each run, failing, and reverted. | Operator ruling 2026-09-06, Option 1 (`architecture/09-update-architecture.md` §3.4). **Nothing was added to the thirteen `compose up -d` call sites** — most are repairs, and a repair that refuses to repair leaves an app down; they were made safe by removing the reason. `pinned_images` is INTENT, `installed_images` is an OBSERVATION — never fed from each other (R-166, one field over). **Known limitations, all recorded rather than fixed:** a frozen app is frozen WHOLE (§8.4); `.felhom.yml` keeps flowing, so a frozen app can get a probe for a newer version — false alarm, never data loss (**R-458**); and **the Update button is still unguarded** (R-448 is slice 4). Closes R-447, R-441, R-438 | | Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad `backup:` block degrades to legacy, loudly) | controller v0.132, catalog | **PROVEN-LIVE** | `CAMPAIGN-2` T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs | | | Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) | agent v0.94.0→**v0.95.0**, controller v0.169.0→v0.171.0→**v0.215.0**, hub v0.73.1 | **PROVEN-LIVE (healthy path + delivery + the severity wire).** **IMPLEMENTED, NOT proven-live: the Hiba-from-counters path** (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (**R-332**) | **2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix):** the card on guest 9201 now shows BOTH real disks with **real verdicts + human model labels** — **„AirDisk 512GB SSD" → Rendben (34°C)** (the system SSD, via LVM/dm resolution) and **„TOSHIBA MQ04ABF100" → Rendben (30°C)** (the USB, via union-path SMART). `/disks` carries `smart.health=PASSED` + `model_name` for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (`SPIKE-smart-coverage-2026-07-25.md` had proven both disks answer `smartctl -a -j` PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. **Notification pipeline PROVEN-LIVE 2026-07-24** — a `disk_health_degraded` POST (the exact `notify.PushEvent` wire call) was **400-rejected by hub v0.73.0** and **200-accepted + „Operator email sent" by hub v0.73.1** | No new smartctl load; feature-detect by payload presence → **MinAgent floor unchanged**; no sudoers/`-d sat` change. **No global banner** (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin `local` on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + `model_name` capture. **2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions** (`audits/DIAG-smart-passed-trap-2026-08-14.md` + two committed fixtures: raw `smartctl -a -j` and 406 `smartd` lines from ST3000VX010 S/N Z6A07P2G). **(1)** `smart_status.passed` is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry `thresh: 0` and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. **(2)** The alert it did produce carried severity `"warn"`, which the hub coerces to `info` and never emails: **the counterfactual is ZERO emails about this drive** (R-328, fixed controller v0.215.0, and the `warning`-vs-`warn` pair proven side by side in `notification_log` on 2026-08-14 — `sent` vs no row at all). **(3)** The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. **The verdict half of that arm remains unit+red-proof covered only** — no live drive has reached Hiba from counters (R-332). **SMART history/trending (hub-side) PARKED** (ROADMAP R-73) | | App crashes → customer notified (one event per transition, no flapping spam) | controller v0.120, hub v0.48 | **IMPLEMENTED** | controller v0.120.0 (dead-app alerting, `app_start_failed`, one-event-per-transition red-proofs); `CAMPAIGN-3` F11 surfaced the gap | End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven | diff --git a/documentation/architecture/02-controller-module-map.md b/documentation/architecture/02-controller-module-map.md index 893f9983..43007616 100644 --- a/documentation/architecture/02-controller-module-map.md +++ b/documentation/architecture/02-controller-module-map.md @@ -335,7 +335,7 @@ own; every caller that is not the customer must decide for itself whether the ap ### `sync/` | File | Class | Reason | Risk | |---|---|---|---| -| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset, copy compose+`.felhom.yml`, never overwrite app.yaml). **It copies into EVERY stack folder, deployed or not — see "the app-definition seam" below.** | clean | +| `sync/sync.go` | **KEEP** | Catalog git-sync (clone/fetch/reset). **Since controller v0.235.0 it RENDERS `docker-compose.yml` rather than copying it** — verbatim while the catalog still offers the app's pinned version, from the app's stored `applied-compose.yml` once the catalog moves past it. `.felhom.yml` is still copied verbatim always, and `app.yaml` is still never touched. Reasoning: `09-update-architecture.md` §5. | clean | ### `system/` — split per-function (not per-file) | File | Class | Reason | Risk | @@ -518,6 +518,14 @@ own; every caller that is not the customer must decide for itself whether the ap Until this was measured, no architecture document said what happens here, and the gap itself is R-438. The three facts below are the ones a reader needs before touching any of it. +> **⚠ SECTIONS 1 AND 2 DESCRIBE THE BEHAVIOUR UP TO CONTROLLER v0.234.0. Controller v0.235.0 +> (2026-09-06) CHANGED IT, on an operator ruling.** They are kept because they are the measured +> account of why it was changed, and because every box below v0.235.0 still behaves this way. **What +> ships now: `09-update-architecture.md` §5.** In one sentence — an app's VERSION is frozen to what the +> customer has and only a deliberate Update moves it, while template CORRECTIONS and the self-healing +> below still arrive on the 15-minute cycle. Nothing was added to the thirteen call sites in §2; they +> were made safe by removing the reason. + ### 1. The catalog syncer rewrites the file under a running app, on a 15-minute cycle `Syncer.copyTemplates` (`sync/sync.go:319`) walks every directory in the catalog cache and copies @@ -526,6 +534,11 @@ the app is deployed.** The only guard is a sha256 content compare (`copyIfChange exclusion is `app.yaml`. Interval is `git.sync_interval`, default `15m` (`config/config.go:351`), plus one immediate sync at controller start (`sync.go:98`). +**SINCE v0.235.0** this walk still happens and `.felhom.yml` is still copied unconditionally, but the +compose file goes through `Syncer.renderSource`, which consults a per-app plan supplied by the stack +manager (`Manager.RenderPlanFor`) through a nil-safe seam. A nil seam is byte-for-byte the behaviour +described above. + **It restarts nothing.** The post-sync hook is `stackMgr.InjectMissingFields(updated)` and nothing else. So from the moment it runs, a deployed app's *definition* and its *running containers* disagree, and they stay that way until something else acts. diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 399d647d..08799842 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -81,10 +81,12 @@ it does not change it. --- -## 3. The three operator decisions (2026-09-02) +## 3. The operator decisions These are rulings, not proposals. Anything specced against a different assumption is wrong. +### 2026-09-02 + 1. **The safety copy is a verified recent backup as a PRECONDITION** — not a new copy invented for the update path. The guest-snapshot alternative is to be **spiked before anything is designed around it**. Context: the existing safety machinery (`Manager.writeSafetyDump`, @@ -98,6 +100,31 @@ These are rulings, not proposals. Anything specced against a different assumptio 3. **Updates are automatic WITHIN a major, never ACROSS one.** The cross-major case needs a human, because §4 says it cannot be undone. +### 2026-09-06 — **Option 1: freeze the version, keep the fixes flowing.** SHIPPED, v0.235.0 + +4. **An app's version is frozen to what the customer has, and only a deliberate Update moves it — + while corrections to its definition keep arriving on the 15-minute cycle exactly as they do + today.** + + **This is the ruling R-447 was blocked on**, and it was blocked for a good reason: §2 establishes + that `RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in + its own comment. Reversing a chosen behaviour is a decision, not a bug fix. + + **What the ruling looked at, and why it is not simply "stop the syncer touching deployed apps".** + The old behaviour had two halves and only one of them was unwanted: + + | half | verdict | + |---|---| + | a restart/repair silently changes the app's VERSION | **unwanted** — nobody chose it, nobody is told, and §4 says it cannot be undone | + | template CORRECTIONS reach a deployed app, and a broken definition heals itself within 15 minutes | **worth keeping** — both measured in the spike §3 | + + So the ruling keeps the second and removes the first. In the operator's own words: *while the + catalog is offering the same version you are running, its fixes flow to you; the moment it moves to + a newer version, you are frozen at what you have until you choose to update.* + + **It does NOT make the Update button safer.** That is slice 4 (R-448), and it is where the backup + precondition goes. Slice 3 only stops the other twelve paths from doing the update's job. + --- ## 4. The vocabulary ruling — "rollback" is struck @@ -122,15 +149,95 @@ half-life because the syncer overwrites it (**R-441**). --- -## 5. The target shape +## 5. The shape, as SHIPPED in v0.235.0 -**The live `docker-compose.yml` becomes DERIVED from a pin recorded in `app.yaml`** — the one file the -syncer never touches (`sync.go:319`'s exclusion). The catalog then proposes; `app.yaml` decides; the -rendered compose file is an output rather than an input, and the thirteen unattended `up -d` paths -stop being able to change a version by accident. +The live `docker-compose.yml` is **DERIVED** from a pin recorded in `app.yaml` — the one file the +syncer never touches. The catalog proposes; `app.yaml` decides; the compose file is an output rather +than an input, and the thirteen unattended `up -d` paths stop being able to change a version by +accident. -**Nothing in slices 1 or 2 implements this.** They make the current state *visible*, which is the -prerequisite for judging how urgent it is. +### 5.1 Nothing was added to the thirteen call sites, and that is deliberate + +**They are made safe by removing the reason, not by gating them.** The most important of them are +REPAIRS — the boot reconciler (`bootrecon.go:269`), the drive-return gate (`intermediary.go:222`), +the app-stop guard (`appstop_marker.go:283`). **A repair path that refuses to repair leaves a +customer's app down, which is worse than the problem this slice solves.** Since the file they act on +no longer changes version, every one of them became safe without being touched. + +### 5.2 The pin, and what it is not + +`AppConfig.PinnedImages` (`app.yaml`, `pinned_images:`), service → image ref. + +**It is NOT `InstalledImages`.** That field is an OBSERVATION — what containers report. This one is a +DECISION — what should run. Letting an observation feed a decision would make a bad reading become a +bad deployment, which is the category error `desired_state` exists to avoid (R-166), one field over. +They will normally agree; when they disagree that is a signal, not a bug to paper over. + +**Absent means UNPINNED, and unpinned means the app behaves exactly as it did before v0.235.0.** + +Beside it, `applied-compose.yml` in the stack directory stores the exact definition that pin came +from. `Syncer.copyTemplates` copies exactly `docker-compose.yml` and `.felhom.yml`, so that name is +safe from the catalog, and keeping it beside the app means it travels with every path that already +moves a stack dir. + +### 5.3 The four writers — the only acts entitled to move a version + +| writer | pin source | +|---|---| +| the deploy path (`runComposeDeploy`) | the template just deployed from | +| **`UpdateStack`** | the catalog's current template, written **BEFORE** the pull | +| the restore (`stackAdapter.RecreateStackDefinitionFromUnit`) | the recovery unit's captured compose — **this closes R-441** | +| `Manager.AdoptPins` | the observation, once, and only when complete AND matching | + +**`UpdateStack`'s ordering is load-bearing, not stylistic.** `compose pull` and `up -d` act on the file +on disk, so the catalog's definition has to BE that file before either runs. A pin set afterwards +would pull the frozen version and change nothing — while reporting success, and a button that lies is +worse than a button that refuses. **A failed pin write REFUSES the update**, which is the opposite of +`recordInstalledImages` and for the same reason `desired_state` refuses: this field is intent. + +### 5.4 The render table, complete + +| app state | result | +|---|---| +| not deployed / protected / seam not wired | the catalog template — today's behaviour | +| deployed, **unpinned** | the catalog template + one DEBUG | +| deployed, pinned, catalog images **equal** | the catalog template — **fixes flow, self-healing works** | +| deployed, pinned, catalog images **differ** | the **stored applied definition** — frozen WHOLE | +| pinned, differ, nothing stored | the catalog template + one WARN. We cannot freeze what we do not have and must not invent it | +| mid-deploy | the compose file is left alone this cycle | + +**`.felhom.yml` is copied verbatim in every case** — it carries no image, and it carries +`catalog_since`, which the badge needs. See §8.5. + +**The frozen branch writes a WHOLE file and never a substitution.** Taking the new template and +putting the old refs back creates a third state nobody chose: `wger 2.6` needs a full DB configuration +the older template cannot supply, so an old image under a new template is broken in a way neither +version is. + +**And this is not "skip deployed apps".** That option was considered and rejected: it also stops +health-check fixes, memory limits and new deploy fields, and it destroys the self-healing measured in +the spike §3 — both halves the ruling explicitly kept. + +### 5.5 Adoption, and why the startup order matters + +`AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every deployed app +to what it is already running. **It reads and writes files only** — no container is started, stopped or +touched. It skips, loudly, when the observation is incomplete or when the app runs something the +current template no longer offers; those apps keep pre-v0.235.0 behaviour rather than receive a +guessed pin. + +**`syncer.Start()` was moved to after adoption.** It fires an immediate sync; at its previous position +that first sync ran while every app was still unpinned, copied the catalog over a deployed app, and +handed the next restart a version change — the exact behaviour this slice removes, once per boot. + +### 5.6 The trap this slice set for the previous one + +`Stack.TemplateImages` is read from the app's **live** compose file — which is now the RENDERED one. +On a frozen app that file names the OLD version, so `web.compareInstalledToTemplate` would find +installed == template and answer **„Naprakész" on exactly the apps that are behind** — with every test +still green, because the new field has the same type and shape. The badge now reads +`Stack.CatalogImages`, taken from the syncer's own git clone. **A feature that silently inverts an +earlier feature is the failure mode to look for whenever a file changes meaning.** --- @@ -141,7 +248,7 @@ prerequisite for judging how urgent it is. | **1** | **The box records what it actually installed** — `app.yaml.installed_images`, per compose service, ref + digest + first-seen. | **SHIPPED, controller v0.233.0 (2026-09-02)** | | **1b** | **Seed the record for apps nobody touches** — a startup backfill, so the label is not restricted to apps that happen to get restarted. | **SHIPPED, controller v0.234.0 (2026-09-03)** | | **2** | **One badge says whether the app is current** — „Naprakész" / „Frissítés elérhető — N napja", from `catalog_since`. No version number. | **SHIPPED, controller v0.233.0 + catalog `69761cf` (2026-09-02)** | -| **3** | **The compose file becomes DERIVED** — stop the syncer overwriting a deployed app's file; the pin in `app.yaml` wins. Needs the operator's ruling on R-438 first. | OPEN — R-447 | +| **3** | **The compose file becomes DERIVED** — the pin in `app.yaml` wins; the syncer renders instead of copying. | **SHIPPED, controller v0.235.0 (2026-09-06)** — operator ruling §3.4 | | **4** | **A guarded update** — verified-backup precondition, abort-on-failure, and the truth at the moment of action rather than 5m16s later (R-443). | OPEN — R-448 | | **5** | **An upgrade test** — prove a real one-major upgrade end to end, including the abort path. | OPEN — R-449 | | **6** | **A version sequence** — updates automatic within a major, a human across one; **an engine change gets its own edge.** | OPEN — R-450 | @@ -235,7 +342,18 @@ Version strings stay in the logs, the API and the hub. startup by reading containers (§7.1). **The residue that stays:** the seed happens at controller START, so a box between upgrade and its next restart still shows nothing — bounded by one restart rather than unbounded. -4. **The hub does not record image tags at all.** Its report's container payload carries name, state, +4. **A frozen app is frozen WHOLE.** While the catalog is ahead, **no** template correction reaches + that app — not even one unrelated to the version. That is the direct consequence of §3.4 and of the + `wger 2.6` hazard, and it is the right trade: a new template around an old image is a third broken + state. Recorded so it is a choice, not a surprise. +5. **`.felhom.yml` keeps flowing while the compose file is frozen** — the deliberate asymmetry in + §5.4. So a frozen app can receive a health check written for a NEWER version and read as degraded. + **The failure direction is a false alarm, never data loss**, and freezing `.felhom.yml` would break + the update badge by withholding `catalog_since`. **R-458.** +6. **The Update button is still unguarded.** It takes no backup, has no rollback, and can still + attempt a multi-major jump the app will refuse (R-40). **Slice 3 did not change that and must not + be read as having done so** — the precondition is slice 4 (R-448). +7. **The hub does not record image tags at all.** Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported. diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index fb05166e..a5ed257a 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,10 @@ --- +| **R-438** | **The catalog sync rewrote a DEPLOYED app's `docker-compose.yml` and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06).** `Syncer.copyTemplates` copied into every stack folder on a 15-minute cycle with **no deployed check**, so a deployed app's file and its running containers disagreed from that moment, and the next `compose up -d` from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s **with a pull**, and a boot reconciliation upgraded it **with nobody pressing anything**. **Reasoning kept — the distinction this row existed to protect:** `RestartStack`'s use of `up -d` to pick up template changes was **CHOSEN and written down in its own comment**, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. **Reasoning kept — one fear was measured SMALLER than stated:** a plain power cut does NOT upgrade anything, because Docker's `restart: unless-stopped` restores the containers on the old image and the reconciler logs `no boot-orphaned apps`; the unattended upgrade needs the narrower precondition *"and the app did not come back"*. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0** | `audits/SPIKE-app-update-2026-09-01.md` §2, §3, §8; `architecture/09-update-architecture.md`; `tests/VALIDATION-update-slice3-2026-09-06.md` | +| **R-447** | **UPDATE ARC SLICE 3 — the live compose file is now DERIVED; the syncer renders instead of copying. SHIPPED controller v0.235.0 (2026-09-06).** The pin lives in `app.yaml` (`pinned_images`), the definition it came from is stored beside the app as `applied-compose.yml`, and `Syncer.renderSource` writes the catalog template verbatim while the catalog still offers the pinned version and the stored definition once it moves past it. **Reasoning kept — the ruling, in the operator's own words:** *while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.* **Reasoning kept — why nothing was added to the thirteen `compose up -d` call sites:** most of them are REPAIRS (the boot reconciler, the drive-return gate, the app-stop guard), and **a repair path that refuses to repair leaves a customer's app down, which is worse than the problem**; they were made safe by removing the reason, not by gating them. **Reasoning kept — why the frozen branch writes a WHOLE file and never a substitution:** `wger 2.6` needs a full DB configuration the older template cannot supply, so an old image under a new template is a third state nobody chose. **Reasoning kept — why this is not "skip deployed apps" (option B, rejected):** that also stops health-check fixes, memory limits and new deploy fields, and destroys the self-healing measured live in the spike §3. **Reasoning kept — `pinned_images` is INTENT and `installed_images` is an OBSERVATION; never feed one from the other** (the R-166 category error, one field over). Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — SHIPPED controller v0.235.0** | `architecture/09-update-architecture.md` §3.4, §5; `tests/VALIDATION-update-slice3-2026-09-06.md`; controller `CHANGELOG.md` v0.235.0 | +| **R-441** | **The restore path and the catalog sync disagreed about the image, and the SYNC WON within 15 minutes. CLOSED controller v0.235.0 (2026-09-06).** `stackAdapter.RecreateStackDefinitionFromUnit` wrote the recovery unit's captured `docker-compose.yml` — carrying the OLD pin — into the live stack dir, and `Syncer.copyIfChanged` overwrote it from the catalog on the next tick, so a restore's image-level recovery had a **<=15-minute half-life**. The restore now PINS to what the unit captured and stores it as the applied definition, so the render obeys the restored file instead. **Reasoning kept:** the overwrite half was MEASURED live 2026-09-01 (`[INFO] [sync] Updated bentopdf/docker-compose.yml` at 18:10:29Z over a locally-modified file); the "the restore writes to that same path" half was READ, and this closure rests on the render behaviour being measured live rather than on the reading. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — SHIPPED controller v0.235.0** | `internal/stacks/pin.go`; `TestGroupF_RestorePinIsReportedToTheSyncer`; `tests/VALIDATION-update-slice3-2026-09-06.md` | +| **R-455** | **DooPlex had no Docker Hub login, and the unauthenticated ceiling blocked a BUILD rather than only a gate. CLOSED 2026-09-06 — a Docker Hub PAT was added to the credentials file (operator).** Measured 2026-09-02: `build.sh 0.233.0 --push` failed at `429 Too Many Requests` resolving `debian:bookworm-slim`, with neither base image in the local store. **Reasoning kept — the workaround and its verification, because the answer outlives the incident:** both bases were pulled from Google's official Docker Hub mirror (`mirror.gcr.io/library/...`) and retagged, and the identity claim was later MEASURED rather than assumed — `docker pull docker.io/library/` answered `Status: Image is up to date` for both, i.e. Hub's own manifest resolved to the images already local, and the `docker manifest inspect` bodies were identical between the registries. **So the mirror is a verified-sound fallback if the PAT is ever unavailable.** Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — credential added by the operator** | `felhom-controller/REPORT.md` (2026-09-03) §5 | | **R-429** | **CORRECTED 2026-09-01 — the snapshots ARE being taken; what was broken was that nobody could tell, and my own probe looked for the wrong name.** Viktor read the control panel on 2026-09-01: **seven automatic daily snapshots** on `storage-box-pool-1` (plan BX11), six days old to ~9 h old, filesystem ~2.5 GB, per-snapshot 0–14 MB, *Display snapshot directory* ON. **The mitigation works.** **MY ERROR, NAMED:** the 2026-09-01 spike probed for a directory called `.snapshots`; the vendor documents the path as **`/.zfs/snapshot`**. The probe's controls were sound and its subject was wrong, so `not found` was true and meant nothing. **The task that set the spike asserted the mechanism without citing the vendor documentation, and I did not check it** — that is how a correct instrument produced a wrong headline. **WHAT REMAINS TRUE, and is the actual finding:** the row claiming it had no R-number so nothing could cite it; its "confirm tomorrow" went **36 days** unanswered; and the `DUE-CHECKS` block built for that exact class (R-341) was **empty**. **The finding was never the snapshots. It was that nobody could tell.** Re-probed at the documented path — see R-432 for what a sub-account can actually reach. Evidence: `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1 and this row. | **CLOSED 2026-09-01 — mitigation CONFIRMED WORKING; the visibility gap is R-432** | panel read 2026-09-01 (Viktor); re-probe at `/.zfs/snapshot` in `audits/SPIKE-r95-offsite-delete-2026-09-01.md` | | **R-431** | **An unexplained fall in a customer's off-site snapshot count is now noticed within a day — SHIPPED hub v0.111.0.** Third signal in `OffsiteChecker`, beside FILL and STALENESS. **It lives on the HUB deliberately:** the event being detected is a box deleting its own backups, so a detector on that box is one the same event can silence; the hub already receives the count and keeps the history. **Threshold: a fall of more than HALF the previous count and at least 5** — reasoned, not invented, because the measurement gave nothing to calibrate against: over 12 898 reports (2026-06-05 → 2026-09-01) every one of the nine decreases lands exactly on ZERO and every one predates `stats_known` (the R-331 shape), and in the 380-report window where `stats_known` is true there are **zero** decreases. Retention keeps 7 daily + 4 weekly + 6 monthly per group, so it **cannot halve a total**; a mass deletion goes to ~0. **Three pre-conditions, each with its scar:** `StatsKnown` (R-331 — a zero is not a zero when unmeasured), the declared `State` (R-204 — the box names its own situation), and run success (R-100 — presence is not success; `incomplete` excluded too). An untrustworthy report neither alarms nor moves the baseline. **ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms**, fixture committed. Severity `error`, operator-only, no customer template. | **CLOSED 2026-09-01 — SHIPPED hub v0.111.0** | `hub/internal/monitor/offsite.go`; `offsite_r431_test.go` incl. 9 009-point real-history replay; hub v0.111.0 | | **R-419** | **`observations_gate.py` accepted an observation whose body merely CONTAINED the string `NOT-A-FINDING`, even in prose disclaiming it.** Found by accident on 2026-09-01 when a planted test observation reading *"it carries no `FILED:` and no `NOT-A-FINDING:` marker"* was reported `OK 1. NOT-A-FINDING` and a real push went green over an unfiled finding. The gate's whole job is to force an explicit choice, and a sentence disclaiming the choice counted as making it. **FIXED 2026-09-01:** a marker must now start a line or follow a sentence boundary, and inline code spans are stripped before matching — a marker inside backticks is being talked about, never used. | **CLOSED — FIXED + PINNED** (2026-09-01, R-421 sweep) | verified in BOTH directions: the decoy and a backticked mention are convicted; a real `**FILED: R-419**` and a real `**NOT-A-FINDING: ...**` still pass. Decoy kept in `felhom.eu/scripts/test_gate_decoys.py` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 11f24360..2673eab4 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -675,16 +675,13 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** | | **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** | | **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** | -| **R-438** | **[P1-HIGH] The catalog sync rewrites a DEPLOYED app's `docker-compose.yml`, and no architecture document records that it does.** `felhom-controller/controller/internal/sync/sync.go`, `Syncer.copyTemplates`, copies `docker-compose.yml` and `.felhom.yml` into EVERY stack folder on a 15-minute cycle (`internal/config/config.go:351`, default `15m`, confirmed 2026-09-01). **The loop does not test whether the app is deployed** — the only guard is a sha256 content compare in `copyIfChanged`, and the only exclusion is `app.yaml`. From the moment it runs, a deployed app's compose file and its running containers disagree, and the next `compose up -d` from ANY source resolves that disagreement without asking anyone. `documentation/architecture/02-controller-module-map.md` describes the syncer accurately (*"copy compose + `.felhom.yml`, never overwrite app.yaml"*) and stops before the consequence; `00-capability-map.md` records the lifecycle actions as PROVEN-LIVE and says nothing about what they do to app data. **The consequence appears in NO architecture document and in no register row until this one.** **This is the mechanism behind R-40.** **NOT CALLED A DEFECT: it may have been chosen** — `Manager.RestartStack` carries an explicit in-code comment saying `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*, which is a stated intent for exactly this behaviour on the RESTART path. Whether that intent extends to the unattended paths is the operator's ruling to make. Evidence attached by Phase 1/2 of `audits/SPIKE-app-update-2026-09-01.md`. **MEASURED LIVE 2026-09-01 — CONFIRMED, and the mechanism is now attributed to an exact symbol.** A real catalog pin change (`bentopdf` v2.8.6 -> v2.8.5, commit `214d448`) travelled the real 15-minute cycle: at **17:45:17Z** `[INFO] [sync] Updated bentopdf/docker-compose.yml` rewrote the DEPLOYED app's file (mtime 17:45:17.646) while the container went on running v2.8.6 (started 17:36:35Z, unchanged). **Nothing told the customer** — no event, no notification, no email, and the customer's own pages carry NO version string at all (searched with ASCII fragments and BOTH controls; a first pass using an unescaped `.` over-counted and was corrected with `grep -F`). **THE CONSEQUENCE IS ALSO MEASURED:** with the file moved, `POST /api/stacks/bentopdf/restart` upgraded the container — **18.3 s and a network PULL** when the target image was absent, 0.5 s when present — and a **boot reconciliation upgraded it with NOBODY PRESSING ANYTHING** (`bootrecon.go:259` -> `StartStack` -> `compose up -d`). **THE DESIGN INTENT IS ALREADY IN THE SOURCE and it narrows this row:** `Manager.RestartStack` comments that `up -d` is used *"so that ... any template changes (new images, healthchecks) are picked up"*. So the RESTART half was chosen and written down; what is recorded nowhere is what the syncer then does to a deployed app, and whether the choice was meant to extend to the 13 UNATTENDED call sites. **AND ONE FEAR IS MEASURED SMALLER THAN FEARED:** a plain power cut does NOT upgrade — Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. The unattended upgrade needs the narrower precondition *"and the app did not come back"*. `audits/SPIKE-app-update-2026-09-01.md` **THE DOCUMENT HALF IS NOW DISCHARGED, 2026-09-02: `documentation/architecture/09-update-architecture.md` exists and is a LIVING document, updated by every slice of this arc.** It records the mechanism as measured (§1), quotes the `RestartStack` comment that proves the restart half was CHOSEN (§2), carries the three operator rulings of 2026-09-02 (§3), strikes the word "rollback" (§4), states the target shape (§5) and lists the seven slices with a status each (§6). **THE ROW STAYS OPEN AND THE REASON IS THE POINT: the mechanism is now DOCUMENTED, not CHANGED.** Whether the syncer should go on overwriting a deployed app's compose file is still the operator's ruling, and acting on it is slice 3 (R-447). | **OPEN — rank P1-HIGH; owner: VIKTOR rules, CC measures** | | **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | -| **R-441** | **[P2-MEDIUM] The restore path and the catalog sync disagree about which image the app should run, and the SYNC WINS within 15 minutes.** `stackAdapter.RecreateStackDefinitionFromUnit` (`felhom-controller/controller/cmd/controller/main.go:2570`) writes the recovery unit's CAPTURED `docker-compose.yml` — carrying the OLD image pin — straight into the live stack dir, and `restore_unit.go:317` states the intent: *"Resolved from the UNIT's compose, because that file is about to BECOME the live one."* But `Syncer.copyIfChanged` overwrites any stack file whose content differs from the catalog, on the next 15-minute tick, with no deployed check (R-438). **So a restore's image-level rollback has a <=15-minute half-life, and the next `compose up -d` from any of the 13 unattended call sites re-applies the catalog pin.** **GRADED HONESTLY — the two halves have different evidence:** the overwrite is **MEASURED** (a locally-modified compose on demo-hp was overwritten by the sync at 18:10:29Z, `[INFO] [sync] Updated bentopdf/docker-compose.yml`); that the restore writes to that same path is **READ, not measured**. Settling it needs one live restore with a stale pin, which is a phase, not a check. **Why it matters more than it reads:** R-361's undo copy plus this is the only route back that exists, and Phase 6 proved putting the old TAG back is not a rollback at all (R-443's sibling finding) — so the data restore is the whole remedy, and it is fighting the syncer. Owner: **CC to measure, Viktor to rule on which wins.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** | | **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** | | **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | | **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** | -| **R-447** | **[P1-HIGH] UPDATE ARC SLICE 3 — make the live compose file DERIVED, so the syncer stops changing a deployed app's version.** The target shape (`architecture/09-update-architecture.md` §5): the pin lives in `app.yaml`, the one file `Syncer.copyTemplates` never touches, and the live `docker-compose.yml` becomes an OUTPUT rather than an input. The catalog then proposes and `app.yaml` decides, and the thirteen unattended `compose up -d` call sites (spike §8) stop being able to change a version by accident. **BLOCKED ON AN OPERATOR RULING, and that is the whole reason this is a separate slice:** R-438 established that the restart half of this behaviour was CHOSEN and written down in `Manager.RestartStack`'s own comment. Changing it is not a bug fix; it is reversing a decision, and the decision-maker is Viktor. **Slices 1 and 2 shipped first ON PURPOSE** — the fleet's real state has to be visible before anyone can judge how urgent this is. Do NOT add a deployed check to any of the thirteen paths ahead of the ruling. `architecture/09-update-architecture.md` §5 | **BLOCKED — rank P1-HIGH; owner: VIKTOR rules, CC implements** | | **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** | | **R-449** | **[P2-MEDIUM] UPDATE ARC SLICE 5 — an upgrade test that proves a real one-major upgrade end to end, INCLUDING the abort path.** Spike §7 performed the pieces by hand on Nextcloud (31.0.14 → 32.0.9 ran the migration; 31.0.14 → 34.0.1 was refused by the app and reported as SUCCESS by the product; the downgrade attempt was refused with the data intact). **None of it is a test that runs again.** A one-off measurement that nothing repeats decays into a claim — this project's most-repeated defect class. The test must assert the CONSEQUENCE (does the app serve after the upgrade? does the abort put it back?), not the mechanism. Venue: a Tier-0 box; `runbooks/target-selection.md` names which. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: CC** | | **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** | @@ -692,9 +689,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** | | **R-453** | **[P3-LOW] `~/.config/credentials` quotes its values with SINGLE quotes, and a half-applied strip cost a session an hour and produced a WRONG diagnosis that was published before the operator corrected it.** MEASURED 2026-09-02: extracting `PASSWORD` with `sed 's/^"//;s/"$//'` — double quotes only — sent the literal `'` characters as part of the password, and `POST /login` returned **HTTP 200 + `Hibás jelszó`** on BOTH demo boxes. **THE DIAGNOSIS THAT FOLLOWED WAS WRONG AND WAS WRITTEN INTO A REGISTER ROW, A STATUS ITEM AND A MEMORY BEFORE IT WAS CHECKED:** "the vaulted password is stale on both boxes". The operator answered in one line — *the value is in single quotes* — and one retry with `sed "s/^['\\"]//;s/['\\"]$//"` returned **302 + `felhom_session`**. **THE INSTRUMENTATION LESSON, which is the actual finding and outlives the typo:** the controller's own log line `auth.go:176 [WARN] Failed login` was quoted as the discriminator, and it IS a true and useful one — **it separates "wrong password" from "wrong Host header", and that is ALL it separates.** It cannot distinguish a wrong password from wrong password HANDLING, and it was read as if it could. A discriminator that rules out one alternative is not a discriminator that rules in the remaining one. **This is the SECOND time this exact file's quoting has produced a confidently wrong verdict** — the memory `credentials-file-values-are-quoted` was minted for the first (`cut -d=` keeps the quotes → a wrong "the password is stale" diagnosis), and this session applied that memory HALF, stripping one quote character and not the other. **The fix is an instrument, not a resolution to be careful:** one shared helper that extracts a value from that file correctly, used everywhere, so the next session cannot get it half right. Evidence: `tests/VALIDATION-update-slice12-2026-09-02.md` §4.0, which states the correction rather than quietly removing the claim. | **READY — rank P3-LOW; owner: CC** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | -| **R-455** | **[P2-MEDIUM] DooPlex has no Docker Hub login, and the unauthenticated ceiling now blocks BUILDS, not just gates.** MEASURED 2026-09-02: `build.sh 0.233.0 --push` failed at `[internal] load metadata for docker.io/library/debian:bookworm-slim: 429 Too Many Requests`, with `golang:1.24-bookworm` cancelled behind it. Neither base image was in the local store, so there was no fallback. The same window also left the catalog's `image-resolvable` gate INCONCLUSIVE (6 of 65 lookups throttled). **This is R-41's known note — *"the full sweep is still OWED — DooPlex is not logged in to Docker Hub"* — arriving with a bigger bill: it is no longer only a gate that cannot reach a verdict, it is a release that cannot be built.** **WORKAROUND USED AND RECORDED RATHER THAN BURIED:** both base images were pulled from Google's official Docker Hub mirror (`mirror.gcr.io/library/...`) and retagged locally, after which `build.sh` ran unmodified; the two digests are written into `felhom-controller/REPORT.md` §5 so the identity check against Hub is one command when the window clears. **THE IDENTITY CLAIM IS NOW MEASURED, not left as an IOU:** the throttle cleared 40 minutes later and `docker pull docker.io/library/debian:bookworm-slim` and `…/golang:1.24-bookworm` both answered **`Status: Image is up to date`** — Docker Hub's own manifest resolved to the images already local, i.e. the ones the mirror supplied and v0.233.0 was built from; and the `docker manifest inspect` bodies are identical between the two registries (`959bc47a76ff713a…` debian, `de3a17b36657e232…` golang). **So the shipped image is byte-for-byte what a Docker-Hub build would have produced, and the mirror is a sound source — which is what makes option (b) below a real option rather than a hope.** **The fix is a credential, and it is a decision:** a Docker Hub account for DooPlex (free tier lifts the ceiling ~6x for authenticated pulls), or an explicit standing ruling that the mirror is the sanctioned source and `build.sh` should name it. **A build that depends on an anonymous third-party quota is not a build you can run when you need to.** Owner: **VIKTOR decides, CC executes.** | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR decides, CC executes** | | **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | +| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** |