R-442 CLOSED (controller v0.236.0): "delete my data too" deletes it or refuses; R-465..467 opened
gates / gates (push) Failing after 18s
gates / gates (push) Failing after 18s
- OPEN-ITEMS: R-442 row removed; R-465 (six remaining Paths.HDDPath readers — audit), R-466 (recovery-unit residue after "delete backups"), R-467 (v0.236.0 owes a golden). - CLOSED-ITEMS: R-442 compressed, reasoning kept, fleet shape now established. - 00-capability-map: lifecycle row narrowed (Campaign 3 proved remove removes the APP, not the data) and re-proven from audits/R442-2026-09-13/. - STATUS: item 13 in plain language; item 7 closed (ruled 2026-09-02, 09 §3). - audits/R442-2026-09-13/: live evidence (A, C, D bodies, controls, log window, teardown). Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -26,6 +26,7 @@
|
||||
|
||||
---
|
||||
|
||||
| **R-442** | **`remove_hdd_data: true` was INERT — the customer's data stayed on the drive while the API reported success (HTTP 200, `hdd_paths_removed: null`, 128 MB of Nextcloud left; demo-hp 2026-09-01).** Shipped in controller **v0.236.0** (2026-09-13). Removal resolved the drive from the GLOBAL `cfg.Paths.HDDPath` (no default, set on NO box — demo-hp AND demo-felhom both measured 0 `hdd_path` / 0 `FELHOM_PATHS_*`, so the fleet shares the shape); it now reads the app's OWN `app.yaml` `HDD_PATH` — the `07-backup-architecture.md` ~L437 rule that deploy, the start gate and the backup destination already implemented — and a data removal it cannot resolve, or whose drive is absent, is REFUSED (409, exact Hungarian sentence, typed `stacks.RemoveRefusedError`) BEFORE `compose down`, app kept. SSD app → `hdd_paths_removed: []`, never `null`, plus `hdd_note`. Missing folders stated in `hdd_paths_missing`. The backup-half refusal reaches the response (`backup_paths_refused`) and its base follows the same rule. **Reasoning kept:** *"declares no drive" ≠ "could not resolve the drive" — the first is a fact, the second a refusal*; *an app gone with its data left behind is unrecoverable from the UI — the customer cannot even re-run the removal*; *no fallback to the global — that silent fallback is the exact path this closes*. **Observation carried:** 8 of the 13 `needs_hdd` catalog apps bind ONLY `${USERDATA_PATH}` (the shared library) and no `${HDD_PATH}` folder, so for them "delete my data" correctly removes nothing on the drive and the modal shows no checkbox. | **CLOSED 2026-09-13 — shipped controller v0.236.0, proven live on demo-hp** | `audits/R442-2026-09-13/` — A: 63 MB written by Nextcloud ITSELF, gone after removal and listed with its size; C: 409 + sentence, all 69 files untouched, app still deployed, `[ERROR] … refused` logged; D: gokapi `[]` + note; ASCII controls (`llap`: C=2 A=0 D=0). 15 tests + two red-proofs in `felhom-controller/REPORT.md`. Full original text: `git show d6837d98ee24:documentation/backlog/OPEN-ITEMS.md`. |
|
||||
| **R-449** | **UPDATE ARC SLICE 5 — an upgrade test that runs again. BUILT AND RUN 2026-09-06:** `app-catalog-felhom.eu/scripts/upgrade-test.py` + `upgrade_fixtures.py`, 7 edges across 3 apps, evidence in `audits/upgrade-spike-2026-09-06/`. **Reasoning kept — success is an APPLICATION-LEVEL READBACK, never file identity:** `survive2.py`'s sha256+inode rule is right for a redeploy and WRONG for an upgrade, because a migration is supposed to rewrite files and that rule would fail every correct upgrade. **Reasoning kept — nothing is ever seeded into a volume by hand (R-156);** an app with no non-browser route is recorded `inconclusive`, which is a result and not a licence to plant a file. **Reasoning kept — run the negative control FIRST:** C3's TO image exits immediately and came back `failed`; a harness that cannot fail a known-broken upgrade proves nothing with its greens. **WHAT IT MEASURED:** all five real catalog upgrades kept the customer's data; and **whether an upgrade can be undone is a property of the individual APP, not of upgrades** — docmost REFUSES (*"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*), privatebin does not, which independently reproduces the Nextcloud finding on a second app by a DIFFERENT mechanism and puts two measurements behind §4's ruling that "rollback" is the wrong word. **It also found a defect in our own catalog (R-459).** **What stays open, as its own rows rather than inside this one:** R-459 (the skipped MariaDB datadir upgrade), R-460 (bookstack's file half is unprovable headlessly), R-462 (the widening, costed). Full original text: `git show 417df06f3529:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — harness built, run, and proven by a red negative control** | `audits/SPIKE-upgrade-test-2026-09-06.md`; `audits/upgrade-spike-2026-09-06/evidence/`; catalog `0474ce387e6f` |
|
||||
| **R-438** | **The catalog sync rewrote a DEPLOYED app's `docker-compose.yml` and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06).** `Syncer.copyTemplates` copied into every stack folder on a 15-minute cycle with **no deployed check**, so a deployed app's file and its running containers disagreed from that moment, and the next `compose up -d` from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s **with a pull**, and a boot reconciliation upgraded it **with nobody pressing anything**. **Reasoning kept — the distinction this row existed to protect:** `RestartStack`'s use of `up -d` to pick up template changes was **CHOSEN and written down in its own comment**, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. **Reasoning kept — one fear was measured SMALLER than stated:** a plain power cut does NOT upgrade anything, because Docker's `restart: unless-stopped` restores the containers on the old image and the reconciler logs `no boot-orphaned apps`; the unattended upgrade needs the narrower precondition *"and the app did not come back"*. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0** | `audits/SPIKE-app-update-2026-09-01.md` §2, §3, §8; `architecture/09-update-architecture.md`; `tests/VALIDATION-update-slice3-2026-09-06.md` |
|
||||
| **R-447** | **UPDATE ARC SLICE 3 — the live compose file is now DERIVED; the syncer renders instead of copying. SHIPPED controller v0.235.0 (2026-09-06).** The pin lives in `app.yaml` (`pinned_images`), the definition it came from is stored beside the app as `applied-compose.yml`, and `Syncer.renderSource` writes the catalog template verbatim while the catalog still offers the pinned version and the stored definition once it moves past it. **Reasoning kept — the ruling, in the operator's own words:** *while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.* **Reasoning kept — why nothing was added to the thirteen `compose up -d` call sites:** most of them are REPAIRS (the boot reconciler, the drive-return gate, the app-stop guard), and **a repair path that refuses to repair leaves a customer's app down, which is worse than the problem**; they were made safe by removing the reason, not by gating them. **Reasoning kept — why the frozen branch writes a WHOLE file and never a substitution:** `wger 2.6` needs a full DB configuration the older template cannot supply, so an old image under a new template is a third state nobody chose. **Reasoning kept — why this is not "skip deployed apps" (option B, rejected):** that also stops health-check fixes, memory limits and new deploy fields, and destroys the self-healing measured live in the spike §3. **Reasoning kept — `pinned_images` is INTENT and `installed_images` is an OBSERVATION; never feed one from the other** (the R-166 category error, one field over). Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — SHIPPED controller v0.235.0** | `architecture/09-update-architecture.md` §3.4, §5; `tests/VALIDATION-update-slice3-2026-09-06.md`; controller `CHANGELOG.md` v0.235.0 |
|
||||
|
||||
@@ -677,7 +677,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
|
||||
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-442** | **[P1-HIGH] `remove_hdd_data: true` is INERT on a box whose `controller.yaml` has no `paths.hdd_path` — the customer's data stays on the drive and the API reports NEITHER removed NOR preserved.** MEASURED on demo-hp 2026-09-01: removing an app with `{"remove_hdd_data":true,"remove_backups":true}` returned **HTTP 200** with `"hdd_paths_removed":null,"hdd_paths_preserved":null` and left **128 MB** at `/mnt/felhom-drives/hdd_1/appdata/nextcloud`. **ROOT CAUSE, with controls:** `Paths.HDDPath` (`internal/config/config.go:117`) has **NO default** — only an env override at `:403` — and demo-hp's `controller.yaml` `paths:` block holds only `data_dir`, `stacks_dir`, `system_data_path`; the container has **no `FELHOM_PATHS_*` variable at all** (measured, count 0). So `cfg.Paths.HDDPath == ""` and `ParseComposeHDDMounts` (`internal/stacks/delete.go:600-603`) returns `nil` on its FIRST line — logging `found 0 HDD mounts` — for a compose that plainly contains `- ${HDD_PATH}/appdata/nextcloud:/var/www/html/data`. **The second half of the same removal ALSO no-op'd:** `[WARN] Refusing to remove backup path outside expected directory: /mnt/felhom-drives/hdd_1/backups/primary/nextcloud/db-dumps`. **Why P1:** a customer who removes an app and asks for the data to be deleted is told it worked, and it was not. This is a privacy answer, not a tidiness one. **NOT ESTABLISHED: whether the fleet shares this config shape** — demo-felhom and any customer box must be checked before sizing it. **The fix needs a test that FAILS when `hdd_path` is empty**, or the guard comes back. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P1-HIGH; owner: CC** |
|
||||
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
|
||||
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
|
||||
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
|
||||
@@ -697,6 +696,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** |
|
||||
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-465** | **[P3-LOW] `cfg.Paths.HDDPath` — the global that R-442 proved is set on NO box (demo-hp AND demo-felhom: 0 `hdd_path`, 0 `FELHOM_PATHS_*`) — still has SIX readers, each reading an always-empty value:** `report/builder.go:69`, `monitor/healthcheck.go:35`, `api/router.go` (system-info), `web/server.go:740`, `cmd/controller/main.go` (auto-discovery seed + metrics HDD path). Removal was silently inert for months on the very same read. Whether any of these is inert the same way — a report field that is always empty, a health check that never fires, a metric never collected — is a one-hour audit: for each reader, name the POSITIVE observable that must appear when it works and check it on the box. R-442 §5 said "do not delete it here"; this row is the audit it deferred. Owner: **CC.** `felhom-controller/REPORT.md` (v0.236.0, Observations 1) | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-466** | **[P3-LOW] Removing an app with „Mentési adatok törlése" ticked leaves its recovery unit's `compose/` + `manifest.json` on the drive.** The router passes only `backup.AppDBDumpPath(nsRoot, name)` to `RemoveStack`, so `backups/primary/<app>/db-dumps/` goes and the unit root keeps `compose/` (the app's `app.yaml` with the portable secret class at 0600) and `manifest.json`. MEASURED 2026-09-13 on demo-hp after the R-442 Scenario A removal — `audits/R442-2026-09-13/teardown-and-log.txt`, residue check: `backups/primary/nextcloud` still listed with the app gone; removed by hand at teardown. A customer who asked for the backups to go is left with the app's definition and a manifest. **Decide:** the button means the WHOLE unit (pass the unit root, under the same `backups/`-prefix guard) or stays db-dumps-only (then the modal must say so). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-467** | **[P3-LOW] Controller v0.236.0 owes a golden** (golden-notice gate, 2026-09-13): newest golden baked is **0.232.0**, so a machine installed now receives 0.232.0 and reaches 0.236.0 only by self-update. Bake + vouch per `runbooks/RUNBOOK-manual-build.md` §4.1 (the THREE-field vouch: `golden_version` + `agent_version` + `min_agent`); bake record to `documentation/tests/golden-<ver>-<date>/`. If the train deliberately skips this version, this row is the waiver the gate asks for. Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user