Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s

Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).

Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-13 12:30:17 +02:00
parent abe567e14d
commit 5ef0f52bcd
46 changed files with 1029 additions and 285 deletions
+3
View File
@@ -258,3 +258,6 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
| **R-459** | **The skipped MariaDB conversion is STABLE but never self-resolving; converting costs 7 s and keeps the abort — RULED YES and shipped 2026-09-13: `MARIADB_AUTO_UPGRADE=1` on `bookstack-db`, `kimai-db`, `nextcloud-db`, `romm-db`** (catalog `eec1228`/`bd32830`/`3525e35`; no image moved, `catalog_since` untouched). Measured `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md`; proven `audits/r459-close-2026-09-13/` — harness E3/E3b `proven` with `engine_state_after` = `already upgraded to 12.3.3-MariaDB [exit=1]`, `skipped due to $MARIADB_AUTO_UPGRADE` 0 lines, C3 still `failed`; landed on demo-hp via the real 15-min cycle with both container IDs unchanged, one deliberate restart → `MariaDB upgrade not required`, `/login` 200. **Reasoning kept:** *ask the engine, not the log* — the entrypoint prints `MariaDB upgrade not required` on an unsupported downgrade too (R-464); `mariadb-upgrade --check-if-upgrade-is-needed` exit 0 = needed, 1 = not. *Not established, unchanged:* whether any MariaDB feature misbehaves on an UNCONVERTED datadir. *The precaution that keeps the setting inert until Slice 4:* the engine-major rule + gate, removal tracked as R-469. | **CLOSED 2026-09-13 — shipped in the catalog, PROVEN by harness and live** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
| **R-467** | **Controller v0.236.0 owed a golden — PAID 2026-09-13:** golden `0.236.0` baked (`GOLDEN_SHA256=58a3cc24…958bf`, 654 115 664 B), round-tripped from the DOWNLOADED bytes, `./etc/felhom-controller-image` says `felhom-controller:0.236.0`, hub dropdown agreed, three-field vouch re-read (`0.236.0` / `agent 0.130.0` / `min_agent 0.129.0`, not the R-216 shape), floor raised 0.232.0 → 0.236.0. Evidence `documentation/tests/golden-0.236.0-2026-09-13/`. **Reasoning kept:** *this bake carried FOUR unbaked releases and is the LAST per-release bake* — goldens are weekly and before any install from today (R-468); *the MinAgent line was missing from four headers* (R-470). | **CLOSED 2026-09-13 — baked, vouched, floor raised** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
| **R-448** | **UPDATE ARC SLICE 4 — a guarded update: verified-backup precondition, abort path, truth at the moment of action.** Shipped controller **v0.237.0** (job) + **v0.238.0** (page) + **v0.238.1** (nightly legs skip an app mid-update, found live). Proven live on demo-hp 2026-09-13, scenarios A/B/E/F/H and the restore walk. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *the precondition is the existing verified backup, not a new copy* (ruling 2026-09-02) — `backup.Tier2UnitRestorePoint`, extracted from the backups page, not copied; *age a copy by its last SUCCESSFUL Tier-2 copy, never the manifest `created_at`* (measured: the manifest moves only on definition changes); *no automatic rollback — measured per-app, the route back is the restore*; *anything that writes a restore point skips an app that is held OR updating*. Open consequences: R-472, R-475, R-476. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
| **R-443** | **The Update button reported success over an app it had just broken.** Closed by slice 4 (v0.237.0): 202 `accepted, not completed`; the outcome exists only as `update_phase` after health. Pinned by `TestR443_UpdateIsNeverReportedCompleteSynchronously`. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a compose exit code is never a success signal* (spike §4: HTTP 200 over a crash loop). | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
| **R-439** | **The restore hold was not honoured by the update path.** Closed by slice 4 (v0.237.0): `update` joined the router's hold check; live Scenario H refused start/restart/update and the boot sweep. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a hold that only one path honours is not a hold* — the audit found the drive-return gate (restart + boot recreate) and the nightly volume dump ignoring any hold; all three fixed and red-proofed. | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
+8 -5
View File
@@ -675,13 +675,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** |
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
| **R-439** | **[P3-LOW] The restore hold is not honoured by the update path.** The R-379/R-380 hold is checked in `felhom-controller/controller/internal/api/router.go`, `Router.actionStack`, under `if action == "start" || action == "restart"` — **`update` is absent from that check** and falls through to `Manager.UpdateStack`, which ends in `compose pull` + `compose up -d --remove-orphans`. The comment above `Manager.RestoreHoldFor` (`internal/backup/offbox_reconstitute.go:323`) states the design intent in terms: *"Every start path consults this — the customer's button, the app-stop Recover() starter, and the boot reconciler — because a hold that only one path honours is not a hold."* Update is a fourth path and does not honour it. **Severity LOW, and the reason is part of the row:** the UI only renders the Frissites button when the app is operational (`internal/web/templates/stacks.html`), and a held app is stopped, so a customer cannot reach this from the page. The API endpoint is ungated. **This is a defence-in-depth gap, not a customer-reachable bug.** One-line fix, taken because the hold's own design comment says so — and it needs a test pinning the invariant, or the comment stays a wish. CONFIRMED BY READING 2026-09-01 (`audits/SPIKE-app-update-2026-09-01.md`). **RE-READ AND CONFIRMED 2026-09-01; the severity argument SURVIVES but its stated reason was imprecise and is corrected here.** The task's reason was *"the UI only renders Frissites when the app is operational, and a held app is stopped"*. Half right: `isOperationalState` (`internal/web/funcmap.go:90`) counts **`StateRestarting` and `StateDegraded` as operational too**, and this was OBSERVED live — the green `Frissites` button rendered over a crash-looping app during the spike's Phase 3b. **So the button is hidden specifically because a held app is `StateStopped`, not because broken apps hide it.** LOW stands; the reason must be stated precisely or the next reader will widen it. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-443** | **[P2-MEDIUM] The Update button reports SUCCESS over an app it has just broken, and the truth arrives 5m16s later by a different road.** MEASURED on demo-hp 2026-09-01: `POST /api/stacks/bentopdf/update` against an image that pulls cleanly and then fails to run returned **HTTP 200 `{"ok":true,"message":"Stack bentopdf update completed"}`** and logged `Stack bentopdf updated successfully (took 3.5s)`, while the container went to `status=restarting RestartCount=9`. **The controller's own post-start line told the truth (`manager.go:1403 ... alpine:3.20 restarting`) — but it runs AFTER the API has already answered.** This is this repo's own `up -d` exits 0 on a crash-loop invariant surfacing at the customer's most consequential button. **What the customer's page then said:** badge **`Ujraindites...`**, `Restarting (0) 15 seconds ago`, and the full green button row — because `isOperationalState` counts `StateRestarting` as operational (see R-439). *"Restarting"* reads as transient, not as failure, and nothing says the update caused it. **THE HONEST OTHER HALF, and it must travel with this row: the customer IS told.** `app_start_failed` fired at 18:05:59Z with severity `warning` (inside the hub's exact vocabulary, so it really delivers) — 5m16s after the update, from `crashLoopAfter = 5 * time.Minute`, a threshold whose own comment argues it well. **So this is NOT the silent-dead-app class; it is a TRUTHFULNESS-AT-THE-MOMENT-OF-ACTION problem.** Also recorded: on a pull FAILURE the product behaves correctly — HTTP 500, and `compose up -d` resolves images before touching a container, so the running app survives (measured twice). Owner: **CC to propose, VIKTOR to rule on whether Update should wait and verify.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules** |
| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** |
| **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** |
| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` | **OPEN — rank P2-MEDIUM; owner: CC** |
| **R-448** | **[P2-MEDIUM] UPDATE ARC SLICE 4 — a guarded update: a verified backup as a precondition, an abort path, and the truth at the moment of action.** Three parts, each already evidenced. (a) **The precondition is a VERIFIED RECENT BACKUP, not a new copy** (operator ruling 2026-09-02); the guest-snapshot alternative must be SPIKED before anything is designed around it. Today's safety machinery is DATABASE-ONLY (`Manager.writeSafetyDump`, `internal/backup/offbox_reconstitute.go:207`) and the file half was never priced (spike §6). (b) **The abort path, never a "rollback"** — spike §7 proved the word is wrong: once a migration has run, the old image refuses to start on the migrated data. The two available shapes are ABORT (before anything migrated) and RESTORE FROM A COPY (after). (c) **Truth at the moment of action** — this subsumes **R-443**: the Update button returns HTTP 200 over an app it has just broken and the alarm arrives 5m16s later. `architecture/09-update-architecture.md` §3, §4 | **READY — rank P2-MEDIUM; owner: CC proposes, VIKTOR rules on (c)** |
| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 | **READY — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 | **READY — rank P3-LOW; owner: CC** |
| **R-452** | **[P3-LOW] Nothing enforces `catalog_since`, so the one number the update badge shows can silently under-report.** `app-catalog-felhom.eu` `CLAUDE.md` now states the rule — any commit that changes an `image:` line must set that app's `catalog_since` to the same day — and all 53 apps were backfilled from git history on 2026-09-02 (`69761cf`). **A rule with no instrument is a wish; that is this project's most-repeated finding and this row exists so it is not repeated silently.** A stale `catalog_since` makes „Frissítés elérhető — N napja" under-report N, which is the single number the badge exists to give. **WHY IT WAS NOT BUILT IN THE SAME SESSION, stated rather than implied:** the gate would have to diff an `image:` line against the PARENT commit, and `catalog_gates.py` runs under a runner that fetches at `--depth 1` — there is no parent to diff against. The gate therefore needs a deeper fetch, which is a change to the CI shape and not to a script. **This is the R-421 class in advance: an enumerated gap becomes a row in the same session it is enumerated.** `architecture/09-update-architecture.md` §8.2 | **READY — rank P3-LOW; owner: CC** |
@@ -698,12 +695,18 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-465** | **[P3-LOW] `cfg.Paths.HDDPath` — the global that R-442 proved is set on NO box (demo-hp AND demo-felhom: 0 `hdd_path`, 0 `FELHOM_PATHS_*`) — still has SIX readers, each reading an always-empty value:** `report/builder.go:69`, `monitor/healthcheck.go:35`, `api/router.go` (system-info), `web/server.go:740`, `cmd/controller/main.go` (auto-discovery seed + metrics HDD path). Removal was silently inert for months on the very same read. Whether any of these is inert the same way — a report field that is always empty, a health check that never fires, a metric never collected — is a one-hour audit: for each reader, name the POSITIVE observable that must appear when it works and check it on the box. R-442 §5 said "do not delete it here"; this row is the audit it deferred. Owner: **CC.** `felhom-controller/REPORT.md` (v0.236.0, Observations 1) | **READY — rank P3-LOW; owner: CC** |
| **R-466** | **[P3-LOW] Removing an app with „Mentési adatok törlése" ticked leaves its recovery unit's `compose/` + `manifest.json` on the drive.** The router passes only `backup.AppDBDumpPath(nsRoot, name)` to `RemoveStack`, so `backups/primary/<app>/db-dumps/` goes and the unit root keeps `compose/` (the app's `app.yaml` with the portable secret class at 0600) and `manifest.json`. MEASURED 2026-09-13 on demo-hp after the R-442 Scenario A removal — `audits/R442-2026-09-13/teardown-and-log.txt`, residue check: `backups/primary/nextcloud` still listed with the app gone; removed by hand at teardown. A customer who asked for the backups to go is left with the app's definition and a manifest. **Decide:** the button means the WHOLE unit (pass the unit root, under the same `backups/`-prefix guard) or stays db-dumps-only (then the modal must say so). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver.yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. | **BLOCKED — on R-448; rank P3-LOW; owner: CC** |
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** |
| **R-470** | **[P3-LOW] Four consecutive controller CHANGELOG headers (v0.233.0 … v0.236.0) carry NO `MinAgent:` line, and `RUNBOOK-manual-build.md` §4.1 step 5 tells the vouch to read it from the header.** Measured 2026-09-13 while baking golden 0.236.0: `sed -n 1,349p CHANGELOG.md | grep -c MinAgent` → **0**; the newest statement is v0.232.0's `**MinAgent: 0.129.0** (unchanged)`. The vouch used 0.129.0 on the ground that no later entry declares a change and `internal/agentapi/features.go` gates per feature, not per release — but a reader following the runbook literally finds nothing to read, which is the R-233 shape (a document pointing at a string that is not there). **Fix:** either every release header carries the line (the v0.232.0 convention), or the runbook says where MinAgent actually lives when a header omits it. One of the two, not both. | **READY — rank P3-LOW; owner: CC** |
| **R-471** | **[P3-LOW] `observations_gate.py` reads the FIRST `Observations` section of `REPORT.md` and nothing after it, so an appended second section with an unmarked item passes — and the decoy that proves it has been reading "LIVE HOLE" at HEAD, unnoticed, because the decoy suite is run by hand.** MEASURED 2026-09-13 on a clean worktree of `4b2e560`: `python3 scripts/test_gate_decoys.py` → `FAIL: observations/R-419: decoy PASSED - LIVE HOLE (rc=0)`; the gate's own output shows it scanned `## 11. Observations — noticed, documented, NOT acted on` (3 items, all marked) and never reached the appended `## Observations` block the decoy planted. Whether the cause is "first heading wins" or a heading-shape filter is NOT established — only that the planted unmarked item was not seen. **Consequence:** a report with two observation sections gets the second one unchecked. **Two fixes, both owed:** scan every section whose heading contains `Observations`, and put `test_gate_decoys.py` where something runs it (it is the instrument for R-421 and nothing in `repo_gates.py` invokes it). | **READY — rank P3-LOW; owner: CC** |
| **R-472** | **[P2-MEDIUM] THE GOLDEN CADENCE RULING AND THE HUB'S FLOOR RULE CONTRADICT EACH OTHER: a floor raise delivers NOTHING until a golden carries the release.** The 2026-09-13 ruling (R-468) rests on *"every release still raises the floor, so both demo boxes keep getting each release in ~20 s; only the golden moves to a cadence"*. MEASURED THE SAME DAY, releasing controller v0.237.0: `POST /configuration/global-floor` → `0.237.0` (303, re-read), and the hub logged for BOTH boxes `managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0, so its agent requirement is unknown — vouch a golden carrying the floor's controller (publish-train rule 1) (controller floor withheld)`; the box logged `SetFloor: floor "0.236.0" → ""`; neither box moved in 8 minutes. Publish-train rule 1 (hub `ResolveManagedFloor`) was built so a controller is never pushed past the agent it needs, and it reads that requirement from the vouched golden. **So under a weekly golden, releases between bakes do not reach the fleet by floor at all** — they need a hand deploy. v0.237.0 and v0.238.0 were hand-deployed to both demo guests by the skill's documented route and the floor was put back to 0.236.0 (a floor held on every box is a standing dashboard flag). Evidence `audits/slice4-2026-09-13/live/00-floor-0.237.0.txt`. **This is the operator's call, not CC's:** (a) bake per release again (the treadmill R-468 ended), or (b) let the floor carry a release above the golden when its CHANGELOG header states an unchanged `MinAgent` — which needs R-470's header line to be reliable first. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
| **R-473** | **[P2-MEDIUM] The `glance` catalog template crash-loops on EVERY fresh install.** MEASURED 2026-09-13 on demo-hp (deployed as a throwaway for the slice-4 live test): `restarts=13 status=restarting`, the log repeating `parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory`, the `glance_config` volume empty. The image does not create a default config and the template seeds none. The install page reports the deploy as started and the card then shows a restarting app. Not a version problem: the template's current tag v0.8.5 has the same config requirement (the catalog walk that pinned it never did a fresh install). **Fix shape:** seed a minimal `glance.yml` the way `gokapi`'s catalog entry seeds `config.json` (memory `gokapi-headless-setup`), then prove a fresh install lands healthy. Evidence `audits/slice4-2026-09-13/live/02-glance-abandoned.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). | **READY — rank P3-LOW; owner: CC** |
| **R-475** | **[P2-MEDIUM] The update precondition is Tier-2-ONLY, so an app with no Tier-2 copy cannot be updated at all — even when its primary unit or off-site copy could restore it.** Slice 4 (v0.237.0) gates the Update button on `backup.Tier2UnitRestorePoint`, exactly as specified. MEASURED on demo-hp 2026-09-13: `gokapi` and `nextcloud` have NO `cross_drive` record, so both are refused with the no-backup sentence. The classes this reaches: a box with no second target (`no_target`), an app whose customer switched Tier 2 off, and a freshly installed app before its first Tier-2 run. **It is a finding about the specification, not a defect in the build:** the primary recovery unit (same drive) and the off-site unit (Tier 3) are both restorable routes the predicate does not consider — the driveless-app class R-356 names. Options: accept Tier 2 as the only update route (and say so in the refusal), or widen the predicate to the primary unit with an explicit "same drive" caveat. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules on the route set, CC implements** |
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
Clearing a row means the check was DONE and its result recorded in that R-row —
+1 -1
View File
@@ -52,7 +52,7 @@
| ID | Item | Size | Status | Notes |
|----|------|------|--------|-------|
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-459 was measured 2026-09-06 and narrowed to a decision, not a defect:** the skipped MariaDB conversion is stable but never self-resolving, and correcting it costs 7 s and does **not** cost the ability to abort — so the trade that row was expected to produce does not exist. Opened R-463 (the PostgreSQL analogue, which fails in the OPPOSITE direction — it refuses to start, and 8 templates sit on `postgres:16-alpine`) and R-464 (an entrypoint line that says `upgrade not required` on an unsupported downgrade). **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. |
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** **2026-09-13 — SLICE 4 COLLAPSED: SHIPPED** (controller v0.237.0/v0.238.0/v0.238.1, R-448 CLOSED, proven live — `audits/slice4-2026-09-13/`). Remaining: slice 6 (R-450), slice 7 (R-451); R-469 unblocked. | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-459 was measured 2026-09-06 and narrowed to a decision, not a defect:** the skipped MariaDB conversion is stable but never self-resolving, and correcting it costs 7 s and does **not** cost the ability to abort — so the trade that row was expected to produce does not exist. Opened R-463 (the PostgreSQL analogue, which fails in the OPPOSITE direction — it refuses to start, and 8 templates sit on `postgres:16-alpine`) and R-464 (an entrypoint line that says `upgrade not required` on an unsupported downgrade). **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. |
| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. |
| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/<share>`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. |
| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** |