SPIKE R-459: the skipped MariaDB conversion is stable, and the trade it implied does not exist
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
Outcome A, qualified. Not B and not C. It does not degrade: 5 of 5 restarts of 12.3 on an 11.6 datadir, readback passed every time, mariadb_upgrade_info unchanged, the entrypoint line never escalated past [Note]. It also never heals - the engine answers 'Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!' on every start and will forever. The trade R-459 was expected to produce is not real. Converting properly SUCCEEDS across the multi-major jump, takes 7 seconds, backs up the system database unasked - and putting 11.6 back afterwards STILL starts and serves the data. So the operator is being handed a cheap correction, not a choice between a correct engine and a reversible one. The exit-code polarity was measured rather than read: 0 means the upgrade IS needed, 1 means it is not. Assuming either the flag name or the polarity would have inverted the headline. And run without credentials the same command returns a confident-looking FATAL ERROR that is an auth failure. R-464: after converting and going back, the entrypoint prints 'MariaDB upgrade not required' on a state the same engine calls an unsupported downgrade. The obvious cheap instrument for R-459 would have been to grep for that line, and it would have reported fine for the broken case. R-463: the PostgreSQL analogue, deliberately NOT measured here. 11 templates, 8 on postgres:16-alpine, register grep for pg_upgrade returns zero. The two engines fail in OPPOSITE directions - MariaDB skips quietly, Postgres refuses to start - so that one cannot hide; it presents as eight apps down at once. No template changed. Teardown all three layers, hub checked rather than asserted, local-lvm 30.53 percent before and after.
This commit is contained in:
@@ -691,10 +691,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-459** | **[P2-MEDIUM] Our own bookstack template moves MariaDB across a MAJOR while the image tells us, in its own words, that the datadir upgrade it needs is being SKIPPED.** MEASURED 2026-09-06 in a throwaway guest, on the catalog's own transition `0b73e5e` (mariadb `11.6` → `12.3`): the moment 12.3 starts on the 11.6 datadir it logs **`[Note] [Entrypoint]: MariaDB upgrade (mariadb-upgrade or creating healthcheck users) required, but skipped due to $MARIADB_AUTO_UPGRADE setting`** — and then serves normally. `templates/bookstack/docker-compose.yml` sets **no `MARIADB_*` environment at all**, so `MARIADB_AUTO_UPGRADE` is unset and the entrypoint declines to run `mariadb-upgrade`. **THE CAUSE IS ASSIGNED, NOT GUESSED:** the edge was decomposed, and the app half alone (`E3a`, engine held at 11.6) produces **no** upgrade line while both edges that move the engine (`E3`, `E3b`) produce it. A bundled edge could never have said which half. **AND IT EXPLAINS A RESULT THAT LOOKED LIKE GOOD NEWS:** the abort of E3 "worked" — 11.6 came back and served the data, logging `MariaDB upgrade not required` — **because the datadir was never converted.** The reversibility is a side-effect of an upgrade that did not fully happen. **WHAT IS NOT ESTABLISHED, and this row must not be read past: whether running 12.3 on an unconverted 11.6 datadir ever actually breaks.** It did not break here. MariaDB calls the upgrade required; we measured that it is skipped and did **not** measure a consequence. **What would settle it, cheaply:** run the E3 edge, then restart the stack several times and exercise the app, watching for the entrypoint's own complaint to become an error — one guest, no new mechanism. **The fix, if the consequence is real, is one line of template env**, but setting `MARIADB_AUTO_UPGRADE` fleet-wide is a change to how every MariaDB app upgrades and is the operator's call, not a quiet edit. Owner: **CC measures, VIKTOR rules on the fleet-wide env.** `audits/SPIKE-upgrade-test-2026-09-06.md` §4 | **READY — rank P2-MEDIUM; owner: CC measures, VIKTOR rules** |
|
||||
| **R-459** | **[P3-LOW, was P2] MEASURED 2026-09-06 — the skipped MariaDB conversion is STABLE but never self-resolving, and fixing it costs 7 SECONDS and does NOT cost the ability to go back.** The consequence R-459 deliberately left unmeasured is now measured: `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md`. **(1) It does not degrade: 5 of 5 restarts of 12.3 on the 11.6 datadir, readback passed every time, `mariadb_upgrade_info` unchanged at `11.6.2-MariaDB`, and the entrypoint's line never escalated past `[Note]`.** **(2) It never heals either** — the engine answers `Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!` on every start, and will forever. **(3) THE TRADE THIS ROW WAS EXPECTED TO PRODUCE DOES NOT EXIST.** The fear was that converting properly would end the ability to abort. Measured: with `MARIADB_AUTO_UPGRADE=1` the conversion **succeeds** across the multi-major jump (so it is not a stepping problem either), takes **7 s**, takes its own system-database backup first (`system_mysql_backup_11.6.2-MariaDB.sql.zst`), and **putting 11.6 back afterwards still starts and serves the data**. **So the choice is no longer a trade; it is a cheap correction.** **WHAT REMAINS OPEN IS THE DECISION, NOT THE MEASUREMENT:** setting `MARIADB_AUTO_UPGRADE` changes how **four** apps upgrade — bookstack (the only one that has moved a major), kimai, nextcloud, romm — and R-459's original owner line reserves a fleet-wide env change for the operator. **Nothing was committed to any template**; the comparison arm used a scratch copy inside a throwaway guest. **NOT ESTABLISHED, and not smuggled in: whether any specific MariaDB FEATURE misbehaves on unconverted system tables.** This run exercised BookStack's normal read/write path only, over minutes, with one seeded record. **Rank dropped P2 → P3** because the failure mode is now bounded by measurement rather than unknown. Raised in `STATUS.md`. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: VIKTOR rules on the fleet-wide env, CC implements** |
|
||||
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-461** | **[P3-LOW] `runbooks/target-selection.md` names a venue that does not exist and fences a fixture that is gone.** MEASURED 2026-09-06 on demo-hp while siting the R-449 guest. (a) The runbook says to put VM disks on a dir storage at **`/mnt/nvme-1tb`, at its root**. **There is no `/mnt/nvme-1tb`** — the 1 TB NVMe is mounted at **`/mnt/hdd_1`**, which is the enrolled user-data drive and the `felhom-backup` target, i.e. the same disk under a different path. The instruction's REASON is still exactly right (`local-lvm` is an over-subscribed thin pool backing the live guest 9201, and this run kept off it — `local-lvm` read **30.50 % before and after**), so the fence held; only its address is stale. (b) The runbook forbids destroying **`drill-r50` (VM 300)**, "the only drift fixture (R-93)". **`qm list` returns nothing on demo-hp** — there are no VMs at all. **The fence currently protects nothing, and R-93's premise that a drift fixture exists is false.** **Why it is a row and not a quiet edit:** a runbook that names a path nobody can find is one a session works around, and working around a safety instruction is how the instruction stops being followed. Both halves need checking against the box before the text is changed — (b) in particular may mean R-93 should be closed or reopened as "the drift fixture is gone", which is a different fact from "do not destroy it". Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §7 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** |
|
||||
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
@@ -52,7 +52,7 @@
|
||||
|
||||
| ID | Item | Size | Status | Notes |
|
||||
|----|------|------|--------|-------|
|
||||
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. |
|
||||
| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-459 was measured 2026-09-06 and narrowed to a decision, not a defect:** the skipped MariaDB conversion is stable but never self-resolving, and correcting it costs 7 s and does **not** cost the ability to abort — so the trade that row was expected to produce does not exist. Opened R-463 (the PostgreSQL analogue, which fails in the OPPOSITE direction — it refuses to start, and 8 templates sit on `postgres:16-alpine`) and R-464 (an entrypoint line that says `upgrade not required` on an unsupported downgrade). **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. |
|
||||
| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. |
|
||||
| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/<share>`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. |
|
||||
| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** |
|
||||
|
||||
Reference in New Issue
Block a user