R-108 CLOSED — D5's precondition is met (controller v0.187.0)

Four-artifact update per the coupling rule, plus the audit.

07-backup-architecture.md: §10.1 retitled CLOSED with the ruling and the D5
sentence; the FileBrowser network-share row flipped YES->NO, closed at the
PLACEMENT rather than at the bind; the exposure chain annotated with the fifth
surface (decommission-with-migrate guarded only its source) and the correction
that the boundary is the deploy POST, not the dropdown; §7.3 retitled UNBLOCKED;
register row collapsed; open question F answered.

00-capability-map.md: new §D row PROVEN-LIVE, with the un-exercised legs named —
the deploy-POST and decommission refusals are unit-tested, not live-fired.

OPEN-ITEMS.md: R-108 dispositioned; D5 given its OWN row as READY/UNBLOCKED (it
had existed only inside other rows' prose — the R-123 thread-loss pattern);
R-126 registered.

ROADMAP.md: R-108 collapsed to a shipped one-liner; R-126 added.

R-126 filed not fixed: a .fab bundle (plaintext secrets, optional password) can
be exported ONTO a NAS. Split out of R-108 rather than folded in — it is an
explicit customer-chosen export destination, not a browsing surface reaching a
backup tree, so it was never part of D5's precondition.

Live evidence: same-box before/after on demo-felhom through the real authenticated
endpoint, the network-specific refusal on demo-hp, non-effect verified in the
registry, and R-67's share-root bind diffed byte-identical across the deploy.
This commit is contained in:
2026-07-30 14:21:44 +02:00
parent 70f84941d4
commit d42d90fed7
5 changed files with 323 additions and 13 deletions
@@ -95,6 +95,7 @@
| Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts | controller, agent v0.87 | **PROVEN-LIVE** | `DISPOSITION-ia-finding2-systemdisks-2026-07-13` (legacy EFI+LVM host, root not offered, byte-identical); enroll/format live in `storage-lifecycle-acceptance-2026-06-15` (E10 re-enroll, data intact); agent fence self-test refuses `/dev/sda` | (Cited `CAMPAIGN-2` T-STG-ENROLL/SEC-FORMAT were auth-hollow CSRF-403.) **Fresh-USB wizard enroll+format through the customer UI PROVEN-LIVE (controller v0.141.0, 2026-07-17):** a 64 GB scratch USB driven through the real `/api/storage/init` endpoints (login+CSRF) → confirm → detached format (~27 s mkfs) → mount → register → mounted+registered at `/mnt/felhom-drives/scratch1`. **F6 (initialize-to-usable) now covered:** the wizard runs the chain as a detached, disconnect-safe, pollable job (3-step progress) with an agent format-status poll for a slow mkfs **2026-07-26 — a SILENT failure class on the channel every agent-backed capability depends on (this row, data migration, USB enrollment, guest RAM, quiesce/PBS) is now DETECTED (controller v0.173.0, R-77). No row status flips.** `controller.yaml` and `bootstrap.json` could disagree on `local_api.endpoint` indefinitely with no signal: the R-50 island migration rewrote the latter, the fleet kept dialling the former, and for 17.5 h the only alert was a generic "agent unreachable" that read as an infrastructure blip. Drift now raises its own event type (`local_api_endpoint_drift`) naming both values. It is DETECTION ONLY — the authority ruling is R-78 — so the class is now loud, not prevented. Evidence: `audits/DIAG-agent-channel-2026-07-26.md`. |
| Data migration between drives (all / per-app), crash-safe | controller | **PROVEN-LIVE** | `CAMPAIGN-6C` 4P-5 (scope=app round-trip, byte-identical); `storage-lifecycle-acceptance-2026-06-15` (two migrate-all runs via dashboard UI, sha256 byte-identical) | (Cited `CAMPAIGN-2` T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS |
| NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated | controller v0.113117, agent v0.81/84/85 | **PROVEN-LIVE** | `SPIKE-nas-verify-2026-07-11`, `SPIKE-nas-dsm-2026-07-11`, `CAMPAIGN-3-2026-07-11` (boot/reassert fixes) | |
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
| Decommission (migrate-first and anyway-paths), eject | agent, controller | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) | (Cited `CAMPAIGN-2` T-STG-DECOM-* were auth-hollow; `SPIKE-decommission` was report-only, button still vestigial.) |
| Boot ordering: automount + networking survive reboot; appliance self-heal watchdog | agent v0.85 | **PROVEN-LIVE** | `CAMPAIGN-4-2026-07-13` (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + `CAMPAIGN-6A-2026-07-14` 1D (re-arm reboot-survival across 9 guest + 1 host reboots) | (`CAMPAIGN-3` F10/F11/F12 were the CRITICAL/HIGH *failures*; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: `skip-active` on `pct reboot` carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A) |
@@ -372,18 +372,28 @@ that.
requires the guest's secrets. So offsite alone cannot rebuild an app onto a fresh guest.
→ **R-107**
### 7.3 What D5 would change — and why it is blocked
### 7.3 What D5 would change — and why it was blocked (PRECONDITION NOW MET, 2026-07-30)
**[DESIGN, TARGET — BLOCKED]** The intended fix is to make app secrets travel with the **local**
**[DESIGN, TARGET — UNBLOCKED]** The intended fix is to make app secrets travel with the **local**
recovery unit, so Tier-1 and Tier-2 restore work **without the guest and without R**. Offsite
already encrypts everything, so secrets travelling offsite would be covered by the escrowed repo
password. R would then be required for **offsite recovery and host identity only** — losing R would
cost the offsite route, not local recovery.
**This is not adopted.** §2 of the task that produced this document required the premise to be
established, not assumed: the backup tree must be unreachable from every browsing, download and
export surface. **It is not.** The verification and the exposure are in §10.1. Until that is closed
(**R-108**), D5 stays a target and §7.1's chain stands as the model.
**The premise D5 rests on is now established.** §2 of the task that produced this document required
it to be proven, not assumed: the backup tree must be unreachable from every browsing, download and
export surface. When this document was written it was **not** — the FileBrowser network-share bind
reached it. **R-108 closed that on 2026-07-30** (controller v0.187.0) by refusing app namespaces on
network storage, so no `backups/` tree can exist under the share-root bind; every other surface was
already clear (§10.1's table). **D5's precondition is therefore MET and D5 may be adopted.**
**Still true, and not part of D5's precondition:** a `.fab` bundle carries plaintext secrets by
design with an optional password, and `storageDriveList()` does not filter network paths, so a bundle
can be **exported onto** a NAS (§5, → **R-126**). That is an export destination the customer chooses
explicitly, not a browsing surface reaching a backup tree, and it is unchanged by D5 — D5 moves
secrets into the local recovery unit, not into `.fab`. It is tracked separately rather than folded in.
Until D5 is actually implemented, §7.1's chain stands as the model.
---
@@ -467,7 +477,39 @@ uplink — **no customer has ever driven a restore**.
Every divergence between the model above and the system as it is, each with an ID.
### 10.1 D5 is BLOCKED — the backup tree is reachable from a browsing surface
### 10.1 ~~D5 is BLOCKED~~ — CLOSED 2026-07-30 by R-108 (controller v0.187.0)
> **D5's PRECONDITION IS MET.** An app's data namespace can no longer be placed on network storage, so
> no `backups/` tree can exist inside FileBrowser's share-root bind, and the browsing surface therefore
> cannot reach the backup tree on ANY storage class. Every other read surface in the table below was
> already NO. **D5 may be adopted** — nothing in this section blocks it.
>
> **The fix inverted the obvious one, and that is the durable lesson here.** The share-root bind was
> not narrowed, because it **cannot** be: (a) the `:rslave` share-ROOT bind is load-bearing — a
> Phase-0 probe (2026-07-22) proved an in-container access through it wakes the idle automount
> trigger, so narrowing it breaks NAS access itself; (b) there is no `userdata/` layer to scope to,
> since apps on a share store at `<share>/<app>`; and (c) creating one would write Felhom's directory
> convention onto a customer's own NAS, which R-67 forbids outright. The browsing surface being
> immovable is precisely *why* the backup tree must never be placed under it. Tier 2 had already
> reached the same conclusion for its own targets (`F-6C-1`); R-108 closes the PRIMARY namespace,
> which was the last remaining route.
>
> **Operator ruling 2026-07-30: refuse the placement, keep the browse bind.** `RefuseAsAppNamespace`
> (`internal/settings/settings.go`) is the single predicate; all placement surfaces consult it. It
> **fails closed**`/mnt/felhom-drives` holds both storage kinds in-guest, so a path prefix cannot
> classify and `Kind` exists only on a REGISTERED path; an unregistered path under that root is
> therefore un-classifiable and is refused rather than assumed to be a drive.
>
> **Nothing was stranded:** zero apps on network storage across all six hub customers including Peti.
> R-67's browse capability is byte-for-byte unchanged (verified by diffing demo-hp's generated
> compose before and after the deploy). This also **supersedes** the controller README's "NAS backup
> locality — decision A" (v0.118.0), which deliberately kept a NAS-resident app's Tier-1 artifacts on
> the NAS: that case can no longer arise.
>
> Evidence: `audits/R108-network-app-namespace-2026-07-30.md`. The pre-fix analysis below is retained
> verbatim as the record of what was wrong.
#### 10.1 (historical) The exposure as it stood before v0.187.0
**[FACT] The verification and its result.** Every surface that can read a file was checked:
@@ -477,7 +519,7 @@ Every divergence between the model above and the system as it is, each with an I
| SMB browse (folder picker) | **NO** | same deny set applied per child (`sharing_handlers.go:521-537`) |
| SMB `ensureImportShare` (the store-direct bypass) | **NO** | writes one controller-generated constant, `GetImportRoot()` = `<system ns>/userdata/import` (`sharing_handlers.go:564-587`) |
| FileBrowser — **local drives** | **NO** | the bind is `appbackup.UserdataDir(sp.Path)` only, and the comment says why (`internal/web/handlers.go:2450-2460`) |
| **FileBrowser — network shares** | **YES** | the bind is the share **ROOT**: `- %s:/srv/%s:rslave` (`handlers.go:2432`), and the path joins the config source list (`:2433`) |
| **FileBrowser — network shares** | ~~**YES**~~**NO** (R-108, v0.187.0) | the bind is still the share **ROOT** (`- %s:/srv/%s:rslave`, `handlers.go:2437`) and deliberately so — but **no app namespace, hence no `backups/` tree, can exist on a share**, so the root bind reaches only the customer's own files. The reachability is closed at the PLACEMENT, not at the bind |
| `.fab` import path validation | **NO** | confined to `<root>/exports` (`handler_export.go:400-408`, `estimate.go:215-217`) |
| `.fab` browser download | **NO** | name-pattern + parent-must-be-the-staging-dir double guard (`handler_export_download.go:36-45,120-140`) |
| `/api/debug/*` | **NO** | no file-serving branch (`handler_debug.go:46-92`) |
@@ -514,6 +556,15 @@ Every divergence between the model above and the system as it is, each with an I
(`recovery_unit.go:73`). **D5 would make it one.** That is exactly the test §2 set, and D5 therefore
does **not** hold as written. → **R-108**
> **CLOSED (v0.187.0).** Links 2, 3 and 4 of the chain above are now guarded, and a **fifth** surface
> the chain did not list was found and guarded too: `handleStorageDecommission` mode=`migrate` checked
> only `req.Where` (the SOURCE) via `refuseNetworkLifecycle`, so a whole namespace could be
> decommissioned ONTO a NAS. Link 2's framing also understated the problem — the deploy **dropdown** is
> only a UI list; the boundary is the deploy **POST** (`internal/api/router.go`), which accepts any
> caller-supplied `HDD_PATH` and whose only other validation is `os.Stat` existence
> (`internal/stacks/deploy.go`). Filtering the list alone would have left the surface open. Links 5 and
> 6 are unchanged and still true — they simply can no longer be reached.
### 10.2 The gap register
| ID | Gap | Consequence |
@@ -524,7 +575,8 @@ does **not** hold as written. → **R-108**
| **R-105** | Three hub-held DR records are empty on the whole live fleet: `hosts.dr_record_json`, `host_escrow.directive_json`, `dr_recipe.host_half.drives` | the Recipe (§4) is incomplete in exactly the fields host-loss recovery reads. Causes may differ per field |
| **R-106** | `dr_recipe.host_half.pbs.namespace` records `"root"` on every box | the recorded restore coordinate is wrong; real namespaces are per-customer |
| **R-107** | No offsite action unpacks the named-volume tars Tier-3 captures on every run | offsite alone cannot rebuild a named-volume app (§7.2) |
| **R-108** | Network storage can host an app's namespace, and FileBrowser binds a network share at its **root** | **blocks D5** (§10.1); today it also lets a `.fab` with plaintext secrets be exported to a NAS (§5) |
| ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` |
| **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) |
| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10) |
| R-86 (open) | Restore-tests are interval-scheduled, not backup-aligned | a tier's proof cadence is unrelated to when its archives are written |
| R-87 (open) | The restic tier is never restore-tested | matrix row 4's route has no unattended proof |
@@ -567,8 +619,11 @@ mountpoints) are both on `/dev/sda3` → VG `pve`. The tier therefore protects a
and operator error only**, never against disk failure. Accept and name it honestly in the customer-
facing description, or move the target.
**F (added by §10.1, not in the original list).** D5 cannot be adopted until R-108 closes. Is
closing R-108 the intended path, or is D5 withdrawn?
**F (added by §10.1, not in the original list). — ANSWERED 2026-07-30.** D5 could not be adopted
until R-108 closed. Closing R-108 *was* the intended path, and it is done (controller v0.187.0): the
operator ruled to refuse app namespaces on network storage rather than narrow the browsing surface,
because the share-root bind is load-bearing and cannot be scoped. **D5 is no longer blocked.** Whether
to now *implement* D5 remains an open scheduling decision, not a blocked one.
---