Backlog triage Part B: 125 finished rows + 20 id-less rows moved to CLOSED-ITEMS (full text at 9e2786c); open rows normalised to one 6-column shape; narratives archived verbatim; closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS (decoys, seen red); rules rehomed to CONTEXT + 07 §11; loose notes triaged; R-814..R-819 filed; register 444 -> 325
gates / gates (push) Successful in 30s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-03 09:17:22 +02:00
parent 9e2786c907
commit 71b8c8c62b
22 changed files with 1080 additions and 815 deletions
@@ -0,0 +1,114 @@
# DIAGNOSIS — post-F9 storage-registration gap (read-only; no system changes made)
**Date:** 2026-06-14. **Scope:** why the HDD (`felhom-usb`) is attached + written-to by RomM yet badged
"Nem regisztrált" and absent from the deploy dropdown. **Method:** live SSH (read-only) + controller
source. **No state changed** — no registration, deploy, restart, or edits. Diagnosis only; fix to be specced next.
## TL;DR
F9 attached the HDD at the **agent/guest layer** (the `pct set -mpN` bind). The controller's **storage
registry** (`settings.json → storage_paths`) is a **separate layer** that F9 never touched. The controller
registers a drive only through its own enroll flows (`runStorageInit` / `runStorageAttach` /
`handleStorageRegister`), each of which calls `registerStoragePath`. In the F9 session the drive was
attached by calling the **agent's `/disks/guest-attach` directly** (and RomM got `HDD_PATH=/mnt/felhom-usb`
via a manual `deploy` value) — both bypass the controller's registry. Auto-discovery can't backfill it
(it's a one-time seed that only scans deployed-app `HDD_PATH`s, never the agent's attached-drive list).
Result: attached + in use, but unregistered → not selectable for a new app. **F9 is NOT closed.**
## Q1 — Real storage topology (what each path maps to)
| Name (UI) | Host path | Device | Size / free | What it is |
|---|---|---|---|---|
| `Tárhely (felhom-data)` ★ (registered) | `/mnt/sys_drive/felhom-data` | `/dev/mapper/pve-vm--9201--disk--0` | 32 G / **~29 G** | the **internal OS-disk volume** (rootfs disk-0). The "28.7 GB szabad" in the dropdown = this. |
| `felhom-usb` (NOT registered) | `/mnt/felhom-usb` | `/dev/sdb1` | **916 G / 870 G** | the **HDD**, guest-attached (mp1), 2.1 MB used |
Evidence: `pct config 9201` → `mp1: /mnt/felhom-usb/felhom-data,mp=/mnt/felhom-usb`; guest `df` → `/mnt/sys_drive`
on `…disk--0` (32 G, 29 G avail) vs `/mnt/felhom-usb` on `/dev/sdb1` (916 G, 870 G avail). The registry
(`settings.json`) lists only `{"path":"/mnt/sys_drive/felhom-data","label":"Tárhely (felhom-data)","is_default":true,"added_at":"2026-06-13T22:22:50Z"}`.
**Confirmed: the "28.7 GB" option is the internal volume, not the HDD.**
## Q2 — Where RomM's data is + how it got an HDD path
- RomM's `app.yaml` (`/opt/docker/stacks/romm/app.yaml`): `env.HDD_PATH: /mnt/felhom-usb`, `deployed: true`,
`locked_fields: [HDD_PATH]`. Its data dirs live under `/mnt/felhom-usb/felhom-data/appdata/romm/…` (the HDD).
- **How it got set:** the `HDD_PATH` deploy field was filled with `/mnt/felhom-usb` when RomM was (re)deployed
onto the HDD during the F9 P3/restore test. The deploy path validates `os.Stat(HDD_PATH)` exists
(`stacks/deploy.go` path-field check) — it does, because F9 had bound `/mnt/felhom-usb` into the guest — but
**the deploy does NOT register the path** in `storage_paths`. So RomM references the HDD path *directly*,
entirely independent of the registry. That's how "RomM uses the HDD" coexists with "HDD unregistered."
## Q3 — Why `felhom-usb` is "Nem regisztrált"
The registry (`settings.json storage_paths`) contains **only** `/mnt/sys_drive/felhom-data`; `/mnt/felhom-usb`
is absent. Nothing registered it, at any layer:
- **F9's attach is agent-layer only.** `ReassertGuestBinds` / `handleDiskGuestAttach` → `guestbind.go AttachBind`
do `pct set -mpN` (the bind). They never call the controller's `registerStoragePath`. In the F9 session the
drive was attached by calling the agent `/disks/guest-attach` **directly**, bypassing the controller enroll flow.
- **Deploy doesn't register.** Setting `HDD_PATH` on RomM only validates existence (above), no registry write.
- **Auto-discovery can't backfill it** — two reasons in `settings.AutoDiscoverStoragePaths` (`internal/settings/settings.go:584`):
1. **One-time seed:** `if len(s.StoragePaths) > 0 { return // already configured }` (≈:592). The registry already
holds `sys_drive` (seeded 2026-06-13), so discovery is now **permanently inert** on every restart.
2. **Wrong source even if it ran:** it's fed by `discoverHDDPaths(StacksDir)` (`cmd/controller/main.go:1337`,
called once at startup, `main.go:126`), which scans **deployed apps' `HDD_PATH`** — never the agent's
`/disks` attached-drive list. A drive attached out-of-band (or whose app was deployed after startup) is invisible.
So: **registration is decoupled from attach; auto-discovery is a one-shot startup seed of app `HDD_PATH`s; the
agent-layer F9 bind and the manual deploy both bypass registration** → the HDD stays unregistered.
## Q4 — Why the deploy dropdown excludes the HDD
The deploy dropdown is populated from `s.settings.GetSchedulableStoragePaths()` (`internal/web/handlers.go:323`),
i.e. **registered paths only**. `/mnt/felhom-usb` is not in the registry → not schedulable → absent. So a new
app (Calibre) can only be pointed at `sys_drive` (the 31 G internal volume). The "Nem regisztrált" badge is the
storage page (`storageWizardPageHandler`) merging the agent `/disks` (attached, `guest_attached=true`) against
`GetStoragePaths()` (registry) and flagging the difference (`storage_handlers.go` `Registered` field;
`settings.html` `window.__registeredPaths`).
## Q5 — Intended registration-vs-attach behavior (per the code)
The controller's design intends **register + attach as one operator-driven enroll operation**, NOT an automatic
side-effect of an agent bind:
- `runStorageInit` (format a blank/new drive) and `runStorageAttach` (mount an existing-fs drive — the "Csatolás"
button) both do: assign (host mount) → **`registerStoragePath`** → `attachIntoGuest`→`agent.GuestAttach`
(`internal/web/storage_handlers.go:83-158`). The comment is explicit: *"the StoragePath **registration is the
durable intent**, so a transient attach failure is logged (not fatal) — P3 self-heal completes it."*
- For an **already-mounted, unregistered** drive (exactly `felhom-usb`'s state), there is a dedicated action:
`POST /api/storage/register` → `handleStorageRegister` (`storage_handlers.go:442`, the "Regisztrálás" button),
commented: *"the natural primary action for a mounted-but-unregistered data drive (e.g. **felhom-usb**): the
customer's intent is to USE the existing data, not wipe it… registering a host-only mount otherwise leaves it
guest-invisible (the exact gap that produced the 'nem elérhető' banner)."* It registers + attaches into the guest.
So per the code, registration is meant to come **through these controller flows**. There is **no reconciliation
that auto-registers a drive the agent attached out-of-band** — which is precisely the F9 case (the agent's
re-assert / a direct guest-attach binds the drive without ever invoking the controller's register step).
## Q6 — The "felhom-data" naming collision
Deliberate convention, but genuinely confusing. `felhom-data` is BOTH:
1. the **path tail of the internal volume's registered path** (`/mnt/sys_drive/felhom-data`), surfaced as the
label `Tárhely (felhom-data)`, and
2. the **per-drive Felhom namespace subdirectory** created on EVERY user-data drive by the agent's `AttachBind`
(`guestbind.go:29 felhomDataNS = "felhom-data"`, "matches `appbackup.FelhomDataDir`") — hence `/mnt/felhom-usb/felhom-data/…`.
So `felhom-data` denotes both a specific volume and the generic namespace. Risk: the label `Tárhely (felhom-data)`
reads as generic but is specifically the OS-disk volume; meanwhile the HDD also has a `felhom-data` dir. A clearer
label (e.g. "Belső SSD" for the internal volume vs the drive's own name for the HDD) would remove the ambiguity.
## Recommended fix direction (to spec next — NOT implemented here)
1. ~~**Auto-register on attach (primary).**~~ **REJECTED (2026-06-14)** — contradicts the new→enrolled
manual-enrollment model; manual enroll is by design. A drive becomes usable through the explicit
enroll/"Regisztrálás" flow (recommendation #2), not by auto-registering whatever the agent attached
out-of-band. (The "make discovery additive, not return-if-any-path-exists" sub-point WAS adopted —
shipped in controller v0.64.0, A1 — but only as additive pickup of paths referenced by *deployed apps*,
NOT as auto-register-on-attach.)
~~Make a drive the agent has attached into the guest become a registered
storage path automatically, so F9's attach is usable end-to-end. Options: (a) the controller's startup/periodic
storage reconcile consults the **agent `/disks`** and registers any `guest_attached=true` user-data drive missing
from the registry (extend beyond the app-`HDD_PATH` scan); and/or (b) the enroll/re-assert path that binds a
drive also performs `registerStoragePath`. Also **remove the one-time-seed gate** (or make discovery additive,
not "return if any path exists") so it can pick up drives added after the first run.~~
2. **Keep `handleStorageRegister` as the explicit fallback** ("Regisztrálás" button) for the mounted-but-unregistered
case, and make the "Nem regisztrált" badge clearly actionable. For the **immediate** demo state, this button
(`POST /api/storage/register {where:/mnt/felhom-usb}`) would register felhom-usb + make it selectable — no fix needed to recover now.
3. **Clarify the `felhom-data` label** (Q6) to disambiguate the internal volume from the per-drive namespace.
Recommended: **both** auto-register-on-attach AND the clearer manual path — auto so the normal flow "just works",
manual for out-of-band/edge cases.
## Verdict on F9 — NOT closed
F9 made the HDD present + writable **inside the guest** (agent layer), but did not make it **usable for new apps
through the UI** (controller registry layer). The deploy dropdown is gated on the registry; the registry is
unaware of the agent-attached drive. **F9 should remain open (or spawn an F9b)** until: attach ⇒ auto-registered
⇒ a NEW app (e.g. Calibre) can be deployed to the HDD through the normal deploy UI with its data landing on
`/dev/sdb1` persistently. Today that round-trip fails at the dropdown.
+51
View File
@@ -0,0 +1,51 @@
> **STATUS: FIXED** in controller v0.62.0 @ `f8afe5c` (2026-06-14) — implemented trunk-based on `main` per the plan below, with a regression test. This note is retained for provenance.
# fix/m18-dump-validation-cache — NOTES (pending review, NOT deployed)
**Verdict:** LIVE @ `controller/internal/appbackup/dbdump.go:469` (commit `6953899`).
**Class:** performance (not correctness/security). **Escape-hatch branch** — the fix crosses the
`settings`↔`appbackup` package boundary and changes an exported signature; not forced unattended.
## The bug (verified mechanism)
`ListDumpFiles(dumpDir)` (dbdump.go:425) unconditionally calls `f.Validation = ValidateDump(fullPath, f.DBType)`
(line 469) for **every** `.sql` file on **every** call. `ValidateDump` (line 320) opens the file and
scans it line-by-line (bufio loop). The caller chain is `backup.RefreshCache` (every ~5 min, scheduler)
→ `listAllDumpFiles` → `ListDumpFiles` for every drive/stack. So every dump is fully re-read every 5
minutes, even when unchanged.
A `settings.DBValidationCache` type exists (`settings.go:159`) and is written in `RunDBDumps`
(`backup.go:447`), but it has only `ValidatedAt/TableCount/HasHeader/Error` — **no size or modtime** —
and `ListDumpFiles` never consults it. `DumpFileInfo` already carries `Size` + `ModTime` (dbdump.go:451-453),
so the inputs for a cheap skip-check are present; they're just not used.
## Impact
On large DB dumps (hundreds of MB) this is wasted disk I/O + CPU every 5 minutes. Negligible on the demo
(small dumps); real on a customer with big databases. No correctness impact — validation results are the
same, just recomputed.
## Fix plan (implementable, low-risk once reviewed)
1. Add `Size int64` and `ModTime string` (RFC3339) to `settings.DBValidationCache`.
2. Change `ListDumpFiles(dumpDir string)` → `ListDumpFiles(dumpDir string, cached func(name string, size int64, mod time.Time) (DBValidationResult, bool))`.
Pass a **plain lookup func**, NOT the `settings` type — `appbackup` must not import `settings` (avoid
an import cycle; keep appbackup leaf-like). The func returns the cached `DBValidationResult` + ok.
3. Before line 469: if `cached(e.Name(), info.Size(), info.ModTime())` returns ok, reuse it and skip
`ValidateDump`; else validate and (caller-side) store the result keyed by name+size+modtime.
4. Update the bridge re-export (`backup/appbackup_bridge.go:66`) and the `backup.go` call sites to build
the lookup from `settings`'s cache (now carrying size+modtime), and to write back fresh validations.
5. Keep a nil-`cached` fast path (validate-always) so other callers/tests don't have to thread it.
## Test plan (regression)
- Unit test in `appbackup`: call `ListDumpFiles` twice with a `cached` func that records calls; assert
`ValidateDump` is NOT re-run for a file whose size+modtime match the cache, and IS run when modtime
changes. (A spy counter on a wrapped validator, or assert via a temp `.sql` whose mtime is bumped.)
- This test fails on the pre-fix code (which always validates).
## Why not tonight
Crosses a package boundary + changes an exported signature with multiple call sites (bridge + backup.go),
and the cache type needs new fields — too entangled to land safely unattended. Hand to the supervised
session: the plan above is complete and mechanical.
+58
View File
@@ -0,0 +1,58 @@
> **STATUS: FIXED** in controller v0.62.0 @ `6bab68b` (2026-06-14) — implemented trunk-based on `main` per the plan below, with a regression test. This note is retained for provenance.
# fix/m19-stackname-crossref — NOTES (pending review, NOT deployed)
**Verdict:** LIVE @ `controller/internal/appbackup/dbdump.go:536-551`, used at `:115` (commit `6953899`).
**Class:** correctness edge — **low real-world incidence** with the current catalog. **Escape-hatch branch**
— the clean fix injects the deployed-stack list into `appbackup` (cross-package), not forced unattended.
## The bug (verified mechanism)
`deriveStackName(containerName)` pure-suffix-strips: it splits on `-` and, if the last segment is in
`{postgres,db,mariadb,mysql,database,redis,cache}`, returns the join of the remaining parts. It never
cross-references actual deployed stack names. `DiscoverDatabases` assigns `StackName: deriveStackName(name)`
directly (line 115).
So a stack whose real name **ends** in one of those tokens is misattributed:
- a stack literally named `my-cache` → its DB container `my-cache` (or `my-cache-postgres`) derives to
`my` / `my-cache`, attributing the dump to the wrong (or a non-existent) stack.
- worse, `a-db` and `a` could collide.
## Impact
A DB dump is filed under the wrong stack name → that stack's backup/restore accounting is wrong, and the
restore-by-stack path could miss or cross-wire the dump. **Incidence is effectively zero in the current
felhom catalog** (stack slugs are `romm`, `nextcloud`, `paperless-ngx`, `immich`, `adventurelog`,
`actualbudget`, `mealie`, `vikunja`, … — none ends in a DB-role token; DB containers are `<stack>-postgres`
etc., which strip correctly). It becomes real only if a future app slug ends in a role token.
## Fix plan (implementable)
1. Thread the set of **known deployed stack names** into discovery:
`DiscoverDatabases(ctx, logger, debug, knownStacks []string)` and
`deriveStackName(containerName string, known map[string]bool)`.
Source the list from the caller in `backup.go` (it holds the `StackDataProvider` — expose/known stack
names via a lookup func to avoid importing `stacks` into `appbackup` and creating a cycle).
2. New `deriveStackName` logic:
- candidate := current suffix-strip result.
- if `known[candidate]` → use it (the suffix was a real DB-role suffix of a real stack).
- else if `known[containerName]` → the container name IS the stack (don't strip).
- else → longest `known` stack name that is a prefix of `containerName` (handles `<stack>_postgres`,
`<stack>-1`, compose-suffixed names); tie-break to the longest match.
- else → fall back to the current suffix-strip (preserve today's behaviour when the stack list is
unavailable/empty, so nothing regresses).
3. Keep a nil/empty-`known` fast path = today's behaviour (back-compat for other callers/tests).
## Test plan (regression)
- Table test for `deriveStackName` with `known = {romm, my-cache}`:
- `romm-postgres` → `romm` (suffix is a role; `romm` is known).
- `my-cache` → `my-cache` (known as-is; must NOT strip to `my`). ← fails on pre-fix code.
- `my-cache-postgres` → `my-cache` (strip role, result known).
- `unknown-db` with no matching known → falls back to `unknown` (today's behaviour).
## Why not tonight
Requires plumbing the deployed-stack list across the `backup`→`appbackup` boundary (cycle-avoidance via a
lookup func) and touching `DiscoverDatabases`'s signature + caller. Low-incidence, so not worth a risky
unattended change. The plan + tests above are complete for the supervised session.
@@ -0,0 +1,40 @@
# FOLLOW-UP — bump the golden's default controller tag + validate the full provision path
**Status:** **RESOLVED 2026-07-03** — `build-golden.sh` v2.0.0 (felhom-agent @ `ceca355`) makes the
controller tag a MANDATORY argument (the hand-bumped default had rotted AGAIN, 0.43.0 → 0.85.1 →
stale vs 0.98.3 — a required arg cannot rot); golden **0.98.3** baked + validated clean-room on the
drill VM (first boot lands the current controller, self-update reports up-to-date, app deploys) +
published + operator-vouched. The full first-boot path was validated WITHOUT a supervised
destroy→re-provision (drill VM virgin snapshot instead — nothing touched 9201/felhom-pve).
Evidence: `../audits/DRILL-golden-098-2026-07-03.md`. Item 3 (provision re-asserting user-data
binds) was shipped separately (agent-startup re-assert + F9 part A). Original note kept below for
history.
**Class:** provisioning correctness / customer-onboarding. **Risk (as queued):** SUPERVISED (golden + a real destroy→provision).
## The problem
`felhom-agent/configs/build-golden.sh:43` bakes the controller image into the golden as:
```
CONTROLLER_IMAGE="${6:-gitea.dooplex.hu/admin/felhom-controller:0.43.0}"
```
`:0.43.0` is ~20 versions stale (current is `0.63.0`). The bootstrap writes it to
`/etc/felhom-controller-image` and `docker run`s that tag on first boot. So the **next time a golden is
actually baked for a real provision, it would stand a customer guest up on an ancient controller** —
missing every fix since 0.43.0 (incl. F1 memory guard, F17 DB restore, the M18/M19 fixes, etc.). It only
hasn't bitten because the running demo guest 9201 had its `/etc/felhom-controller-image` updated in place
post-provision; a fresh provision would not.
This is why the F9 session chose **attach-to-existing** over destroy+re-provision (a re-provision would
have regressed 9201's controller 0.62.0 → 0.43.0).
## The fix (separate task)
1. Bump the `build-golden.sh` default `CONTROLLER_IMAGE` to the current released controller tag (and
establish a convention so it tracks releases — e.g. read a `LATEST_CONTROLLER` pin, or pass it
explicitly from the build pipeline).
2. **Validate the full path end-to-end**, which *does* warrant a supervised destroy + re-provision (it is
the real customer-onboarding flow, not coverable by attach-to-existing): bake golden → provision a
scratch guest → first boot stands up the **current** controller → controller ONLINE on the hub →
enrolled user-data drive is bound (F9 `ReassertGuestBinds` / provision bind) → an app deploys.
3. While there: confirm the provision path itself asserts known user-data binds (F9 part A for the
re-provision case, complementing the agent-startup re-assert already shipped in v0.31.0).
@@ -0,0 +1,361 @@
# OPEN-ITEMS — the narrative sections, as they stood on 2026-10-03
> **History, not a register.** These sections sat between the register tables of
> `documentation/backlog/OPEN-ITEMS.md` until the 2026-10-03 triage. They are moved here word for word
> (`git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` holds them in place). Every row they discuss is in
> `OPEN-ITEMS.md` (open) or `CLOSED-ITEMS.md` (finished). A status word below is the status ON THE DATE OF
> ITS SECTION and may be stale — the registers are the authority. The ranking paragraphs ("Why the TOP
> READY rows rank this way", "The 2026-08-02 intake, ranked") are superseded by
> `documentation/audits/backlog-triage-2026-10-03/RECOMMENDATION.md`.
## Operator rulings — 2026-08-04
Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule).
1. **Run the recovery drill, after R-198.** R-198 shipped in hub v0.93.0; **the drill is the next
session.** Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10 — demo-hp, ~3–4 h, a recovery
code created and KEPT, a sentinel file, wipe, reinstall, recover, and **pass = a byte-identical
sha256, not "the repository opened"**. → R-201
2. **Delete the orphaned ciphertext** (~1.2 GB across the two demo boxes, in set-aside restic stores
nothing prunes). **STILL OWED — not done in v0.93.0.** It is a destructive act on a protected
endpoint and belongs to a session that is scoped for it, not to a release that ships a schema
change. → R-193
3. **Accept the risk on R-193 candidate (c)** — no repository password retained on the Proxmox host.
**This is what makes R-198 load-bearing rather than tidy:** with no host-retained copy, the
customer-present recovery path is the ONLY way back from a rebuild, and that path runs entirely
through the retained identity blob. → R-193, R-199, R-200, R-201
~~**Still open and untouched by v0.93.0:** R-199, R-200, R-201.~~ **UPDATED 2026-08-04 (evening), after
hub v0.94.0 + agent v0.125.0 + controller v0.195.0:**
- **R-199 — CLOSED, proven on hardware.** Chain links 6–8 are assembled and walked. The offsite
repository password came back out of the sealed bundle **byte-identical** to the one on disk
(`c60c8bc737a6…` from three independent sources: the box's file, the recovered bundle, and the hash
the hub already stored).
- **R-200 — the plumbing half shipped**, the customer-facing form did not, deliberately.
- **R-201 — PASSED 2026-08-04 (night run).** A customer's file survived a machine rebuild and came
back byte-identical, through the customer's own restore flow. **It passed only because a person was
there:** four manual interventions stood between the recovered key and the restored file, none of
them in any design document → R-204.
- **R-202 — untouched.** The orphan card still promises recoverability unconditionally.
- **The orphaned ciphertext deletion (~1.2 GB) is STILL OWED** — ruling 2, above.
**UPDATED 2026-08-05, after controller v0.198.0 + hub v0.95.0 (R-204):**
- **R-204 items 1–3 — CLOSED.** The reset code works without a restart; a re-issue no longer marks a
healthy escrow stale (→ **R-196 CLOSED**); a unit restore states what it did NOT restore.
- **R-204 item 4 — OPEN and unstarted:** a rebuilt box cannot obtain an off-site credential unaided.
It needs an operator ruling on the one-shot credential design → **R-193**.
- **Still open and untouched by this session, stated so nothing is presumed closed by association:**
**R-202** (the orphan card's unconditional promise), **the ~1.2 GB orphaned-ciphertext deletion**
(ruling 2 — still owed, still needs its own scoped session), and **R-198's retention, which remains
UNIT-PROVEN ONLY** — nothing has superseded a key in production, and proving it needs a SECOND
deliberate wipe. That retention drill is the next item, and it is not this session's.
v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.197.0 got the file into
the snapshot and the drill got it out again. **v0.198.0/v0.95.0 remove three of the four crutches the
drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.**
## R-201 — THE RE-WALK, 2026-08-06 (attended)
**The question was asked a second time, on the fixed build, on a brand-new appliance built from the
published ISO. The answer is still no — but it is a nearer no.**
| half | verdict |
|---|---|
| **the data** | **PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also byte-identical** (verified as hex, not as rendered text). Restored in **23 s** out of the pre-destruction snapshot `a7bc23bd`, through the customer's own two-step full-restore flow |
| **the journey** | **FAIL — two dead ends**, against Phase 1's four. One needed a command line **inside the guest**, one a **Proxmox-host** action |
**RTO: the unaided figure is STILL UNDEFINED**, because the unaided journey still does not complete.
Attended: login 11:42:22 → key placed **+45 s** → tier up **+24 m 12 s** (after intervention 1) →
all three sentinels restored and verified **+30 m 13 s** (after intervention 2). **The 30 m figure must
not be quoted as the customer number.** The only segment that reflects the product working alone is
**23 seconds** to pull 12.8 MB back once everything was in place.
**The two dead ends:** **R-218's consume half** (row corrected above) and **R-220** (drives
unenrollable after a rebuild — reproduced and red-proved again; without it no app can be redeployed,
and without a redeployed app the restore page is empty, which is R-213's territory and follows from
R-220 rather than being separate).
**What PASSED and is worth keeping:** the recovery screen **appeared without being sought**
(`/` → `/launcher` → `/recovery`); it answered all three of its questions and its seal date matched the
hub's `created_at` exactly; the emailed reset code worked **first try**; the unlock took **1.528 s** —
a real unseal — and placed the key; and **R-225's fix was seen working in the wild** (the store read
„a pillanatképek száma még ismeretlen" rather than a false zero).
**R-216 part 4 reproduced live:** the reinstall **downgraded the agent 0.126.0 → 0.125.0**, back to the
vouched version — an operator's hand-fix undone by the very event that makes recovery necessary.
⚠ **THE DELIVERY GAP, and it is owed.** A fresh install landed on controller **0.201.0** / agent
**0.125.0** — the vouched versions, **neither carrying the fixes**. They were installed **by hand**.
Fleet delivery needs a golden carrying 0.202.0 **and** a vouched agent 0.126.0. **Nothing was vouched;
that is the operator's act.** **This re-walk proves the JOURNEY on the fixed build; it does NOT prove a
customer would receive that build.**
Evidence: `tests/rewalk-r201-2026-08-06/journal.md`.
## CAMPAIGN 11 — the recovery journey, 2026-08-05
**The whole journey was walked end to end for the first time, on a throwaway appliance built from the
published ISO. The data came back byte-identical; the journey did not exist.** **Ten findings from
Phases 1 and 3, R-214 … R-223** (seven fixed in controller **v0.201.0** + hub **v0.97.0/0.97.1**;
**three deliberately still open**, each blocking a real flow), **plus five from Phase 2's injected
faults, R-224 … R-228.** Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` (Phases 0/1/3)
and `journal-phase24.md` (Phases 2/4). Campaign document:
`audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`.
> **Phase 2's verdict in one line.** The cryptography, the retention and the transport all work and
> are now proven live. **What fails is being told the truth:** a mistyped code, a hub outage, a
> stopped agent and a correct code for a retained earlier package all produce **one** message, and
> three of the four are wrong.
### Phase 2 — the injected faults, 2026-08-05/06 (unattended)
> **ALL FIVE CLOSED 2026-08-06** in controller **v0.202.0** + agent **v0.126.0**. The rule they now
> enforce, stated so it outlives them: **on the unlock path the customer is blamed only after a real
> attempt REFUSED their code; every other outcome, including an unclassifiable one, says something
> else.**
>
> **STILL OPEN AND DELIBERATELY UNTOUCHED BY THAT WORK — said explicitly rather than left to
> inference:** **R-214** (the console never stops showing a stale pairing code), **R-220** (a rebuilt
> box's drives cannot be re-enrolled — *currently worked around BY HAND on the campaign venue*, which
> is the only reason an app could be deployed there at all), **R-221** (a rebuilt box cannot run the
> escrow ceremony), **R-213** (putting files back), **R-202** (the orphan card's unconditional
> promise — now the **last** place on that surface still promising recoverability, two doors from
> where R-228 removed the same promise).
Eleven faults, each judged on **the message, not the outcome**, each with a positive control proving
the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/journal-phase24.md`.
## Instruction files — deferred half, 2026-08-06
**Recorded against existing rows by Phase 2:**
- **R-216 — §4.1 is now MEASURED, not deduced.** The previous session could only offer two absences.
The box's own `/settings` renders „Minimális verzió (üzemeltető) **0.200.0**" (`GetFloor()`, whose
only writer is the report-ACK handler; **both hold branches serve `Floor=""`**, pinned by
`managed_floor_test.go:94`), and a **cold-started** controller logs
`settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)` against the same line reading
`floor still unknown after 1m30s` while the hold was in force. The hub's HELD lines ran every
15 min to 22:12:06 and stopped, with a liveness control proving the hub kept logging. **The floor
is served.**
- **⚠ A correction to how that positive was to be taken.** `SetFloor`'s line is `u.dbg(...)`, gated on
`cfg.Logging.Level == "debug"` and written to the **logger** — it can **never** reach the logx debug
ring, so it cannot appear in `/api/debug/logs` at any level. A controller restart alone would not
have produced it. Confirmed with a level census on the ring first (1196 DEBUG / 2802 INFO / 2 WARN),
so the absence was known to be structural rather than evidential.
- **R-218 — the live half is STILL NOT MEASURED, deliberately.** The fix is present and correct
(`needsOffsiteCredential` now retires on the **target**, not the key), but the venue **has** a
target, so the box correctly does not declare; declaring here would be the bug. The state that
exercises it is shape (a), which the venue no longer holds. **Recorded as not measured rather than
inferred from the unit test.**
- **R-217 — its fix HELD under exactly its fault** (F5): with the store blocked after a successful
unlock, the page rendered M3 and **no listing block at all**; the three false-claim strings are
absent, verified in UTF-8 with accented positive controls present.
- **R-215 — its fix is present** and the `GET /recovery` gate consults the same predicate as the POST
sibling.
- **R-199's back-pointer in `architecture/00-capability-map.md` was already added** — the brief lists
it as owed; it is present on the escrow-recovery row, explicitly labelled as the omitted
back-pointer. **No action taken; the brief's assumption was stale.**
**Untouched by this session, stated so nothing is presumed closed by association:** **R-213** (putting
files back — the half the recovery screen deliberately does not do) and **R-202** (the orphan card's
unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of).
## CAMPAIGN 12 — the class sweep, 2026-08-08 (unattended)
Eight rows, **grouped by class so the classes are visible as classes**. Full method, controls, blind
spots and the Part-4 gating ranking: `audits/CAMPAIGN-12-class-sweep-2026-08-08.md`. Gating candidates
are in `ROADMAP.md`, not here. **C1 produced no new instance** and has no row, deliberately.
**⚠ A correction the campaign owed to its own brief:** the task described C5's `escrow_stale` instance
as "closed individually". **It is not — it is R-247, `READY`.** The live repo is the source.
## CAMPAIGN 12 follow-through — G-1 built, R-260 and R-247 closed, 2026-08-08
**The gate was built BEFORE the fixes and was seen failing on 40 fields**, captured verbatim in
`documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md`. That order was the method, not
bureaucracy: the night before, `deadcode` was rejected for class C6 precisely because it was made to
prove itself first and found neither of the two defects it was meant for.
**⚠ A COUNT THIS SESSION'S OWN PROMPT GOT WRONG, corrected against the repo rather than quoted.** The
prompt said *"465 emitted tags, eight unreachable"*. R-260's wording was "at least eight
DECISION-BEARING facts", never eight tags in total, and its own census already listed more. Measured
on the three declared wires: **210 tags checked, 51 skipped, 40 convicted.** Two prompt claims were
wrong this week and both were caught the same way.
**Two things the gate's CONTROL caught before it was trusted**, each a defect in the instrument:
1. **A substring false negative.** `grep -F healed_at` also matches `privsep_healed_at`, so a
genuinely dropped field read as received. R-260 named `healed_at`, so its absence from the output
was the tell. Now a whole-token regex.
2. **`dr_recipe` is not wholly opaque.** The hub stores each half as `json.RawMessage` and re-emits
nested shapes verbatim — but the TOP-LEVEL section keys are decoded by `hostHalfShape` /
`appHalfShape`, and **those are allow-lists**: a section an emitter adds is silently dropped until
named in both. That already cost `offsite_restic` (R-122). The gate is therefore opaque BELOW
depth 1, so the sections are checked; treating the whole subtree as opaque would have put R-122's
shape back outside its reach.
**THE FORTY, BY DISPOSITION.** Full per-field reasons live in the gate's own `ALLOWLIST`, where each
entry is a claim someone can re-check.
| # | field(s) | direction | decision | what changed |
|---|---|---|---|---|
| 1 | `oob.operator_key_configured` | agent → hub | **RECEIVE AND ACT ON IT** | decoded as a pointer; `oobDegraded` now fails when the key is absent, and the alert NAMES it. Hub v0.99.0 |
| 2 | `oob.wg_handshake_age_s`, `oob.healed_at` | agent → hub | **RECEIVE, for the message only** | decoded into `HostOOBRow` and put in the event payload; deliberately NOT in the predicate — widening the check beyond the fact that is now arriving is how a check stops being read |
| 3 | `escrow_stale` | hub → controller | **RECEIVE AND ACT ON IT** | `report.EscrowStatus.Stale`; the box tells a withheld hash from a hash-less one. Controller v0.209.0. **This is R-247** |
| 4 | `host.{cpu_temp_c,loadavg,memory_total_bytes,memory_used_bytes,uptime_seconds}`, `system.{load_avg_1,5,15, memory_total_mb, memory_used_mb, temperature_celsius, uptime_seconds}` | agent/controller → hub | **NO CONSUMER WANTED — redundant** | allowlisted: the hub decodes `cpu_percent` / `memory_percent` / `disk_percent` from the same stanzas and every threshold is expressed on those |
| 5 | `guests.spec.{disk_bytes,memory_bytes}` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: guest sizing is hub-owned INTENT (the manifest), not mirrored reality |
| 6 | `storage_targets.smart.model_name` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: a display label with no threshold on it; `smart.health` and every counter the hub bands on ARE decoded |
| 7 | `wireguard.last_handshake_age_s` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: wgsync reconciles from its own state, and the OOB path's own handshake age is now decoded |
| 8 | `guest_net` + its 7 children, `selfupdate_pending(_version)`, `mgmt_plane.healed_recently`, `pbs_dr.applied_at`, `restore_tests.mount_{parity,inventory}`, `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check` | agent/controller → hub | **NO CONSUMER TODAY, AND ONE IS ARGUABLY OWED** | allowlisted **against R-264, which stays OPEN**. Allowlisting is not deciding, and the entries say so |
**A C7 instance found while doing it, and corrected.** `HostReport.SelfUpdatePending`'s own comment
claimed *"The hub reads an absent field as pending=false, the correct default."* The hub has no field
for it and reads nothing either way. Comment corrected in the agent (no version bump — comment only).
**The hub's OOB test fixture was part of why this survived.** `oobReport()` omitted
`operator_key_configured` entirely, so every pre-existing scenario ran against a report shape **no
released agent produces**. Fixed, and the new tests drive the raw JSON decode boundary — a test that
builds the receiving struct by hand cannot see a field that never decodes, which is the whole class.
**Explicitly still open, untouched by this session:** R-246 (the wrong stale flag on `demo-hp` —
clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and
**C7's test-comment half**, which Campaign 12 recorded as *owed, not done* (60 of 2652 production
invariant comments sampled; none of the 1440 test comments).
## The seed that never ran twice, and three pictures that were not true — 2026-08-08
Four defects of one family: something the box already knows, either thrown away or drawn as its
opposite. Agent **v0.128.0**, controller **v0.210.0**, `gates.yml` (no hub change, no hub bump).
**R-221's writer, ESTABLISHED at `file:line` rather than assumed** — the prompt asked for this and it
was owed. `step_agent_config` (`felhom.eu/scripts/felhom-host-install.sh:2396`) renders `agent.json`
from `base = {}` unless an explicit `--preserve-from` is passed (flag `:1246`, defaulting empty at
`:256`), and writes it with `O_TRUNC` (`:2579`). **The render never writes an `escrow` section at
all** — grep over the whole heredoc returns zero hits. The pbsdr marker lives host-side
(`<agent-state>/pbsdr/marker.json`) and survives. So a rebuild keeps the marker and takes the key:
same descriptor, same hash, early return, seed never re-runs. **The attribution in R-221 was
correct.** A rebuild is nonetheless only the case that was measured — the same hole opens for a
hand-edited or restored config, which is the honest reason the fix is at the seam and not in the
installer.
**The idempotent early return was KEPT**, and that is load-bearing: it stops a converged box
re-running Proxmox operations every 60 s.
`TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner calls on
that tick, so a "fix" that simply deleted the return fails the test. Verified by mutation.
**The §7.3 truth table as implemented** (R-258):
| this app's own most recent dump result | restore point | verdict |
| any of its databases failed | yes | `error` — cross |
| all clean | yes | `ok` — tick |
| none recorded (no database / no run yet) | yes | **no icon**, time only, title „Erről a mentésről nincs eredményünk." |
| any | no | no tier-1 row at all, unchanged |
**An existing test was asserting the defect and was corrected, not deleted.**
`TestBuildAppBackupRows_Tier1FromRestorePoints` expected `"ok"` for a `FullBackupStatus` with **no
`LastDBDump` at all** — a green tick derived from nothing but a file's existence, i.e. Scenario G.
Its real subject, the `Tier1LastRun` time, is unchanged.
**The convention is now ruled** (§7.2, `CONTEXT.md` S-39): a `…Known bool` companion beside the
figures. `ROADMAP.md` G-3 was blocked on that decision and is unblocked.
**Six red-proofs, every one demonstrated failing and restored, each with the mutation asserted
applied.** The one that matters: Scenario A **fails against today's tree** with the intended message
— so the test tests the defect.
**Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262,
R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243,
R-202, R-213, R-244, R-214/R-235, and **C7's test-comment half**, which Campaign 12 recorded as
*owed, not done*. **G-8's other half** (a hub-side check that notices a *vouch* has been forgotten)
was deliberately not built: it is hub work whose payoff is a daily email, and this session already
ends with a bake-and-vouch cycle in front of the operator.
## Why the TOP READY rows rank this way
This covers the next few only — it is deliberately **not** a full ordering of the table above, so that
there is one ranking to maintain rather than two.
1. **R-95** — the largest *data* exposure: the tier holding the customer's documents and photos is the
one whose credential can delete. **THIS PARAGRAPH WAS STALE UNTIL 2026-09-01 AND ITS OLD TEXT IS
NAMED SO THE CORRECTION IS NOT MISTAKEN FOR A RE-RANK.** It said the snapshot mitigation was
*"armed (daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong:**
seven daily snapshots do exist (R-429), and the word "armed" was withdrawn by the spike that same
day. **What is true now:** the snapshots exist and the box cannot write into them, but **no account
we hold can read anything out of one** (R-433, measured — 777,600 names, zero hits, controlled), so
they do not yet bound this exposure. The root cause is untouched either way — the box can still
`forget --prune` its own repo, from two call sites. **The ORDER of this list is unchanged and is
Viktor's**; only the facts under item 1 were corrected.
2. **R-94** — **de-ranked 2026-07-29.** The prior rationale ("until it moves every hub-driven
install gets the pre-R-82 default") was false: the constant selects no script and every install
already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment —
a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing;
not high-consequence, and it blocks nothing.
3. ~~**R-86**~~ — **CLOSED 2026-08-03**, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.
4. **R-87** — **SPIKED 2026-08-31 and now a DECISION, not work.** It also spent 2026-08-22..31 in
`CLOSED-ITEMS.md` by mistake while this paragraph ranked it fourth and pointed at nothing
(R-405). The spike measured it rather than designing it: a scratch restore of all 8 apps costs
25 s against the 40.3 s weekly check, but it would have caught ONE of the five drill-found
restore defects. **Recommendation: build the NARROW version — prove the snapshot still
CONTAINS a recoverable unit — or close the row.** Viktor's call; see
`audits/SPIKE-restic-restore-test-2026-08-31.md`. *The 2026-08-03 rationale, kept:*
R-86 built most of what it was waiting for (per-archive due-ness, a
proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is
restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is
no longer waiting on a scheduling model that did not exist.
5. ~~**R-185**~~ — **CLOSED 2026-08-03**, agent v0.123.0 + installer 1.24.0, proven live on both demo
boxes. The silence was fixed as well as the grant: the box now asks whether it may READ each tier
it depends on, because an empty listing cannot distinguish forbidden from newborn.
6. ~~**R-189**~~ — **CLOSED 2026-08-03** with **R-188** and **R-186**, agent v0.122.0. The three
reporting/release signals that misreported their own work are fixed; **R-185 is the one that
remains open from that group** and is untouched by this — it is a missing storage ACL on
demo-felhom, not a reporting defect.
7. **R-110** — last **because it is not a READY row**: the ruling is the operator's, not CC's, and
there is nothing for CC to build until it lands. Ranked here rather than omitted because it is the
only item on this page about the *publish channel* of the most privileged artifact Felhom ships,
and today's exposure is zero — which makes now the cheapest moment it will ever be to decide.
### The 2026-08-02 intake (R-156 … R-164), ranked
Filed in one pass from Campaign 10, its two spikes, and the 53-template catalog persistence sweep.
**R-156 and R-157 had lived only in audit documents** — the identical *"minted in a spike doc and
never carried across"* failure the register already records for R-153/R-154/R-155, caught by the
sweep's own §8.0 while it was happening. R-158 was minted by a second session on the same day for an
unrelated finding, which is why the sweep's proposals were renumbered to R-159…R-162 at filing time.
1. **R-157** — highest: a `deployed: true` app can stay **down indefinitely** after a power cut or
hard reset, and in mechanism **B** nothing reports it on any channel (`0 currently down`). It is the
only row here where the customer loses service and has no signal at all.
2. **R-156** — the class is now detectable and two of three apps are fixed; what remains is **papra's
referral**, one app, well understood. *(Promoted 2026-08-02: R-161 was ranked here because nothing
ran the gate; it now has a mandated entry point, so R-156's residue is the larger remaining item.)*
4. **R-163** — a real ceiling that silently caps local backup once an app outgrows `mp1`, and it gates
Tier-2 and Tier-3 as well. Ranked below the above only because **overflow itself is safe today** —
it refuses per app and preserves the last good unit byte-identical. **RE-FRAMED 2026-08-02:** no
longer waiting on a ratio — decision **D-a** merges `mp1` away, so the row is now the record of the
constraint and the work moves to **R-165** (with **R-167** shipping in the same step). R-165 inherits
this rank; it is the highest-ranked item that must land **before any external install**.
5. **R-158** — the gap that makes R-163 dangerous: cross the size line and **one page** tells you. On
its own it is a notification gap, not a silent failure, which is why it sits here and not higher.
6. **R-164** — blocked on a predicate, no customer impact today; it only becomes urgent if the unit
size in R-163 is judged unacceptable, since the tar-drop is the cheapest way to halve it.
7. **R-161** — **de-ranked 2026-08-02, ruled and shipped at reduced scope.** The gate now has one
mandated entry point (`catalog_gates.py`), which is the shape that actually gets run here. What is
left is the automatic half, and that is sufficient while **one** person touches templates — so it
ranks low by design, not by neglect. Revisit when a second does.
8. **R-162** — `WATCHING` only. A limitation that fails closed; revisit if a non-overlay driver ships.
**R-159 and R-160 are SHIPPED** and are not ranked; they are filed to record the class, and R-159's
class (an image `VOLUME` at an unmounted path) is still live — `immich-server` has one today.
<!-- RESTORED 2026-08-22: these rows are NOT fully closed and were moved to CLOSED-ITEMS.md in
error by this session's compressor, whose status regex matched the word CLOSED inside
'PARTLY CLOSED' and the word FIXED inside 'OPEN - NOT FIXED'. They are restored VERBATIM
from commit fddfe00ce268, not from the compressed form: an open row keeps its detail. -->
<!-- ── MIGRATED FROM ROADMAP.md, 2026-08-22 (R-369 ruling: ONE register) ──────────────────
16 rows moved: every ROADMAP row that asserted something checkable about the shipped
product, plus one owed operator decision. Ideas and proposals stayed in ROADMAP, which
is their home. Each row below keeps its ORIGINAL identifier and filing date — the age is
the point. The roadmap keeps its copy as history, marked moved, with a pointer here. -->