Backlog triage Part B: 125 finished rows + 20 id-less rows moved to CLOSED-ITEMS (full text at 9e2786c); open rows normalised to one 6-column shape; narratives archived verbatim; closed_register_gate RULE 3 refuses a finished row in OPEN-ITEMS (decoys, seen red); rules rehomed to CONTEXT + 07 §11; loose notes triaged; R-814..R-819 filed; register 444 -> 325
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,114 @@
|
||||
# DIAGNOSIS — post-F9 storage-registration gap (read-only; no system changes made)
|
||||
|
||||
**Date:** 2026-06-14. **Scope:** why the HDD (`felhom-usb`) is attached + written-to by RomM yet badged
|
||||
"Nem regisztrált" and absent from the deploy dropdown. **Method:** live SSH (read-only) + controller
|
||||
source. **No state changed** — no registration, deploy, restart, or edits. Diagnosis only; fix to be specced next.
|
||||
|
||||
## TL;DR
|
||||
F9 attached the HDD at the **agent/guest layer** (the `pct set -mpN` bind). The controller's **storage
|
||||
registry** (`settings.json → storage_paths`) is a **separate layer** that F9 never touched. The controller
|
||||
registers a drive only through its own enroll flows (`runStorageInit` / `runStorageAttach` /
|
||||
`handleStorageRegister`), each of which calls `registerStoragePath`. In the F9 session the drive was
|
||||
attached by calling the **agent's `/disks/guest-attach` directly** (and RomM got `HDD_PATH=/mnt/felhom-usb`
|
||||
via a manual `deploy` value) — both bypass the controller's registry. Auto-discovery can't backfill it
|
||||
(it's a one-time seed that only scans deployed-app `HDD_PATH`s, never the agent's attached-drive list).
|
||||
Result: attached + in use, but unregistered → not selectable for a new app. **F9 is NOT closed.**
|
||||
|
||||
## Q1 — Real storage topology (what each path maps to)
|
||||
| Name (UI) | Host path | Device | Size / free | What it is |
|
||||
|---|---|---|---|---|
|
||||
| `Tárhely (felhom-data)` ★ (registered) | `/mnt/sys_drive/felhom-data` | `/dev/mapper/pve-vm--9201--disk--0` | 32 G / **~29 G** | the **internal OS-disk volume** (rootfs disk-0). The "28.7 GB szabad" in the dropdown = this. |
|
||||
| `felhom-usb` (NOT registered) | `/mnt/felhom-usb` | `/dev/sdb1` | **916 G / 870 G** | the **HDD**, guest-attached (mp1), 2.1 MB used |
|
||||
|
||||
Evidence: `pct config 9201` → `mp1: /mnt/felhom-usb/felhom-data,mp=/mnt/felhom-usb`; guest `df` → `/mnt/sys_drive`
|
||||
on `…disk--0` (32 G, 29 G avail) vs `/mnt/felhom-usb` on `/dev/sdb1` (916 G, 870 G avail). The registry
|
||||
(`settings.json`) lists only `{"path":"/mnt/sys_drive/felhom-data","label":"Tárhely (felhom-data)","is_default":true,"added_at":"2026-06-13T22:22:50Z"}`.
|
||||
**Confirmed: the "28.7 GB" option is the internal volume, not the HDD.**
|
||||
|
||||
## Q2 — Where RomM's data is + how it got an HDD path
|
||||
- RomM's `app.yaml` (`/opt/docker/stacks/romm/app.yaml`): `env.HDD_PATH: /mnt/felhom-usb`, `deployed: true`,
|
||||
`locked_fields: [HDD_PATH]`. Its data dirs live under `/mnt/felhom-usb/felhom-data/appdata/romm/…` (the HDD).
|
||||
- **How it got set:** the `HDD_PATH` deploy field was filled with `/mnt/felhom-usb` when RomM was (re)deployed
|
||||
onto the HDD during the F9 P3/restore test. The deploy path validates `os.Stat(HDD_PATH)` exists
|
||||
(`stacks/deploy.go` path-field check) — it does, because F9 had bound `/mnt/felhom-usb` into the guest — but
|
||||
**the deploy does NOT register the path** in `storage_paths`. So RomM references the HDD path *directly*,
|
||||
entirely independent of the registry. That's how "RomM uses the HDD" coexists with "HDD unregistered."
|
||||
|
||||
## Q3 — Why `felhom-usb` is "Nem regisztrált"
|
||||
The registry (`settings.json storage_paths`) contains **only** `/mnt/sys_drive/felhom-data`; `/mnt/felhom-usb`
|
||||
is absent. Nothing registered it, at any layer:
|
||||
- **F9's attach is agent-layer only.** `ReassertGuestBinds` / `handleDiskGuestAttach` → `guestbind.go AttachBind`
|
||||
do `pct set -mpN` (the bind). They never call the controller's `registerStoragePath`. In the F9 session the
|
||||
drive was attached by calling the agent `/disks/guest-attach` **directly**, bypassing the controller enroll flow.
|
||||
- **Deploy doesn't register.** Setting `HDD_PATH` on RomM only validates existence (above), no registry write.
|
||||
- **Auto-discovery can't backfill it** — two reasons in `settings.AutoDiscoverStoragePaths` (`internal/settings/settings.go:584`):
|
||||
1. **One-time seed:** `if len(s.StoragePaths) > 0 { return // already configured }` (≈:592). The registry already
|
||||
holds `sys_drive` (seeded 2026-06-13), so discovery is now **permanently inert** on every restart.
|
||||
2. **Wrong source even if it ran:** it's fed by `discoverHDDPaths(StacksDir)` (`cmd/controller/main.go:1337`,
|
||||
called once at startup, `main.go:126`), which scans **deployed apps' `HDD_PATH`** — never the agent's
|
||||
`/disks` attached-drive list. A drive attached out-of-band (or whose app was deployed after startup) is invisible.
|
||||
|
||||
So: **registration is decoupled from attach; auto-discovery is a one-shot startup seed of app `HDD_PATH`s; the
|
||||
agent-layer F9 bind and the manual deploy both bypass registration** → the HDD stays unregistered.
|
||||
|
||||
## Q4 — Why the deploy dropdown excludes the HDD
|
||||
The deploy dropdown is populated from `s.settings.GetSchedulableStoragePaths()` (`internal/web/handlers.go:323`),
|
||||
i.e. **registered paths only**. `/mnt/felhom-usb` is not in the registry → not schedulable → absent. So a new
|
||||
app (Calibre) can only be pointed at `sys_drive` (the 31 G internal volume). The "Nem regisztrált" badge is the
|
||||
storage page (`storageWizardPageHandler`) merging the agent `/disks` (attached, `guest_attached=true`) against
|
||||
`GetStoragePaths()` (registry) and flagging the difference (`storage_handlers.go` `Registered` field;
|
||||
`settings.html` `window.__registeredPaths`).
|
||||
|
||||
## Q5 — Intended registration-vs-attach behavior (per the code)
|
||||
The controller's design intends **register + attach as one operator-driven enroll operation**, NOT an automatic
|
||||
side-effect of an agent bind:
|
||||
- `runStorageInit` (format a blank/new drive) and `runStorageAttach` (mount an existing-fs drive — the "Csatolás"
|
||||
button) both do: assign (host mount) → **`registerStoragePath`** → `attachIntoGuest`→`agent.GuestAttach`
|
||||
(`internal/web/storage_handlers.go:83-158`). The comment is explicit: *"the StoragePath **registration is the
|
||||
durable intent**, so a transient attach failure is logged (not fatal) — P3 self-heal completes it."*
|
||||
- For an **already-mounted, unregistered** drive (exactly `felhom-usb`'s state), there is a dedicated action:
|
||||
`POST /api/storage/register` → `handleStorageRegister` (`storage_handlers.go:442`, the "Regisztrálás" button),
|
||||
commented: *"the natural primary action for a mounted-but-unregistered data drive (e.g. **felhom-usb**): the
|
||||
customer's intent is to USE the existing data, not wipe it… registering a host-only mount otherwise leaves it
|
||||
guest-invisible (the exact gap that produced the 'nem elérhető' banner)."* It registers + attaches into the guest.
|
||||
|
||||
So per the code, registration is meant to come **through these controller flows**. There is **no reconciliation
|
||||
that auto-registers a drive the agent attached out-of-band** — which is precisely the F9 case (the agent's
|
||||
re-assert / a direct guest-attach binds the drive without ever invoking the controller's register step).
|
||||
|
||||
## Q6 — The "felhom-data" naming collision
|
||||
Deliberate convention, but genuinely confusing. `felhom-data` is BOTH:
|
||||
1. the **path tail of the internal volume's registered path** (`/mnt/sys_drive/felhom-data`), surfaced as the
|
||||
label `Tárhely (felhom-data)`, and
|
||||
2. the **per-drive Felhom namespace subdirectory** created on EVERY user-data drive by the agent's `AttachBind`
|
||||
(`guestbind.go:29 felhomDataNS = "felhom-data"`, "matches `appbackup.FelhomDataDir`") — hence `/mnt/felhom-usb/felhom-data/…`.
|
||||
So `felhom-data` denotes both a specific volume and the generic namespace. Risk: the label `Tárhely (felhom-data)`
|
||||
reads as generic but is specifically the OS-disk volume; meanwhile the HDD also has a `felhom-data` dir. A clearer
|
||||
label (e.g. "Belső SSD" for the internal volume vs the drive's own name for the HDD) would remove the ambiguity.
|
||||
|
||||
## Recommended fix direction (to spec next — NOT implemented here)
|
||||
1. ~~**Auto-register on attach (primary).**~~ **REJECTED (2026-06-14)** — contradicts the new→enrolled
|
||||
manual-enrollment model; manual enroll is by design. A drive becomes usable through the explicit
|
||||
enroll/"Regisztrálás" flow (recommendation #2), not by auto-registering whatever the agent attached
|
||||
out-of-band. (The "make discovery additive, not return-if-any-path-exists" sub-point WAS adopted —
|
||||
shipped in controller v0.64.0, A1 — but only as additive pickup of paths referenced by *deployed apps*,
|
||||
NOT as auto-register-on-attach.)
|
||||
~~Make a drive the agent has attached into the guest become a registered
|
||||
storage path automatically, so F9's attach is usable end-to-end. Options: (a) the controller's startup/periodic
|
||||
storage reconcile consults the **agent `/disks`** and registers any `guest_attached=true` user-data drive missing
|
||||
from the registry (extend beyond the app-`HDD_PATH` scan); and/or (b) the enroll/re-assert path that binds a
|
||||
drive also performs `registerStoragePath`. Also **remove the one-time-seed gate** (or make discovery additive,
|
||||
not "return if any path exists") so it can pick up drives added after the first run.~~
|
||||
2. **Keep `handleStorageRegister` as the explicit fallback** ("Regisztrálás" button) for the mounted-but-unregistered
|
||||
case, and make the "Nem regisztrált" badge clearly actionable. For the **immediate** demo state, this button
|
||||
(`POST /api/storage/register {where:/mnt/felhom-usb}`) would register felhom-usb + make it selectable — no fix needed to recover now.
|
||||
3. **Clarify the `felhom-data` label** (Q6) to disambiguate the internal volume from the per-drive namespace.
|
||||
Recommended: **both** auto-register-on-attach AND the clearer manual path — auto so the normal flow "just works",
|
||||
manual for out-of-band/edge cases.
|
||||
|
||||
## Verdict on F9 — NOT closed
|
||||
F9 made the HDD present + writable **inside the guest** (agent layer), but did not make it **usable for new apps
|
||||
through the UI** (controller registry layer). The deploy dropdown is gated on the registry; the registry is
|
||||
unaware of the agent-attached drive. **F9 should remain open (or spawn an F9b)** until: attach ⇒ auto-registered
|
||||
⇒ a NEW app (e.g. Calibre) can be deployed to the HDD through the normal deploy UI with its data landing on
|
||||
`/dev/sdb1` persistently. Today that round-trip fails at the dropdown.
|
||||
@@ -0,0 +1,51 @@
|
||||
> **STATUS: FIXED** in controller v0.62.0 @ `f8afe5c` (2026-06-14) — implemented trunk-based on `main` per the plan below, with a regression test. This note is retained for provenance.
|
||||
|
||||
# fix/m18-dump-validation-cache — NOTES (pending review, NOT deployed)
|
||||
|
||||
**Verdict:** LIVE @ `controller/internal/appbackup/dbdump.go:469` (commit `6953899`).
|
||||
**Class:** performance (not correctness/security). **Escape-hatch branch** — the fix crosses the
|
||||
`settings`↔`appbackup` package boundary and changes an exported signature; not forced unattended.
|
||||
|
||||
## The bug (verified mechanism)
|
||||
|
||||
`ListDumpFiles(dumpDir)` (dbdump.go:425) unconditionally calls `f.Validation = ValidateDump(fullPath, f.DBType)`
|
||||
(line 469) for **every** `.sql` file on **every** call. `ValidateDump` (line 320) opens the file and
|
||||
scans it line-by-line (bufio loop). The caller chain is `backup.RefreshCache` (every ~5 min, scheduler)
|
||||
→ `listAllDumpFiles` → `ListDumpFiles` for every drive/stack. So every dump is fully re-read every 5
|
||||
minutes, even when unchanged.
|
||||
|
||||
A `settings.DBValidationCache` type exists (`settings.go:159`) and is written in `RunDBDumps`
|
||||
(`backup.go:447`), but it has only `ValidatedAt/TableCount/HasHeader/Error` — **no size or modtime** —
|
||||
and `ListDumpFiles` never consults it. `DumpFileInfo` already carries `Size` + `ModTime` (dbdump.go:451-453),
|
||||
so the inputs for a cheap skip-check are present; they're just not used.
|
||||
|
||||
## Impact
|
||||
|
||||
On large DB dumps (hundreds of MB) this is wasted disk I/O + CPU every 5 minutes. Negligible on the demo
|
||||
(small dumps); real on a customer with big databases. No correctness impact — validation results are the
|
||||
same, just recomputed.
|
||||
|
||||
## Fix plan (implementable, low-risk once reviewed)
|
||||
|
||||
1. Add `Size int64` and `ModTime string` (RFC3339) to `settings.DBValidationCache`.
|
||||
2. Change `ListDumpFiles(dumpDir string)` → `ListDumpFiles(dumpDir string, cached func(name string, size int64, mod time.Time) (DBValidationResult, bool))`.
|
||||
Pass a **plain lookup func**, NOT the `settings` type — `appbackup` must not import `settings` (avoid
|
||||
an import cycle; keep appbackup leaf-like). The func returns the cached `DBValidationResult` + ok.
|
||||
3. Before line 469: if `cached(e.Name(), info.Size(), info.ModTime())` returns ok, reuse it and skip
|
||||
`ValidateDump`; else validate and (caller-side) store the result keyed by name+size+modtime.
|
||||
4. Update the bridge re-export (`backup/appbackup_bridge.go:66`) and the `backup.go` call sites to build
|
||||
the lookup from `settings`'s cache (now carrying size+modtime), and to write back fresh validations.
|
||||
5. Keep a nil-`cached` fast path (validate-always) so other callers/tests don't have to thread it.
|
||||
|
||||
## Test plan (regression)
|
||||
|
||||
- Unit test in `appbackup`: call `ListDumpFiles` twice with a `cached` func that records calls; assert
|
||||
`ValidateDump` is NOT re-run for a file whose size+modtime match the cache, and IS run when modtime
|
||||
changes. (A spy counter on a wrapped validator, or assert via a temp `.sql` whose mtime is bumped.)
|
||||
- This test fails on the pre-fix code (which always validates).
|
||||
|
||||
## Why not tonight
|
||||
|
||||
Crosses a package boundary + changes an exported signature with multiple call sites (bridge + backup.go),
|
||||
and the cache type needs new fields — too entangled to land safely unattended. Hand to the supervised
|
||||
session: the plan above is complete and mechanical.
|
||||
@@ -0,0 +1,58 @@
|
||||
> **STATUS: FIXED** in controller v0.62.0 @ `6bab68b` (2026-06-14) — implemented trunk-based on `main` per the plan below, with a regression test. This note is retained for provenance.
|
||||
|
||||
# fix/m19-stackname-crossref — NOTES (pending review, NOT deployed)
|
||||
|
||||
**Verdict:** LIVE @ `controller/internal/appbackup/dbdump.go:536-551`, used at `:115` (commit `6953899`).
|
||||
**Class:** correctness edge — **low real-world incidence** with the current catalog. **Escape-hatch branch**
|
||||
— the clean fix injects the deployed-stack list into `appbackup` (cross-package), not forced unattended.
|
||||
|
||||
## The bug (verified mechanism)
|
||||
|
||||
`deriveStackName(containerName)` pure-suffix-strips: it splits on `-` and, if the last segment is in
|
||||
`{postgres,db,mariadb,mysql,database,redis,cache}`, returns the join of the remaining parts. It never
|
||||
cross-references actual deployed stack names. `DiscoverDatabases` assigns `StackName: deriveStackName(name)`
|
||||
directly (line 115).
|
||||
|
||||
So a stack whose real name **ends** in one of those tokens is misattributed:
|
||||
- a stack literally named `my-cache` → its DB container `my-cache` (or `my-cache-postgres`) derives to
|
||||
`my` / `my-cache`, attributing the dump to the wrong (or a non-existent) stack.
|
||||
- worse, `a-db` and `a` could collide.
|
||||
|
||||
## Impact
|
||||
|
||||
A DB dump is filed under the wrong stack name → that stack's backup/restore accounting is wrong, and the
|
||||
restore-by-stack path could miss or cross-wire the dump. **Incidence is effectively zero in the current
|
||||
felhom catalog** (stack slugs are `romm`, `nextcloud`, `paperless-ngx`, `immich`, `adventurelog`,
|
||||
`actualbudget`, `mealie`, `vikunja`, … — none ends in a DB-role token; DB containers are `<stack>-postgres`
|
||||
etc., which strip correctly). It becomes real only if a future app slug ends in a role token.
|
||||
|
||||
## Fix plan (implementable)
|
||||
|
||||
1. Thread the set of **known deployed stack names** into discovery:
|
||||
`DiscoverDatabases(ctx, logger, debug, knownStacks []string)` and
|
||||
`deriveStackName(containerName string, known map[string]bool)`.
|
||||
Source the list from the caller in `backup.go` (it holds the `StackDataProvider` — expose/known stack
|
||||
names via a lookup func to avoid importing `stacks` into `appbackup` and creating a cycle).
|
||||
2. New `deriveStackName` logic:
|
||||
- candidate := current suffix-strip result.
|
||||
- if `known[candidate]` → use it (the suffix was a real DB-role suffix of a real stack).
|
||||
- else if `known[containerName]` → the container name IS the stack (don't strip).
|
||||
- else → longest `known` stack name that is a prefix of `containerName` (handles `<stack>_postgres`,
|
||||
`<stack>-1`, compose-suffixed names); tie-break to the longest match.
|
||||
- else → fall back to the current suffix-strip (preserve today's behaviour when the stack list is
|
||||
unavailable/empty, so nothing regresses).
|
||||
3. Keep a nil/empty-`known` fast path = today's behaviour (back-compat for other callers/tests).
|
||||
|
||||
## Test plan (regression)
|
||||
|
||||
- Table test for `deriveStackName` with `known = {romm, my-cache}`:
|
||||
- `romm-postgres` → `romm` (suffix is a role; `romm` is known).
|
||||
- `my-cache` → `my-cache` (known as-is; must NOT strip to `my`). ← fails on pre-fix code.
|
||||
- `my-cache-postgres` → `my-cache` (strip role, result known).
|
||||
- `unknown-db` with no matching known → falls back to `unknown` (today's behaviour).
|
||||
|
||||
## Why not tonight
|
||||
|
||||
Requires plumbing the deployed-stack list across the `backup`→`appbackup` boundary (cycle-avoidance via a
|
||||
lookup func) and touching `DiscoverDatabases`'s signature + caller. Low-incidence, so not worth a risky
|
||||
unattended change. The plan + tests above are complete for the supervised session.
|
||||
@@ -0,0 +1,40 @@
|
||||
# FOLLOW-UP — bump the golden's default controller tag + validate the full provision path
|
||||
|
||||
**Status:** **RESOLVED 2026-07-03** — `build-golden.sh` v2.0.0 (felhom-agent @ `ceca355`) makes the
|
||||
controller tag a MANDATORY argument (the hand-bumped default had rotted AGAIN, 0.43.0 → 0.85.1 →
|
||||
stale vs 0.98.3 — a required arg cannot rot); golden **0.98.3** baked + validated clean-room on the
|
||||
drill VM (first boot lands the current controller, self-update reports up-to-date, app deploys) +
|
||||
published + operator-vouched. The full first-boot path was validated WITHOUT a supervised
|
||||
destroy→re-provision (drill VM virgin snapshot instead — nothing touched 9201/felhom-pve).
|
||||
Evidence: `../audits/DRILL-golden-098-2026-07-03.md`. Item 3 (provision re-asserting user-data
|
||||
binds) was shipped separately (agent-startup re-assert + F9 part A). Original note kept below for
|
||||
history.
|
||||
**Class:** provisioning correctness / customer-onboarding. **Risk (as queued):** SUPERVISED (golden + a real destroy→provision).
|
||||
|
||||
## The problem
|
||||
`felhom-agent/configs/build-golden.sh:43` bakes the controller image into the golden as:
|
||||
|
||||
```
|
||||
CONTROLLER_IMAGE="${6:-gitea.dooplex.hu/admin/felhom-controller:0.43.0}"
|
||||
```
|
||||
|
||||
`:0.43.0` is ~20 versions stale (current is `0.63.0`). The bootstrap writes it to
|
||||
`/etc/felhom-controller-image` and `docker run`s that tag on first boot. So the **next time a golden is
|
||||
actually baked for a real provision, it would stand a customer guest up on an ancient controller** —
|
||||
missing every fix since 0.43.0 (incl. F1 memory guard, F17 DB restore, the M18/M19 fixes, etc.). It only
|
||||
hasn't bitten because the running demo guest 9201 had its `/etc/felhom-controller-image` updated in place
|
||||
post-provision; a fresh provision would not.
|
||||
|
||||
This is why the F9 session chose **attach-to-existing** over destroy+re-provision (a re-provision would
|
||||
have regressed 9201's controller 0.62.0 → 0.43.0).
|
||||
|
||||
## The fix (separate task)
|
||||
1. Bump the `build-golden.sh` default `CONTROLLER_IMAGE` to the current released controller tag (and
|
||||
establish a convention so it tracks releases — e.g. read a `LATEST_CONTROLLER` pin, or pass it
|
||||
explicitly from the build pipeline).
|
||||
2. **Validate the full path end-to-end**, which *does* warrant a supervised destroy + re-provision (it is
|
||||
the real customer-onboarding flow, not coverable by attach-to-existing): bake golden → provision a
|
||||
scratch guest → first boot stands up the **current** controller → controller ONLINE on the hub →
|
||||
enrolled user-data drive is bound (F9 `ReassertGuestBinds` / provision bind) → an app deploys.
|
||||
3. While there: confirm the provision path itself asserts known user-data binds (F9 part A for the
|
||||
re-provision case, complementing the agent-startup re-assert already shipped in v0.31.0).
|
||||
@@ -0,0 +1,361 @@
|
||||
# OPEN-ITEMS — the narrative sections, as they stood on 2026-10-03
|
||||
|
||||
> **History, not a register.** These sections sat between the register tables of
|
||||
> `documentation/backlog/OPEN-ITEMS.md` until the 2026-10-03 triage. They are moved here word for word
|
||||
> (`git show 9e2786c:documentation/backlog/OPEN-ITEMS.md` holds them in place). Every row they discuss is in
|
||||
> `OPEN-ITEMS.md` (open) or `CLOSED-ITEMS.md` (finished). A status word below is the status ON THE DATE OF
|
||||
> ITS SECTION and may be stale — the registers are the authority. The ranking paragraphs ("Why the TOP
|
||||
> READY rows rank this way", "The 2026-08-02 intake, ranked") are superseded by
|
||||
> `documentation/audits/backlog-triage-2026-10-03/RECOMMENDATION.md`.
|
||||
|
||||
## Operator rulings — 2026-08-04
|
||||
|
||||
Recorded here because a ruling that lives only in a conversation binds nobody (the R-96 standing rule).
|
||||
|
||||
1. **Run the recovery drill, after R-198.** R-198 shipped in hub v0.93.0; **the drill is the next
|
||||
session.** Design: `audits/RECON-offsite-dr-chain-2026-08-04.md` §10 — demo-hp, ~3–4 h, a recovery
|
||||
code created and KEPT, a sentinel file, wipe, reinstall, recover, and **pass = a byte-identical
|
||||
sha256, not "the repository opened"**. → R-201
|
||||
2. **Delete the orphaned ciphertext** (~1.2 GB across the two demo boxes, in set-aside restic stores
|
||||
nothing prunes). **STILL OWED — not done in v0.93.0.** It is a destructive act on a protected
|
||||
endpoint and belongs to a session that is scoped for it, not to a release that ships a schema
|
||||
change. → R-193
|
||||
3. **Accept the risk on R-193 candidate (c)** — no repository password retained on the Proxmox host.
|
||||
**This is what makes R-198 load-bearing rather than tidy:** with no host-retained copy, the
|
||||
customer-present recovery path is the ONLY way back from a rebuild, and that path runs entirely
|
||||
through the retained identity blob. → R-193, R-199, R-200, R-201
|
||||
|
||||
~~**Still open and untouched by v0.93.0:** R-199, R-200, R-201.~~ **UPDATED 2026-08-04 (evening), after
|
||||
hub v0.94.0 + agent v0.125.0 + controller v0.195.0:**
|
||||
|
||||
- **R-199 — CLOSED, proven on hardware.** Chain links 6–8 are assembled and walked. The offsite
|
||||
repository password came back out of the sealed bundle **byte-identical** to the one on disk
|
||||
(`c60c8bc737a6…` from three independent sources: the box's file, the recovered bundle, and the hash
|
||||
the hub already stored).
|
||||
- **R-200 — the plumbing half shipped**, the customer-facing form did not, deliberately.
|
||||
- **R-201 — PASSED 2026-08-04 (night run).** A customer's file survived a machine rebuild and came
|
||||
back byte-identical, through the customer's own restore flow. **It passed only because a person was
|
||||
there:** four manual interventions stood between the recovered key and the restored file, none of
|
||||
them in any design document → R-204.
|
||||
- **R-202 — untouched.** The orphan card still promises recoverability unconditionally.
|
||||
- **The orphaned ciphertext deletion (~1.2 GB) is STILL OWED** — ruling 2, above.
|
||||
|
||||
**UPDATED 2026-08-05, after controller v0.198.0 + hub v0.95.0 (R-204):**
|
||||
|
||||
- **R-204 items 1–3 — CLOSED.** The reset code works without a restart; a re-issue no longer marks a
|
||||
healthy escrow stale (→ **R-196 CLOSED**); a unit restore states what it did NOT restore.
|
||||
- **R-204 item 4 — OPEN and unstarted:** a rebuilt box cannot obtain an off-site credential unaided.
|
||||
It needs an operator ruling on the one-shot credential design → **R-193**.
|
||||
- **Still open and untouched by this session, stated so nothing is presumed closed by association:**
|
||||
**R-202** (the orphan card's unconditional promise), **the ~1.2 GB orphaned-ciphertext deletion**
|
||||
(ruling 2 — still owed, still needs its own scoped session), and **R-198's retention, which remains
|
||||
UNIT-PROVEN ONLY** — nothing has superseded a key in production, and proving it needs a SECOND
|
||||
deliberate wipe. That retention drill is the next item, and it is not this session's.
|
||||
|
||||
v0.93.0 made the key survive. v0.94.0/v0.125.0/v0.195.0 make it come back. v0.197.0 got the file into
|
||||
the snapshot and the drill got it out again. **v0.198.0/v0.95.0 remove three of the four crutches the
|
||||
drill needed — the fourth is R-193, and until it goes the recovery is still operator-assisted.**
|
||||
|
||||
## R-201 — THE RE-WALK, 2026-08-06 (attended)
|
||||
|
||||
**The question was asked a second time, on the fixed build, on a brand-new appliance built from the
|
||||
published ISO. The answer is still no — but it is a nearer no.**
|
||||
|
||||
| half | verdict |
|
||||
|---|---|
|
||||
| **the data** | **PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also byte-identical** (verified as hex, not as rendered text). Restored in **23 s** out of the pre-destruction snapshot `a7bc23bd`, through the customer's own two-step full-restore flow |
|
||||
| **the journey** | **FAIL — two dead ends**, against Phase 1's four. One needed a command line **inside the guest**, one a **Proxmox-host** action |
|
||||
|
||||
**RTO: the unaided figure is STILL UNDEFINED**, because the unaided journey still does not complete.
|
||||
Attended: login 11:42:22 → key placed **+45 s** → tier up **+24 m 12 s** (after intervention 1) →
|
||||
all three sentinels restored and verified **+30 m 13 s** (after intervention 2). **The 30 m figure must
|
||||
not be quoted as the customer number.** The only segment that reflects the product working alone is
|
||||
**23 seconds** to pull 12.8 MB back once everything was in place.
|
||||
|
||||
**The two dead ends:** **R-218's consume half** (row corrected above) and **R-220** (drives
|
||||
unenrollable after a rebuild — reproduced and red-proved again; without it no app can be redeployed,
|
||||
and without a redeployed app the restore page is empty, which is R-213's territory and follows from
|
||||
R-220 rather than being separate).
|
||||
|
||||
**What PASSED and is worth keeping:** the recovery screen **appeared without being sought**
|
||||
(`/` → `/launcher` → `/recovery`); it answered all three of its questions and its seal date matched the
|
||||
hub's `created_at` exactly; the emailed reset code worked **first try**; the unlock took **1.528 s** —
|
||||
a real unseal — and placed the key; and **R-225's fix was seen working in the wild** (the store read
|
||||
„a pillanatképek száma még ismeretlen" rather than a false zero).
|
||||
|
||||
**R-216 part 4 reproduced live:** the reinstall **downgraded the agent 0.126.0 → 0.125.0**, back to the
|
||||
vouched version — an operator's hand-fix undone by the very event that makes recovery necessary.
|
||||
|
||||
⚠ **THE DELIVERY GAP, and it is owed.** A fresh install landed on controller **0.201.0** / agent
|
||||
**0.125.0** — the vouched versions, **neither carrying the fixes**. They were installed **by hand**.
|
||||
Fleet delivery needs a golden carrying 0.202.0 **and** a vouched agent 0.126.0. **Nothing was vouched;
|
||||
that is the operator's act.** **This re-walk proves the JOURNEY on the fixed build; it does NOT prove a
|
||||
customer would receive that build.**
|
||||
|
||||
Evidence: `tests/rewalk-r201-2026-08-06/journal.md`.
|
||||
|
||||
## CAMPAIGN 11 — the recovery journey, 2026-08-05
|
||||
|
||||
**The whole journey was walked end to end for the first time, on a throwaway appliance built from the
|
||||
published ISO. The data came back byte-identical; the journey did not exist.** **Ten findings from
|
||||
Phases 1 and 3, R-214 … R-223** (seven fixed in controller **v0.201.0** + hub **v0.97.0/0.97.1**;
|
||||
**three deliberately still open**, each blocking a real flow), **plus five from Phase 2's injected
|
||||
faults, R-224 … R-228.** Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` (Phases 0/1/3)
|
||||
and `journal-phase24.md` (Phases 2/4). Campaign document:
|
||||
`audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`.
|
||||
|
||||
> **Phase 2's verdict in one line.** The cryptography, the retention and the transport all work and
|
||||
> are now proven live. **What fails is being told the truth:** a mistyped code, a hub outage, a
|
||||
> stopped agent and a correct code for a retained earlier package all produce **one** message, and
|
||||
> three of the four are wrong.
|
||||
|
||||
### Phase 2 — the injected faults, 2026-08-05/06 (unattended)
|
||||
|
||||
> **ALL FIVE CLOSED 2026-08-06** in controller **v0.202.0** + agent **v0.126.0**. The rule they now
|
||||
> enforce, stated so it outlives them: **on the unlock path the customer is blamed only after a real
|
||||
> attempt REFUSED their code; every other outcome, including an unclassifiable one, says something
|
||||
> else.**
|
||||
>
|
||||
> **STILL OPEN AND DELIBERATELY UNTOUCHED BY THAT WORK — said explicitly rather than left to
|
||||
> inference:** **R-214** (the console never stops showing a stale pairing code), **R-220** (a rebuilt
|
||||
> box's drives cannot be re-enrolled — *currently worked around BY HAND on the campaign venue*, which
|
||||
> is the only reason an app could be deployed there at all), **R-221** (a rebuilt box cannot run the
|
||||
> escrow ceremony), **R-213** (putting files back), **R-202** (the orphan card's unconditional
|
||||
> promise — now the **last** place on that surface still promising recoverability, two doors from
|
||||
> where R-228 removed the same promise).
|
||||
|
||||
Eleven faults, each judged on **the message, not the outcome**, each with a positive control proving
|
||||
the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/journal-phase24.md`.
|
||||
|
||||
## Instruction files — deferred half, 2026-08-06
|
||||
|
||||
**Recorded against existing rows by Phase 2:**
|
||||
|
||||
- **R-216 — §4.1 is now MEASURED, not deduced.** The previous session could only offer two absences.
|
||||
The box's own `/settings` renders „Minimális verzió (üzemeltető) **0.200.0**" (`GetFloor()`, whose
|
||||
only writer is the report-ACK handler; **both hold branches serve `Floor=""`**, pinned by
|
||||
`managed_floor_test.go:94`), and a **cold-started** controller logs
|
||||
`settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)` against the same line reading
|
||||
`floor still unknown after 1m30s` while the hold was in force. The hub's HELD lines ran every
|
||||
15 min to 22:12:06 and stopped, with a liveness control proving the hub kept logging. **The floor
|
||||
is served.**
|
||||
- **⚠ A correction to how that positive was to be taken.** `SetFloor`'s line is `u.dbg(...)`, gated on
|
||||
`cfg.Logging.Level == "debug"` and written to the **logger** — it can **never** reach the logx debug
|
||||
ring, so it cannot appear in `/api/debug/logs` at any level. A controller restart alone would not
|
||||
have produced it. Confirmed with a level census on the ring first (1196 DEBUG / 2802 INFO / 2 WARN),
|
||||
so the absence was known to be structural rather than evidential.
|
||||
- **R-218 — the live half is STILL NOT MEASURED, deliberately.** The fix is present and correct
|
||||
(`needsOffsiteCredential` now retires on the **target**, not the key), but the venue **has** a
|
||||
target, so the box correctly does not declare; declaring here would be the bug. The state that
|
||||
exercises it is shape (a), which the venue no longer holds. **Recorded as not measured rather than
|
||||
inferred from the unit test.**
|
||||
- **R-217 — its fix HELD under exactly its fault** (F5): with the store blocked after a successful
|
||||
unlock, the page rendered M3 and **no listing block at all**; the three false-claim strings are
|
||||
absent, verified in UTF-8 with accented positive controls present.
|
||||
- **R-215 — its fix is present** and the `GET /recovery` gate consults the same predicate as the POST
|
||||
sibling.
|
||||
- **R-199's back-pointer in `architecture/00-capability-map.md` was already added** — the brief lists
|
||||
it as owed; it is present on the escrow-recovery row, explicitly labelled as the omitted
|
||||
back-pointer. **No action taken; the brief's assumption was stale.**
|
||||
|
||||
**Untouched by this session, stated so nothing is presumed closed by association:** **R-213** (putting
|
||||
files back — the half the recovery screen deliberately does not do) and **R-202** (the orphan card's
|
||||
unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing cost of).
|
||||
|
||||
## CAMPAIGN 12 — the class sweep, 2026-08-08 (unattended)
|
||||
|
||||
Eight rows, **grouped by class so the classes are visible as classes**. Full method, controls, blind
|
||||
spots and the Part-4 gating ranking: `audits/CAMPAIGN-12-class-sweep-2026-08-08.md`. Gating candidates
|
||||
are in `ROADMAP.md`, not here. **C1 produced no new instance** and has no row, deliberately.
|
||||
|
||||
**⚠ A correction the campaign owed to its own brief:** the task described C5's `escrow_stale` instance
|
||||
as "closed individually". **It is not — it is R-247, `READY`.** The live repo is the source.
|
||||
|
||||
## CAMPAIGN 12 follow-through — G-1 built, R-260 and R-247 closed, 2026-08-08
|
||||
|
||||
**The gate was built BEFORE the fixes and was seen failing on 40 fields**, captured verbatim in
|
||||
`documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md`. That order was the method, not
|
||||
bureaucracy: the night before, `deadcode` was rejected for class C6 precisely because it was made to
|
||||
prove itself first and found neither of the two defects it was meant for.
|
||||
|
||||
**⚠ A COUNT THIS SESSION'S OWN PROMPT GOT WRONG, corrected against the repo rather than quoted.** The
|
||||
prompt said *"465 emitted tags, eight unreachable"*. R-260's wording was "at least eight
|
||||
DECISION-BEARING facts", never eight tags in total, and its own census already listed more. Measured
|
||||
on the three declared wires: **210 tags checked, 51 skipped, 40 convicted.** Two prompt claims were
|
||||
wrong this week and both were caught the same way.
|
||||
|
||||
**Two things the gate's CONTROL caught before it was trusted**, each a defect in the instrument:
|
||||
|
||||
1. **A substring false negative.** `grep -F healed_at` also matches `privsep_healed_at`, so a
|
||||
genuinely dropped field read as received. R-260 named `healed_at`, so its absence from the output
|
||||
was the tell. Now a whole-token regex.
|
||||
2. **`dr_recipe` is not wholly opaque.** The hub stores each half as `json.RawMessage` and re-emits
|
||||
nested shapes verbatim — but the TOP-LEVEL section keys are decoded by `hostHalfShape` /
|
||||
`appHalfShape`, and **those are allow-lists**: a section an emitter adds is silently dropped until
|
||||
named in both. That already cost `offsite_restic` (R-122). The gate is therefore opaque BELOW
|
||||
depth 1, so the sections are checked; treating the whole subtree as opaque would have put R-122's
|
||||
shape back outside its reach.
|
||||
|
||||
**THE FORTY, BY DISPOSITION.** Full per-field reasons live in the gate's own `ALLOWLIST`, where each
|
||||
entry is a claim someone can re-check.
|
||||
|
||||
| # | field(s) | direction | decision | what changed |
|
||||
|---|---|---|---|---|
|
||||
| 1 | `oob.operator_key_configured` | agent → hub | **RECEIVE AND ACT ON IT** | decoded as a pointer; `oobDegraded` now fails when the key is absent, and the alert NAMES it. Hub v0.99.0 |
|
||||
| 2 | `oob.wg_handshake_age_s`, `oob.healed_at` | agent → hub | **RECEIVE, for the message only** | decoded into `HostOOBRow` and put in the event payload; deliberately NOT in the predicate — widening the check beyond the fact that is now arriving is how a check stops being read |
|
||||
| 3 | `escrow_stale` | hub → controller | **RECEIVE AND ACT ON IT** | `report.EscrowStatus.Stale`; the box tells a withheld hash from a hash-less one. Controller v0.209.0. **This is R-247** |
|
||||
| 4 | `host.{cpu_temp_c,loadavg,memory_total_bytes,memory_used_bytes,uptime_seconds}`, `system.{load_avg_1,5,15, memory_total_mb, memory_used_mb, temperature_celsius, uptime_seconds}` | agent/controller → hub | **NO CONSUMER WANTED — redundant** | allowlisted: the hub decodes `cpu_percent` / `memory_percent` / `disk_percent` from the same stanzas and every threshold is expressed on those |
|
||||
| 5 | `guests.spec.{disk_bytes,memory_bytes}` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: guest sizing is hub-owned INTENT (the manifest), not mirrored reality |
|
||||
| 6 | `storage_targets.smart.model_name` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: a display label with no threshold on it; `smart.health` and every counter the hub bands on ARE decoded |
|
||||
| 7 | `wireguard.last_handshake_age_s` | agent → hub | **NO CONSUMER WANTED — redundant** | allowlisted: wgsync reconciles from its own state, and the OOB path's own handshake age is now decoded |
|
||||
| 8 | `guest_net` + its 7 children, `selfupdate_pending(_version)`, `mgmt_plane.healed_recently`, `pbs_dr.applied_at`, `restore_tests.mount_{parity,inventory}`, `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check` | agent/controller → hub | **NO CONSUMER TODAY, AND ONE IS ARGUABLY OWED** | allowlisted **against R-264, which stays OPEN**. Allowlisting is not deciding, and the entries say so |
|
||||
|
||||
**A C7 instance found while doing it, and corrected.** `HostReport.SelfUpdatePending`'s own comment
|
||||
claimed *"The hub reads an absent field as pending=false, the correct default."* The hub has no field
|
||||
for it and reads nothing either way. Comment corrected in the agent (no version bump — comment only).
|
||||
|
||||
**The hub's OOB test fixture was part of why this survived.** `oobReport()` omitted
|
||||
`operator_key_configured` entirely, so every pre-existing scenario ran against a report shape **no
|
||||
released agent produces**. Fixed, and the new tests drive the raw JSON decode boundary — a test that
|
||||
builds the receiving struct by hand cannot see a field that never decodes, which is the whole class.
|
||||
|
||||
**Explicitly still open, untouched by this session:** R-246 (the wrong stale flag on `demo-hp` —
|
||||
clearing it is an operator act hub-side), R-255, R-256, R-257, R-258, R-259, R-261, R-262, R-263, and
|
||||
**C7's test-comment half**, which Campaign 12 recorded as *owed, not done* (60 of 2652 production
|
||||
invariant comments sampled; none of the 1440 test comments).
|
||||
|
||||
## The seed that never ran twice, and three pictures that were not true — 2026-08-08
|
||||
|
||||
Four defects of one family: something the box already knows, either thrown away or drawn as its
|
||||
opposite. Agent **v0.128.0**, controller **v0.210.0**, `gates.yml` (no hub change, no hub bump).
|
||||
|
||||
**R-221's writer, ESTABLISHED at `file:line` rather than assumed** — the prompt asked for this and it
|
||||
was owed. `step_agent_config` (`felhom.eu/scripts/felhom-host-install.sh:2396`) renders `agent.json`
|
||||
from `base = {}` unless an explicit `--preserve-from` is passed (flag `:1246`, defaulting empty at
|
||||
`:256`), and writes it with `O_TRUNC` (`:2579`). **The render never writes an `escrow` section at
|
||||
all** — grep over the whole heredoc returns zero hits. The pbsdr marker lives host-side
|
||||
(`<agent-state>/pbsdr/marker.json`) and survives. So a rebuild keeps the marker and takes the key:
|
||||
same descriptor, same hash, early return, seed never re-runs. **The attribution in R-221 was
|
||||
correct.** A rebuild is nonetheless only the case that was measured — the same hole opens for a
|
||||
hand-edited or restored config, which is the honest reason the fix is at the seam and not in the
|
||||
installer.
|
||||
|
||||
**The idempotent early return was KEPT**, and that is load-bearing: it stops a converged box
|
||||
re-running Proxmox operations every 60 s.
|
||||
`TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner calls on
|
||||
that tick, so a "fix" that simply deleted the return fails the test. Verified by mutation.
|
||||
|
||||
**The §7.3 truth table as implemented** (R-258):
|
||||
|
||||
| this app's own most recent dump result | restore point | verdict |
|
||||
| any of its databases failed | yes | `error` — cross |
|
||||
| all clean | yes | `ok` — tick |
|
||||
| none recorded (no database / no run yet) | yes | **no icon**, time only, title „Erről a mentésről nincs eredményünk." |
|
||||
| any | no | no tier-1 row at all, unchanged |
|
||||
|
||||
**An existing test was asserting the defect and was corrected, not deleted.**
|
||||
`TestBuildAppBackupRows_Tier1FromRestorePoints` expected `"ok"` for a `FullBackupStatus` with **no
|
||||
`LastDBDump` at all** — a green tick derived from nothing but a file's existence, i.e. Scenario G.
|
||||
Its real subject, the `Tier1LastRun` time, is unchanged.
|
||||
|
||||
**The convention is now ruled** (§7.2, `CONTEXT.md` S-39): a `…Known bool` companion beside the
|
||||
figures. `ROADMAP.md` G-3 was blocked on that decision and is unblocked.
|
||||
|
||||
**Six red-proofs, every one demonstrated failing and restored, each with the mutation asserted
|
||||
applied.** The one that matters: Scenario A **fails against today's tree** with the intended message
|
||||
— so the test tests the defect.
|
||||
|
||||
**Explicitly still open, untouched by this session:** R-246, R-255, R-256, R-257, R-261, R-262,
|
||||
R-263, **R-264** (the twenty-one undecided facts — a design session of its own), R-240, R-243,
|
||||
R-202, R-213, R-244, R-214/R-235, and **C7's test-comment half**, which Campaign 12 recorded as
|
||||
*owed, not done*. **G-8's other half** (a hub-side check that notices a *vouch* has been forgotten)
|
||||
was deliberately not built: it is hub work whose payoff is a daily email, and this session already
|
||||
ends with a bake-and-vouch cycle in front of the operator.
|
||||
|
||||
## Why the TOP READY rows rank this way
|
||||
|
||||
This covers the next few only — it is deliberately **not** a full ordering of the table above, so that
|
||||
there is one ranking to maintain rather than two.
|
||||
|
||||
1. **R-95** — the largest *data* exposure: the tier holding the customer's documents and photos is the
|
||||
one whose credential can delete. **THIS PARAGRAPH WAS STALE UNTIL 2026-09-01 AND ITS OLD TEXT IS
|
||||
NAMED SO THE CORRECTION IS NOT MISTAKEN FOR A RE-RANK.** It said the snapshot mitigation was
|
||||
*"armed (daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong:**
|
||||
seven daily snapshots do exist (R-429), and the word "armed" was withdrawn by the spike that same
|
||||
day. **What is true now:** the snapshots exist and the box cannot write into them, but **no account
|
||||
we hold can read anything out of one** (R-433, measured — 777,600 names, zero hits, controlled), so
|
||||
they do not yet bound this exposure. The root cause is untouched either way — the box can still
|
||||
`forget --prune` its own repo, from two call sites. **The ORDER of this list is unchanged and is
|
||||
Viktor's**; only the facts under item 1 were corrected.
|
||||
2. **R-94** — **de-ranked 2026-07-29.** The prior rationale ("until it moves every hub-driven
|
||||
install gets the pre-R-82 default") was false: the constant selects no script and every install
|
||||
already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment —
|
||||
a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing;
|
||||
not high-consequence, and it blocks nothing.
|
||||
3. ~~**R-86**~~ — **CLOSED 2026-08-03**, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.
|
||||
4. **R-87** — **SPIKED 2026-08-31 and now a DECISION, not work.** It also spent 2026-08-22..31 in
|
||||
`CLOSED-ITEMS.md` by mistake while this paragraph ranked it fourth and pointed at nothing
|
||||
(R-405). The spike measured it rather than designing it: a scratch restore of all 8 apps costs
|
||||
25 s against the 40.3 s weekly check, but it would have caught ONE of the five drill-found
|
||||
restore defects. **Recommendation: build the NARROW version — prove the snapshot still
|
||||
CONTAINS a recoverable unit — or close the row.** Viktor's call; see
|
||||
`audits/SPIKE-restic-restore-test-2026-08-31.md`. *The 2026-08-03 rationale, kept:*
|
||||
R-86 built most of what it was waiting for (per-archive due-ness, a
|
||||
proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is
|
||||
restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is
|
||||
no longer waiting on a scheduling model that did not exist.
|
||||
5. ~~**R-185**~~ — **CLOSED 2026-08-03**, agent v0.123.0 + installer 1.24.0, proven live on both demo
|
||||
boxes. The silence was fixed as well as the grant: the box now asks whether it may READ each tier
|
||||
it depends on, because an empty listing cannot distinguish forbidden from newborn.
|
||||
6. ~~**R-189**~~ — **CLOSED 2026-08-03** with **R-188** and **R-186**, agent v0.122.0. The three
|
||||
reporting/release signals that misreported their own work are fixed; **R-185 is the one that
|
||||
remains open from that group** and is untouched by this — it is a missing storage ACL on
|
||||
demo-felhom, not a reporting defect.
|
||||
7. **R-110** — last **because it is not a READY row**: the ruling is the operator's, not CC's, and
|
||||
there is nothing for CC to build until it lands. Ranked here rather than omitted because it is the
|
||||
only item on this page about the *publish channel* of the most privileged artifact Felhom ships,
|
||||
and today's exposure is zero — which makes now the cheapest moment it will ever be to decide.
|
||||
|
||||
### The 2026-08-02 intake (R-156 … R-164), ranked
|
||||
|
||||
Filed in one pass from Campaign 10, its two spikes, and the 53-template catalog persistence sweep.
|
||||
**R-156 and R-157 had lived only in audit documents** — the identical *"minted in a spike doc and
|
||||
never carried across"* failure the register already records for R-153/R-154/R-155, caught by the
|
||||
sweep's own §8.0 while it was happening. R-158 was minted by a second session on the same day for an
|
||||
unrelated finding, which is why the sweep's proposals were renumbered to R-159…R-162 at filing time.
|
||||
|
||||
1. **R-157** — highest: a `deployed: true` app can stay **down indefinitely** after a power cut or
|
||||
hard reset, and in mechanism **B** nothing reports it on any channel (`0 currently down`). It is the
|
||||
only row here where the customer loses service and has no signal at all.
|
||||
2. **R-156** — the class is now detectable and two of three apps are fixed; what remains is **papra's
|
||||
referral**, one app, well understood. *(Promoted 2026-08-02: R-161 was ranked here because nothing
|
||||
ran the gate; it now has a mandated entry point, so R-156's residue is the larger remaining item.)*
|
||||
4. **R-163** — a real ceiling that silently caps local backup once an app outgrows `mp1`, and it gates
|
||||
Tier-2 and Tier-3 as well. Ranked below the above only because **overflow itself is safe today** —
|
||||
it refuses per app and preserves the last good unit byte-identical. **RE-FRAMED 2026-08-02:** no
|
||||
longer waiting on a ratio — decision **D-a** merges `mp1` away, so the row is now the record of the
|
||||
constraint and the work moves to **R-165** (with **R-167** shipping in the same step). R-165 inherits
|
||||
this rank; it is the highest-ranked item that must land **before any external install**.
|
||||
5. **R-158** — the gap that makes R-163 dangerous: cross the size line and **one page** tells you. On
|
||||
its own it is a notification gap, not a silent failure, which is why it sits here and not higher.
|
||||
6. **R-164** — blocked on a predicate, no customer impact today; it only becomes urgent if the unit
|
||||
size in R-163 is judged unacceptable, since the tar-drop is the cheapest way to halve it.
|
||||
7. **R-161** — **de-ranked 2026-08-02, ruled and shipped at reduced scope.** The gate now has one
|
||||
mandated entry point (`catalog_gates.py`), which is the shape that actually gets run here. What is
|
||||
left is the automatic half, and that is sufficient while **one** person touches templates — so it
|
||||
ranks low by design, not by neglect. Revisit when a second does.
|
||||
8. **R-162** — `WATCHING` only. A limitation that fails closed; revisit if a non-overlay driver ships.
|
||||
|
||||
**R-159 and R-160 are SHIPPED** and are not ranked; they are filed to record the class, and R-159's
|
||||
class (an image `VOLUME` at an unmounted path) is still live — `immich-server` has one today.
|
||||
|
||||
<!-- RESTORED 2026-08-22: these rows are NOT fully closed and were moved to CLOSED-ITEMS.md in
|
||||
error by this session's compressor, whose status regex matched the word CLOSED inside
|
||||
'PARTLY CLOSED' and the word FIXED inside 'OPEN - NOT FIXED'. They are restored VERBATIM
|
||||
from commit fddfe00ce268, not from the compressed form: an open row keeps its detail. -->
|
||||
|
||||
<!-- ── MIGRATED FROM ROADMAP.md, 2026-08-22 (R-369 ruling: ONE register) ──────────────────
|
||||
16 rows moved: every ROADMAP row that asserted something checkable about the shipped
|
||||
product, plus one owed operator decision. Ideas and proposals stayed in ROADMAP, which
|
||||
is their home. Each row below keeps its ORIGINAL identifier and filing date — the age is
|
||||
the point. The roadmap keeps its copy as history, marked moved, with a pointer here. -->
|
||||
Reference in New Issue
Block a user