Decisions 71-73 recorded before the work (ep0 copy keeps 8 weekly; tester-1 keys removed via registrar; token not rotated now) + R-831 token row, R-832 roadmap P4
gates / gates (push) Successful in 30s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 06:57:19 +02:00
parent 710a2505f9
commit 697c2a7b10
4 changed files with 16 additions and 2 deletions
+4 -2
View File
@@ -191,7 +191,7 @@ stopping line that lies.
| **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC |
| **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator |
## Backup & restore — 58 rows (P2 12, P3 26, P4 20)
## Backup & restore — 59 rows (P2 12, P3 26, P4 21)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -253,6 +253,7 @@ stopping line that lies.
| **R-691** | Backup & restore | P4 | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). **-- 2026-09-28 afternoon: LIVE-PROVEN on demo-hp 9201 (controller 0.278.0), endpoint level.** A throwaway nextcloud, seeded through its own front door, joined the off-site copy; the off-site run-now pushed snapshot `6cb379a8`. (a) The full off-site restore (prepare → download → reconstitute): 3 volumes + the database replayed, the seed read back, a marker user written after the snapshot read ABSENT, same versions. (b) Remove keeping the data + deleting the local copies → reinstall: the choice and the kept list both named "távoli mentés, 2026-09-28 15:40" / "the off-site copy, 2026-09-28 15:40"; "use my kept data" downloaded the unit alone and loaded 3/3 volumes + 1/1 database in 55 s; the seed and the kept files read back. App removed with its data; demo-hp's app list equals the list before. `audits/kept-offsite-2026-09-28/E/`. | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the second-drive (Tier 2) path of Use/Load is not proven live — 9202 has one drive) — **CLOSED — controller v0.277.0, live 2026-09-28** | — | — | CC |
| **R-815** | Backup & restore | P4 | First-ever **GC** on `felhom-offsite` (armed today 13:11 UTC, never run) | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-815. WATCHING — no completion record found; schedule `sun 04:30` still present 2026-09-30.) — WATCHING | schedule | **Sun 2026-08-02 04:30 UTC** — confirm it completes | CC |
| **R-816** | Backup & restore | P4 | **No off-site failure class has ever been seen live.** F-DIAG (controller v0.182.0, 2026-07-28) split off-site failures into six causes — quota, orphaned, no_repo, no_units, transport, unknown — each with its own Hungarian message. None of the six has been exercised by a real failure on a box; the recovery inventory records it only as a known limit (`documentation/architecture/_recovery-inventory-2026-07-28.md:955`). Filed 2026-10-03 from the F-DIAG row's residue when that row moved to `CLOSED-ITEMS.md`. | **READY — filed 2026-10-03 (triage); owner: CC.** Exercise each class once on a scratch guest (a full quota, a missing repository, a blocked transport) and read the message the household sees. | — | — | CC |
| **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator |
## Storage & devices — 12 rows (P3 7, P4 5)
@@ -271,7 +272,7 @@ stopping line that lies.
| **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks/<n>/deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC |
| **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC |
## Security & access — 30 rows (P2 3, P3 24, P4 3)
## Security & access — 31 rows (P2 3, P3 25, P4 3)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -302,6 +303,7 @@ stopping line that lies.
| **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC |
| **R-783** | Security & access | P3 | **[P3-LOW] SparkyFitness: three wrong sign-ins by anyone shut EVERY visitor out of sign-in for ~10 s — a stranger retrying every 10 s keeps the household out.** MEASURED 2026-10-01 on 9202 through the simulated tunnel (`audits/visitors-2026-10-01/C/box/sparky-box.txt`): better-auth's sign-in limit (3 per 10 s) is keyed on one address — it reads the LEFTMOST X-Forwarded-For, whose chain the R-753 router reset removes, so its frontend nginx hands it traefik's address; the household from another address got 429 at 3 s and 10 s, in at 16 s. Not forgeable (a rotating forged address did not escape). better-auth's header setting is not exposed as an env by SparkyFitness. **Needs:** accept (10 s), or an upstream setting for better-auth's `ipAddressHeaders` + a right-walking reader. | **OPEN — rank P3-LOW; owner: CC** | — | — | CC |
| **R-826** | Security & access | P3 | **Storage Box sub-accounts kept every earlier box's key, unpinned — each a route to delete history.** The registrar's first install dropped 4 stale unpinned lines on demo-felhom's sub-account and 5 on demo-hp's (every reinstall had added a key; nothing removed one). tester-1's sub-account still holds 3 (its box is gone), so the daily key check raises `offsite_key_unlocked` for it every day until they go. Fix: the next tester-1 box's registration removes them, or the operator lets CC rewrite that file through the registrar. | **READY — tester-1 alarm stands until then** | — | — | CC |
| **R-831** | Security & access | P3 | **The Hetzner storage API token (`HETZNER_TOKEN`, the storage project's token in `Secret/storagebox`) was printed into the 2026-10-03 session transcript** — CC read the gitignored `manifests/storagebox.secret.yaml` and its redaction pattern missed the quoted value. It can create, reset and delete Storage Box sub-accounts. Not rotated by the operator's choice (decision 73). **Rotation, whenever chosen (3 steps):** create a new token in the storage project in the Hetzner console → patch `Secret/storagebox` key `HETZNER_TOKEN` in `felhom-system` and `kubectl rollout restart deployment/hub` → delete the old token in the console. Rule for sessions: never print a file that holds secrets — read the one field needed. | **WAITING-ON-OPERATOR — rotation is his call** | — | rotate when chosen | operator |
| **R-134** | Security & access | P4 | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC |
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |