diff --git a/CONTEXT.md b/CONTEXT.md index 6dbb9e7f..e794dfd8 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -1608,6 +1608,8 @@ Two corollaries, both of which cost something to learn: ## Standing rulings +### S-39 + **S-39 — "WE DO NOT KNOW" IS NEVER DRAWN AS "FINE", AND THE CODEBASE HAS ONE WAY OF SAYING IT (2026-08-08, R-259 / R-258; controller v0.210.0).** @@ -1647,6 +1649,8 @@ block drifts and a drifted copy passes while the page it claims to cover has cha fixture-is-not-the-wire mistake, hit twice already (R-262's golden, and the OOB fixture in hub v0.99.0). +### S-38 + **S-38 — A FACT ONE SIDE EMITS AND THE OTHER CANNOT RECEIVE IS A DEFECT, AND A CHECK NOW SAYS SO (2026-08-08, G-1 / R-260 / R-247).** @@ -1706,6 +1710,8 @@ Longhorn's target is DooPlex itself over NFS, and the only outbound-looking cron **offline operator-held copy**, so disk loss is recoverable — it is not system-held escrow, which is the only residual. Full survey: `audits/RECON-dooplex-backup-2026-08-06.md`. +### S-37 (2026-08-06) + **S-37 — A CLAIM IN AN INSTRUCTION FILE IS CHECKED, NOT TRUSTED (2026-08-06, R-229/R-230 close-out).** 1. **The workspace-root `CLAUDE.md` is a SYMLINK** to `documentation/runbooks/workspace-CLAUDE.md`. @@ -1733,6 +1739,8 @@ the only residual. Full survey: `audits/RECON-dooplex-backup-2026-08-06.md`. which machine may be destroyed, so a wrong statement about what a machine holds is how a drill lands somewhere it should not. +### S-36 + **S-36 — THE AUTO-MEMORY STORE IS BACKED UP, NOT COMMITTED; AND THE WORKSPACE IS INSTALLABLE (2026-08-06, R-229 part 2).** @@ -1763,6 +1771,8 @@ the only residual. Full survey: `audits/RECON-dooplex-backup-2026-08-06.md`. that should have matched them produced no hook line at all. Verify new rules from a fresh session (`claude -p`), never from the frontmatter. +### S-35 (2026-08-06) + **S-35 — INSTRUCTION FILES ARE A SHORT CORE PLUS PATH-SCOPED RULES (2026-08-06, R-229).** Decided while rightsizing the four `CLAUDE.md` files. The mechanisms were verified before being @@ -1797,6 +1807,8 @@ Enforced by `scripts/instructions_gate.py`, registered in `controller_gates.py` `agent_gates.py`. **`felhom.eu/CLAUDE.md` is knowingly still over the ceiling (227 effective lines)** and is therefore not yet gated — closing it needs the restructure R-229 defers. +### S-34 (2026-08-05) + **S-34 — UNLOCKING AND RESTORING ARE SEPARATE. The recovery screen shipped (2026-08-05, controller v0.200.0, R-193 CLOSED). Read with S-32 and S-33; together they close the whole customer journey up to the listing.** @@ -1839,6 +1851,8 @@ the listing.** orphaned history (`RECON-offsite-dr-chain-2026-08-04.md` §12.3) and demo-hp's is operator-held. The live run exercised handler → agent → hub fetch → age KDF and stopped at the unseal. +### S-33 + **S-33 — THE BOX DECLARES, THE HUB ANSWERS. R-204 item 4 / R-193's credential half closed (2026-08-05, controller v0.199.0 + hub v0.96.0). Read with S-32; together they close all four of the drill's manual interventions.** @@ -1885,6 +1899,8 @@ off-site backups — declaring on it would make the whole fleet ask for credenti double-issue, and that one can only mint. - **NEVER widen this to the ceremony. Credential automatic, key customer-present.** +### S-32 (2026-08-05) + **S-32 — THREE OF S-31's FOUR MANUAL STEPS ARE CLOSED (2026-08-05, R-204 items 1–3 / R-196). controller v0.198.0 + hub v0.95.0. Read this BEFORE S-31 — it supersedes S-31's steps 2–5.** @@ -1924,6 +1940,8 @@ controller v0.198.0 + hub v0.95.0. Read this BEFORE S-31 — it supersedes S-31' size-gate reveal on demo-hp 9201, using `privatebin` so the drill's `calibre-web` scratch was not touched. **The Part 2 change was NOT fired live on demo-hp** — a Re-issue there was out of scope. +### S-31 + **S-31 — THE DRILL PASSED: a customer's file survives a machine rebuild and comes back byte-identical. The capability is proven; the customer JOURNEY is four undocumented manual steps (2026-08-04, R-201/R-204).** @@ -1960,6 +1978,8 @@ endpoint. **The local escape hatch does not work unaided:** `--print-reset-code` - **Never run a ceremony while a recovery is in flight** — it supersedes the identity blob. Under v0.93.0 the old blob is retained, but nothing serves a superseded blob back (R-199). +### S-30 (2026-08-04) + **S-30 — the R-201 drill was PREPARED and HALTED BEFORE THE WIPE (2026-08-04). Nothing was wiped.** It stopped at step 4 because the sentinel file was **not in the off-site snapshot** while the run @@ -1981,6 +2001,8 @@ Wiping would have destroyed the only copy of the sentinel and proven nothing. *Still not established, unchanged:* **no file has ever been restored from an off-site backup after a wipe**, and Part 0's install path (controller v0.196.0) has never run against a live recovery. +### S-29 + **S-29 — a box may fetch its OWN sealed recovery blob with its OWN credential; the operator-driven DR path is a separate thing and stays gated (2026-08-04, R-199; hub v0.94.0 + agent v0.125.0 + controller v0.195.0).** @@ -2016,6 +2038,8 @@ serves that ONE object to its authenticated owner. - **The chain today: links 1–8 walked, 9–11 not.** The KEY comes back. Nothing installs it, reopens a repository with it, or restores a file — R-200's remaining half and R-201. +### S-28 + **S-28 — the escrow retention now covers the OFF-SITE data key, and customer-present recovery is the accepted design, which makes that retention load-bearing (2026-08-04, R-198/R-197; hub v0.93.0).** @@ -2048,6 +2072,8 @@ the retained identity blob. A session that touches escrow custody is touching th both hashes are known and differ) is the evidential signal that a box's off-site data key moved. It carries **no hash value**. `MarkEscrowStale` is **precautionary**, not evidential — see S-26(a). +### S-27 + **S-27 — a customer with NO machine ever bound is UNKNOWN, silently; one that was bound and went quiet still alarms (2026-08-04, R-195; hub v0.92.0).** Operator ruling, implemented as `store.HasEverBoundHost` (live `hosts` row OR `host_deletions` tombstone) consulted once at the top of @@ -2064,6 +2090,8 @@ customer it would most obviously cover.** `david` (created 2026-08-01, no machin 2026-07-15 — does not, because its 482 old reports make it `down`. Generalise it: **a "skip the dead" guard built on evidence of life cannot see something that was never alive.** +### S-26 + **S-26 — the one-shot secret is the recoverable one; the irreplaceable one is minted fresh on every guest rebuild (2026-08-04, R-193 spike — `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`). No code shipped for it; the decision is the operator's.** @@ -2106,6 +2134,8 @@ the ceremony. Retaining and serving it back needs no new seam. Its price is one the data key at rest on the Proxmox host. **That trade is the operator's to make and the spike does not make it.** +### S-24 (2026-08-04) + **S-24 — offsite retention is ep0's, and the box asks for none (2026-08-04, R-191; installer 1.25.0).** R-89 moved offsite pruning server-side and box tokens stay write-only. The 2026-07-26 "two weeks" ruling was not reversed — **where it is ENFORCED moved, and the installer's `keep_last: 2` did not @@ -2116,6 +2146,8 @@ both namespaces have a prune job at 03:30 keep-last 2 that has run daily since 2 all OK). **If that ever stops, `keep_last: 0` is unbounded growth** — check ep0's prune jobs before assuming the offsite tier is retained. +### S-25 (2026-08-04) + **S-25 — a lost storage grant repairs itself, and the repair is REPORTED (2026-08-04, R-190; agent v0.124.1).** On a missing grant the agent runs the existing root wrapper (`felhom-backup-target-apply grant `, already sudoers-permitted for any id) and re-reads once — the pbsdr R-22 shape. Bounded @@ -2133,6 +2165,8 @@ hub's existing ok→degraded→ok edge is the channel. *Caveat measured live:* **PVE caches permissions** (~40 s and ~16 min observed), so detection lags the loss and a single permission read is a lagging indicator → R-194. +### S-23 + **S-23 — the host (on-box) whole-guest tier is restore-PROVEN, unattended, on both demo boxes (2026-08-04). Scope: those two boxes, not the fleet.** @@ -2154,6 +2188,8 @@ carrying a host-tier entry for the first time. *The asymmetry worth remembering:* a host-tier restore is **83–109 s**; an offsite one is **300–540 s**. The tier that matters for an ordinary recovery is also the cheapest to prove. +### S-21 + **S-21 — an empty listing cannot distinguish FORBIDDEN from NEWBORN, so the box asks the permission question directly (2026-08-03, R-185; agent v0.123.0 + installer 1.24.0).** @@ -2182,6 +2218,8 @@ does not page — turning an ordinary documented setup into an alert is how a si an operator archives unread. It never consults content, so it cannot alarm on a newborn tier by construction, and it never reports ok when it could not ask. +### S-22 + **S-22 — the installer's Scenario-F arm must finish the job, not just leave the definition alone (2026-08-03, R-185).** `configure_backup_target` has two arms. Case A creates the storage and grants in the same breath. The reuse arm — *"the target already exists"* — returned **without granting**, and @@ -2194,6 +2232,8 @@ retargeting the box. `$BACKUP_TARGET_ID` stays OUT of `PVE_STORAGES`: that list before the target is resolved, and `--acl-storages` entries are preflight-checked for existence. A gate asserts every arm that resolves the target also grants on it. +### S-19 + **S-19 — a restore-test PROOF is durable and reportable; a FAILURE is neither, and that asymmetry is the design (2026-08-03, R-189; agent v0.122.0).** @@ -2217,6 +2257,8 @@ VMID, duration) are not re-invented — an absent duration is not a claim, a fab **Migration consequence, seen live:** a pre-R-189 record has no tier, so upgrading does not retroactively make an old proof visible to the hub; the tier's next real proof fills it in. +### S-20 + **S-20 — the release order is build → tag LOCALLY → publish → push tag, and every step protects something (2026-08-03, R-188 + R-186).** @@ -2236,6 +2278,8 @@ yields the same bytes with or without the tag; the verification command lives in `felhom-agent/CLAUDE.md`. Both build paths (`release-agent.sh` and `publish-agent.sh`'s fallback) use identical flags: they differed by `CGO_ENABLED=0` and produced binaries 74 KB apart for one version. +### S-17 + **S-17 — restore-testing is PER ARCHIVE GENERATION, and the hub's staleness window follows each tier's own rhythm (2026-08-03, R-86; agent v0.121.0 + hub v0.91.0).** @@ -2268,6 +2312,8 @@ observed archive interval × 4 generations, floored at 7 days, capped at 12 days `offsiteBackupStaleAfter` 8 d — the backup-freshness checker's own thresholds) when history is too short to observe one. Shipping Part 1 alone would have produced a nightly false alarm. +### S-18 + **S-18 — `ep0` is Tier 2, PROTECTED (operator ruling, 2026-08-03).** D-d named two protected machines and did not name ep0 either way; `runbooks/target-selection.md` carried the question in writing for two days. The ruling **extends D-d's protected list to three machines**: DooPlex, Peti's cluster, @@ -2275,6 +2321,8 @@ two days. The ruling **extends D-d's protected list to three machines**: DooPlex tunnel config or nftables rules was already forbidden by what it would destroy, and the ordinary off-site READ a restore-test performs remains permitted. +### S-13 (2026-08-03) + **S-13 — the `mp1` merge landed, and the variant was chosen on measurement (2026-08-03, R-165 / D-a).** The appliance's two data volumes are one. **Variant V-c**: the volume mounts at the NEUTRAL path `/var/lib/felhom`, and both `/var/lib/docker` and `/mnt/sys_drive` are binds of subdirectories of it. @@ -2297,6 +2345,8 @@ so pruning could only mean deleting a **different** app's only local recovery un boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed shortly. So R-176's in-place migration rehearsal is **withdrawn**, not deferred. +### S-14 (2026-08-03) + **S-14 — prove first, then vouch (2026-08-03) — SPENT, and the ordering did not survive contact.** The rule was: golden **0.192.0** stays UNVOUCHED until a box has been proven from it, because vouching is what makes a fresh install pick a golden up. **In the event the golden was vouched at 07:23:26 CEST on @@ -2310,6 +2360,8 @@ and reminders do not hold. If prove-then-vouch is to be a rule it needs the shap a refusal at `handleSetArtifacts`, the sole path to `SetArtifactManifest`, which runs without anyone choosing to run it. +### S-15 (2026-08-03) + **S-15 — the merged layout is proven live, by two different supply paths (2026-08-03, R-178).** Both demo boxes were wiped and reinstalled from golden 0.192.0 and taken through claim → deploy → back up → **restore**. **demo-hp** was installed with `--golden ` (the layout proof) and @@ -2323,6 +2375,8 @@ the floor guarded `captureAllRecoveryUnits` and not `runVolumeDumps`, the leg th and its refusal's "the previous unit is untouched" was measured false. **R-181 CLOSED the same day (controller v0.193.0 + v0.193.1), so R-165 is now PROVEN-LIVE in both halves** — see S-14. +### S-14 (2026-08-03) + **S-14 — the reserve is a per-app, per-run ADMISSION decision, not a capture check (2026-08-03, R-181; controller v0.193.0 + v0.193.1).** B2 as first shipped was consulted in exactly one place — `captureAllRecoveryUnits`, a few KB — while `RunDBDumps`' database leg and `runVolumeDumps` wrote the @@ -2355,6 +2409,8 @@ not provide, and the fourth of those found on live hardware rather than by revie no run scope, so a refused app re-alerts on every status refresh (measured: a second identical alert pair 13 s after the run's). Pre-existing in v0.192.0; R-181 changed neither caller. +### S-15 (2026-08-03) + **S-15 — publishing is an act, not a side-effect of pushing (2026-08-03, R-110 + R-115 + R-183).** Two rulings, one shape: something became live because someone pushed, not because anyone decided. @@ -2381,6 +2437,8 @@ Two rulings, one shape: something became live because someone pushed, not becaus that bumps a version, before publishing — and a gate that fails on the normal path is one people learn to ignore. +### S-16 + **S-16 — a backup run NOTIFIES ONCE and RECORDS ALWAYS, and those are different things (2026-08-03, R-182; controller v0.194.0 + hub v0.90.0/.1).** Measured: nine per-app capture failures reached the hub, two were mailed, seven were dropped by a cooldown whose key carries no app @@ -2408,6 +2466,8 @@ identifier — *before* `LogNotification`, so they left no row anywhere. the 4 GiB swapfile survived. The 40 GB local disk is UNCHANGED** — a CPU/RAM resize only, so no disk figure in any runbook needed correcting. That closed **R-90** and unblocked **R-86**. +### S-11 (2026-08-02) + **S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).** Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the CUSTOMER's** (they can free space, delete files, add a drive) and **a per-app backup capture failure @@ -2430,6 +2490,8 @@ label and the free-space figures — the same reason `offbox_enlarge_blocked` an have none. `notify.IsOperatorOnly` was added so ONE test pins both registers; checked separately, an allowlisted-but-not-operator-only type is invisible. +### S-12 (2026-08-02) + **S-12 — the monitoring landed BEFORE the merge, not with it (2026-08-02).** D-a's condition (2) says R-167 ships in the same step as the `mp1`→`mp0` merge and never after, because the merge removes a wall that currently fails safely. **This session landed it FIRST**, which @@ -2442,6 +2504,8 @@ load-bearing findings for anyone picking that up: **"the layout" is not one thin is also a BULKHEAD**, not only a ceiling — today an overflow cannot reach `/var/lib/docker`, and after the merge it can. +### S-8 (2026-08-02) + **S-8 — CI detects; it does not block, and that is structural (2026-08-02, R-168).** A Gitea Actions runner in `gitea-system` re-runs every repo's gate entry point on every push, independent of who pushed and of what they typed. It **cannot refuse a push**: every felhom repo @@ -2452,6 +2516,8 @@ skipped or was never armed. Making CI blocking needs branch protection plus a PR changes how the operator works and is **their** call → R-169. Do not "fix" this by adding branch protection. +### S-9 (2026-08-02) + **S-9 — a detector that tells no one is not finished (2026-08-02, R-168 probe P5).** Probe P5 measured that a failed run produces **no mail, no notification row and no log line** from Gitea. So the workflow sends its own alarm on the project's existing Resend path and **prints the @@ -2462,6 +2528,8 @@ the runner image has **no `curl`** (deliberately — python3 and git only, so us `api.resend.com` sits behind **Cloudflare, which 403s the default `Python-urllib` User-Agent with error 1010** — a failure that looks exactly like an auth failure and is not one. +### S-10 (2026-08-02) + **S-10 — the runner is unprivileged, and the reason is the host (2026-08-02).** The usual `act_runner` recipe pairs it with a `docker:dind` sidecar and `privileged: true`. Rejected: DooPlex is **Tier 2** and *is* the recovery chain — Gitea, the hub, the registry, PBS and @@ -2471,6 +2539,8 @@ job sees exactly the runner image's tools**, which is why `python3` had to be ba stock `act_runner` carries git but not python3). If a future job genuinely needs Docker, that is a conversation, not a patch. +### S-11 + **S-11 — CI reproduces the workspace's sibling layout, because two entry points depend on it (2026-08-02).** `controller_gates.py` and `agent_gates.py` invoke the shared `reuse_refs_check.py` that lives in the `felhom.eu` clone next door and is deliberately never copied, and both repos' @@ -2478,6 +2548,8 @@ that lives in the `felhom.eu` clone next door and is deliberately never copied, sibling; without it the gate fails **closed** — correctly, but for the wrong reason. Verified that CI and the local hook then agree exactly (controller 126 exact / 6 suffix / 1 cross-repo). +### S-6 (2026-08-02) + **S-6 — the hub renders no host-install version, and the gate pins its absence (2026-08-02, R-94).** The Setup tab's *"host-install 1.19.0"* label is **deleted, not derived**. Deriving it is not achievable honestly: the Option-1 command downloads `felhom-host-install.sh` from the website **at @@ -2492,6 +2564,8 @@ tautological `render_test.go` assertion (`html contains hostInstallVersion`, whe put it there) **passed at `9.9.9`** — an assertion that compares a value to itself tests the plumbing, never the claim. +### S-7 + **S-7 — gates run from ONE entry point per repo, and `reuse_refs_check` was fixed rather than the convention it polices (2026-08-02, R-29).** Two rulings from the same census. @@ -2519,6 +2593,8 @@ exact → suffix → ambiguous → sibling repo → FAIL, **prints every non-exa resolution attempted on a failure. It stays in **one** place and is invoked across the workspace — never copied, which would recreate the drift it detects. +### S-1 (2026-07-26) + **S-1 — N.5 gains a third leg: architecture docs are same-session coupled (2026-07-26, R-81).** Any task that changes an **architectural contract** — tiers, targets, cadences, trust boundaries — updates the owning `documentation/architecture/*.md` in the **same session**, under exactly the same @@ -2527,6 +2603,8 @@ coupling rule that already binds the capability map and the ROADMAP. Origin: R-8 (single target, single cadence), while being cited as authoritative. A stale architecture doc is worse than a missing one, because it is trusted. +### S-2 (2026-07-26) + **S-2 — architecture docs carry an honest status header (2026-07-26, R-81).** Every `documentation/architecture/*.md` opens with the version it was **verified against** and the date. A doc more than a few trains behind its subject is marked **STALE** *in that header*, so a @@ -2535,6 +2613,8 @@ reader meets the warning before the content, not after acting on it. Origin: versions stale (live v0.173.0), and cited as authoritative throughout the R-80 diagnostic. Ratifying or retiring it is → **R-83**. +### S-3 + **S-3 — the recovery model: six decisions, 2026-07-28.** Taken in an architecture discussion and expressed in the `07-backup-architecture.md` full rewrite (which replaces the 2026-07-14 DRAFT entirely — that doc was verified against controller v0.132.0, **51 versions stale**, while being @@ -2569,6 +2649,8 @@ cited as authoritative). They are **decisions, not observations**; the rewrite l compromised hub yields blobs nobody can open, *provided the operator's key is never stored in the hub*. That proviso is why escrow custody is an open decision (`07` §11-A). +### S-4 (2026-07-31) + **S-4 — the hub session password alone now unlocks console root on every managed box (2026-07-31, hub v0.84.0).** Retrieving a host's vaulted break-glass `root@pam` credential previously required the **global operator API key**, a secret distinct from the hub login and kept out-of-band. The `Console access` card on the @@ -2589,6 +2671,8 @@ Five decisions were deliberately **left open for the operator** and are recorded stated**) · Hetzner as a single failure domain · and `local` vzdump sharing a physical device with the guest it backs up. Gaps minted the same session: **R-102 … R-108**. +### S-13 + **S-13 — boot recovery finished, and the lesson is about the DIAGNOSIS ORDER (controller v0.190.0, 2026-08-02, R-157 A · R-170 · R-171).** @@ -2639,6 +2723,8 @@ you are watching the cache settle, not the system.** six), window settle times 10/40/10/10/15/15 s — routinely 2–8× the old fixed 5 s. The sharpest evidence is a same-app before/after on one box: missed at 18:08:35, recovered at 18:18:50. +### S-12 + **S-12 — D-b is BUILT (controller v0.189.0, 2026-08-02, R-166).** The desired/in-flight/observed split now exists; the S-1 contract lives in `architecture/02-controller-module-map.md` §0a. @@ -2682,6 +2768,8 @@ behaviour changes under one live validation is one too many. **IMPLEMENTED, not PROVEN-LIVE** — unit-proven and red-proofed, but nobody killed the controller mid-backup on real hardware; the capability map says so rather than rounding it up. +### S-5 + **S-5 — four operator decisions taken in discussion on 2026-08-02, recorded before anything is built.** They existed only in conversation, which is the condition the standing rules were written against. Labels are the ones used in the discussion (**D-a … D-d**) and are deliberately kept diff --git a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md index 63969537..d29f7d35 100644 --- a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md @@ -35,3 +35,4 @@ no reboot. | R-31 | **race half fixed on felhom.eu main** (hub, unreleased): a second off-site Save while one provisions → 409, nothing saved or created; red-proof `hub/R-31-red.txt`. LEFT: the async save + status card | 35 | (this batch) | | R-30 | needs a design — wait-channel presence is indirect (~240 s + grace, unmeasured) and gates host-delete before RESET | 15 | — | | R-336 | no hub half — the pollers run on the boxes (pvestatd, proxmox-backup-client); a design question | 5 | — | +| R-377 | **closed** — 44 `### S-n` sub-headings in CONTEXT.md, no ruling text edited (88 lines added, 0 removed) | 15 | (this batch) | diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 9c7b7d07..b3b88f79 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,14 @@ --- +## 2026-10-06 (night) — the second burn-down night + +The full text of every row below: `git show e866a66b56:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-377** | **`CONTEXT.md`'s standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them.** (P4) | CLOSED 2026-10-06 night — DONE AS THE ROW SAID | 44 `### S-n (date)` sub-headings added above the rulings in `CONTEXT.md` § Standing rulings; no ruling's text edited, reordered or compressed (the diff is 88 added lines, 0 removed). Five ids are used twice in the log (S-11..S-15) — kept as written, the log is never edited. | + ## 2026-10-06 (evening) — the design build The full text of every row below: `git show 8db4425d98:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 8e82b5de..3b97cd10 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -278,7 +278,7 @@ stopping line that lies. | **R-89** | Business & legal | P4 | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 23 rows (P3 2, P4 21) +## Process & tooling — 22 rows (P3 2, P4 20) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -291,7 +291,6 @@ stopping line that lies. | **R-288** | Process & tooling | P4 | **The capability map is too long to be read, and that is why it stops being true.** `architecture/00-capability-map.md` is **134 642 bytes / 19 456 words across 99 table rows in only 159 lines** — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is **3 024 words**, the offsite-password-recovery row **1 087**, the unattended-restore-proof row **971**, the app/guest-network-failure row **904**. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads **2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5`** (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. **This is the project's memory, so restructuring it is surgery and wants daylight** — filed, deliberately not attempted in the 2026-08-09 session. **What the shape should probably be:** one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from | **READY (M) — NEW 2026-08-09** | — | Do not fold this into another session; it needs its own. **SECOND CONCRETE COST, 2026-08-10 — and it is a different failure mode from the first.** An *execution record* — 33 package deletions on the operator's rule — was undiscoverable for two days because it lives **inside the row about the Configuration page being slow**. Two sessions searched for it: one reported "no register row records a package prune", the other exhausted the Gitea logs, the activity feed and the schema before concluding it might be unestablishable. It was in `OPEN-ITEMS.md` the whole time. **The first cost (2026-08-09) was two records that looked contradictory and were not; this one is a record that could not be found at all.** Illegibility now has two measured costs and they are different in kind: prose rows make claims ambiguous, and rows-about-other-things make facts unfindable. **The rule this earns is in CONTEXT.md: a record that lives inside a row about something else has not been recorded** | Viktor | | **R-290** | Process & tooling | P4 | **Most capability-map rows that back a green dot cite no evidence document at all — measured, 20 of 28 probed.** The page's *Walked* means *"done end to end on real hardware, evidence on file"*. Extracting the evidence column for the 28 rows behind the page's claims found a `tests/` or `audits/` path in **8**; the other 20 carry prose only. **Consequence, applied this session:** of 32 claims the page drew as Walked, **12 were downgraded to Built** because no walk document exists for them — `install.installer-by-tag`, `use.lifecycle`, `drives.enrol`, `drives.migrate`, `backup.tier1`, `backup.whole-machine`, `backup.restore-proof`, `fault.selfheal`, `fault.operator-email`, `fail.drive-filling`, `fail.lost-recovery-code`, `fail.hub-down`. **This is not a claim that those twelve are false** — several are near-certainly fine — it is a claim that nothing on file distinguishes them from an opinion, which is exactly what the status word promises. **The gate now enforces it going forward:** `scripts/check_stands.py` fails on `status: walked` with no `evidence:` source. **What is owed:** either a walk document per row, or an honest demotion in the map itself (the map is the source; the dataset only follows it) | **READY (M) — NEW 2026-08-09** | R-288 | The dataset was corrected; **the capability map itself still says PROVEN-LIVE for these rows** and is the thing to fix | Viktor | | **R-327** | Process & tooling | P4 | **The standing picture still describes a defect that has been fixed twice over.** Found by the first run of `unproven.py` (R-326), which is the argument for having built it. `where-felhom-stands.yaml`'s `claim.code-naming` is `status: partial` and its title reads *"The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show"* — **both halves of which are now false.** The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new `reenroll` mail kind), and the third near-homograph on 2026-08-13 (R-323). **NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule:** *"A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it."* Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met | **READY (S) — NEW 2026-08-13, RANK 4** | R-295, R-323, R-326 | Decide the capability-map status for the naming arc, then let the dataset follow it. **Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet**, so `walked` would be an over-claim; `built` is probably right, and the title needs rewriting either way because it describes a defect rather than a capability | operator + CC | -| **R-377** | Process & tooling | P4 | **`CONTEXT.md`'s standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them.** Measured 2026-08-22: `CONTEXT.md` is 217 KB, of which **187,913 bytes — 86% — is a single `## Standing rulings` section** carrying 39 `S-` ids and 153 bullets under **one** heading. **This session deliberately did NOT compress or split it**, and the reason is the ruling itself: `PROMPT-TEMPLATE.md` §3.4 and this file's own contract say the decision log is *dated, never edited afterwards*, and it is the only place that answers *"has this been proposed before, and why did we say no?"* — compressing it destroys exactly that. With no per-ruling delimiter, any mechanical split risks cutting a live ruling from its reason, which is the failure this whole arc is correcting. **So the problem is navigational, not volumetric, and the fix is structural: give each ruling a sub-heading with its `S-` id and date.** Then it can be linked, cited and found without a single word being edited. **13 mentions of SUPERSEDED already sit inside that blob** and cannot be separated from live text safely today. | **OPEN — LOW** | R-369 | Add per-ruling sub-headings only. Do not compress, do not reorder, do not edit any ruling's text. | CC | | **R-392** | Process & tooling | P4 | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC | | **R-394** | Process & tooling | P4 | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | | **R-421** | Process & tooling | P4 | **THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident.** R-410 (a `mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). **The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked.** The 2026-09-01 decoy sweep read all 29 scripts and fooled **16**. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. **The shapes, so the next one is cheap to recognise:** (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. **The single largest cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. `decoy_coverage_gate.py` now refuses a new gate that ships without a decoy. | **OPEN — the class row; it stays open as the place the next instance is recorded** | — | — | CC |