R-377 closed: CONTEXT.md standing rulings get S-id sub-headings (no ruling text edited)
gates / gates (push) Successful in 2m44s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 21:08:30 +02:00
parent e866a66b56
commit 93b942f136
4 changed files with 98 additions and 2 deletions
+88
View File
@@ -1608,6 +1608,8 @@ Two corollaries, both of which cost something to learn:
## Standing rulings
### S-39
**S-39 — "WE DO NOT KNOW" IS NEVER DRAWN AS "FINE", AND THE CODEBASE HAS ONE WAY OF SAYING IT
(2026-08-08, R-259 / R-258; controller v0.210.0).**
@@ -1647,6 +1649,8 @@ block drifts and a drifted copy passes while the page it claims to cover has cha
fixture-is-not-the-wire mistake, hit twice already (R-262's golden, and the OOB fixture in hub
v0.99.0).
### S-38
**S-38 — A FACT ONE SIDE EMITS AND THE OTHER CANNOT RECEIVE IS A DEFECT, AND A CHECK NOW SAYS SO
(2026-08-08, G-1 / R-260 / R-247).**
@@ -1706,6 +1710,8 @@ Longhorn's target is DooPlex itself over NFS, and the only outbound-looking cron
**offline operator-held copy**, so disk loss is recoverable — it is not system-held escrow, which is
the only residual. Full survey: `audits/RECON-dooplex-backup-2026-08-06.md`.
### S-37 (2026-08-06)
**S-37 — A CLAIM IN AN INSTRUCTION FILE IS CHECKED, NOT TRUSTED (2026-08-06, R-229/R-230 close-out).**
1. **The workspace-root `CLAUDE.md` is a SYMLINK** to `documentation/runbooks/workspace-CLAUDE.md`.
@@ -1733,6 +1739,8 @@ the only residual. Full survey: `audits/RECON-dooplex-backup-2026-08-06.md`.
which machine may be destroyed, so a wrong statement about what a machine holds is how a drill
lands somewhere it should not.
### S-36
**S-36 — THE AUTO-MEMORY STORE IS BACKED UP, NOT COMMITTED; AND THE WORKSPACE IS INSTALLABLE
(2026-08-06, R-229 part 2).**
@@ -1763,6 +1771,8 @@ the only residual. Full survey: `audits/RECON-dooplex-backup-2026-08-06.md`.
that should have matched them produced no hook line at all. Verify new rules from a fresh session
(`claude -p`), never from the frontmatter.
### S-35 (2026-08-06)
**S-35 — INSTRUCTION FILES ARE A SHORT CORE PLUS PATH-SCOPED RULES (2026-08-06, R-229).**
Decided while rightsizing the four `CLAUDE.md` files. The mechanisms were verified before being
@@ -1797,6 +1807,8 @@ Enforced by `scripts/instructions_gate.py`, registered in `controller_gates.py`
`agent_gates.py`. **`felhom.eu/CLAUDE.md` is knowingly still over the ceiling (227 effective lines)**
and is therefore not yet gated — closing it needs the restructure R-229 defers.
### S-34 (2026-08-05)
**S-34 — UNLOCKING AND RESTORING ARE SEPARATE. The recovery screen shipped (2026-08-05, controller
v0.200.0, R-193 CLOSED). Read with S-32 and S-33; together they close the whole customer journey up to
the listing.**
@@ -1839,6 +1851,8 @@ the listing.**
orphaned history (`RECON-offsite-dr-chain-2026-08-04.md` §12.3) and demo-hp's is operator-held. The
live run exercised handler → agent → hub fetch → age KDF and stopped at the unseal.
### S-33
**S-33 — THE BOX DECLARES, THE HUB ANSWERS. R-204 item 4 / R-193's credential half closed
(2026-08-05, controller v0.199.0 + hub v0.96.0). Read with S-32; together they close all four of the
drill's manual interventions.**
@@ -1885,6 +1899,8 @@ off-site backups — declaring on it would make the whole fleet ask for credenti
double-issue, and that one can only mint.
- **NEVER widen this to the ceremony. Credential automatic, key customer-present.**
### S-32 (2026-08-05)
**S-32 — THREE OF S-31's FOUR MANUAL STEPS ARE CLOSED (2026-08-05, R-204 items 1–3 / R-196).
controller v0.198.0 + hub v0.95.0. Read this BEFORE S-31 — it supersedes S-31's steps 2–5.**
@@ -1924,6 +1940,8 @@ controller v0.198.0 + hub v0.95.0. Read this BEFORE S-31 — it supersedes S-31'
size-gate reveal on demo-hp 9201, using `privatebin` so the drill's `calibre-web` scratch was not
touched. **The Part 2 change was NOT fired live on demo-hp** — a Re-issue there was out of scope.
### S-31
**S-31 — THE DRILL PASSED: a customer's file survives a machine rebuild and comes back byte-identical.
The capability is proven; the customer JOURNEY is four undocumented manual steps (2026-08-04, R-201/R-204).**
@@ -1960,6 +1978,8 @@ endpoint. **The local escape hatch does not work unaided:** `--print-reset-code`
- **Never run a ceremony while a recovery is in flight** — it supersedes the identity blob. Under
v0.93.0 the old blob is retained, but nothing serves a superseded blob back (R-199).
### S-30 (2026-08-04)
**S-30 — the R-201 drill was PREPARED and HALTED BEFORE THE WIPE (2026-08-04). Nothing was wiped.**
It stopped at step 4 because the sentinel file was **not in the off-site snapshot** while the run
@@ -1981,6 +2001,8 @@ Wiping would have destroyed the only copy of the sentinel and proven nothing.
*Still not established, unchanged:* **no file has ever been restored from an off-site backup after a
wipe**, and Part 0's install path (controller v0.196.0) has never run against a live recovery.
### S-29
**S-29 — a box may fetch its OWN sealed recovery blob with its OWN credential; the operator-driven DR
path is a separate thing and stays gated (2026-08-04, R-199; hub v0.94.0 + agent v0.125.0 + controller
v0.195.0).**
@@ -2016,6 +2038,8 @@ serves that ONE object to its authenticated owner.
- **The chain today: links 1–8 walked, 9–11 not.** The KEY comes back. Nothing installs it, reopens a
repository with it, or restores a file — R-200's remaining half and R-201.
### S-28
**S-28 — the escrow retention now covers the OFF-SITE data key, and customer-present recovery is the
accepted design, which makes that retention load-bearing (2026-08-04, R-198/R-197; hub v0.93.0).**
@@ -2048,6 +2072,8 @@ the retained identity blob. A session that touches escrow custody is touching th
both hashes are known and differ) is the evidential signal that a box's off-site data key moved. It
carries **no hash value**. `MarkEscrowStale` is **precautionary**, not evidential — see S-26(a).
### S-27
**S-27 — a customer with NO machine ever bound is UNKNOWN, silently; one that was bound and went quiet
still alarms (2026-08-04, R-195; hub v0.92.0).** Operator ruling, implemented as
`store.HasEverBoundHost` (live `hosts` row OR `host_deletions` tombstone) consulted once at the top of
@@ -2064,6 +2090,8 @@ customer it would most obviously cover.** `david` (created 2026-08-01, no machin
2026-07-15 — does not, because its 482 old reports make it `down`. Generalise it: **a "skip the dead"
guard built on evidence of life cannot see something that was never alive.**
### S-26
**S-26 — the one-shot secret is the recoverable one; the irreplaceable one is minted fresh on every
guest rebuild (2026-08-04, R-193 spike — `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`).
No code shipped for it; the decision is the operator's.**
@@ -2106,6 +2134,8 @@ the ceremony. Retaining and serving it back needs no new seam. Its price is one
the data key at rest on the Proxmox host. **That trade is the operator's to make and the spike does not
make it.**
### S-24 (2026-08-04)
**S-24 — offsite retention is ep0's, and the box asks for none (2026-08-04, R-191; installer 1.25.0).**
R-89 moved offsite pruning server-side and box tokens stay write-only. The 2026-07-26 "two weeks"
ruling was not reversed — **where it is ENFORCED moved, and the installer's `keep_last: 2` did not
@@ -2116,6 +2146,8 @@ both namespaces have a prune job at 03:30 keep-last 2 that has run daily since 2
all OK). **If that ever stops, `keep_last: 0` is unbounded growth** — check ep0's prune jobs before
assuming the offsite tier is retained.
### S-25 (2026-08-04)
**S-25 — a lost storage grant repairs itself, and the repair is REPORTED (2026-08-04, R-190; agent
v0.124.1).** On a missing grant the agent runs the existing root wrapper (`felhom-backup-target-apply
grant <id>`, already sudoers-permitted for any id) and re-reads once — the pbsdr R-22 shape. Bounded
@@ -2133,6 +2165,8 @@ hub's existing ok→degraded→ok edge is the channel.
*Caveat measured live:* **PVE caches permissions** (~40 s and ~16 min observed), so detection lags the
loss and a single permission read is a lagging indicator → R-194.
### S-23
**S-23 — the host (on-box) whole-guest tier is restore-PROVEN, unattended, on both demo boxes
(2026-08-04). Scope: those two boxes, not the fleet.**
@@ -2154,6 +2188,8 @@ carrying a host-tier entry for the first time.
*The asymmetry worth remembering:* a host-tier restore is **83–109 s**; an offsite one is
**300–540 s**. The tier that matters for an ordinary recovery is also the cheapest to prove.
### S-21
**S-21 — an empty listing cannot distinguish FORBIDDEN from NEWBORN, so the box asks the permission
question directly (2026-08-03, R-185; agent v0.123.0 + installer 1.24.0).**
@@ -2182,6 +2218,8 @@ does not page — turning an ordinary documented setup into an alert is how a si
an operator archives unread. It never consults content, so it cannot alarm on a newborn tier by
construction, and it never reports ok when it could not ask.
### S-22
**S-22 — the installer's Scenario-F arm must finish the job, not just leave the definition alone
(2026-08-03, R-185).** `configure_backup_target` has two arms. Case A creates the storage and grants
in the same breath. The reuse arm — *"the target already exists"* — returned **without granting**, and
@@ -2194,6 +2232,8 @@ retargeting the box. `$BACKUP_TARGET_ID` stays OUT of `PVE_STORAGES`: that list
before the target is resolved, and `--acl-storages` entries are preflight-checked for existence.
A gate asserts every arm that resolves the target also grants on it.
### S-19
**S-19 — a restore-test PROOF is durable and reportable; a FAILURE is neither, and that asymmetry is
the design (2026-08-03, R-189; agent v0.122.0).**
@@ -2217,6 +2257,8 @@ VMID, duration) are not re-invented — an absent duration is not a claim, a fab
**Migration consequence, seen live:** a pre-R-189 record has no tier, so upgrading does not
retroactively make an old proof visible to the hub; the tier's next real proof fills it in.
### S-20
**S-20 — the release order is build → tag LOCALLY → publish → push tag, and every step protects
something (2026-08-03, R-188 + R-186).**
@@ -2236,6 +2278,8 @@ yields the same bytes with or without the tag; the verification command lives in
`felhom-agent/CLAUDE.md`. Both build paths (`release-agent.sh` and `publish-agent.sh`'s fallback) use
identical flags: they differed by `CGO_ENABLED=0` and produced binaries 74 KB apart for one version.
### S-17
**S-17 — restore-testing is PER ARCHIVE GENERATION, and the hub's staleness window follows each
tier's own rhythm (2026-08-03, R-86; agent v0.121.0 + hub v0.91.0).**
@@ -2268,6 +2312,8 @@ observed archive interval × 4 generations, floored at 7 days, capped at 12 days
`offsiteBackupStaleAfter` 8 d — the backup-freshness checker's own thresholds) when history is too
short to observe one. Shipping Part 1 alone would have produced a nightly false alarm.
### S-18
**S-18 — `ep0` is Tier 2, PROTECTED (operator ruling, 2026-08-03).** D-d named two protected machines
and did not name ep0 either way; `runbooks/target-selection.md` carried the question in writing for
two days. The ruling **extends D-d's protected list to three machines**: DooPlex, Peti's cluster,
@@ -2275,6 +2321,8 @@ two days. The ruling **extends D-d's protected list to three machines**: DooPlex
tunnel config or nftables rules was already forbidden by what it would destroy, and the ordinary
off-site READ a restore-test performs remains permitted.
### S-13 (2026-08-03)
**S-13 — the `mp1` merge landed, and the variant was chosen on measurement (2026-08-03, R-165 / D-a).**
The appliance's two data volumes are one. **Variant V-c**: the volume mounts at the NEUTRAL path
`/var/lib/felhom`, and both `/var/lib/docker` and `/mnt/sys_drive` are binds of subdirectories of it.
@@ -2297,6 +2345,8 @@ so pruning could only mean deleting a **different** app's only local recovery un
boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed shortly.
So R-176's in-place migration rehearsal is **withdrawn**, not deferred.
### S-14 (2026-08-03)
**S-14 — prove first, then vouch (2026-08-03) — SPENT, and the ordering did not survive contact.** The
rule was: golden **0.192.0** stays UNVOUCHED until a box has been proven from it, because vouching is
what makes a fresh install pick a golden up. **In the event the golden was vouched at 07:23:26 CEST on
@@ -2310,6 +2360,8 @@ and reminders do not hold. If prove-then-vouch is to be a rule it needs the shap
a refusal at `handleSetArtifacts`, the sole path to `SetArtifactManifest`, which runs without anyone
choosing to run it.
### S-15 (2026-08-03)
**S-15 — the merged layout is proven live, by two different supply paths (2026-08-03, R-178).** Both
demo boxes were wiped and reinstalled from golden 0.192.0 and taken through claim → deploy → back up →
**restore**. **demo-hp** was installed with `--golden <local volid>` (the layout proof) and
@@ -2323,6 +2375,8 @@ the floor guarded `captureAllRecoveryUnits` and not `runVolumeDumps`, the leg th
and its refusal's "the previous unit is untouched" was measured false. **R-181 CLOSED the same day
(controller v0.193.0 + v0.193.1), so R-165 is now PROVEN-LIVE in both halves** — see S-14.
### S-14 (2026-08-03)
**S-14 — the reserve is a per-app, per-run ADMISSION decision, not a capture check (2026-08-03, R-181;
controller v0.193.0 + v0.193.1).** B2 as first shipped was consulted in exactly one place —
`captureAllRecoveryUnits`, a few KB — while `RunDBDumps`' database leg and `runVolumeDumps` wrote the
@@ -2355,6 +2409,8 @@ not provide, and the fourth of those found on live hardware rather than by revie
no run scope, so a refused app re-alerts on every status refresh (measured: a second identical alert
pair 13 s after the run's). Pre-existing in v0.192.0; R-181 changed neither caller.
### S-15 (2026-08-03)
**S-15 — publishing is an act, not a side-effect of pushing (2026-08-03, R-110 + R-115 + R-183).**
Two rulings, one shape: something became live because someone pushed, not because anyone decided.
@@ -2381,6 +2437,8 @@ Two rulings, one shape: something became live because someone pushed, not becaus
that bumps a version, before publishing — and a gate that fails on the normal path is one people
learn to ignore.
### S-16
**S-16 — a backup run NOTIFIES ONCE and RECORDS ALWAYS, and those are different things
(2026-08-03, R-182; controller v0.194.0 + hub v0.90.0/.1).** Measured: nine per-app capture failures
reached the hub, two were mailed, seven were dropped by a cooldown whose key carries no app
@@ -2408,6 +2466,8 @@ identifier — *before* `LogNotification`, so they left no row anywhere.
the 4 GiB swapfile survived. The 40 GB local disk is UNCHANGED** — a CPU/RAM resize only, so no disk
figure in any runbook needed correcting. That closed **R-90** and unblocked **R-86**.
### S-11 (2026-08-02)
**S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).**
Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the
CUSTOMER's** (they can free space, delete files, add a drive) and **a per-app backup capture failure
@@ -2430,6 +2490,8 @@ label and the free-space figures — the same reason `offbox_enlarge_blocked` an
have none. `notify.IsOperatorOnly` was added so ONE test pins both registers; checked separately, an
allowlisted-but-not-operator-only type is invisible.
### S-12 (2026-08-02)
**S-12 — the monitoring landed BEFORE the merge, not with it (2026-08-02).**
D-a's condition (2) says R-167 ships in the same step as the `mp1`→`mp0` merge and never after,
because the merge removes a wall that currently fails safely. **This session landed it FIRST**, which
@@ -2442,6 +2504,8 @@ load-bearing findings for anyone picking that up: **"the layout" is not one thin
is also a BULKHEAD**, not only a ceiling — today an overflow cannot reach `/var/lib/docker`, and after
the merge it can.
### S-8 (2026-08-02)
**S-8 — CI detects; it does not block, and that is structural (2026-08-02, R-168).**
A Gitea Actions runner in `gitea-system` re-runs every repo's gate entry point on every push,
independent of who pushed and of what they typed. It **cannot refuse a push**: every felhom repo
@@ -2452,6 +2516,8 @@ skipped or was never armed. Making CI blocking needs branch protection plus a PR
changes how the operator works and is **their** call → R-169. Do not "fix" this by adding branch
protection.
### S-9 (2026-08-02)
**S-9 — a detector that tells no one is not finished (2026-08-02, R-168 probe P5).**
Probe P5 measured that a failed run produces **no mail, no notification row and no log line** from
Gitea. So the workflow sends its own alarm on the project's existing Resend path and **prints the
@@ -2462,6 +2528,8 @@ the runner image has **no `curl`** (deliberately — python3 and git only, so us
`api.resend.com` sits behind **Cloudflare, which 403s the default `Python-urllib` User-Agent with
error 1010** — a failure that looks exactly like an auth failure and is not one.
### S-10 (2026-08-02)
**S-10 — the runner is unprivileged, and the reason is the host (2026-08-02).**
The usual `act_runner` recipe pairs it with a `docker:dind` sidecar and `privileged: true`. Rejected:
DooPlex is **Tier 2** and *is* the recovery chain — Gitea, the hub, the registry, PBS and
@@ -2471,6 +2539,8 @@ job sees exactly the runner image's tools**, which is why `python3` had to be ba
stock `act_runner` carries git but not python3). If a future job genuinely needs Docker, that is a
conversation, not a patch.
### S-11
**S-11 — CI reproduces the workspace's sibling layout, because two entry points depend on it
(2026-08-02).** `controller_gates.py` and `agent_gates.py` invoke the shared `reuse_refs_check.py`
that lives in the `felhom.eu` clone next door and is deliberately never copied, and both repos'
@@ -2478,6 +2548,8 @@ that lives in the `felhom.eu` clone next door and is deliberately never copied,
sibling; without it the gate fails **closed** — correctly, but for the wrong reason. Verified that CI
and the local hook then agree exactly (controller 126 exact / 6 suffix / 1 cross-repo).
### S-6 (2026-08-02)
**S-6 — the hub renders no host-install version, and the gate pins its absence (2026-08-02, R-94).**
The Setup tab's *"host-install 1.19.0"* label is **deleted, not derived**. Deriving it is not
achievable honestly: the Option-1 command downloads `felhom-host-install.sh` from the website **at
@@ -2492,6 +2564,8 @@ tautological `render_test.go` assertion (`html contains hostInstallVersion`, whe
put it there) **passed at `9.9.9`** — an assertion that compares a value to itself tests the
plumbing, never the claim.
### S-7
**S-7 — gates run from ONE entry point per repo, and `reuse_refs_check` was fixed rather than the
convention it polices (2026-08-02, R-29).** Two rulings from the same census.
@@ -2519,6 +2593,8 @@ exact → suffix → ambiguous → sibling repo → FAIL, **prints every non-exa
resolution attempted on a failure. It stays in **one** place and is invoked across the workspace —
never copied, which would recreate the drift it detects.
### S-1 (2026-07-26)
**S-1 — N.5 gains a third leg: architecture docs are same-session coupled (2026-07-26, R-81).**
Any task that changes an **architectural contract** — tiers, targets, cadences, trust boundaries —
updates the owning `documentation/architecture/*.md` in the **same session**, under exactly the same
@@ -2527,6 +2603,8 @@ coupling rule that already binds the capability map and the ROADMAP. Origin: R-8
(single target, single cadence), while being cited as authoritative. A stale architecture doc is
worse than a missing one, because it is trusted.
### S-2 (2026-07-26)
**S-2 — architecture docs carry an honest status header (2026-07-26, R-81).**
Every `documentation/architecture/*.md` opens with the version it was **verified against** and the
date. A doc more than a few trains behind its subject is marked **STALE** *in that header*, so a
@@ -2535,6 +2613,8 @@ reader meets the warning before the content, not after acting on it. Origin:
versions stale (live v0.173.0), and cited as authoritative throughout the R-80 diagnostic. Ratifying
or retiring it is → **R-83**.
### S-3
**S-3 — the recovery model: six decisions, 2026-07-28.** Taken in an architecture discussion and
expressed in the `07-backup-architecture.md` full rewrite (which replaces the 2026-07-14 DRAFT
entirely — that doc was verified against controller v0.132.0, **51 versions stale**, while being
@@ -2569,6 +2649,8 @@ cited as authoritative). They are **decisions, not observations**; the rewrite l
compromised hub yields blobs nobody can open, *provided the operator's key is never stored in the
hub*. That proviso is why escrow custody is an open decision (`07` §11-A).
### S-4 (2026-07-31)
**S-4 — the hub session password alone now unlocks console root on every managed box (2026-07-31, hub v0.84.0).**
Retrieving a host's vaulted break-glass `root@pam` credential previously required the **global operator
API key**, a secret distinct from the hub login and kept out-of-band. The `Console access` card on the
@@ -2589,6 +2671,8 @@ Five decisions were deliberately **left open for the operator** and are recorded
stated**) · Hetzner as a single failure domain · and `local` vzdump sharing a physical device with
the guest it backs up. Gaps minted the same session: **R-102 … R-108**.
### S-13
**S-13 — boot recovery finished, and the lesson is about the DIAGNOSIS ORDER (controller v0.190.0,
2026-08-02, R-157 A · R-170 · R-171).**
@@ -2639,6 +2723,8 @@ you are watching the cache settle, not the system.**
six), window settle times 10/40/10/10/15/15 s — routinely 2–8× the old fixed 5 s. The sharpest
evidence is a same-app before/after on one box: missed at 18:08:35, recovered at 18:18:50.
### S-12
**S-12 — D-b is BUILT (controller v0.189.0, 2026-08-02, R-166).** The desired/in-flight/observed
split now exists; the S-1 contract lives in `architecture/02-controller-module-map.md` §0a.
@@ -2682,6 +2768,8 @@ behaviour changes under one live validation is one too many.
**IMPLEMENTED, not PROVEN-LIVE** — unit-proven and red-proofed, but nobody killed the controller
mid-backup on real hardware; the capability map says so rather than rounding it up.
### S-5
**S-5 — four operator decisions taken in discussion on 2026-08-02, recorded before anything is
built.** They existed only in conversation, which is the condition the standing rules were written
against. Labels are the ones used in the discussion (**D-a … D-d**) and are deliberately kept