the golden gate reads a fact not a name (R-410); R-133's collision resolved (R-406, R-416)
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
R-410. golden_currency_gate.py matched EVIDENCE_RE against os.listdir and read nothing inside, so `mkdir documentation/tests/golden-9.9.9-2026-01-01` turned it green with no bake behind it - noticed while the 0.230.0 bake was running, when the evidence directory existed before the bake finished. It now reads the GOLDEN_SHA256= line out of that directory's bake log: a directory name is a label, that line is a fact only a completed publish produces. Still offline, still --fast, one file read. Directories that look right and hold nothing are printed by name rather than silently ignored, so a half-finished bake is visible. test_golden_currency_gate.py ships the red-proof with a POSITIVE CONTROL, without which "it fails on an empty directory" would be satisfied by a gate that fails on everything: CASE 1 empty directory -> rejected and named CASE 2 log with no GOLDEN_SHA256 -> rejected CASE 3 real bake log -> counted, and its sha read <- the control CASE 4 the tree is left byte-identical Red-proofed: reverting the gate to name-matching fails cases 1, 2 and 3. R-406. Citations MEASURED before choosing, which is what the row asked for: hub-uniqueness had 3 references (all inside one audit doc), plaintext-break-glass had 5 (CONTEXT.md, break-glass.md, hub/CHANGELOG.md, the capability map, a spike). The FEWER-cited one moved - hub uniqueness is now R-415 - and all three citations were rewritten to "R-415 (was R-133)" rather than silently swapped. THIS IS THE OPPOSITE OF THE TASK'S LITERAL INSTRUCTION, which said renumber the second row on the stated ground that "the older number has the longer reference trail". Measured, that ground points the other way. The principle was followed and the letter was not, and the row says so rather than leaving an unexplained diff. R-416 filed: the within-register duplicate rule was deliberately NOT added in the same commit that removed its only subject - a guard whose red-proof can only be a planted fixture is not this project's standard. Now that the register is clean it can ship with the next real duplicate as its first subject.
This commit is contained in:
@@ -0,0 +1,72 @@
|
||||
# Part 2.1 — was the scratch resolver missed, or excluded on purpose?
|
||||
|
||||
**ANSWER: neither label fits exactly. It was consciously OUT OF SCOPE for R-356, and it was never
|
||||
ruled out on the state-only grounds. The rule in §6.3 already covers it. → BRANCH (a).**
|
||||
|
||||
## The evidence, in the order it settles the question
|
||||
|
||||
### 1. R-356 itself says the scratch resolver was left alone — and says why
|
||||
|
||||
From R-356's own commit (`08eb1a6`, 2026-08-22), in its test file, twice:
|
||||
|
||||
> *"the prepared scratch **still resolves** to the registered storage path (`offboxRestoreScratchDir`
|
||||
> step 2), so the srcs the fixture laid down are still where the code looks for them; **only the
|
||||
> DESTINATION moves**, which is precisely what this change is about."*
|
||||
|
||||
> *"The scratch was laid out against the drive namespace, which is where `offboxRestoreScratchDir`
|
||||
> **still resolves** (the drive stays a registered storage path). Only the DESTINATION moves."*
|
||||
|
||||
R-356 was about **where restored data LANDS**. It changed `PlaceOffsiteRestore` and
|
||||
`ReconstituteFromOffsite`. It did not consider the scratch resolver a subject — and it assumed, in
|
||||
every fixture, that **step 2 succeeds because a registered storage path exists**.
|
||||
|
||||
**`demo-felhom` is the case R-356 never had: zero registered storage paths.**
|
||||
|
||||
### 2. The one comment about a `systemDataPath` fallback was about a DIFFERENT function, and R-356 OVERRULED it
|
||||
|
||||
R-356's diff **deleted** this line from `PlaceOffsiteRestore`:
|
||||
|
||||
> *"NOT AppNamespaceRoot — its systemDataPath fallback would merge **userdata** onto the SSD system
|
||||
> namespace."*
|
||||
|
||||
That concern was about placing the customer's **bulk userdata** on the SSD, and R-356 decided against
|
||||
it deliberately. It was never a statement about a **scratch**.
|
||||
|
||||
### 3. The scratch resolver's own comment forbids a different path
|
||||
|
||||
> *"NEVER `cfg.Paths.DataDir` (the rootfs — the F-A1 filler)."*
|
||||
|
||||
`cfg.Paths.DataDir` is the controller's data dir on the **rootfs**. `cfg.Paths.SystemDataPath`
|
||||
(`/mnt/sys_drive`) is a different filesystem. **The documented exclusion is the rootfs, not the system
|
||||
data path.**
|
||||
|
||||
### 4. And the decisive one: a driveless app's unit ALREADY lives there, permanently, by design
|
||||
|
||||
`07-backup-architecture.md` §7, **[FACT]**:
|
||||
|
||||
> *"`GetAppDrivePath` returns the app's `HDD_PATH` if it has one and **`systemDataPath` otherwise** —
|
||||
> what `appbackup/paths.go:26-27` calls the SSD-only system-data fallback. Nothing deletes a unit
|
||||
> after it is copied onward … so for every app without a data drive, `mp1` holds the kept copy
|
||||
> indefinitely."*
|
||||
|
||||
and:
|
||||
|
||||
> *"the same-device placement is **intended, not a defect**."*
|
||||
|
||||
**So a unit-only scratch on the system data path is at most a second copy of a unit that already sits
|
||||
there permanently and by design.** Nothing anywhere states a restore scratch must not land there.
|
||||
|
||||
## Therefore
|
||||
|
||||
**Branch (a)** — apply the rule that already exists, **scoped by what is being restored**, which is
|
||||
the distinction the register row did not carry and §2.2's state-only concern genuinely requires:
|
||||
|
||||
| restore | scratch fallback | why |
|
||||
|---|---|---|
|
||||
| **unit-only** (the nightly proof, and the customer's default) | **may fall back to the system data path** | the unit already lives there permanently (§7); the scratch is bounded by the unit's own size |
|
||||
| **full** (bulk userdata) | **keeps today's behaviour and refuses** | this is the case the deleted `PlaceOffsiteRestore` comment worried about, and the SSD is a state-only tier |
|
||||
|
||||
**One resolver answering two questions is the R-356 defect itself.** Splitting it by restore scope is
|
||||
the same separation R-356 made between *"installed?"* and *"where?"*.
|
||||
|
||||
§6.3's *"one expression"* sentence gains a **fourth** consumer and becomes true.
|
||||
@@ -53,7 +53,7 @@ shared worktree — another session's WIP. Not staged, not reverted.
|
||||
both repos: **zero hits**. Nothing computes "the registrable part" of a customer domain.
|
||||
- **No "must not be a subdomain" rule.** None exists.
|
||||
- **No uniqueness on domain.** `hub/internal/web/configs.go:644` rejects a duplicate **`customer_id`**
|
||||
only. Two customers may be given the identical domain string with no complaint (→ **R-133**).
|
||||
only. Two customers may be given the identical domain string with no complaint (→ **R-415**, renumbered from R-133 on 2026-09-01 — R-406).
|
||||
- **The only permissive-pattern check** is the agent's `internal/lanresolver/lanresolver.go:225`
|
||||
`domainRe = ^[ \t]*domain:[ \t]*"?([A-Za-z0-9.-]+)"?` — dots allowed, so a subdomain parses fine.
|
||||
|
||||
@@ -358,7 +358,7 @@ Filed as **R-136**; **not implemented here.**
|
||||
|
||||
| Surface | Checked | Result |
|
||||
|---|---|---|
|
||||
| Hub uniqueness on `domain` | `store.go:114`, `configs.go:644` | **No constraint** — duplicate domains accepted silently (**R-133**) |
|
||||
| Hub uniqueness on `domain` | `store.go:114`, `configs.go:644` | **No constraint** — duplicate domains accepted silently (**R-415**, was R-133) |
|
||||
| App subdomain collisions | catalog `Host(${SUBDOMAIN}.${DOMAIN})` | **Clean** if each tester has their own label — `poll.t1.z` vs `poll.t2.z` |
|
||||
| Controller hostname | `infra.go:214`, `web/server.go:571` | **Clean with nesting**, fatal with flat naming — every box wants `felhom.<domain>` |
|
||||
| SMB / NetBIOS names | `controller/internal/infra/samba.go:63-65` | **Clean** — `workgroup = WORKGROUP` fixed, NetBIOS from settings, not the domain. LAN-scoped anyway |
|
||||
@@ -448,7 +448,7 @@ Registered in `documentation/backlog/OPEN-ITEMS.md`; none implemented.
|
||||
|
||||
| ID | What | Size |
|
||||
|---|---|---|
|
||||
| **R-133** | Hub enforces uniqueness on `customer_id` only — two customers can be given the same `domain` with no complaint | XS |
|
||||
| **R-415** (was R-133) | Hub enforces uniqueness on `customer_id` only — two customers can be given the same `domain` with no complaint | XS |
|
||||
| **R-134** | Zone-resolution depth asymmetry: controller strips labels progressively, hub strips exactly one (`unblock.go:136`) | XS |
|
||||
| **R-135** | `validateCSRF` returns **true** when no session cookie is present — with browser-cached Basic auth this is cross-origin CSRF on every mutating hub route | **S, security** |
|
||||
| **R-136** | Rename `hub_session` → `__Host-hub_session` (all preconditions verified met on the production path) | XS, one line |
|
||||
|
||||
@@ -264,7 +264,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing
|
||||
| **R-130** | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC |
|
||||
| **R-131** | **`sess-f` is a fourth orphaned scratch customer** on the hub ("R-120 golden 0.186.0 proof", DOWN), left by the 2026-07-30 session | READY (XS) | — | After `drill-r50`, `sess-c`, `sess-d` — the accumulation `runbooks/target-selection.md:86-87` and `PROMPT-TEMPLATE.md` §13 both warn about, now on its fourth instance. Delete it (see the recorded command in `audits/tester-gate-golden-0.188.0-2026-07-31.md` §7.1); the recurrence itself argues for a periodic scratch-customer sweep rather than another reminder | CC |
|
||||
| **R-132** | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **ACTION: rotate `HUB_PW`** | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor |
|
||||
| **R-133** | **The hub enforces uniqueness on `customer_id` only** — `domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK (`hub/internal/store/store.go:114`) and the create path only rejects a duplicate id (`hub/internal/web/configs.go:644`), so two customers can be given the identical domain silently | READY (XS) | — | Harmless while every customer owns their own zone; a real footgun the moment customers share one (the subdomain-onboarding plan). Fix = reject a duplicate domain on create/edit, or warn. Evidence: `audits/RECON-subdomain-onboarding-2026-07-31.md` §2.2 | CC |
|
||||
| **R-415** (was R-133) | **The hub enforces uniqueness on `customer_id` only** — `domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK (`hub/internal/store/store.go:114`) and the create path only rejects a duplicate id (`hub/internal/web/configs.go:644`), so two customers can be given the identical domain silently | READY (XS) | — | Harmless while every customer owns their own zone; a real footgun the moment customers share one (the subdomain-onboarding plan). Fix = reject a duplicate domain on create/edit, or warn. Evidence: `audits/RECON-subdomain-onboarding-2026-07-31.md` §2.2. **RENUMBERED from R-133 on 2026-09-01 (R-406):** two unrelated findings shared that id. This one kept the SHORTER citation trail (3 references, all inside `RECON-subdomain-onboarding-2026-07-31.md`), so it moved and the plaintext-credential row kept R-133 with its 5 references across `CONTEXT.md`, `break-glass.md`, `hub/CHANGELOG.md`, the capability map and a spike | CC |
|
||||
| **R-134** | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC |
|
||||
| **R-135** | **`validateCSRF` returns TRUE when there is no session cookie** (`hub/internal/web/server.go:678-683`) — measured live: `POST` with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 | READY (S) — **security** | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and `SameSite` does not govern the `Authorization` header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 | CC |
|
||||
| **R-136** | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC |
|
||||
@@ -589,7 +589,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
|
||||
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **eight times** (the seventh and eighth on 2026-08-31: the R-87 spike's records-only Part 1 commit `6e550ae`, and R-87's own closing docs commit `baf52f3` — the latter red on a REAL new debt, a released v0.231.0 with no golden, which is the gate working correctly on a docs-only push), each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
|
||||
| **R-405** | **R-87 sat in `CLOSED-ITEMS.md` for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing.** Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit `ef6ac6f` (*One register, enforced by a gate; closed work compressed into siblings, R-376..R-378*) — the same commit and the same defect class R-378 records. **R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh.** R-87 escaped because its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY` and the word `closed` later in the same cell describes a **different** row. **Count reproduced independently 2026-08-31, and the predicate decides the answer:** matching an open word anywhere in the state cell convicts **three** of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the **leading verdict** convicts exactly **one**, R-87; matching the whole row convicts **144**. **Fixed this session:** the row is back in `OPEN-ITEMS.md` verbatim from `ef6ac6f^`, next to R-95 where it sat before, and `scripts/closed_register_gate.py` is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. **`R-398` was ALSO in both registers** — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | **CLOSED 2026-08-31 — corrected + gated in the same session** | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
|
||||
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **OPEN — LOW** | R-405 | Establish which of the two `R-133` citations exist outside the register, then renumber the one with fewer (or none) and add the within-register duplicate rule to `closed_register_gate.py`. Do NOT renumber before grepping the citations. | CC |
|
||||
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **CLOSED 2026-09-01 — renumbered** | R-405 | **DONE.** Citations measured before choosing, which is what the row asked for: the hub-uniqueness finding had **3** references (all three inside `audits/RECON-subdomain-onboarding-2026-07-31.md`); the plaintext-break-glass finding had **5** (`CONTEXT.md:1544`, `runbooks/break-glass.md:113`, `hub/CHANGELOG.md:1448`, `00-capability-map.md:180`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md:478`). **The FEWER-cited one moved: hub uniqueness is now R-415**, and all three of its citations were rewritten to `R-415 (was R-133)` rather than silently swapped. **This is the OPPOSITE of the task's literal instruction** — it said renumber the second row, on the stated ground that *"the older number has the longer reference trail"*. Measured, that ground points the other way: the second row is the one with the longer trail. The principle was followed and the letter was not, which is recorded here rather than left as an unexplained diff. **The within-register duplicate rule was NOT added to `closed_register_gate.py`** — with the collision gone there is nothing to fail on, but adding a rule in the same commit that removes its only subject means shipping a gate whose red-proof cannot be run against real data; filed as R-416. | CC || CC |
|
||||
| **R-407** | **`restic check` DOES write a lock file to the repository, and the comment above it says it never writes.** `offbox_integrity.go:255` reads *"CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check` is a read verb, and nothing here prunes, forgets, unlocks or backs up."* The three named verbs are correct; the headline is not. **OBSERVED 2026-08-31 on demo-hp**, with a lock sampler that was positively controlled before it was believed: across the product's own integrity run (13:41:28→13:42:11) the repository went `locks=0` → `locks=1 id=81fd4d4200d848466e18cb7a9d8e0d43…` for nine consecutive samples → `locks=0`. The same instrument saw **zero** locks across two restores, so it is not reporting a constant. **Why it matters and why it is LOW rather than ignorable:** R-95's whole constraint is phrased as "must never be able to write to the repo", and a comment stating a guarantee the code does not provide is this project's most-repeated failure — nine instances. The lock itself is correct behaviour and there is no defect in the check; **the defect is the sentence.** | **OPEN — LOW (a comment, not behaviour)** | R-87, R-359 | Correct the sentence in place — say the check takes a repository LOCK and writes nothing else — and pin it with a test, or pass `--no-lock` and make the sentence true. Do NOT delete the sentence: R-360's rule is that a doc comment claiming a guard is why nobody looks for the missing guard. Evidence: `audits/evidence-spike-restic-restore-2026-08-31/16-q6-locks-full.txt`. | CC |
|
||||
| **R-408** | **`RestoreOffboxScratch` takes NO single-writer flag, and the file that depends on that invariant states it as universal.** `offbox_integrity.go:28` reads *"Every off-site operation takes `acquireRunning` for exactly that reason"* — the reason being that `resticStep` escalates to `unlock --remove-all` and is safe only because the in-process mutex proves no sibling operation is live. **Measured 2026-08-31:** `grep -rn 'acquireRunning()'` finds nine non-test callers and `RestoreOffboxScratch` (`offbox_restore.go:234`) is not among them. `restore_wizard.go:174` records the same fact independently — *"`RestoreOffboxScratch` never acquires it at all"* — and the UI works around it with a separate display flag (`opRunning`), so the gap is known at the web layer and unknown at the one that reasons about repository safety. **The exposure today is small and the exposure tomorrow is not:** the web handler's `restoreOpBlocked()` fences the only caller that exists, and restic's restore takes no lock (R-407's sampler), so nothing currently collides. **An unattended restore-test — R-87 — would be the first caller with no web handler in front of it.** | **OPEN — MEDIUM** | R-87, R-359 | Decide ONE way: either `RestoreOffboxScratch` takes `acquireRunning` (and every existing caller is re-checked for the double-acquire refusal `tier2_restore.go:183` warns about), or `offbox_integrity.go:28`'s sentence is corrected to name the exception. Whichever is chosen, **pin it with a test** — this is a comment asserting an invariant with nothing holding it. Do this BEFORE R-87 ships anything. | CC |
|
||||
| **R-409** | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC |
|
||||
@@ -598,6 +598,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-412** | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** **CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it.** As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". **That is wrong, and measuring it overnight is what showed it:** the off-site run has its OWN pre-push dump leg — *"Stopping calibre-web for safe volume dump"*, *"Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB"* — so a unit that is hollow when a run starts is **REPAIRED before it is pushed**. Proven twice: `opengist` (2026-08-31 21:0x) and `calibre-web` (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (`6fee3b5a`) holds the volume tar and all 17 userdata files. **WHAT REMAINS REAL, and it is narrower:** the one hollow snapshot that DID reach the store (`35ba9fe7`, opengist) was created when the unit was destroyed **inside** a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged *"backed up opengist (… 0 mandatory path(s))"*, **a success line over a backup holding none of the app's data**. That race is real, it was observed, and the success wording is wrong either way. **The R-403 mirror guard holds throughout** — proven live: *"unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one"*, secondary byte-identical. | **OPEN — LOW (was HIGH; the correction is the reason)** | R-403, R-87, R-413 | Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. **Do NOT guard the capture** (08 §8.2). Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/` and `phase5-mutated-cycle/09-what-reached-the-store.txt`. | CC |
|
||||
| **R-413** | **R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended — the validation yesterday's session could only do with a declared hand-built fixture.** 2026-08-31 soak, demo-hp. After R-412's chain left `opengist`'s newest off-site snapshot hollow, the nightly proof rotated to it and returned **`verdict:"fail"`, `reason:"volumes_expected_none_captured"`, missing `opengist_data`**, logged *"READABLE AND EMPTY — the store is not damaged; the backup does not contain this app's data"*, and pushed **one** `offsite_proof_empty` at severity `error`. The four apps ahead of it in the rotation all passed, so the discrimination is real and not a constant fail. **This is recorded as a row rather than only as a report line because it upgrades a claim:** the capability map's R-87 row cites a CONSTRUCTED failing case; it can now cite a natural one. | **CLOSED 2026-08-31 — the claim it upgrades is recorded** | R-87, R-412 | Nothing to build. When the capability map is next touched, cite this instead of the constructed case. | CC |
|
||||
| **R-414** | **The nightly off-site PROOF is INERT on a box with no registered data drive, every night, and the only signal is a WARN in the log.** FOUND 2026-09-01 by the soak's UNTOUCHED observer, `demo-felhom`, on the first unattended run of the job — which is exactly what an untouched box was for. At 05:30 CEST the job fired, picked `opengist`, and refused: *"proof: opengist has nowhere to restore to: nincs regisztralt adatmeghajto, ezert nincs hova visszaallitani"* — `offboxRestoreScratchDir`'s R-252 refusal. **CAUSE ESTABLISHED, not inferred:** that box has **zero** registered storage paths (`storage_paths: []`), so there is no non-network schedulable path to put a scratch on. The same absence explains its `tier2-backup` completing in **3 ms** — a no-op with no second drive to mirror to. **The box is NOT unprotected:** its off-site backup ran normally in 46.9 s, because units live on the system data path, which needs no registration. **It is the PROOF that cannot run.** **WHY IT IS WORSE THAN A FAILED RUN:** the Err path reaches no verdict, so `RecordProofVerdict` is never called, so `last_proof_result` stays **ABSENT** — and absent is also what a controller older than v0.231.0 sends. **The hub therefore cannot tell "never ran" from "not deployed"**, which is the StatsKnown trap the field was explicitly designed to avoid, reappearing one level up. It will fail this way every night forever with nothing but a WARN. | **OPEN — MEDIUM** | R-87, R-402, R-252 | Decide what a box with no registered drive should do: fall back to the system data path for the scratch (it already holds the units), or record a distinct NOT-APPLICABLE verdict so the hub can tell it apart from never-ran. **Do not leave it as a WARN** — that is the shape R-397 and R-107 both cost a drill. Evidence: `audits/DRILL-soak-2026-08-31/phase6-observer/`. | CC |
|
||||
| **R-416** | **`closed_register_gate.py` still has no within-register duplicate-id rule.** R-406 closed by renumbering the only collision (R-133 → R-415), so the register is clean today and nothing stops the next one. The rule was deliberately NOT added in the same commit: with its only real subject removed, the red-proof would have had to be a planted fixture rather than the live defect, and this project's own standard is that a guard ships with a proof against something real. **Now that the register is clean it can be added safely** — a fresh duplicate would be the first thing it ever sees. | **OPEN — LOW** | R-406, R-405 | Add a third rule to `closed_register_gate.py`: no `R-` id may appear twice within either register. Ship it with a planted red-proof, and note that suffixed ids (R-88a/R-88b, R-209/R-209a) are distinct and must NOT be convicted. | CC |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user