docs: R-168 SHIPPED, R-29 CLOSED on the demonstrated alarm, R-169 minted
gates / gates (push) Successful in 7s
gates / gates (push) Successful in 7s
SPIKE-ci-runner-2026-08-02.md: all six probes with method, measurement and ruling; none STOPped. P2 (stock image has git but no python3) and P6 (a runner that loses its state re-registers and orphans the old record) changed the design; P5 (a failed run signals NOTHING) is why the alarm exists at all. R-168 SHIPPED with its evidence. R-29 CLOSED — on the demonstrated alarm and not on a green run, as required: the class it opened is answered at both ends, the hook refusing locally and CI catching a --no-verify bypass and emailing. R-161 noted: its automatic half now exists for the STATIC gate, while its original scope, the runtime gate, is deliberately still not automatic and should stay that way. NEW R-169 (grep established R-168 was the highest in use): CI can only report, because there is no gate in the road. Making it blocking needs branch protection plus a PR workflow, both of which change how the operator works — so it is theirs to decide, and the row states the cost honestly rather than recommending it. CONTEXT gains S-8 (CI detects, does not block, and why that is structural), S-9 (a detector that tells no one is not finished, plus the curl and Cloudflare-1010 traps), S-10 (the runner is unprivileged because DooPlex is Tier 2), S-11 (CI reproduces the sibling layout). CLAUDE.md gains the rule earned by red-proofing: a go test -run pattern that matches no test prints ok and exits 0, and an instrument that can silently drop results is not a measurement.
This commit is contained in:
@@ -56,6 +56,13 @@ UI. Package map, helpers, seams, extension points: **`REUSE.md`** (e.g. new even
|
|||||||
(`{{if .Flag}}` around a button/form/script) ships with a render test per branch of the gate —
|
(`{{if .Flag}}` around a button/form/script) ships with a render test per branch of the gate —
|
||||||
handler tests that POST directly prove nothing about reachability. The v0.70.0 ghost-delete was
|
handler tests that POST directly prove nothing about reachability. The v0.70.0 ghost-delete was
|
||||||
fully implemented server-side and fully dead UI because the button sat inside the wrong gate.
|
fully implemented server-side and fully dead UI because the button sat inside the wrong gate.
|
||||||
|
- **A `go test -run` pattern that matches no test prints `ok` and exits 0.** Found 2026-08-02 while
|
||||||
|
red-proofing: `-run TestCustomerUnified` matched nothing in the target file and reported
|
||||||
|
`ok … 0.062s`, which was read as a passing red-proof. **A red-proof that uses `-run` must first
|
||||||
|
prove the filter matched something** (`-v` and look for `=== RUN`). This is the "an absent line is
|
||||||
|
not evidence" rule aimed at the one place a false green costs most — the proof itself. The same
|
||||||
|
class bit twice that day: a `| tail -5` inside a census query silently dropped rows and looked
|
||||||
|
exactly like a real finding. **An instrument that can drop results silently is not a measurement.**
|
||||||
- **A health check issues no block I/O.** A probe that touches a wedged device enters uninterruptible
|
- **A health check issues no block I/O.** A probe that touches a wedged device enters uninterruptible
|
||||||
sleep, survives `SIGKILL`, and cannot be recovered until the device returns or the host reboots — so
|
sleep, survives `SIGKILL`, and cannot be recovered until the device returns or the host reboots — so
|
||||||
`systemctl restart` hangs too. A timeout protects the caller's control flow and nothing else: the
|
`systemctl restart` hangs too. A timeout protects the caller's control flow and nothing else: the
|
||||||
|
|||||||
+36
@@ -17,6 +17,42 @@
|
|||||||
|
|
||||||
## Standing rulings
|
## Standing rulings
|
||||||
|
|
||||||
|
**S-8 — CI detects; it does not block, and that is structural (2026-08-02, R-168).**
|
||||||
|
A Gitea Actions runner in `gitea-system` re-runs every repo's gate entry point on every push,
|
||||||
|
independent of who pushed and of what they typed. It **cannot refuse a push**: every felhom repo
|
||||||
|
pushes straight to `main` with no pull request, so there is no merge for a status check to stand at.
|
||||||
|
That is not a gap in the runner — there is no gate in the road. The refusing half is
|
||||||
|
`.githooks/pre-push` (per-clone, `--no-verify`-able); this half is what notices when that hook was
|
||||||
|
skipped or was never armed. Making CI blocking needs branch protection plus a PR workflow, which
|
||||||
|
changes how the operator works and is **their** call → R-169. Do not "fix" this by adding branch
|
||||||
|
protection.
|
||||||
|
|
||||||
|
**S-9 — a detector that tells no one is not finished (2026-08-02, R-168 probe P5).**
|
||||||
|
Probe P5 measured that a failed run produces **no mail, no notification row and no log line** from
|
||||||
|
Gitea. So the workflow sends its own alarm on the project's existing Resend path and **prints the
|
||||||
|
provider's accepted id**, which makes "a message left the machine" an observable rather than an
|
||||||
|
assumption. The acceptance test for this work was never a green pipeline — it was a red run with a
|
||||||
|
message in hand (`RESEND-ACCEPTED id=…`). Two traps found while building it, both worth keeping:
|
||||||
|
the runner image has **no `curl`** (deliberately — python3 and git only, so use `urllib`), and
|
||||||
|
`api.resend.com` sits behind **Cloudflare, which 403s the default `Python-urllib` User-Agent with
|
||||||
|
error 1010** — a failure that looks exactly like an auth failure and is not one.
|
||||||
|
|
||||||
|
**S-10 — the runner is unprivileged, and the reason is the host (2026-08-02).**
|
||||||
|
The usual `act_runner` recipe pairs it with a `docker:dind` sidecar and `privileged: true`. Rejected:
|
||||||
|
DooPlex is **Tier 2** and *is* the recovery chain — Gitea, the hub, the registry, PBS and
|
||||||
|
k3s + Longhorn all live on it and it cannot be rebuilt from anything else. Host execution mode costs
|
||||||
|
nothing here because every CI job is one Python command. Consequence to remember: **in host mode the
|
||||||
|
job sees exactly the runner image's tools**, which is why `python3` had to be baked in (probe P2:
|
||||||
|
stock `act_runner` carries git but not python3). If a future job genuinely needs Docker, that is a
|
||||||
|
conversation, not a patch.
|
||||||
|
|
||||||
|
**S-11 — CI reproduces the workspace's sibling layout, because two entry points depend on it
|
||||||
|
(2026-08-02).** `controller_gates.py` and `agent_gates.py` invoke the shared `reuse_refs_check.py`
|
||||||
|
that lives in the `felhom.eu` clone next door and is deliberately never copied, and both repos'
|
||||||
|
`REUSE.md` files cite a path that lives in the hub. Their workflows therefore clone `felhom.eu` as a
|
||||||
|
sibling; without it the gate fails **closed** — correctly, but for the wrong reason. Verified that CI
|
||||||
|
and the local hook then agree exactly (controller 126 exact / 6 suffix / 1 cross-repo).
|
||||||
|
|
||||||
**S-6 — the hub renders no host-install version, and the gate pins its absence (2026-08-02, R-94).**
|
**S-6 — the hub renders no host-install version, and the gate pins its absence (2026-08-02, R-94).**
|
||||||
The Setup tab's *"host-install 1.19.0"* label is **deleted, not derived**. Deriving it is not
|
The Setup tab's *"host-install 1.19.0"* label is **deleted, not derived**. Deriving it is not
|
||||||
achievable honestly: the Option-1 command downloads `felhom-host-install.sh` from the website **at
|
achievable honestly: the Option-1 command downloads `felhom-host-install.sh` from the website **at
|
||||||
|
|||||||
@@ -1,274 +1,237 @@
|
|||||||
# REPORT — Gate enforcement: fix the reds, one entry point per repo, pre-push hook (2026-08-02)
|
# REPORT — CI: a Gitea Actions runner, and a red run that reaches a person (R-168, 2026-08-02)
|
||||||
|
|
||||||
**Overwritten** per the standing rule. The prior contents (`STATUS.md` created + the 2026-08-02
|
**Overwritten** per the standing rule. The prior contents (session 1's gate enforcement, same day)
|
||||||
operator decisions, same day) have their durable record in `STATUS.md` itself and in
|
have their durable record in `scripts/CHANGELOG.md`, `hub/CHANGELOG.md` and `CONTEXT.md` S-6/S-7.
|
||||||
`documentation/backlog/OPEN-ITEMS.md`; nothing was lost by this overwrite.
|
|
||||||
|
|
||||||
**Outcome: complete.** Nothing was left out. Two deviations from the spec and one census addition are
|
**Outcome: complete, and the half that matters is demonstrated.** All six probes answered; none
|
||||||
recorded in §13; both deviations were made deliberately, not by accident.
|
STOPped. The acceptance test was never a green pipeline — it was a **red run with a message in
|
||||||
|
hand**, and that is `RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 1. Confirmed baselines — all four matched §1 on arrival
|
## 1. Confirmed baselines — all matched §1 on arrival
|
||||||
|
|
||||||
| Repo | `main` on arrival | matched §1? | version on arrival | version now |
|
| Repo | `main` on arrival | matched? | version |
|
||||||
|---|---|---|---|---|
|
|
||||||
| `felhom.eu` | `e994bf35d290` | yes | hub `v0.86.0` | **hub `v0.87.0`** |
|
|
||||||
| `felhom-controller` | `4115e88f688b` | yes | `v0.188.0` | `v0.188.0` (no bump — tooling only) |
|
|
||||||
| `felhom-agent` | `4663df7ff3cc` | yes | `v0.119.0` | `v0.119.0` (no bump — tooling only) |
|
|
||||||
| `app-catalog-felhom.eu` | `fd7747d1293c` | yes | n/a | n/a |
|
|
||||||
|
|
||||||
All four trees were clean and `HEAD == origin/main` on arrival.
|
|
||||||
|
|
||||||
## 2. Files created / modified
|
|
||||||
|
|
||||||
**`felhom.eu`**
|
|
||||||
- `hub/internal/web/configs.go` — `hostInstallVersion` const, `pageData.ScriptVersion` field and its
|
|
||||||
assignment DELETED; a NOTE records why there is deliberately no constant here.
|
|
||||||
- `hub/internal/web/templates/customer_unified.html` — Setup Command copy rewritten, no version.
|
|
||||||
- `hub/internal/web/render_test.go` — tautological version assertion deleted (not replaced).
|
|
||||||
- `scripts/hostinstall_gates.py` — gate 1's third assertion inverted; docstring rewritten.
|
|
||||||
- `scripts/felhom-host-install.sh` — **comment only**; `SCRIPT_VERSION` untouched (deviation, §13).
|
|
||||||
- `scripts/reuse_refs_check.py` — rewritten resolution (exact/suffix/ambiguous/cross-repo/FAIL).
|
|
||||||
- `scripts/test_reuse_refs_check.py` — NEW, 13 tests.
|
|
||||||
- `scripts/repo_gates.py` — NEW, the entry point.
|
|
||||||
- `scripts/test_repo_gates.py` — NEW, 3 tests (seam).
|
|
||||||
- `.githooks/pre-push` — NEW.
|
|
||||||
- `CLAUDE.md`, `CONTEXT.md` (S-6, S-7), `.gitignore`, `hub/CHANGELOG.md`, `scripts/CHANGELOG.md`,
|
|
||||||
`manifests/hub.yaml`, `documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/ROADMAP.md`,
|
|
||||||
`STATUS.md`, this file.
|
|
||||||
|
|
||||||
**`felhom-controller`** — `controller/scripts/docker_run_volume_path_gate.py` (one ALLOWLIST entry),
|
|
||||||
`controller/scripts/controller_gates.py` (NEW), `controller/scripts/test_controller_gates.py` (NEW),
|
|
||||||
`.githooks/pre-push` (NEW), `CLAUDE.md`, `controller/README.md`, `CHANGELOG.md`, `REPORT.md`.
|
|
||||||
|
|
||||||
**`felhom-agent`** — `scripts/agent_gates.py` (NEW), `scripts/test_agent_gates.py` (NEW),
|
|
||||||
`.githooks/pre-push` (NEW), `CLAUDE.md`, `CHANGELOG.md`, `REPORT.md`.
|
|
||||||
|
|
||||||
**`app-catalog-felhom.eu`** — `scripts/catalog_gates.py` (`--fast`),
|
|
||||||
`scripts/test_catalog_gates.py` (NEW), `.githooks/pre-push` (NEW), `CLAUDE.md`, `CHANGELOG.md`,
|
|
||||||
`REPORT.md`.
|
|
||||||
|
|
||||||
## 3. Commits pushed to `main`
|
|
||||||
|
|
||||||
**`felhom-controller`** — Part 1.1 is its own commit, with nothing else in it, as specified.
|
|
||||||
|
|
||||||
| Hash | What |
|
|
||||||
|---|---|
|
|
||||||
| `c432f70` | **gate: allowlist appexport/estimate.go named-volume `-v` mount (R-29 leg a)** — the reviewed diff, alone |
|
|
||||||
| `8cb3d7a` | gates: one entry point (`controller/scripts/controller_gates.py`) + pre-push hook |
|
|
||||||
| `7c32c74` | docs: CHANGELOG + README for the gate entry point (no version bump) |
|
|
||||||
|
|
||||||
**`felhom.eu`**
|
|
||||||
|
|
||||||
| Hash | What |
|
|
||||||
|---|---|
|
|
||||||
| `d319ae5` | hub: delete the host-install version label (R-94) + invert hostinstall gate 1 |
|
|
||||||
| `2137094` | scripts: `reuse_refs_check` resolves package shorthand and sibling repos |
|
|
||||||
| `9bd1a54` | gates: one entry point (`scripts/repo_gates.py`) + pre-push hook |
|
|
||||||
| `4707be7` | docs: R-94 closed, R-29 leg (a) closed + leg (b) half, R-168 minted |
|
|
||||||
| `8d9b78c` | manifests: hub 0.86.0 → 0.87.0 |
|
|
||||||
| (final) | REPORT + STATUS |
|
|
||||||
|
|
||||||
**`felhom-agent`** — `054e85a` (entry point + hook), `75245a4` (CHANGELOG).
|
|
||||||
**`app-catalog-felhom.eu`** — `c3e4bb1` (`--fast` + hook), `340ff2a` (CHANGELOG).
|
|
||||||
|
|
||||||
## 4. Per-test results and EVERY red-proof
|
|
||||||
|
|
||||||
| # | Red-proof | Method | Outcome |
|
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| 1 | **Part 1.1 — the docker-`-v` gate** | renamed the new ALLOWLIST suffix to a non-existent file | gate **exit 1**, naming `estimate.go:179`; restored → **exit 0** |
|
| `homelab-manifests` | `420e8197769b` | yes | n/a |
|
||||||
| 2 | **Part 1.2 — the rewritten gate 1** | re-inserted `const hostInstallVersion = "1.19.0"` into `configs.go` | gate **exit 1**, convicting on **3 of its 6 shapes** (const / assignment / renamed-const); restored → **exit 0** |
|
| `felhom.eu` | `af2d10388021` | yes | hub `v0.87.0` — **no bump** |
|
||||||
| 3 | **Part 1.2 — the OLD `render_test.go` assertion tested nothing** | set the const to `"9.9.9"` while the served script was 1.22.0, ran `TestTemplates_InstallGenerator -v` | **`--- PASS`**. It compared the constant to itself. *(First attempt used `-run TestCustomerUnified`, which matches no test in that file and printed a vacuous `ok` — caught, corrected, re-run against the real test name.)* |
|
| `felhom-controller` | `eaded79b1894` | yes | `v0.188.0` — **no bump** |
|
||||||
| 4 | **Part 2 — Scenario E, the checker can still fail** | made `resolve()` return `exact` for an unresolvable token — the "made green by being made blind" mutation | **4 of 13 tests went red** (`test_absent_path_fails`, `test_failure_lists_every_resolution_attempted`, `test_evidence_copy_does_not_satisfy_a_citation`, `test_non_git_sibling_is_not_searched`); restored → 13/13 green |
|
| `felhom-agent` | `2c4efed5de32` | yes | `v0.119.0` — **no bump** |
|
||||||
| 5 | **Part 3 — the inert runner, all four entry points** | replaced each `run_gate` body with `return 0` | every runner still printed its summary and **exited 0**; the seam tests went red — felhom.eu 1 failure, controller 1, agent 2, catalog 1. All restored green |
|
| `app-catalog-felhom.eu` | `f16f29757e95` | yes | n/a |
|
||||||
|
|
||||||
**Integration scenarios**
|
Live Gitea re-confirmation, **all matched**: version **1.26.2**; Actions enabled on **all five**
|
||||||
|
repos; **0 runners**; **0 workflow runs ever**; **0 branch protections**; **no `.gitea/`** in any
|
||||||
|
repo. ArgoCD `gitea` app: `path: gitea-system`, auto-sync **off**, Synced at `420e819`, Healthy.
|
||||||
|
The only new artifact is the runner image, at its own `0.1.0`.
|
||||||
|
|
||||||
- **A — all four entry points green in `--fast`:** exit 0 each. Summary lines quoted in §6.
|
## 2. All six probe results
|
||||||
- **B — a red member fails the runner:** with `hostInstallVersion` restored, `repo_gates.py --fast`
|
|
||||||
exited **1**, `hostinstall_gates`' own finding text appeared on stdout, and the summary read
|
|
||||||
`hostinstall FAILED (exit 1)` / `CONVICTED: hostinstall`. Reverted. The companion inert-runner
|
|
||||||
proof is #5 above.
|
|
||||||
- **C — the hook refuses:** §7.
|
|
||||||
- **D — 13 findings resolve, none by blindness:** all four roots exit 0.
|
|
||||||
`appbackup/userdata.go` resolves **by suffix** inside `felhom-controller`;
|
|
||||||
`wgsync/reconciler.go` is reported as **cross-repo → `felhom.eu/hub/internal/wgsync/reconciler.go`**,
|
|
||||||
not as an exact match.
|
|
||||||
- **E — the kill condition:** pinned by `test_reuse_refs_check.py` and red-proofed (#4 above).
|
|
||||||
- **F — runtime catalogue gates stay out of the hook:** `catalog_gates.py --fast` ran
|
|
||||||
`check-image-pins.py` only (`image-pin gate OK — 53 templates, 0 unpinned images`); the other two
|
|
||||||
did not run and the skip was announced with its reason. Pinned by `test_catalog_gates.py`.
|
|
||||||
|
|
||||||
## 5. Test count before / after
|
Full method and measurement: `documentation/audits/SPIKE-ci-runner-2026-08-02.md`.
|
||||||
|
|
||||||
| Repo | Before | After | Delta |
|
| Probe | Ruling | Measurement |
|
||||||
|---|---|---|---|
|
|
||||||
| `felhom.eu` — Go (`hub/`) | 568 test funcs, 17 packages `ok` | 568, 17 `ok` | 0 (one *assertion* deleted, no test funcs) |
|
|
||||||
| `felhom.eu` — Python | 0 | **16** (13 `test_reuse_refs_check` + 3 `test_repo_gates`) | +16 |
|
|
||||||
| `felhom-controller` — Python | 0 | **4** | +4 |
|
|
||||||
| `felhom-agent` — Python | 0 | **4** | +4 |
|
|
||||||
| `app-catalog-felhom.eu` — Python | existing `test_check_volume_persistence.py` | **+5** (`test_catalog_gates`) | +5 |
|
|
||||||
|
|
||||||
All green. No failures.
|
|
||||||
|
|
||||||
## 6. The four entry points' summary lines, verbatim (`--fast`)
|
|
||||||
|
|
||||||
```
|
|
||||||
site OK (exit 0) template-id OK (exit 0)
|
|
||||||
hostinstall OK (exit 0) emoji OK (exit 0)
|
|
||||||
hub-confirm OK (exit 0) native-confirm OK (exit 0)
|
|
||||||
manifest-bearer OK (exit 0) offbox-rename OK (exit 0)
|
|
||||||
reuse-refs OK (exit 0) app-row-dedup OK (exit 0)
|
|
||||||
mojibake OK (exit 0)
|
|
||||||
all felhom.eu gates OK docker-v OK (exit 0)
|
|
||||||
reuse-refs OK (exit 0)
|
|
||||||
|
|
||||||
all controller gates OK
|
|
||||||
|
|
||||||
reuse-refs OK (exit 0) image-pins OK (exit 0)
|
|
||||||
|
|
||||||
all agent gates OK all catalog gates OK
|
|
||||||
```
|
|
||||||
|
|
||||||
Each run also prints its per-root reuse tally, which is the positive observable — e.g.
|
|
||||||
`OK [felhom-controller]: 133 cited paths — exact 126, suffix 6, ambiguous 0, cross-repo 1, FAILED 0`.
|
|
||||||
|
|
||||||
## 7. Part 4.0 probe — all four measured outcomes
|
|
||||||
|
|
||||||
Throwaway clone, git **2.47.3**, relative `core.hooksPath .githooks`:
|
|
||||||
|
|
||||||
| Step | Expected | **Measured** |
|
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| 1. hook + `core.hooksPath .githooks` + `chmod +x` | — | set; `git config --get core.hooksPath` → `.githooks` |
|
| **P1** runner picks up a job | **PASS** | run #1 claimed by `felhom-gates-runner`, `P1-OK`, `🏁 Job succeeded` |
|
||||||
| 2. `git push` from the **repo root** | refused | **REFUSED, rc=1** |
|
| **P2** stock image runs our checks | **known branch → build an image** | stock `act_runner:0.6.1` (Alpine 3.23): `git 2.52.0` PRESENT, `python3` **ABSENT**. Host mode itself fine → privileged/dind never reached for |
|
||||||
| 3. `git push` from a **subdirectory** (`hub/`) | refused | **REFUSED, rc=1** |
|
| **P3** source without JS actions | **PASS** | `pushed sha == checked-out sha` (`bbd62319…`), plain `git clone`/`fetch` from the in-cluster Service, no `uses:` |
|
||||||
| 4. `git push --no-verify` | succeeds | **ACCEPTED, rc=0** |
|
| **P4** one registration, four repos | **PASS** | tasks 7,8,9,10 across four repos, all `runner_id=2 felhom-gates-runner` |
|
||||||
|
| **P5** does a failure signal anything | **NOTHING SENT → build the alarm** | 0 mail/SMTP/notification log lines, **0** notification rows, after a real `failure` run |
|
||||||
|
| **P6** persistent state needed | **PVC is load-bearing** | PVC intact → same registration; `/data/.runner` removed → **second record** appears and the first is orphaned |
|
||||||
|
|
||||||
**Step 3 did not fail, so there is no STOP and no `SPIKE-githooks-*.md` was written.** The form used
|
**P5's honest limit, stated rather than glossed:** I did not independently prove Gitea's SMTP path
|
||||||
is therefore the **relative** `core.hooksPath .githooks`.
|
can *deliver*, so "no mail" cannot be split with certainty between *no action-failure notification in
|
||||||
|
1.26.2* and *a broken mailer*. It does not change the design — the alarm rides Resend, a different
|
||||||
|
path — and Scenario C proves that path end to end.
|
||||||
|
|
||||||
Step 3 was then re-measured with a hook that prints a distinctive line, because "the push was
|
## 3. Files created / modified
|
||||||
refused" alone does not say *what* refused it. The hook printed `PROBE-HOOK-RAN` **with cwd = the
|
|
||||||
repo root** when `git push` was issued from the root, from `hub/`, and from a two-level-deeper
|
|
||||||
subdirectory. So git resolves a relative `core.hooksPath` against the repo root and runs the hook
|
|
||||||
there regardless of where the push is issued. The shipped hook does not rely on that — it resolves
|
|
||||||
the root explicitly with `git rev-parse --show-toplevel`.
|
|
||||||
|
|
||||||
**Scenario C**, in a throwaway clone of `felhom.eu` with `origin` re-pointed at a scratch bare repo
|
**`homelab-manifests`** — `gitea-system/act-runner.yaml` (NEW: PVC + Deployment),
|
||||||
so nothing real could be pushed:
|
`gitea-system/act-runner/Dockerfile` (NEW), `CHANGELOG.md`, `REPORT.md`.
|
||||||
|
|
||||||
- **unarmed clone** — a manual `repo_gates.py --fast` printed
|
**`felhom.eu`** — `.gitea/workflows/gates.yml` (NEW),
|
||||||
`WARNING: this clone is UNARMED — core.hooksPath is unset…`
|
`documentation/audits/SPIKE-ci-runner-2026-08-02.md` (NEW), `CLAUDE.md`, `CONTEXT.md` (S-8…S-11),
|
||||||
- **C0, clean tree** — push **ACCEPTED**, and the hook said so:
|
`scripts/CHANGELOG.md`, `documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/ROADMAP.md`,
|
||||||
`pre-push [felhom.eu]: gates OK - push proceeding.`
|
`STATUS.md`, `REPORT.md`. *(`hub/internal/web/configs.go` was broken and restored within the session
|
||||||
- **C1, deliberately broken gate** — push **REFUSED**; `hostinstall_gates`' own three findings and
|
for Scenario B; it is byte-identical to its starting state — verified by the gate that objects to it.)*
|
||||||
`CONVICTED: hostinstall` appeared, then
|
|
||||||
`pre-push [felhom.eu]: PUSH REFUSED - gates exited 1.`
|
|
||||||
- **C2, `git push --no-verify` on the same commit** — **ACCEPTED** (`9bd1a54..19e8cbb`), and the
|
|
||||||
broken commit was confirmed present in the scratch bare repo. The escape hatch works.
|
|
||||||
|
|
||||||
All five real pushes in this session ran their hook and announced it — see §3.
|
**`felhom-controller`**, **`felhom-agent`**, **`app-catalog-felhom.eu`** —
|
||||||
|
`.gitea/workflows/gates.yml` (NEW), `CHANGELOG.md`, `REPORT.md`.
|
||||||
|
|
||||||
## 8. `git config --get core.hooksPath` in the four live workspace clones
|
## 4. Commits pushed to `main`
|
||||||
|
|
||||||
|
| Repo | Hash | What |
|
||||||
|
|---|---|---|
|
||||||
|
| `homelab-manifests` | `c51b77c` | the runner manifest + Dockerfile |
|
||||||
|
| `felhom.eu` | `bbd6231` | the temporary Part 0 probe workflow |
|
||||||
|
| | `666a34d` | real `gates.yml`, probe workflow **deleted** |
|
||||||
|
| | `3252d51` | **Scenario B** deliberate breakage (`--no-verify`) |
|
||||||
|
| | `dd13f63` | alarm step v1 — died on `curl: command not found` |
|
||||||
|
| | `f7dbc33` | alarm rewritten in python3/urllib |
|
||||||
|
| | `9530de7` | alarm User-Agent (Cloudflare 1010) — **the run that produced the accepted-id** |
|
||||||
|
| | `4cc1238` | **breakage reverted**, main green |
|
||||||
|
| `felhom-controller` | `dcc400e` | `gates.yml` (with sibling clone) |
|
||||||
|
| `felhom-agent` | `eb99144` | `gates.yml` (with sibling clone) |
|
||||||
|
| `app-catalog-felhom.eu` | `aa57588` | `gates.yml` (`--fast`, no sibling needed) |
|
||||||
|
|
||||||
|
Plus the documentation commits listed at the end of this session.
|
||||||
|
|
||||||
|
## 5. Scenario results A-G
|
||||||
|
|
||||||
|
| # | Result | Positive observable |
|
||||||
|
|---|---|---|
|
||||||
|
| **A** | **PASS** | run #2 (`666a34d`) SUCCESS. **Every member gate's own stdout in the CI log**, not just the summary: `site gates OK — BOM, emoji=0…`, `hostinstall gates: ALL PASS`, `hub confirm gate OK…`, `manifest bearer gate OK…`, `OK [hostexecutor]: 61 cited paths — exact 60, suffix 1…`, then `all felhom.eu gates OK` |
|
||||||
|
| **B** | **PASS** | run #3 (`3252d51`) **FAILURE**. The failing gate's own text: `FAIL: hub/internal/web/configs.go:34 carries const hostInstallVersion…`, `hostinstall gates: 3 FAILURE(S)`, `CONVICTED: hostinstall`, `🏁 Job failed` |
|
||||||
|
| **C** | **PASS** | run #6 FAILURE → **`RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`**. See §6 |
|
||||||
|
| **D** | **PASS** | four repos, four runs, **one runner**: tasks 7/8/9/10 all `runner_id=2 felhom-gates-runner`; conclusions `success` ×4 |
|
||||||
|
| **E** | **PASS** | live pod spec (§8) **and** the refusal in a job log: `/root/.cache/act/…/3.sh: line 2: docker: command not found` → `E-OK docker refused` |
|
||||||
|
| **F** | **PASS** | catalogue run: `image-pin gate OK — 53 templates, 0 unpinned images`; `--fast SKIPPED: image-resolvable, volume-persistence…`; resolvability/volume-persistence/canary output lines **0**, image-pull lines **0** |
|
||||||
|
| **G** | **PASS** | P6: PVC intact → runner list unchanged (`1 felhom-gates-runner`); state removed → **`1` and `2`**, i.e. a second registration. Orphan deleted, `204` |
|
||||||
|
|
||||||
|
**Scenario A note worth keeping:** the controller and agent runs prove more than "it ran" — their
|
||||||
|
reuse tallies in CI (**126 exact / 6 suffix / 1 cross-repo**; **88 / 1 / 1**) are *identical* to the
|
||||||
|
local pre-push hook's. **CI and the hook agree.** Per §12 a disagreement would have outranked this
|
||||||
|
whole task; there is none.
|
||||||
|
|
||||||
|
## 6. Scenario C evidence in full
|
||||||
|
|
||||||
|
The run was genuinely red for a genuinely broken gate, and the message is the observable — not a
|
||||||
|
tick in a UI:
|
||||||
|
|
||||||
```
|
```
|
||||||
felhom.eu core.hooksPath = .githooks
|
CONVICTED: hostinstall
|
||||||
felhom-controller core.hooksPath = .githooks
|
RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab
|
||||||
felhom-agent core.hooksPath = .githooks
|
🏁 Job failed
|
||||||
app-catalog-felhom.eu core.hooksPath = .githooks
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## 9. Hub deployment
|
The Resend id is the **provider's accepted-id**, returned by `api.resend.com` for a message
|
||||||
|
`monitoring@felhom.eu → admin@felhom.eu`, subject `[felhom CI] gates FAILED in admin/felhom.eu`,
|
||||||
|
carrying the repo, the commit, a link to the run, and the line that matters most: *if the local
|
||||||
|
pre-push hook was green for this commit, CI and the hook disagree, and that outranks whatever the
|
||||||
|
push was for.*
|
||||||
|
|
||||||
- **Version:** `gitea.dooplex.hu/admin/felhom-hub:0.87.0`, built and pushed from
|
**The key never appeared in the log** — Gitea masked it (`RESEND_API_KEY: ***`) and a grep for the
|
||||||
`/mnt/5_hdd/felhom.eu/build/felhom-hub` after the clean-tree gate (empty `git status --porcelain`,
|
key prefix across the run log returns **0 lines**.
|
||||||
`HEAD == origin/main == 4707be7`) and a green `go build ./... && go vet ./... && go test ./...`.
|
|
||||||
- **ArgoCD:** `manifests/hub.yaml` bumped in git (`8d9b78c`), hard-refresh → `OutOfSync`, then a
|
## 7. Runner deployment
|
||||||
**deliberate** sync. No `kubectl set image` at any point.
|
|
||||||
Final: `sync=Synced health=Healthy`, revision `8d9b78c153ee…`.
|
- **GitOps only.** `act-runner.yaml` committed to `homelab-manifests` `main` (`c51b77c`), pushed,
|
||||||
- **Rollout:** `deployment "hub" successfully rolled out`; pod `hub-67774ccf4f-74wwx` `1/1 Running`.
|
then a **deliberate ArgoCD sync** of the `gitea` Application. No `kubectl apply`, no
|
||||||
- **Running image:** `gitea.dooplex.hu/admin/felhom-hub:0.87.0`.
|
`kubectl set image` at any point.
|
||||||
- **Pod log excerpt:**
|
- **ArgoCD:** `sync=Synced health=Healthy`, revision `c51b77c68f10…`.
|
||||||
|
- **Pod:** `act-runner-…` `1/1 Running`; PVC `act-runner-data` **Bound**, 5Gi, `longhorn`.
|
||||||
|
- **Running image:** `gitea.dooplex.hu/admin/felhom-act-runner:0.1.0`, verified by deleting the local
|
||||||
|
copy and **re-pulling from the registry** (`Python 3.12.13`, `git version 2.52.0`) rather than
|
||||||
|
trusting the push's own output.
|
||||||
|
- **Online in Gitea's runner list**, with its label:
|
||||||
```
|
```
|
||||||
[INFO] Storage fill checker initialized: warn=90% crit=95%, 8 ok seeded, 0 already-breached left unseeded, 3 root-backed excluded
|
id=2 name=felhom-gates-runner status=online labels=[felhom-gates]
|
||||||
[INFO] Offsite checker initialized: fill warn=90% crit=95%, stale after 48h0m0s, 3 ok-seeded
|
|
||||||
[INFO] Listening on :8080
|
|
||||||
```
|
```
|
||||||
- **Live read of the Setup tab — method: endpoint-level** (`curl` over the hub ClusterIP with Basic
|
Owner-scoped (`owner_id=1`, `repo_id=0`) — which is why one registration serves four repos.
|
||||||
auth; browser automation is not available on DooPlex). `GET /customers/peti-felhom`, 105 164 bytes.
|
|
||||||
The Setup Command card renders:
|
|
||||||
|
|
||||||
> Day-0 host bootstrap. The command always fetches the **current** felhom-host-install.sh from
|
## 8. Security posture — shown, not asserted
|
||||||
> felhom.eu — there is no version to pick here. Run it on a freshly-PVE-installed Proxmox
|
|
||||||
> **host** as root …
|
|
||||||
|
|
||||||
**Zero occurrences of `1.19.0` or `1.22.0` anywhere on the page.** The page also served
|
From the **live** pod spec:
|
||||||
`style.css?v=0.87.0`, independently confirming the new image is the one rendering.
|
|
||||||
|
```
|
||||||
|
privileged : False
|
||||||
|
allowPrivilegeEscalation : False
|
||||||
|
capabilities.drop : ['ALL']
|
||||||
|
automountServiceAccountToken : False
|
||||||
|
hostNetwork/hostPID/hostIPC : None None None
|
||||||
|
nodeSelector : None
|
||||||
|
volumes : [('data', ['persistentVolumeClaim'])]
|
||||||
|
hostPath volumes : NONE
|
||||||
|
docker socket mounts : NONE
|
||||||
|
serviceaccount token mounts : NONE
|
||||||
|
resources : requests 50m/128Mi, limits 500m/512Mi
|
||||||
|
```
|
||||||
|
|
||||||
|
And the negative proven rather than inferred — from a job log:
|
||||||
|
|
||||||
|
```
|
||||||
|
/root/.cache/act/187d434f3fdbb3b4/act/workflow/3.sh: line 2: docker: command not found
|
||||||
|
E-OK docker refused (output above is the refusal)
|
||||||
|
```
|
||||||
|
|
||||||
|
Resources are **half of Gitea's limits** (Gitea: 1 CPU / 1Gi) and a fifth of its requests,
|
||||||
|
deliberately: Gitea shares this node and holds every repository. The heaviest entry point was
|
||||||
|
measured at **11.9 MiB RSS, 0.98 s wall**.
|
||||||
|
|
||||||
|
## 9. Teardown — all three layers
|
||||||
|
|
||||||
|
1. **Workload** — nothing to remove: the runner is the deliverable and stays. No scratch Deployment
|
||||||
|
or PVC was created beyond it.
|
||||||
|
2. **Host** — no scratch Longhorn volume was created beyond the runner's own 5Gi PVC, which is
|
||||||
|
intended and in git. Local Docker images (`gitea/act_runner:0.6.1` base and the built
|
||||||
|
`felhom-act-runner:0.1.0`) remain in DooPlex's local Docker store; they are ordinary build
|
||||||
|
artifacts and **no global prune was run** (forbidden on this host).
|
||||||
|
3. **The Gitea-side registration records — the layer this project has missed four times.**
|
||||||
|
- The **P6b orphan** (`id=1`) was deleted **immediately when it was created**, not left for the
|
||||||
|
end: `DELETE /api/v1/admin/actions/runners/1 → 204`.
|
||||||
|
- **Final runner list: exactly one, the intended live runner** — see §12 for the closing check.
|
||||||
|
- The **temporary probe workflow** `.gitea/workflows/probe.yml` was deleted from `main` in
|
||||||
|
`666a34d`, the same commit that added the real one.
|
||||||
|
- The **admin API access token** minted for this session (used to read run conclusions and delete
|
||||||
|
the orphan) is revoked at the end of the session; it is stored out-of-band and appears in no
|
||||||
|
file.
|
||||||
|
|
||||||
## 10. `OPEN-ITEMS.md` rows opened / closed / re-ranked
|
## 10. `OPEN-ITEMS.md` rows opened / closed / re-ranked
|
||||||
|
|
||||||
- **R-94 → CLOSED**, all three legs. Leg (a) closed **by deletion, not derivation**, with the reason
|
- **R-168 → SHIPPED**, with the evidence in the row: the runner, the one-registration-four-repos
|
||||||
recorded in the row: the Setup command fetches the installer at run time from a website that
|
measurement, the two probes that changed the design (P2, P6), and the demonstrated alarm with its
|
||||||
git-syncs `main` every 30 s (R-110), so no build-time value in the hub can be true.
|
accepted-id.
|
||||||
- **R-29 → leg (a) CLOSED** (`c432f70`, its own reviewed diff); **leg (b) HALF-SHIPPED**. The census
|
- **R-29 → CLOSED**, both halves — and **closed on the demonstrated alarm, not on a green run**, as
|
||||||
is written into the row: thirteen gates; every gate a `CLAUDE.md` names was green, and two of the
|
the task required. The class it opened (*a gate that exists, is red, and is invoked by nothing*) is
|
||||||
four nobody names were red. Stays open for the automatic half.
|
answered at both ends.
|
||||||
- **R-161 → annotated**: `--fast` + hook shipped; the automatic half now points at its successor row.
|
- **R-161 → updated**: its automatic half now exists for the static gate. Its *original* scope, the
|
||||||
- **R-168 → NEW.** `grep -oE "R-[0-9]+" OPEN-ITEMS.md | sort -u -t- -k2 -n | tail` established
|
runtime volume-persistence gate, is deliberately still **not** automatic and should stay that way.
|
||||||
**R-167 as the highest in use**, so R-168 was free. Gitea Actions runner for CI; owner CC; state
|
- **R-169 → NEW.** `grep -oE "R-[0-9]+" | sort -u -t- -k2 -n | tail` established **R-168 as the
|
||||||
**BLOCKED** on its own spike. It records what was measured on 2026-08-02 — Gitea **1.26.2**,
|
highest in use**, so R-169 was free. *CI can only report, because there is no gate in the road* —
|
||||||
Actions enabled on all four repos, **0 runners, 0 workflow runs, 0 branch protections** — and the
|
making it blocking needs branch protection plus a PR workflow, both of which change how the
|
||||||
consequence: with trunk-based direct-to-`main` pushes there is no merge for a status check to gate,
|
operator works. **WAITING-ON-OPERATOR**, owner: operator. The row states the cost honestly (for a
|
||||||
so **CI here can detect but not block**. Its two spike questions: host-mode runner versus a
|
single-operator project a PR workflow may be worse than the disease) and names the real gap: the
|
||||||
privileged Docker-in-Docker sidecar on DooPlex (Tier 2 — it *is* the recovery chain), and whether
|
window between a `--no-verify` push landing and the operator reading the alarm.
|
||||||
the workflow can avoid JavaScript actions.
|
- `ROADMAP.md`: R-168 collapsed to its one-liner, R-29 closed, R-169 added.
|
||||||
- `ROADMAP.md`: R-94 collapsed to its one-liner, R-29 updated, R-168 added.
|
|
||||||
- **`architecture/00-capability-map.md` — N/A, stated rather than silently skipped.** No
|
- **`architecture/00-capability-map.md` — N/A, stated rather than silently skipped.** No
|
||||||
customer-visible or platform capability changed: this task alters tooling, one operator-facing
|
customer-visible or platform capability changed; this is developer tooling.
|
||||||
sentence, and documentation.
|
|
||||||
|
|
||||||
## 11. Teardown
|
## 11. Was `--no-verify` used?
|
||||||
|
|
||||||
**Provisioned nothing.** No VM, no guest, no container and no host was created or modified. The only
|
**Yes — three times, all on `felhom.eu`, all deliberate and all disclosed.**
|
||||||
scratch artifacts were three throwaway git repos under the session scratchpad (the `core.hooksPath`
|
|
||||||
probe, the Scenario C clone, and its bare origin) — disposable, and touching nothing real.
|
|
||||||
|
|
||||||
## 12. Was `--no-verify` used?
|
| Commit | Why |
|
||||||
|
|---|---|
|
||||||
|
| `3252d51` | **Scenario B itself.** The commit had to carry a broken gate, and pushing it with `--no-verify` *is* the bypass CI exists to catch. The local hook would have refused it — that is the point. |
|
||||||
|
| `dd13f63` | The tree was still deliberately broken; pushing the alarm step required bypassing the hook that was correctly objecting. |
|
||||||
|
| `f7dbc33`, `9530de7` | Same, iterating on the alarm while the tree was still red on purpose. |
|
||||||
|
|
||||||
**Once, deliberately, and only inside the throwaway Scenario C clone** whose `origin` was a scratch
|
Every other push in this session ran its pre-push hook and passed. The breakage was reverted in
|
||||||
bare repo — that use *was* the test of the escape hatch (§7, C2). **It was never used against a real
|
`4cc1238`, and the hook passed on that push.
|
||||||
repository.** All five real pushes ran their pre-push hook, and it passed each time.
|
|
||||||
|
|
||||||
## 13. Observations — noticed, recorded, NOT acted on
|
## 12. Observations — noticed, recorded, NOT acted on
|
||||||
|
|
||||||
1. **The census held, with one thing worth stating plainly.** Everything §4.9 said was reproduced:
|
1. **CI and the pre-push hook never disagreed.** §12 said a disagreement would outrank the CI work;
|
||||||
`hostinstall_gates` RED, `reuse_refs_check` RED on all four repos with exactly 13 findings,
|
there was none. The controller's and agent's reuse tallies in CI match the local run exactly. The
|
||||||
`docker_run_volume_path_gate` RED with exactly one finding, and every gate named by a `CLAUDE.md`
|
one place they *would* have disagreed was designed out rather than discovered: without a sibling
|
||||||
green. Since the "two of the four" headline can be read as "two red gates", the arrival state was
|
`felhom.eu` clone, both entry points fail closed with `gate is MISSING`.
|
||||||
in fact **three** red scripts — and they were exactly the two named nowhere plus the one named
|
2. **A near-miss that outranks a probe result.** My first census query reported Actions enabled on
|
||||||
only in `REUSE.md:284`. The correlation is unchanged and slightly stronger than the headline.
|
five *unrelated* repos — a baseline drift big enough to change the task. It was a `| tail -5`
|
||||||
2. **DEVIATION — `scripts/felhom-host-install.sh` was edited, which §12 fenced.** Comment only;
|
inside my own helper truncating the output. The measurement was never wrong; the instrument was.
|
||||||
`SCRIPT_VERSION="1.22.0"` is untouched and no behaviour changed. Its comment asserted that
|
This is the same class as the `go test -run` filter from session 1, and it is now a `CLAUDE.md`
|
||||||
`hostinstall_gates.py` "asserts the two stay equal" — an invariant this task deleted. Leaving it
|
rule: **an instrument that can silently drop results is not a measurement.**
|
||||||
would have shipped exactly the class `CLAUDE.md` names ("a comment asserting an invariant needs a
|
3. **Two failures on the way to the alarm, both worth their write-up** — `curl: command not found`
|
||||||
test pinning it, or it is a wish"), so the fence — whose stated reason is that `SCRIPT_VERSION` is
|
(the image is minimal on purpose; fixed with `urllib`, not a bigger image) and **Cloudflare 403
|
||||||
the single source and stays — was read as fencing the act, not the object. Recorded here so the
|
error 1010** on `api.resend.com`, which blocks the default `Python-urllib` User-Agent. The second
|
||||||
decision is visible rather than absorbed silently.
|
matters because **it looks exactly like an auth failure**: the next person to see a 403 from
|
||||||
3. **DEVIATION — `.gitignore` gained `__pycache__/`.** Not in the spec. The new fixture tests import
|
Resend should check the User-Agent before rotating a key. Both are in the spike doc and in
|
||||||
the checker by path, which writes bytecode into `scripts/`, and an untracked directory there would
|
`CONTEXT.md` S-9.
|
||||||
have failed the clean-tree gate before every future build.
|
4. **A full clone is slow enough to notice** — the P3 probe's `git clone` of `felhom.eu` took ~48 s.
|
||||||
4. **A `go test -run` pattern that matches nothing prints `ok` and exits 0.** My first attempt at
|
The shipped workflows use `git fetch --depth 1 <sha>` instead, which is both faster and pinned to
|
||||||
red-proof #3 used `-run TestCustomerUnified`, which matches no test in `render_test.go`, and
|
the exact pushed commit. Not a finding, but the reason the shipped form differs from the probe's.
|
||||||
`go test` reported `ok … 0.062s`. Caught and re-run against the real name. This is another
|
5. **Not acted on, out of scope:** Gitea's own mailer is configured and enabled but was never proven
|
||||||
instance of the "an absent line is not evidence" class already in `CLAUDE.md`, and it is worth its
|
to deliver (§2). If it *is* broken, nothing else in Gitea that relies on email works either —
|
||||||
own line because it makes a **red-proof itself hollow** — the one place a false green costs most.
|
worth a five-minute test email from the admin UI at some point. It does not affect the CI alarm,
|
||||||
5. **`hub/CHANGELOG.md` and `scripts/CHANGELOG.md` still contain historical `hostInstallVersion`
|
which does not use it.
|
||||||
mentions.** Correct and left alone: the rewritten gate matches code shapes, never prose, so the
|
6. **Not acted on, out of scope:** `homelab-manifests` has no gate entry point and no pre-push hook,
|
||||||
history is safe to keep and the gate stays green.
|
unlike the four product repos, and now no CI workflow either. It is the repo that deploys the
|
||||||
6. **Not acted on, out of scope.** `manifest_bearer_gate.py` prints
|
cluster, so that asymmetry may deserve a look — but adding one was not in this task, and the repo
|
||||||
`manifests/felhom.secret.yaml:39 KNOWN-BACKLOG committed secret …` on every run and still exits
|
has no gates to run yet.
|
||||||
0. It is a tracked backlog item (`documentation/runbooks/secrets.md`) and the gate deliberately
|
|
||||||
does not fail on it — but it means every green `repo_gates.py` run now carries a line that reads
|
|
||||||
like a finding, in front of a reader who did not previously see it. Left untouched, per "do not
|
|
||||||
tighten or refactor a currently-green gate".
|
|
||||||
|
|||||||
@@ -37,11 +37,14 @@ lost: the last good copy is kept intact. *(R-163)*
|
|||||||
**When that happens, only one page says so** — no email, no alert. The page that answers "is this app
|
**When that happens, only one page says so** — no email, no alert. The page that answers "is this app
|
||||||
backed up?" is the one that stays silent. *(R-158)*
|
backed up?" is the one that stays silent. *(R-158)*
|
||||||
|
|
||||||
**The checks that catch this still can't stop a change on their own — but they now try.** Every
|
**The checks now have two nets, and the second one emails you.** Every repository has one command
|
||||||
repository has one command that runs all of its checks, and it runs by itself before every push, so a
|
that runs all of its checks; it runs by itself before every push and refuses a push that fails. That
|
||||||
push carrying a broken check is refused rather than reported. Two limits are real: it has to be
|
one lives on the workstation and can be skipped. So the build server now runs the same checks again,
|
||||||
switched on once in each copy of the code, and it can be skipped on purpose. The version that can be
|
on a machine that does not care who pushed or what they typed — and **when they fail it sends you an
|
||||||
neither is a build server, which we do not have yet. *(R-29, R-161, R-168)*
|
email**, because a red mark on a page nobody watches is not a warning. Proven with a real broken
|
||||||
|
change, not assumed. The one thing it still cannot do is *stop* the change: every change here goes
|
||||||
|
straight to the main copy with no review step, so there is no point in the road for it to stand at.
|
||||||
|
It notices, quickly, and tells you. *(R-29, R-161, R-168, R-169)*
|
||||||
|
|
||||||
## What we're working on
|
## What we're working on
|
||||||
|
|
||||||
@@ -62,12 +65,13 @@ neither is a build server, which we do not have yet. *(R-29, R-161, R-168)*
|
|||||||
|
|
||||||
## Changed since last update
|
## Changed since last update
|
||||||
|
|
||||||
- **2026-08-02** — Fixed: thirteen mechanical checks had built up across the four repositories and
|
- **2026-08-02** — Fixed and then doubled: thirteen mechanical checks had built up across the four
|
||||||
nothing ran most of them. Two were failing quietly, one since 14 July; neither did harm, but nothing
|
repositories and nothing ran most of them; two were failing quietly, one since 14 July. Both fixed.
|
||||||
would have said so. Both fixed; every repository now has one command that runs all its checks, and
|
Every repository now has one command that runs all its checks, it runs automatically before every
|
||||||
it runs automatically before every push. One of the two was the setup page claiming the installer
|
push, and the build server re-runs it after every push and **emails you when it fails** — proven
|
||||||
was version 1.19.0 while machines got 1.22.0 — the number is now **gone** rather than corrected,
|
with a real broken change. One of the two failures was the setup page claiming the installer was
|
||||||
because the page cannot know it: the installer is fetched fresh each run, so any number there is a
|
version 1.19.0 while machines got 1.22.0; that number is now **gone** rather than corrected,
|
||||||
|
because the page cannot know it — the installer is fetched fresh each run, so any number there is a
|
||||||
guess. A check refuses to let one be added back.
|
guess. A check refuses to let one be added back.
|
||||||
- **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That
|
- **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That
|
||||||
changes the disk layout, so it happens before any machine is installed outside the house.
|
changes the disk layout, so it happens before any machine is installed outside the house.
|
||||||
|
|||||||
@@ -0,0 +1,220 @@
|
|||||||
|
# SPIKE — a Gitea Actions runner for the felhom gate entry points (R-168, 2026-08-02)
|
||||||
|
|
||||||
|
**Verdict: the mechanism works and is SHIPPED.** All six probes answered; none produced a STOP.
|
||||||
|
The runner is unprivileged, host-mode, one registration for all four repos, and a failed run now
|
||||||
|
sends its own alarm — measured, not assumed.
|
||||||
|
|
||||||
|
**Context.** Session 1 (same day) gave every felhom repo one gate entry point and a
|
||||||
|
`.githooks/pre-push` that runs it and refuses a failing push. That hook is per-clone and
|
||||||
|
`git push --no-verify` skips it, so nothing independent of the person pushing ever saw whether the
|
||||||
|
gates passed. This spike built the independent half.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Arrival state, re-confirmed live (all matched the anchor)
|
||||||
|
|
||||||
|
| Fact | Expected | Measured |
|
||||||
|
|---|---|---|
|
||||||
|
| Gitea version | 1.26.2 | **1.26.2** |
|
||||||
|
| Actions enabled | all five repos | **all five** (`repo_unit` type 10 present on each) |
|
||||||
|
| Runners registered | 0 | **0** (`action_runner` empty) |
|
||||||
|
| Workflow runs, ever | 0 | **0** (`action_run` empty) |
|
||||||
|
| Branch protections | 0 | **0** (`protected_branch` empty) |
|
||||||
|
| `.gitea/` directory | none | **none in any of the five** |
|
||||||
|
| ArgoCD `gitea` app | `path: gitea-system`, auto-sync off | **confirmed**, Synced at `420e819`, Healthy |
|
||||||
|
|
||||||
|
Repo visibility, which decided P3's design: the four product repos are **public**;
|
||||||
|
`homelab-manifests` is private.
|
||||||
|
|
||||||
|
> **A near-miss worth recording.** My first census query reported Actions enabled on only five
|
||||||
|
> *unrelated* repos, which looked like a baseline drift big enough to change the task. It was a
|
||||||
|
> `| tail -5` inside my own query helper truncating the result. The measurement was never wrong;
|
||||||
|
> the instrument was. Re-run without the pipe, all five target repos had the Actions unit. **A tool
|
||||||
|
> that silently drops rows is indistinguishable from a finding** — the same class as the
|
||||||
|
> `go test -run` filter that matches nothing and prints `ok`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P1 — does a registered runner pick up a job at all?
|
||||||
|
|
||||||
|
**Method.** Registered one runner at **owner scope** for `admin` (`repo_id=0`), label `felhom-gates`,
|
||||||
|
via a token minted with `gitea actions generate-runner-token --scope admin` and stored out-of-band as
|
||||||
|
`Secret/act-runner-registration`. Pushed a temporary `.gitea/workflows/probe.yml` with one
|
||||||
|
`run: echo` step to `felhom.eu`.
|
||||||
|
|
||||||
|
**Measured.** Run #1 appeared against the pushed commit, was claimed by `felhom-gates-runner`, and
|
||||||
|
finished **success**:
|
||||||
|
|
||||||
|
```
|
||||||
|
felhom-gates-runner(version:v0.6.1) received task 1 of job probe, be triggered by event: push
|
||||||
|
P1-OK runner picked up the job
|
||||||
|
🏁 Job succeeded
|
||||||
|
```
|
||||||
|
|
||||||
|
**Ruling: PASS.** No STOP.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P2 — can the stock image run our checks without a container runtime?
|
||||||
|
|
||||||
|
**Method.** Inspected `gitea/act_runner:0.6.1` directly, then re-measured *inside a job* — because
|
||||||
|
what matters is what the **job** sees, not what the image contains, and in host mode those are the
|
||||||
|
same thing only if host mode is really in effect.
|
||||||
|
|
||||||
|
**Measured.** Stock image (Alpine Linux v3.23):
|
||||||
|
|
||||||
|
```
|
||||||
|
python3: sh: python3: not found ABSENT
|
||||||
|
git: git version 2.52.0 PRESENT
|
||||||
|
```
|
||||||
|
|
||||||
|
In-job, after building the minimal image:
|
||||||
|
|
||||||
|
```
|
||||||
|
P2 python3: Python 3.12.13
|
||||||
|
P2 git: git version 2.52.0
|
||||||
|
```
|
||||||
|
|
||||||
|
**Ruling: the known branch, not a failure.** Host mode works; only `python3` was missing. Built
|
||||||
|
`gitea.dooplex.hu/admin/felhom-act-runner:0.1.0` = stock (pinned) + `python3`, nothing else, from
|
||||||
|
`homelab-manifests/gitea-system/act-runner/Dockerfile`. Verified by deleting the local copy and
|
||||||
|
**re-pulling from the registry** rather than trusting the push's own output.
|
||||||
|
|
||||||
|
**Host mode was never in doubt, so the privileged/dind pattern was never reached for** (§5/§12 of the
|
||||||
|
task, and the reason is written into `act-runner.yaml`: DooPlex is Tier 2 and *is* the recovery
|
||||||
|
chain).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P3 — can a workflow obtain the source without JavaScript actions?
|
||||||
|
|
||||||
|
**Method.** No `uses:` anywhere. A plain `run:` step clones from the **in-cluster** Gitea Service
|
||||||
|
(`http://gitea.gitea-system.svc.cluster.local:3000`) — no ingress, no TLS hop, no geo rule — and
|
||||||
|
checks out `$GITHUB_SHA`. The four product repos are public, so no credential is needed at all.
|
||||||
|
|
||||||
|
**Measured.**
|
||||||
|
|
||||||
|
```
|
||||||
|
P3 pushed sha = bbd62319096a1fb92f72218a28756e6fa101e87b
|
||||||
|
P3 checked-out sha = bbd62319096a1fb92f72218a28756e6fa101e87b
|
||||||
|
P3-OK checkout equals pushed commit
|
||||||
|
```
|
||||||
|
|
||||||
|
**Ruling: PASS.** The shipped workflows use the tighter form — `git init` + `git fetch --depth 1
|
||||||
|
origin $GITHUB_SHA` + `checkout FETCH_HEAD` — which is both shallow and pinned to the **exact pushed
|
||||||
|
commit** rather than the branch tip, so two racing pushes cannot make a run test the wrong tree. The
|
||||||
|
full clone in the probe took ~48 s for `felhom.eu`; the shallow fetch is materially faster.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P4 — does one registration serve all four repos?
|
||||||
|
|
||||||
|
**Method.** The registration is owner-scoped (`owner_id=1`, `repo_id=0`). Pushed the real workflow to
|
||||||
|
all four repos and read which runner claimed each task.
|
||||||
|
|
||||||
|
**Measured.**
|
||||||
|
|
||||||
|
```
|
||||||
|
task repo runner_id runner
|
||||||
|
7 felhom.eu 2 felhom-gates-runner
|
||||||
|
8 felhom-controller 2 felhom-gates-runner
|
||||||
|
9 felhom-agent 2 felhom-gates-runner
|
||||||
|
10 app-catalog-felhom.eu 2 felhom-gates-runner
|
||||||
|
```
|
||||||
|
|
||||||
|
**Ruling: PASS.** One registration, four repos. No per-repo registration needed.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P5 — does a failed run signal anything outside the UI? *(the probe that decided Part 4)*
|
||||||
|
|
||||||
|
**Method.** Broke a gate deliberately and pushed it **with `--no-verify`**, which is precisely the
|
||||||
|
bypass CI exists to catch. Run #3 concluded `failure`. Then looked for any outbound signal in the
|
||||||
|
ten minutes around it: Gitea pod logs filtered for mail/SMTP/notification activity, and the
|
||||||
|
`notification` table.
|
||||||
|
|
||||||
|
**Measured — nothing left the machine.**
|
||||||
|
|
||||||
|
```
|
||||||
|
gitea pod logs, mail/smtp/notif lines since the failure : (none)
|
||||||
|
notification rows created in the last 10 minutes : 0
|
||||||
|
```
|
||||||
|
|
||||||
|
Gitea's mailer is *configured and enabled* (`[mailer] ENABLED = true`, Gmail SMTP, FROM is the
|
||||||
|
operator's own address) and the sole user `admin` is active with
|
||||||
|
`email_notifications_preference = enabled` — so this is not a disabled-mailer artefact at the
|
||||||
|
config level.
|
||||||
|
|
||||||
|
**Honest limit on this measurement.** I did not independently prove that Gitea's SMTP path can
|
||||||
|
*deliver*, so "no mail" cannot be split with certainty between *"1.26.2 has no action-failure
|
||||||
|
notification"* and *"the mailer is broken"*. That distinction does not change the design: the alarm
|
||||||
|
is built on **Resend**, a different path entirely, and Scenario C proves that path end to end. It is
|
||||||
|
recorded here so nobody later reads this probe as a clean bill of health for Gitea's mailer.
|
||||||
|
|
||||||
|
**Ruling: Part 4 applies — build the alarm.** A red tick in a web UI nobody watches is exactly the
|
||||||
|
defect R-29 filed, rebuilt one layer up.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## P6 — does the runner need persistent state?
|
||||||
|
|
||||||
|
**Method.** Two measurements, not one, because "it survived a restart" and "the PVC is load-bearing"
|
||||||
|
are different claims.
|
||||||
|
|
||||||
|
- **P6a — restart with the PVC intact:** deleted the pod, let it come back.
|
||||||
|
- **P6b — restart with state lost:** removed `/data/.runner` (exactly what ephemeral storage would
|
||||||
|
do), deleted the pod, let it come back.
|
||||||
|
|
||||||
|
**Measured.**
|
||||||
|
|
||||||
|
| | runner list before | after | runner log |
|
||||||
|
|---|---|---|---|
|
||||||
|
| P6a | `1 felhom-gates-runner` | `1 felhom-gates-runner` | no re-registration; straight to `Starting runner daemon` |
|
||||||
|
| P6b | `1 felhom-gates-runner` | `1 …` **and** `2 …` | `Registering runner…` → `Runner registered successfully` |
|
||||||
|
|
||||||
|
**Ruling: the PVC is load-bearing.** Without persistence, every restart mints a new registration and
|
||||||
|
leaves the previous one behind as a permanently-offline record — the runner list would silently fill
|
||||||
|
with corpses. Shipped with a 5 Gi Longhorn PVC (largest repo checked out is ~125 MiB).
|
||||||
|
|
||||||
|
**The P6b casualty was cleaned up immediately**, not left for the teardown section:
|
||||||
|
`DELETE /api/v1/admin/actions/runners/1 → 204`, leaving exactly one runner.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What the probes changed about the design
|
||||||
|
|
||||||
|
1. **A custom image** (P2) — stock + `python3` only, base pinned by tag.
|
||||||
|
2. **Shallow fetch of the exact SHA** rather than a clone of the branch tip (P3).
|
||||||
|
3. **A PVC** (P6), with the reason recorded in the manifest.
|
||||||
|
4. **A self-sent alarm** (P5) — the task's whole second half.
|
||||||
|
5. **A sibling clone for two of the four repos** — discovered while writing the workflows, not by a
|
||||||
|
probe. `controller_gates.py` and `agent_gates.py` invoke the shared `reuse_refs_check.py` that
|
||||||
|
lives in the `felhom.eu` clone next door and is deliberately never copied, and both repos'
|
||||||
|
`REUSE.md` files cite a path that lives in the hub. Without the sibling, CI would have failed
|
||||||
|
**closed** — correctly, but for the wrong reason. CI now reproduces the workspace's sibling
|
||||||
|
layout, and the resulting tallies match the local hook exactly (controller 126 exact / 6 suffix /
|
||||||
|
1 cross-repo; agent 88 / 1 / 1). **CI and the hook agree.**
|
||||||
|
|
||||||
|
## Two things that failed on the way, and why they are recorded
|
||||||
|
|
||||||
|
Both were caught because the step failed **loudly**; either would have shipped as a silent
|
||||||
|
non-alarm if the step had swallowed its exit code.
|
||||||
|
|
||||||
|
- **`curl: command not found`** — the first alarm used `curl`, which the deliberately minimal image
|
||||||
|
does not carry. Fixed by using `python3` + `urllib` rather than by growing the image: reaching for
|
||||||
|
a bigger base to send one HTTP request is the wrong trade, and every added tool becomes something
|
||||||
|
the next person assumes is load-bearing.
|
||||||
|
- **Cloudflare 403, error 1010** — `api.resend.com` sits behind Cloudflare, which blocks the default
|
||||||
|
`Python-urllib/3.x` User-Agent. **This failure looks exactly like an auth failure and is not one**,
|
||||||
|
which is the reason it is written down: the next person to see a 403 from Resend should check the
|
||||||
|
agent before rotating a key. Verified the fix from inside the runner image with a deliberately
|
||||||
|
invalid payload — with a User-Agent set, Resend answers `422 missing_required_field`, i.e. the
|
||||||
|
request reaches the API rather than the CDN.
|
||||||
|
|
||||||
|
## Standing limit, stated so it is not mistaken for something it is not
|
||||||
|
|
||||||
|
**CI here DETECTS. It does not BLOCK.** Every felhom repo pushes straight to `main` with no pull
|
||||||
|
request, so there is no merge for a status check to gate. This is not a gap in the runner — there is
|
||||||
|
no gate in the road. The refusing half is the local pre-push hook; this half notices when that hook
|
||||||
|
was skipped. Making CI blocking requires branch protection and a pull-request workflow, which is a
|
||||||
|
change to how the operator works and is **their** decision, not this task's.
|
||||||
@@ -24,7 +24,7 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha
|
|||||||
| **R-113** | ~~**The drive-absent gate CANNOT FIRE on device loss — E-2b's alarm is wired to an unreachable condition.**~~ `planDriveGates` (`intermediary.go:216-262`) computes presence by OR-ing `d.BoundUnderParent` into `present[GuestPath]`, and the agent derives `BoundUnderParent` from `GuestSeesMount()` — *"is this path a mount target in the guest's `/proc/<pid>/mountinfo`"* (`localapi/disks.go:210`). The raw drive mount is a **device-bound systemd unit** and dies with the device; **the agent's own bind under the shared parent is NOT device-bound, so its mountinfo entry outlives the device**. The gate reads the surviving bind as "present" ⇒ no `Stop` action ⇒ `notifyDriveAbsent` never called | **SHIPPED + PROVEN-LIVE** (agent v0.114.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md`.** Target drive hot-detached on a fresh box running the SHIPPED agent 0.114.0 (from the Day-0 manifest, not a hand build). **The gate fired in 4 seconds** — E-2d measured ZERO over 4½ minutes — and `SetDisconnected` was reached. It fired on exactly the shape that defeated it: raw `/mnt/mentes` NOT mounted while the bind `/mnt/felhom-drives/mentes` still read `/dev/sdb[/felhom-data]`. **Over-correction guard PASSED:** with both drives present, 0 ABSENT lines and the target stayed healthy. **Note: C5 still FAILED — but on a different defect (→ R-116), not on this one.** **SESSION C IS NOW UNBLOCKED.** Agent **0.114.0 published and vouched** 2026-07-29 (sha `5e4c15ebee2d7583…`, round-trip GET verified; hub manifest read back showing it) — a fresh drill box now installs the fix under test instead of the bug. The golden was NOT re-baked and `min_agent` was NOT raised: the golden bakes the controller, not the agent, and controller v0.185.0 declares MinAgent 0.113.0, which 0.114.0 already satisfies. **FIX: `BoundUnderParent` is now a CONJUNCTION — bound under the parent AND the drive's raw host mount still mounted** (`devicePresent`, new `deviceCheck` seam), at BOTH `/disks` construction sites. The raw mount is the device-bound systemd unit that dies with the device; the agent's bind is not — so the raw mount IS the device signal, grounded in E-2d's measurement rather than inference. **Conjunction, deliberately:** the device half alone would regress boot ordering (raw mounts early, bind lands ~18 s later — that window must keep reading absent), so existing behaviour is byte-identical and only the unreachable case is closed. **Unknown is never absent** (`devicePresent("")` = true) — a false absent stops a working customer's apps. **Controller UNCHANGED, no MinAgent bump:** `BoundUnderParent` has exactly one functional consumer (`planDriveGates:226`); a new `DevicePresent` bool was rejected because absent-from-JSON decodes to `false`, which would have made every drive on an older agent read ABSENT. +6 tests (208→214), 4 red-proofs run and reverted. **Deployed to demo-felhom; over-correction guard verified in production** — raw mount present, drive still reads present, 10/10 apps untouched, no gate action, no false alarm. **demo-hp deliberately left on 0.113.0** (the spec scoped deploy to felhom-pve) — it still carries the bug. **⚠️ SESSION C BLOCKER: the hub Day-0 manifest vouches agent 0.113.0**, so a fresh drill box would install WITHOUT this fix and validate nothing — publish + vouch 0.114.0 first (R-111's trap, same shape). **Awaiting live proof of exactly one leg:** device loss → gate `Stop` → `SetDisconnected` → `backup_target_absent` on the wire. Original finding: **PROVEN LIVE 2026-07-29 on a fresh box.** Target drive hot-detached under a running agent; over 4½ min (budget was 60 s): the **agent** said `enrolled drive absent by UUID` every 20 s, the controller logged **0** `[gate]` lines, and the hub received **zero** events — no `backup_target_absent` **and no generic `storage_disconnected`**. Measured with the device gone: `/mnt/mentes2` NOT mounted, `/mnt/felhom-drives/mentes2` still `/dev/sdb[/felhom-data]`. **Not a virtualisation artefact** — the asymmetry is device-bound-mount vs manual-bind, identical on metal (caveat: proven on SCSI hot-detach; physical unplug not staged). **Mirror scenario reasoned, not observed:** both the specific and generic events come from the same `a.Stop` branch, so the generic one is equally unreachable. **Sixth instance of seam-built-but-never-wired** — E-2b wired the seam, to a condition that cannot occur. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.2 **Session C scope UNCHANGED by Session B** — R-114/R-112 shipped as controller v0.186.0 and neither touches the agent; the leg awaiting proof is still device loss → gate `Stop` → `SetDisconnected` → `backup_target_absent` on the wire. One rebuild now validates all three. | CC |
|
| **R-113** | ~~**The drive-absent gate CANNOT FIRE on device loss — E-2b's alarm is wired to an unreachable condition.**~~ `planDriveGates` (`intermediary.go:216-262`) computes presence by OR-ing `d.BoundUnderParent` into `present[GuestPath]`, and the agent derives `BoundUnderParent` from `GuestSeesMount()` — *"is this path a mount target in the guest's `/proc/<pid>/mountinfo`"* (`localapi/disks.go:210`). The raw drive mount is a **device-bound systemd unit** and dies with the device; **the agent's own bind under the shared parent is NOT device-bound, so its mountinfo entry outlives the device**. The gate reads the surviving bind as "present" ⇒ no `Stop` action ⇒ `notifyDriveAbsent` never called | **SHIPPED + PROVEN-LIVE** (agent v0.114.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md`.** Target drive hot-detached on a fresh box running the SHIPPED agent 0.114.0 (from the Day-0 manifest, not a hand build). **The gate fired in 4 seconds** — E-2d measured ZERO over 4½ minutes — and `SetDisconnected` was reached. It fired on exactly the shape that defeated it: raw `/mnt/mentes` NOT mounted while the bind `/mnt/felhom-drives/mentes` still read `/dev/sdb[/felhom-data]`. **Over-correction guard PASSED:** with both drives present, 0 ABSENT lines and the target stayed healthy. **Note: C5 still FAILED — but on a different defect (→ R-116), not on this one.** **SESSION C IS NOW UNBLOCKED.** Agent **0.114.0 published and vouched** 2026-07-29 (sha `5e4c15ebee2d7583…`, round-trip GET verified; hub manifest read back showing it) — a fresh drill box now installs the fix under test instead of the bug. The golden was NOT re-baked and `min_agent` was NOT raised: the golden bakes the controller, not the agent, and controller v0.185.0 declares MinAgent 0.113.0, which 0.114.0 already satisfies. **FIX: `BoundUnderParent` is now a CONJUNCTION — bound under the parent AND the drive's raw host mount still mounted** (`devicePresent`, new `deviceCheck` seam), at BOTH `/disks` construction sites. The raw mount is the device-bound systemd unit that dies with the device; the agent's bind is not — so the raw mount IS the device signal, grounded in E-2d's measurement rather than inference. **Conjunction, deliberately:** the device half alone would regress boot ordering (raw mounts early, bind lands ~18 s later — that window must keep reading absent), so existing behaviour is byte-identical and only the unreachable case is closed. **Unknown is never absent** (`devicePresent("")` = true) — a false absent stops a working customer's apps. **Controller UNCHANGED, no MinAgent bump:** `BoundUnderParent` has exactly one functional consumer (`planDriveGates:226`); a new `DevicePresent` bool was rejected because absent-from-JSON decodes to `false`, which would have made every drive on an older agent read ABSENT. +6 tests (208→214), 4 red-proofs run and reverted. **Deployed to demo-felhom; over-correction guard verified in production** — raw mount present, drive still reads present, 10/10 apps untouched, no gate action, no false alarm. **demo-hp deliberately left on 0.113.0** (the spec scoped deploy to felhom-pve) — it still carries the bug. **⚠️ SESSION C BLOCKER: the hub Day-0 manifest vouches agent 0.113.0**, so a fresh drill box would install WITHOUT this fix and validate nothing — publish + vouch 0.114.0 first (R-111's trap, same shape). **Awaiting live proof of exactly one leg:** device loss → gate `Stop` → `SetDisconnected` → `backup_target_absent` on the wire. Original finding: **PROVEN LIVE 2026-07-29 on a fresh box.** Target drive hot-detached under a running agent; over 4½ min (budget was 60 s): the **agent** said `enrolled drive absent by UUID` every 20 s, the controller logged **0** `[gate]` lines, and the hub received **zero** events — no `backup_target_absent` **and no generic `storage_disconnected`**. Measured with the device gone: `/mnt/mentes2` NOT mounted, `/mnt/felhom-drives/mentes2` still `/dev/sdb[/felhom-data]`. **Not a virtualisation artefact** — the asymmetry is device-bound-mount vs manual-bind, identical on metal (caveat: proven on SCSI hot-detach; physical unplug not staged). **Mirror scenario reasoned, not observed:** both the specific and generic events come from the same `a.Stop` branch, so the generic one is equally unreachable. **Sixth instance of seam-built-but-never-wired** — E-2b wired the seam, to a condition that cannot occur. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.2 **Session C scope UNCHANGED by Session B** — R-114/R-112 shipped as controller v0.186.0 and neither touches the agent; the leg awaiting proof is still device loss → gate `Stop` → `SetDisconnected` → `backup_target_absent` on the wire. One rebuild now validates all three. | CC |
|
||||||
| **R-112** | **E-2's degraded banner and offer have NO UI CONSUMER — the endpoint is correct and the customer never sees it.** `GET /api/storage/backup-target` returns byte-exact copy (proven live), and **nothing fetches it**: `grep 'backup-target'` across all `*.html`/`*.js`/`*.css` → **0 hits**; no template references `OfferPath`/`Degraded`/the copy; `resolveBackupTargetState` + `degradedMessageFor` are consumed **only** by the JSON handler — **no page handler injects the state** | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md` — the banner reached a customer's page for the first time.** Never-configured box: banner element 1, never-configured copy 1. After the wizard: offer block 1 with `data-path="/mnt/felhom-drives/mentes"`. Healthy after assign: all four markers 0, **proven POSITIVELY** — idle delta 0 `/backup/tiers` calls, page-load delta +1, single caller ⇒ the seam ran and chose silence. **FIXED: the state now has a consumer.** Server-rendered on `/backups` via `backupsHandler` → `backupTargetView` → `backups.html`, following the existing `SingleCopyWarning` banner pattern — NOT a 19th JS fetch, because a banner that needs JavaScript to appear is one more thing that can silently not happen. `backupTargetView` returns **nil** for healthy AND unknown, so those render nothing at all. **Scenario-E seam test drives `backupsHandler` over httptest and asserts the RENDERED HTML** — deleting the one line that sets `data["BackupTarget"]` reproduces the old state and fails every render assertion. **SEAM PROVEN LIVE on demo-felhom by a DIFFERENTIAL positive observable, not by an absent banner:** idle 8 s → 0 new `/backup/tiers` agent calls; each `/backups` load → exactly +1, and that call has only one caller (`resolveBackupTargetState`). The box is healthy (`degraded:false, target:felhom-backup`) and the page correctly rendered **nothing** — which matches its real state but, being a negative, is NOT by itself proof of wiring. **Still unproven live:** that a customer sees actual copy — impossible on a healthy box. **The decisive contrast: templates fetch 18 distinct `/api/storage/*` endpoints; `backup-target` and `backup-target/assign` are the only two with zero references.** The handler's own comment calls itself *"the dashboard's source for the degraded banner and the offer"* — an invariant comment asserting a consumer that does not exist (7th instance of that class). v0.185.1 shipped as *"the offer endpoints were mounted where nothing routed to them"* — it fixed the **router mount** and stopped one layer short of the **render**; its test `TestBackupTargetRoutesLiveUnderTheStorageAPIMount` pins dispatch, not reachability, which is exactly what `CLAUDE.md`'s seam rule warns about. **Fifth instance of seam-built-but-never-wired.** **Fix R-114 FIRST** — wiring this alone would start showing customers the wrong message. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.1 | CC |
|
| **R-112** | **E-2's degraded banner and offer have NO UI CONSUMER — the endpoint is correct and the customer never sees it.** `GET /api/storage/backup-target` returns byte-exact copy (proven live), and **nothing fetches it**: `grep 'backup-target'` across all `*.html`/`*.js`/`*.css` → **0 hits**; no template references `OfferPath`/`Degraded`/the copy; `resolveBackupTargetState` + `degradedMessageFor` are consumed **only** by the JSON handler — **no page handler injects the state** | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md` — the banner reached a customer's page for the first time.** Never-configured box: banner element 1, never-configured copy 1. After the wizard: offer block 1 with `data-path="/mnt/felhom-drives/mentes"`. Healthy after assign: all four markers 0, **proven POSITIVELY** — idle delta 0 `/backup/tiers` calls, page-load delta +1, single caller ⇒ the seam ran and chose silence. **FIXED: the state now has a consumer.** Server-rendered on `/backups` via `backupsHandler` → `backupTargetView` → `backups.html`, following the existing `SingleCopyWarning` banner pattern — NOT a 19th JS fetch, because a banner that needs JavaScript to appear is one more thing that can silently not happen. `backupTargetView` returns **nil** for healthy AND unknown, so those render nothing at all. **Scenario-E seam test drives `backupsHandler` over httptest and asserts the RENDERED HTML** — deleting the one line that sets `data["BackupTarget"]` reproduces the old state and fails every render assertion. **SEAM PROVEN LIVE on demo-felhom by a DIFFERENTIAL positive observable, not by an absent banner:** idle 8 s → 0 new `/backup/tiers` agent calls; each `/backups` load → exactly +1, and that call has only one caller (`resolveBackupTargetState`). The box is healthy (`degraded:false, target:felhom-backup`) and the page correctly rendered **nothing** — which matches its real state but, being a negative, is NOT by itself proof of wiring. **Still unproven live:** that a customer sees actual copy — impossible on a healthy box. **The decisive contrast: templates fetch 18 distinct `/api/storage/*` endpoints; `backup-target` and `backup-target/assign` are the only two with zero references.** The handler's own comment calls itself *"the dashboard's source for the degraded banner and the offer"* — an invariant comment asserting a consumer that does not exist (7th instance of that class). v0.185.1 shipped as *"the offer endpoints were mounted where nothing routed to them"* — it fixed the **router mount** and stopped one layer short of the **render**; its test `TestBackupTargetRoutesLiveUnderTheStorageAPIMount` pins dispatch, not reachability, which is exactly what `CLAUDE.md`'s seam rule warns about. **Fifth instance of seam-built-but-never-wired.** **Fix R-114 FIRST** — wiring this alone would start showing customers the wrong message. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.1 | CC |
|
||||||
| **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that just vanished.** With the assigned target absent, the endpoint returned `degraded:true, target:"felhom-backup"` **plus** the *"a rendszermentés ugyanazon a lemezen van, mint a rendszer"* message — false, the target is a drive that has disappeared, not the system disk — **and** `offer_path` pointing at the missing drive as the remedy | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md`.** With the target absent the page rendered the ABSENT copy (1), the system-disk copy 0, the offer block 0 — both of E-2d's falsehoods gone. API carried `message:"A rendszermentés meghajtója nem érhető el…"` with `target:felhom-backup`. **FIXED: the third state exists.** New `BackupTargetState.TargetAbsent` separates *configured-and-gone* from *never-configured*. `Degraded` keeps its meaning (is there a problem) so the wire contract is unchanged for every consumer; `TargetAbsent` answers which problem, because the remedies are OPPOSITE — attach any second drive vs reconnect *that* one. Copy routed through `degradedMessageFor` (still one decision point) and taken **verbatim** from the hub's `backup_target_absent` email so banner and mail tell one story. **Offer suppressed on the branch itself**, deliberately not left to `firstOfferableDrive`'s `Disconnected` skip — that flag is set by R-113 in another repo, and this state must be right without it. Red-proof: deleting the branch reproduces E-2d's exact payload, offering `/mnt/felhom-drives/mentes2`, the drive that had vanished. **MinAgent unchanged 0.113.0** — R-114 reads `BackupTarget`/`MountPath`/`GuestPath`/`Role`, none of which R-113 altered, so demo-hp is not held. **NOT live-validated: Scenario C cannot occur on a healthy box.** `resolveBackupTargetState` falls through to the generic degraded branch whenever no disk satisfies `d.BackupTarget && d.MountPath != ""`, never distinguishing **never configured** from **configured and now missing**. Shares R-113's root cause — two disagreeing presence signals — but is a different code path with a different fix. **Currently invisible ONLY because of R-112; fix this before wiring that.** Also seen: after reattach the drive returned as `/dev/sdc` while the stable bind still recorded `/dev/sdb`, and the state read healthy. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.3 | CC |
|
| **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that just vanished.** With the assigned target absent, the endpoint returned `degraded:true, target:"felhom-backup"` **plus** the *"a rendszermentés ugyanazon a lemezen van, mint a rendszer"* message — false, the target is a drive that has disappeared, not the system disk — **and** `offer_path` pointing at the missing drive as the remedy | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md`.** With the target absent the page rendered the ABSENT copy (1), the system-disk copy 0, the offer block 0 — both of E-2d's falsehoods gone. API carried `message:"A rendszermentés meghajtója nem érhető el…"` with `target:felhom-backup`. **FIXED: the third state exists.** New `BackupTargetState.TargetAbsent` separates *configured-and-gone* from *never-configured*. `Degraded` keeps its meaning (is there a problem) so the wire contract is unchanged for every consumer; `TargetAbsent` answers which problem, because the remedies are OPPOSITE — attach any second drive vs reconnect *that* one. Copy routed through `degradedMessageFor` (still one decision point) and taken **verbatim** from the hub's `backup_target_absent` email so banner and mail tell one story. **Offer suppressed on the branch itself**, deliberately not left to `firstOfferableDrive`'s `Disconnected` skip — that flag is set by R-113 in another repo, and this state must be right without it. Red-proof: deleting the branch reproduces E-2d's exact payload, offering `/mnt/felhom-drives/mentes2`, the drive that had vanished. **MinAgent unchanged 0.113.0** — R-114 reads `BackupTarget`/`MountPath`/`GuestPath`/`Role`, none of which R-113 altered, so demo-hp is not held. **NOT live-validated: Scenario C cannot occur on a healthy box.** `resolveBackupTargetState` falls through to the generic degraded branch whenever no disk satisfies `d.BackupTarget && d.MountPath != ""`, never distinguishing **never configured** from **configured and now missing**. Shares R-113's root cause — two disagreeing presence signals — but is a different code path with a different fix. **Currently invisible ONLY because of R-112; fix this before wiring that.** Also seen: after reattach the drive returned as `/dev/sdc` while the stable bind still recorded `/dev/sdb`, and the state read healthy. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.3 | CC |
|
||||||
| **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** This is the **class**, not an instance: a gate that exists, asserts something true, is red, and is invoked by nothing reads as coverage it is not providing. `controller/scripts/docker_run_volume_path_gate.py` failed continuously from **2026-07-14 (v0.129.0)** until R-7b's close-out ran it by hand at v0.145.0 — sixteen releases in which every REPORT said "green" | **leg (a) CLOSED; leg (b) HALF-SHIPPED** (2026-08-02) | — | **This item has existed at `ROADMAP.md:158` since before the register was rebuilt (2026-07-27) and was never carried across — that omission is itself part of the finding**, because it is an open item *about work not getting done* that then went missing from the page that decides what gets done. Two separable parts, per R-29's own analysis: **(a)** the `docker_run_volume_path_gate` finding is benign and the fix is a 3-line ALLOWLIST addition with its why — **not** a rewrite of the flagged call — and it gets its own reviewed diff, never bundled into a feature commit; **(b)** the systemic half, the real item: decide where gates run (pre-push hook, `build.sh` step, or CI) and make a red gate block the train the way the Go green gate does. **Two further orphans confirmed 2026-07-29** by repo-wide grep across all file types + sibling repos + `~/.claude` settings/skills/hooks + `.git/hooks` (none non-sample) + Makefile/justfile/Taskfile find (only `hub/Makefile`, zero `gate` occurrences) + CI-directory find (**this repo has no CI at all**) — every one of the 19 hits is a docstring, a code comment or prose, and **not one is an invocation**: `scripts/hostinstall_gates.py` — **RED today** (`hub Setup-tab hostInstallVersion=1.19.0 != SCRIPT_VERSION=1.22.0`, exit 1), the same finding as **R-94 leg (b)** — and `scripts/hub_confirm_gate.py`. Of the four gates in `scripts/`, only `site_gates.py` is mandated anywhere (`CLAUDE.md:153`) and `manifest_bearer_gate.py` is named in `runbooks/secrets.md:76`. **In R-29's own words, carried forward deliberately: do not mint a new ID for a new instance** — the 2026-07-18 rehearsal independently re-raised this item and no second ID was minted then either **UPDATE 2026-08-02 — leg (a) CLOSED** (`felhom-controller` `c432f70`, its own reviewed diff as specified): `appexport/estimate.go`'s `-v` is a NAMED VOLUME mounted read-only into a throwaway container, no host path, structurally identical to the allowlisted `backup/backup.go` entry — allowlisted with its why; `realVolumeSize` untouched. **Leg (b) HALF-SHIPPED:** the 'decide where gates run' ruling is now made and half-implemented — **every repo has ONE entry point** (`felhom.eu/scripts/repo_gates.py`, `felhom-controller/controller/scripts/controller_gates.py`, `felhom-agent/scripts/agent_gates.py`, `app-catalog-felhom.eu/scripts/catalog_gates.py`), each mandated in its `CLAUDE.md` and each wired to `.githooks/pre-push` via `--fast`. **THE CENSUS, which is the finding:** thirteen gate scripts across four repos; **every gate a `CLAUDE.md` names was GREEN, and two of the four nobody names were RED** — `hostinstall_gates.py` (red since 2026-07-14) and `reuse_refs_check.py` (red on all four repos); a third, `docker_run_volume_path_gate.py`, was named only in `REUSE.md:284` and was also red. Correlation with 'named in a CLAUDE.md' was exact. **STAYS OPEN for the automatic half** — a hook is per-clone and `--no-verify` skips it; the unbypassable half is CI → **R-168** | CC |
|
| **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** This is the **class**, not an instance: a gate that exists, asserts something true, is red, and is invoked by nothing reads as coverage it is not providing. `controller/scripts/docker_run_volume_path_gate.py` failed continuously from **2026-07-14 (v0.129.0)** until R-7b's close-out ran it by hand at v0.145.0 — sixteen releases in which every REPORT said "green" | **CLOSED — both halves shipped** (2026-08-02) | — | **This item has existed at `ROADMAP.md:158` since before the register was rebuilt (2026-07-27) and was never carried across — that omission is itself part of the finding**, because it is an open item *about work not getting done* that then went missing from the page that decides what gets done. Two separable parts, per R-29's own analysis: **(a)** the `docker_run_volume_path_gate` finding is benign and the fix is a 3-line ALLOWLIST addition with its why — **not** a rewrite of the flagged call — and it gets its own reviewed diff, never bundled into a feature commit; **(b)** the systemic half, the real item: decide where gates run (pre-push hook, `build.sh` step, or CI) and make a red gate block the train the way the Go green gate does. **Two further orphans confirmed 2026-07-29** by repo-wide grep across all file types + sibling repos + `~/.claude` settings/skills/hooks + `.git/hooks` (none non-sample) + Makefile/justfile/Taskfile find (only `hub/Makefile`, zero `gate` occurrences) + CI-directory find (**this repo has no CI at all**) — every one of the 19 hits is a docstring, a code comment or prose, and **not one is an invocation**: `scripts/hostinstall_gates.py` — **RED today** (`hub Setup-tab hostInstallVersion=1.19.0 != SCRIPT_VERSION=1.22.0`, exit 1), the same finding as **R-94 leg (b)** — and `scripts/hub_confirm_gate.py`. Of the four gates in `scripts/`, only `site_gates.py` is mandated anywhere (`CLAUDE.md:153`) and `manifest_bearer_gate.py` is named in `runbooks/secrets.md:76`. **In R-29's own words, carried forward deliberately: do not mint a new ID for a new instance** — the 2026-07-18 rehearsal independently re-raised this item and no second ID was minted then either **UPDATE 2026-08-02 — leg (a) CLOSED** (`felhom-controller` `c432f70`, its own reviewed diff as specified): `appexport/estimate.go`'s `-v` is a NAMED VOLUME mounted read-only into a throwaway container, no host path, structurally identical to the allowlisted `backup/backup.go` entry — allowlisted with its why; `realVolumeSize` untouched. **Leg (b) HALF-SHIPPED:** the 'decide where gates run' ruling is now made and half-implemented — **every repo has ONE entry point** (`felhom.eu/scripts/repo_gates.py`, `felhom-controller/controller/scripts/controller_gates.py`, `felhom-agent/scripts/agent_gates.py`, `app-catalog-felhom.eu/scripts/catalog_gates.py`), each mandated in its `CLAUDE.md` and each wired to `.githooks/pre-push` via `--fast`. **THE CENSUS, which is the finding:** thirteen gate scripts across four repos; **every gate a `CLAUDE.md` names was GREEN, and two of the four nobody names were RED** — `hostinstall_gates.py` (red since 2026-07-14) and `reuse_refs_check.py` (red on all four repos); a third, `docker_run_volume_path_gate.py`, was named only in `REUSE.md:284` and was also red. Correlation with 'named in a CLAUDE.md' was exact. **STAYS OPEN for the automatic half** — a hook is per-clone and `--no-verify` skips it; the unbypassable half is CI → **R-168** **CLOSED 2026-08-02, on the demonstrated ALARM and not on a green run.** Leg (b)'s automatic half is now live: a Gitea Actions runner re-runs every repo's entry point on every push, independent of who pushed and of what they typed (→ R-168). The class this row opened — *a gate that exists, asserts something true, is red, and is invoked by nothing* — is answered at both ends: the pre-push hook refuses locally, and CI catches a `--no-verify` bypass and **emails the operator**, proven with a real red run and a provider accepted-id. What remains is not this row's finding but a working-style choice — CI reports rather than blocks because there is no merge to gate (→ R-169) | CC |
|
||||||
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
|
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
|
||||||
| **R-86** | Restore-tests are interval-scheduled, not backup-aligned | **READY** | R-90 (ep0 headroom) informs cadence | Trigger a tier ~24 h after **its own** newest archive | CC |
|
| **R-86** | Restore-tests are interval-scheduled, not backup-aligned | **READY** | R-90 (ep0 headroom) informs cadence | Trigger a tier ~24 h after **its own** newest archive | CC |
|
||||||
| **R-87** | The restic tier is never restore-tested | **READY** | — | Design a controller-side test (no scratch-guest analogue transfers) | CC |
|
| **R-87** | The restic tier is never restore-tested | **READY** | — | Design a controller-side test (no scratch-guest analogue transfers) | CC |
|
||||||
@@ -89,14 +89,15 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha
|
|||||||
| **R-158** | **A local Tier-1 app-data backup failure reaches no hub channel — `NotifyBackupFailed` exists, the hub allowlists `backup_failed`, and its only production caller is the off-box/NAS leg** (`cmd/controller/main.go:659`). The backup manager has `tier2Notify`/`offboxNotify`/`offboxEnlargeBlockedNotify` seams (`internal/backup/backup.go:33,37,58`) and **none for the recovery-unit capture**. Fifth instance of *seam built but never wired*; R-97's defect one tier over. | **READY (S)** | — | **Ranked BELOW R-157 — it is a notification GAP, not silent failure.** Measured: with `mp1` full, `/backups` DOES render `✗ Adatmentés sikertelen`, the marker **persists** across a second failed run and **clears** on recovery, and `/backups/apps` honestly shows the last good unit's real mtime — no surface claims a fresh backup over a stale unit. **The half worth fixing: `/backups/apps` is where you ask whether one app is backed up, and it is the one page that never says.** Proposed shape: a `unitNotify` seam wired in `main()` like `SetOffboxNotify`, emitting the existing `backup_failed`. Evidence: `audits/SPIKE-recovery-unit-space-2026-08-02.md` §5, `audits/CAMPAIGN-10-closeout-2026-08-02.md` Q1 | CC |
|
| **R-158** | **A local Tier-1 app-data backup failure reaches no hub channel — `NotifyBackupFailed` exists, the hub allowlists `backup_failed`, and its only production caller is the off-box/NAS leg** (`cmd/controller/main.go:659`). The backup manager has `tier2Notify`/`offboxNotify`/`offboxEnlargeBlockedNotify` seams (`internal/backup/backup.go:33,37,58`) and **none for the recovery-unit capture**. Fifth instance of *seam built but never wired*; R-97's defect one tier over. | **READY (S)** | — | **Ranked BELOW R-157 — it is a notification GAP, not silent failure.** Measured: with `mp1` full, `/backups` DOES render `✗ Adatmentés sikertelen`, the marker **persists** across a second failed run and **clears** on recovery, and `/backups/apps` honestly shows the last good unit's real mtime — no surface claims a fresh backup over a stale unit. **The half worth fixing: `/backups/apps` is where you ask whether one app is backed up, and it is the one page that never says.** Proposed shape: a `unitNotify` seam wired in `main()` like `SetOffboxNotify`, emitting the existing `backup_failed`. Evidence: `audits/SPIKE-recovery-unit-space-2026-08-02.md` §5, `audits/CAMPAIGN-10-closeout-2026-08-02.md` Q1 | CC |
|
||||||
| **R-159** | **wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy.** The image declares `VOLUME /usr/src/app/data`; the template mounted `wishlist_data:/data`, a path the app never writes. `ResolveDockerVolumeNames` returns `<project>_<name>` only for volumes **declared in the compose file**, so `DumpAppVolumes` never sees an anonymous one. Survives a restart, loses on redeploy, never in a backup — harder to notice than papra's. | **SHIPPED** (`templates/wishlist/docker-compose.yml`, 2026-08-02) — filed to record the CLASS | — | **The class is open even though the app is fixed:** any image `VOLUME` at a path the template does not mount creates unbacked-up storage silently. **`immich-server` has one today** at `/data` — empty when measured, so nothing is at risk now. Proposed `REUSE.md` rule: *a template must mount every path in its image's `Config.Volumes`, or state why not.* Checkable only with the image pulled, hence the runtime gate | CC |
|
| **R-159** | **wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy.** The image declares `VOLUME /usr/src/app/data`; the template mounted `wishlist_data:/data`, a path the app never writes. `ResolveDockerVolumeNames` returns `<project>_<name>` only for volumes **declared in the compose file**, so `DumpAppVolumes` never sees an anonymous one. Survives a restart, loses on redeploy, never in a backup — harder to notice than papra's. | **SHIPPED** (`templates/wishlist/docker-compose.yml`, 2026-08-02) — filed to record the CLASS | — | **The class is open even though the app is fixed:** any image `VOLUME` at a path the template does not mount creates unbacked-up storage silently. **`immich-server` has one today** at `/data` — empty when measured, so nothing is at risk now. Proposed `REUSE.md` rule: *a template must mount every path in its image's `Config.Volumes`, or state why not.* Checkable only with the image pulled, hence the runtime gate | CC |
|
||||||
| **R-160** | **gramps-web persisted three paths and wrote to none of them.** `/app/data` appears nowhere in the image's environment; the accounts DB (`GRAMPSWEB_USER_DB_URI`) and **the family tree** (`GRAMPS_DATABASE_PATH=/root/.gramps/grampsdb`) both landed in the writable layer. Upstream persists **eight** paths; the template persisted three, one a phantom. | **SHIPPED** (`templates/gramps-web/docker-compose.yml`, 2026-08-02) | — | **Severity above papra's, and worth keeping visible:** papra loses documents the customer may hold elsewhere; gramps-web loses **the family tree — the artefact built inside the app, of which no other copy exists by construction.** Evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` | CC |
|
| **R-160** | **gramps-web persisted three paths and wrote to none of them.** `/app/data` appears nowhere in the image's environment; the accounts DB (`GRAMPSWEB_USER_DB_URI`) and **the family tree** (`GRAMPS_DATABASE_PATH=/root/.gramps/grampsdb`) both landed in the writable layer. Upstream persists **eight** paths; the template persisted three, one a phantom. | **SHIPPED** (`templates/gramps-web/docker-compose.yml`, 2026-08-02) | — | **Severity above papra's, and worth keeping visible:** papra loses documents the customer may hold elsewhere; gramps-web loses **the family tree — the artefact built inside the app, of which no other copy exists by construction.** Evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` | CC |
|
||||||
| **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run | operator |
|
| **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator |
|
||||||
| **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** The gate's power comes from `docker diff` excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where `docker diff` is unsupported or lies, the gate degrades to the mount-occupancy and writability legs **and would not say so**. | **WATCHING** — a limitation, not a defect | — | It **fails closed**: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships | CC |
|
| **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** The gate's power comes from `docker diff` excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where `docker diff` is unsupported or lies, the gate degrades to the mount-occupancy and writability legs **and would not say so**. | **WATCHING** — a limitation, not a defect | — | It **fails closed**: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships | CC |
|
||||||
| **R-163** | **`mp1` is RETENTION, not staging — and it is sized as if it were neither.** A recovery unit is the KEPT copy on the app's **own** drive (`GetAppDrivePath`, `internal/backup/backup.go:245-255`); for an app with no `HDD_PATH` the namespace falls back to the system SSD — *"the SSD-only system-data fallback"* (`internal/appbackup/paths.go:26-27`). There is **no post-copy deletion**: the only prune is F5 (`backup.go:1053-1112`), residue on OLD drives when an app MOVES. So `mp1` (**20 G**) retains the units of every driveless app, while `mp0` permits **50 G** of volumes — and a DB app's unit is up to **~2×** its data (volume tar **plus** SQL dump; measured 21.1 GB → 40.2 GB). `--sysdata-grow` defaults to **0** (`felhom-agent/cmd/felhom-agent/main.go:178`) and is **not** derived from the physical drive; demo-hp's real guest 9201 ships `mp0 50G / mp1 20G`. | **RE-FRAMED 2026-08-02 — open, no longer waiting on a ratio** | — (the sizing question is answered; the work is **R-165**) | **RE-FRAMED, NOT CLOSED (operator decision D-a, 2026-08-02 — `CONTEXT.md` S-5).** The row asked *what ratio should `mp1` be?* and that question is **withdrawn rather than answered**: `mp1` is merged into `mp0` so local recovery units share the app-data area and the ceiling stops existing — a bigger number is the same wall further away. **This row stays open as the record of the constraint** (what `mp1` is for, what it gates, and the measured 2× DB-app unit size) **until the merge lands**, because until then every consequence below is still live on every box. **The work is R-165; the warning that must ship with it is R-167.** Original finding, unchanged, follows. **No number is proposed here deliberately.** What is recorded is the constraint and its blast radius: **`mp1` gates the whole app-data chain**, because Tier-2 mirrors the unit *"(always)"* from `RecoveryUnitPath` (`internal/backup/tier2.go:302,368`) and Tier-3 carries it too — a unit that cannot be written has nothing for either to copy. Bounded on the other side: a unit holds **volume tars + DB dumps only, never `mp8` userdata** (`internal/backup/recovery_unit.go:20-25`), so a 1 TB photo library is never in one. **This bounds D5's Lane-1 independence** — see `architecture/07-backup-architecture.md` §7.5. Overflow itself is SAFE (R-158's measurement: refuses per app, last good unit preserved byte-identical) — what is missing is the warning, which is R-158 (widened to R-167) | CC |
|
| **R-163** | **`mp1` is RETENTION, not staging — and it is sized as if it were neither.** A recovery unit is the KEPT copy on the app's **own** drive (`GetAppDrivePath`, `internal/backup/backup.go:245-255`); for an app with no `HDD_PATH` the namespace falls back to the system SSD — *"the SSD-only system-data fallback"* (`internal/appbackup/paths.go:26-27`). There is **no post-copy deletion**: the only prune is F5 (`backup.go:1053-1112`), residue on OLD drives when an app MOVES. So `mp1` (**20 G**) retains the units of every driveless app, while `mp0` permits **50 G** of volumes — and a DB app's unit is up to **~2×** its data (volume tar **plus** SQL dump; measured 21.1 GB → 40.2 GB). `--sysdata-grow` defaults to **0** (`felhom-agent/cmd/felhom-agent/main.go:178`) and is **not** derived from the physical drive; demo-hp's real guest 9201 ships `mp0 50G / mp1 20G`. | **RE-FRAMED 2026-08-02 — open, no longer waiting on a ratio** | — (the sizing question is answered; the work is **R-165**) | **RE-FRAMED, NOT CLOSED (operator decision D-a, 2026-08-02 — `CONTEXT.md` S-5).** The row asked *what ratio should `mp1` be?* and that question is **withdrawn rather than answered**: `mp1` is merged into `mp0` so local recovery units share the app-data area and the ceiling stops existing — a bigger number is the same wall further away. **This row stays open as the record of the constraint** (what `mp1` is for, what it gates, and the measured 2× DB-app unit size) **until the merge lands**, because until then every consequence below is still live on every box. **The work is R-165; the warning that must ship with it is R-167.** Original finding, unchanged, follows. **No number is proposed here deliberately.** What is recorded is the constraint and its blast radius: **`mp1` gates the whole app-data chain**, because Tier-2 mirrors the unit *"(always)"* from `RecoveryUnitPath` (`internal/backup/tier2.go:302,368`) and Tier-3 carries it too — a unit that cannot be written has nothing for either to copy. Bounded on the other side: a unit holds **volume tars + DB dumps only, never `mp8` userdata** (`internal/backup/recovery_unit.go:20-25`), so a 1 TB photo library is never in one. **This bounds D5's Lane-1 independence** — see `architecture/07-backup-architecture.md` §7.5. Overflow itself is SAFE (R-158's measurement: refuses per app, last good unit preserved byte-identical) — what is missing is the warning, which is R-158 (widened to R-167) | CC |
|
||||||
| **R-164** | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC |
|
| **R-164** | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC |
|
||||||
| **R-165** | **Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.** Operator decision **D-a**, 2026-08-02 (`CONTEXT.md` S-5). Local recovery units share the app-data area instead of holding their own fixed ceiling, so the wall R-163 describes is removed rather than moved further away. Guest 9201 on demo-hp ships `mp0 50G / mp1 20G` today | **READY (M) — NEW 2026-08-02** | — | **Two conditions travel WITH the decision and are not optional.** **(1) Before any external install.** It changes the **disk layout**, so it is a fresh-install shape while there are no external boxes and a per-box migration after — and the decision's cheapness is entirely a function of that ordering. **(2) It removes a wall that currently fails safely**, so **R-167** (D-c: fill warning + failure alert) lands in the same step, never after: today an app that outgrows `mp1` is refused per app with the last good unit preserved byte-identical (R-158's measurement), and after the merge the same overflow consumes the space the app itself is using. Touches the installer/agent guest shape (`--sysdata-grow` defaults to **0** and is not derived from the physical drive, `felhom-agent/cmd/felhom-agent/main.go:178`) and the golden. **Does NOT close R-163** — that row is the record of the constraint and stays open until this lands | CC |
|
| **R-165** | **Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.** Operator decision **D-a**, 2026-08-02 (`CONTEXT.md` S-5). Local recovery units share the app-data area instead of holding their own fixed ceiling, so the wall R-163 describes is removed rather than moved further away. Guest 9201 on demo-hp ships `mp0 50G / mp1 20G` today | **READY (M) — NEW 2026-08-02** | — | **Two conditions travel WITH the decision and are not optional.** **(1) Before any external install.** It changes the **disk layout**, so it is a fresh-install shape while there are no external boxes and a per-box migration after — and the decision's cheapness is entirely a function of that ordering. **(2) It removes a wall that currently fails safely**, so **R-167** (D-c: fill warning + failure alert) lands in the same step, never after: today an app that outgrows `mp1` is refused per app with the last good unit preserved byte-identical (R-158's measurement), and after the merge the same overflow consumes the space the app itself is using. Touches the installer/agent guest shape (`--sysdata-grow` defaults to **0** and is not derived from the physical drive, `felhom-agent/cmd/felhom-agent/main.go:178`) and the golden. **Does NOT close R-163** — that row is the record of the constraint and stays open until this lands | CC |
|
||||||
| **R-166** | **App state gets a desired/observed model with its own store.** Operator decision **D-b**, 2026-08-02 (`CONTEXT.md` S-5). Three kinds of information, deliberately separated: **desired** (running, or stopped because the customer said so) stays in `app.yaml`; **in-flight operations** (a backup started and never reported finishing) are **persisted** — written before the operation and after it — so an interruption is visible after a reboot; **observed** (unhealthy, restarting, drive gone) is **not** persisted and is rebuilt by looking. The last two live in a store **separate from `app.yaml`**, so frequent writes never touch a file holding deploy settings and secrets | **BLOCKED — NEW 2026-08-02** | two facts to establish before it is specced | **NOT a spec yet, deliberately.** **Establish first:** (a) whether the crash-safe journal already in the backup code covers the in-flight case — *it may already exist and merely be unwired, which has been the answer four times in this project* (`CLAUDE.md` "seam built but never wired"); (b) whether the existing SQLite store is reachable from this code path. **Binding safety rule, verbatim from the decision:** *losing the state store must never cause an app to be deleted, restarted wrongly, or reported healthy when it is not — the worst acceptable outcome is re-running a backup that already ran.* The controller must also check **every container of an app**, so a partly-dead app cannot read as healthy. **This is the durable fix for R-157 mechanism B** (zero containers reading as a deliberate stop) and for F-CRIT-1's class; R-157 keeps its own row because it wants a bounded fix now, not this rework | CC |
|
| **R-166** | **App state gets a desired/observed model with its own store.** Operator decision **D-b**, 2026-08-02 (`CONTEXT.md` S-5). Three kinds of information, deliberately separated: **desired** (running, or stopped because the customer said so) stays in `app.yaml`; **in-flight operations** (a backup started and never reported finishing) are **persisted** — written before the operation and after it — so an interruption is visible after a reboot; **observed** (unhealthy, restarting, drive gone) is **not** persisted and is rebuilt by looking. The last two live in a store **separate from `app.yaml`**, so frequent writes never touch a file holding deploy settings and secrets | **BLOCKED — NEW 2026-08-02** | two facts to establish before it is specced | **NOT a spec yet, deliberately.** **Establish first:** (a) whether the crash-safe journal already in the backup code covers the in-flight case — *it may already exist and merely be unwired, which has been the answer four times in this project* (`CLAUDE.md` "seam built but never wired"); (b) whether the existing SQLite store is reachable from this code path. **Binding safety rule, verbatim from the decision:** *losing the state store must never cause an app to be deleted, restarted wrongly, or reported healthy when it is not — the worst acceptable outcome is re-running a backup that already ran.* The controller must also check **every container of an app**, so a partly-dead app cannot read as healthy. **This is the durable fix for R-157 mechanism B** (zero containers reading as a deliberate stop) and for F-CRIT-1's class; R-157 keeps its own row because it wants a bounded fix now, not this rework | CC |
|
||||||
| **R-167** | **Storage monitoring and backup alerts.** Operator decision **D-c**, 2026-08-02 (`CONTEXT.md` S-5). Two pieces: a warning to the **customer** as a drive fills, **before** anything fails; and a failure alert to the **hub/operator** when any backup fails for **any** reason, lack of space included | **READY (M) — NEW 2026-08-02** | — | **Prerequisite for R-165** — D-a removes a wall that currently fails safely, so this ships in the same step, not after. **Subsumes and widens R-158**, which is the same alert one tier down (`NotifyBackupFailed` exists and the hub allowlists `backup_failed`; only the off-box/NAS leg calls it, `cmd/controller/main.go:659`). R-158 stays as the named seam and its proposed `unitNotify` shape is the concrete first leg — **do not file a second row for the same wire.** The customer half is new: nothing today warns before a fill, and R-158's measurement is that when it happens exactly one page says so | CC |
|
| **R-167** | **Storage monitoring and backup alerts.** Operator decision **D-c**, 2026-08-02 (`CONTEXT.md` S-5). Two pieces: a warning to the **customer** as a drive fills, **before** anything fails; and a failure alert to the **hub/operator** when any backup fails for **any** reason, lack of space included | **READY (M) — NEW 2026-08-02** | — | **Prerequisite for R-165** — D-a removes a wall that currently fails safely, so this ships in the same step, not after. **Subsumes and widens R-158**, which is the same alert one tier down (`NotifyBackupFailed` exists and the hub allowlists `backup_failed`; only the off-box/NAS leg calls it, `cmd/controller/main.go:659`). R-158 stays as the named seam and its proposed `unitNotify` shape is the concrete first leg — **do not file a second row for the same wire.** The customer half is new: nothing today warns before a fill, and R-158's measurement is that when it happens exactly one page says so | CC |
|
||||||
| **R-168** | **CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK.** The pre-push hook shipped by R-29 is local to a clone and `--no-verify` skips it; the unbypassable half is continuous integration, and it has no execution mechanism here yet | **BLOCKED** on its own spike | a spike deciding the runner's execution mode | **Measured 2026-08-02:** Gitea **1.26.2**; Actions **enabled on all four repos**; **0 runners registered, 0 workflow runs, 0 branch protections**. The consequence is the load-bearing part: **we push directly to `main` with no branches, so there is no merge for a status check to gate** — CI here reports after the fact unless the workflow itself is given teeth, which is a separate decision from standing one up. Two open questions for the spike, both about the execution mechanism (the reason this is not bundled into R-29): **(1)** host-mode runner versus a privileged Docker-in-Docker sidecar on DooPlex — DooPlex is Tier 2, it *is* the recovery chain (Gitea, registry, k3s+Longhorn, PBS, hub), and a privileged DinD sidecar there needs its blast radius argued, not assumed; **(2)** whether the workflow can avoid JavaScript actions entirely (act_runner's node runtime is the usual reason a host-mode runner fails). Successor to R-161's automatic half | CC |
|
| **R-168** | ~~CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK~~ | **SHIPPED — and the alarm is DEMONSTRATED** (2026-08-02) | — | **Runner live**: `homelab-manifests/gitea-system/act-runner.yaml`, an unprivileged host-mode `act_runner` in `gitea-system`, one owner-scoped registration serving all four repos (measured: tasks 7-10 all claimed by `felhom-gates-runner`). `.gitea/workflows/gates.yml` in each repo runs that repo's entry point with `--fast` and nothing else; no `uses:` step anywhere. **Six probes, all answered, none STOPped** — `audits/SPIKE-ci-runner-2026-08-02.md`. The two that changed the design: **P2** (stock image has git but NO python3 → custom image `felhom-act-runner:0.1.0`, base pinned, python3 and nothing else) and **P6** (a runner that loses `/data/.runner` re-registers and leaves a dead record behind → the PVC is load-bearing, measured both ways). **P5 is the one that mattered**: a failed run produced NO mail, NO notification row and NO log line from Gitea, so the run now sends its own alarm via Resend and prints the provider's accepted id. **Proven end to end, not asserted**: a deliberately broken commit pushed with `--no-verify` → run #6 `failure` → `RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`. Posture shown from the live pod spec: `privileged: false`, all caps dropped, no docker socket, no hostPath, `automountServiceAccountToken: false`, sized at half Gitea's limits so it cannot crowd out the service holding every repository on the same node. **The standing limit stays true and is written into the manifest and every workflow: it DETECTS, it does not BLOCK** — making it block is → R-169 | — |
|
||||||
|
| **R-169** | **CI can only report, because there is no gate in the road.** Every felhom repo pushes straight to `main` with no pull request, so there is no merge for a status check to stand at. R-168's runner therefore notices a broken push *after* it has landed | **WAITING-ON-OPERATOR** (a working-style decision, not a defect) | an operator ruling | Making CI *blocking* requires two things this task deliberately did NOT do, because both change how the operator works and that is not a task's call: **(a)** branch protection on `main`, and **(b)** a pull-request workflow instead of direct-to-`main` pushes. The cost is real — every change would need a PR, which for a single-operator project may be worse than the disease. **The current arrangement is two nets, and it is not nothing**: `.githooks/pre-push` REFUSES locally, and R-168's runner NOTICES when that hook was skipped or was never armed in a clone, and emails. The honest gap is the window between a `--no-verify` push landing and the operator reading the alarm. Decide only if that window ever actually costs something | operator |
|
||||||
|
|
||||||
## Why the TOP READY rows rank this way
|
## Why the TOP READY rows rank this way
|
||||||
|
|
||||||
|
|||||||
File diff suppressed because one or more lines are too long
@@ -1,3 +1,52 @@
|
|||||||
|
## CI — a Gitea Actions runner, and a red run that reaches a person (2026-08-02, R-168)
|
||||||
|
|
||||||
|
**No version bump anywhere: nothing in the product repos is compiled, built or deployed by this.**
|
||||||
|
Recorded explicitly so the omission reads as a decision rather than a miss.
|
||||||
|
|
||||||
|
**What this closes.** Session 1 (same day) gave every repo one gate entry point and a
|
||||||
|
`.githooks/pre-push` that refuses a failing push. That hook is per-clone and `--no-verify` skips it,
|
||||||
|
so nothing independent of the person pushing ever saw whether the gates passed. This is the
|
||||||
|
independent half, and with it **R-29 CLOSES** — on the demonstrated alarm, not on a green run.
|
||||||
|
|
||||||
|
**`.gitea/workflows/gates.yml` (new)** — triggers on `push`, `runs-on: felhom-gates`, obtains the
|
||||||
|
source with a shallow `git fetch` of the **exact pushed SHA** from the in-cluster Gitea Service, and
|
||||||
|
runs `scripts/repo_gates.py --fast` and nothing else. **No `uses:` step anywhere** — JavaScript
|
||||||
|
actions need a node runtime the host-mode runner does not have, and probe P3 measured that a plain
|
||||||
|
`git fetch` is sufficient and lands on the pushed commit. No `|| true`; the entry point's exit code
|
||||||
|
IS the job's result.
|
||||||
|
|
||||||
|
**The alarm, which is the half that matters.** Probe P5 measured that a failed run produces **no
|
||||||
|
mail, no notification row and no log line** from Gitea. A red tick in a web UI nobody watches is
|
||||||
|
exactly the defect R-29 filed, rebuilt one layer up — so the workflow sends its own email on failure
|
||||||
|
via Resend (the hub's existing transactional path) and **prints the provider's accepted id**, making
|
||||||
|
"a message left the machine" an observable. **Demonstrated, not asserted:** a deliberately broken
|
||||||
|
commit pushed with `--no-verify` produced run #6 `failure` and
|
||||||
|
`RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`.
|
||||||
|
|
||||||
|
Two traps found while building it, both worth keeping because each looks like something else:
|
||||||
|
the runner image has **no `curl`** on purpose (python3 and git only — so the step uses `urllib`
|
||||||
|
rather than growing the image), and `api.resend.com` sits behind **Cloudflare, which 403s the default
|
||||||
|
`Python-urllib` User-Agent with error 1010** — a failure that reads exactly like an auth failure and
|
||||||
|
is not one.
|
||||||
|
|
||||||
|
**The standing limit, written into the workflow itself: it REPORTS, it cannot REFUSE.** Every repo
|
||||||
|
pushes straight to `main` with no pull request, so there is no merge for a status check to stand at.
|
||||||
|
That is not a gap in the runner; there is no gate in the road. Making it blocking needs branch
|
||||||
|
protection plus a PR workflow, which changes how the operator works → **R-169**, waiting on them.
|
||||||
|
|
||||||
|
**`documentation/audits/SPIKE-ci-runner-2026-08-02.md` (new)** — all six probes, method, measurement
|
||||||
|
and ruling; none produced a STOP. Also records a near-miss worth more than the probes: a `| tail -5`
|
||||||
|
inside my own census query silently dropped rows and looked exactly like a baseline drift big enough
|
||||||
|
to change the task. **An instrument that can drop results silently is not a measurement.**
|
||||||
|
|
||||||
|
**`CLAUDE.md`** gains the matching rule from session 1's red-proofing: a `go test -run` pattern that
|
||||||
|
matches no test prints `ok` and exits 0, so a red-proof using `-run` must first prove the filter
|
||||||
|
matched something.
|
||||||
|
|
||||||
|
**`CONTEXT.md`** gains S-8 (CI detects, does not block, and why that is structural), S-9 (a detector
|
||||||
|
that tells no one is not finished), S-10 (the runner is unprivileged because DooPlex is Tier 2), and
|
||||||
|
S-11 (CI reproduces the workspace's sibling layout, because two entry points depend on it).
|
||||||
|
|
||||||
## Gate enforcement — one entry point per repo, and a pre-push hook (2026-08-02)
|
## Gate enforcement — one entry point per repo, and a pre-push hook (2026-08-02)
|
||||||
|
|
||||||
**No version bump: `scripts/` carries no version, and this is tooling.** Recorded explicitly so the
|
**No version bump: `scripts/` carries no version, and this is tooling.** Recorded explicitly so the
|
||||||
|
|||||||
Reference in New Issue
Block a user