R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s

The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.

The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.

Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.

08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.

R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.

Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
2026-08-23 07:59:52 +02:00
parent 1eb64bec51
commit 55274d5ef3
39 changed files with 4986 additions and 65 deletions
+2 -2
View File
@@ -136,8 +136,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
| ID | What | State |
|---|---|---|
| **R-383** | **The double-failure message tells the customer their previous state was saved, and names a file that is not there.** The sentence ends *"a korábbi állapot mentése megvan: <file>"* — "the backup of the previous state EXISTS" — built from the path `writeSafetyDump` returned, WITHOUT asking whether it is still on disk. But one of the two ways a rollback can fail is that the undo copy is missing or unreadable, and in exactly that case the sentence is FALSE. **Measured live twice, on v0.220.2 (2026-08-22 21:11:19) and again on v0.221.1 (22:20:32):** rollback failed with `stat …pre-restore-…sql: no such file or directory`, and the customer message named that same file as existing. 369 bytes, `offbox_reconstitute.go` (the double-failure branch). **This is R-361's own class** — a sentence asserting a property the code does not check — one surface over. | **OPEN — MEDIUM** | — | Say what is true: name the undo copy only when it is verifiably on disk, and say plainly when it is not. **Do not simply drop the filename** — an operator needs it, and R-351's lesson was that a refusal which names nothing forces someone to remember what the product already knew. Evidence: `audits/DRILL-r361-2026-08-22/evidence/03-observed-false-sentence.txt`, `audits/DRILL-r361-2026-08-22/evidence/16-part4-message.txt`. | CC |
| **R-384** | **An app whose DATABASE has died raises no dead-app alarm — `unhealthy` masks the mixed state.** `aggregateState` (`internal/stacks/manager.go`) checks `if unhealthy > 0 → StateUnhealthy` BEFORE the mixed-case degraded branch, and `IsDownState` (`manager.go:54`) is `{stopped, exited, degraded}` — `unhealthy` is absent. So a multi-container app whose database container dies goes `degraded` for a moment and then `unhealthy` as its own healthcheck fails, and stops being a fault. **Measured live 2026-08-22:** `bookstack-db` stopped out-of-band at 21:27:01; `bookstack` read `unhealthy`; the dead-app heartbeat reported **`0 currently down`** across the whole window (scans 600 and 620), with 8 apps evaluated. **NOT invisible everywhere** — the health report counts it (`cr.Unhealthy++`, `internal/report/builder.go:251`) and that reaches the hub — but it raises no banner and no customer e-mail. **This is the F-CRIT-1 class the `classifyRunStates` comment says was closed:** it WAS closed for `StateStopped`+failedRestart, and `unhealthy` was never in scope. Found while building a positive control for a different question. | **OPEN — MEDIUM** | — | Decide whether a SUSTAINED `unhealthy` is a fault (it is not a brief one — that is why it is excluded), on the `crashLoopAfter` model: a threshold above the deploy/health windows rather than a state test. **Do not simply add `unhealthy` to `IsDownState`** — it has other callers and would alarm on every deploy, which is the over-correction F-A1 nearly cost. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`. | CC |
| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC |
| **R-386** | **A single-container app stopped OUT OF BAND raises no alarm at all — and a comment states the opposite as settled fact.** `aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**. `classifyRunStates` then whitelists `StateStopped` as a deliberate user stop unless the quiesce loop reports a failed restart. The comment at `cmd/controller/main.go` says: *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable.)"* — **measured FALSE.** The neighbouring I2 claim (*"a CRASHING app never comes to rest at `stopped` — faults surface as StateExited"*) is false in the same way. **Measured live on `demo-hp` 2026-08-23 (controller v0.222.0):** `privatebin` (1 container, `unless-stopped`) stopped out of band at 05:47:35Z; at 05:51:53Z it read `state=stopped`, **9 dead-app scans had run, and there were ZERO `app_start_failed` events and ZERO banner lines.** **The absence is trustworthy — positive control from the same box 17 minutes earlier:** `app_start_failed` fired for BookStack at 05:30:14Z, so the detector demonstrably works there. **SCOPE, stated so it is not overclaimed:** a genuine crash under `unless-stopped` is RESTARTED by Docker and surfaces as `restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member. **This is case #10 of "a comment asserting an invariant the code does not provide".** Found by §4 of the R-384 task, which asked for a measurement and explicitly forbade a fix in that session. | **OPEN — MEDIUM** | — | Decide whether an out-of-band stop is distinguishable from a customer stop at all — they are byte-identical on the Docker side, exactly as invariant I1 says, so the answer is probably NOT a state test but a recorded intent (`DesiredStateOf` already exists and `bootrecon` already consumes it). **Do NOT simply un-whitelist `StateStopped`** — that re-alarms every genuine customer stop, which is the over-correction F-CRIT-1's fix was careful to avoid. **And fix the comment either way:** it is load-bearing and it is false. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/live-17-sec4-stop.txt`, `live-18-sec4-verdict.txt`, `live-19-scenarioD-sec4-full-log.txt`. | CC |
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |