docs: R-86 closed and proven live; ep0 recorded as protected; R-185/186/187 filed
gates / gates (push) Successful in 8s
gates / gates (push) Successful in 8s
- OPEN-ITEMS: R-86 CLOSED with the trap in its own wording recorded (the literal reading is never true on a daily tier); R-87 re-ranked UP because R-86 built most of what it waited for; R-185 (the agent cannot list demo-felhom's host backup tier — a missing storage ACL, pre-existing), R-186 (a released binary's sha is not reproducible from its tag), R-187 (R-115's publish leg had never actually run) filed. R-184 was the highest ID in use. - ROADMAP: R-86 collapsed, keeping the reasoning and correcting the shape the row itself proposed — which would have been the never-fires version. - 07-backup-architecture: new contract section — restore-testing is per ARCHIVE GENERATION, with the trap and what did not change (S-1). - 00-capability-map: the unattended restore-proof row upgraded to PROVEN-LIVE on the 635 s due-triggered offsite run, with the restart and teardown evidence. - CONTEXT: S-17 (the rule, the trap, the config key, the hub's derivation) and S-18 (ep0 is Tier 2 — extends D-d's protected list to three machines). Numbered 17/18 because S-14 and S-15 were already duplicated in the file. - STATUS: rewritten for the operator, trimmed back to one screen.
This commit is contained in:
@@ -26,8 +26,11 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha
|
||||
| **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that just vanished.** With the assigned target absent, the endpoint returned `degraded:true, target:"felhom-backup"` **plus** the *"a rendszermentés ugyanazon a lemezen van, mint a rendszer"* message — false, the target is a drive that has disappeared, not the system disk — **and** `offer_path` pointing at the missing drive as the remedy | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | — | **PROVEN LIVE `audits/SESSION-C-2026-07-29.md`.** With the target absent the page rendered the ABSENT copy (1), the system-disk copy 0, the offer block 0 — both of E-2d's falsehoods gone. API carried `message:"A rendszermentés meghajtója nem érhető el…"` with `target:felhom-backup`. **FIXED: the third state exists.** New `BackupTargetState.TargetAbsent` separates *configured-and-gone* from *never-configured*. `Degraded` keeps its meaning (is there a problem) so the wire contract is unchanged for every consumer; `TargetAbsent` answers which problem, because the remedies are OPPOSITE — attach any second drive vs reconnect *that* one. Copy routed through `degradedMessageFor` (still one decision point) and taken **verbatim** from the hub's `backup_target_absent` email so banner and mail tell one story. **Offer suppressed on the branch itself**, deliberately not left to `firstOfferableDrive`'s `Disconnected` skip — that flag is set by R-113 in another repo, and this state must be right without it. Red-proof: deleting the branch reproduces E-2d's exact payload, offering `/mnt/felhom-drives/mentes2`, the drive that had vanished. **MinAgent unchanged 0.113.0** — R-114 reads `BackupTarget`/`MountPath`/`GuestPath`/`Role`, none of which R-113 altered, so demo-hp is not held. **NOT live-validated: Scenario C cannot occur on a healthy box.** `resolveBackupTargetState` falls through to the generic degraded branch whenever no disk satisfies `d.BackupTarget && d.MountPath != ""`, never distinguishing **never configured** from **configured and now missing**. Shares R-113's root cause — two disagreeing presence signals — but is a different code path with a different fix. **Currently invisible ONLY because of R-112; fix this before wiring that.** Also seen: after reattach the drive returned as `/dev/sdc` while the stable bind still recorded `/dev/sdb`, and the state read healthy. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.3 | CC |
|
||||
| **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** This is the **class**, not an instance: a gate that exists, asserts something true, is red, and is invoked by nothing reads as coverage it is not providing. `controller/scripts/docker_run_volume_path_gate.py` failed continuously from **2026-07-14 (v0.129.0)** until R-7b's close-out ran it by hand at v0.145.0 — sixteen releases in which every REPORT said "green" | **CLOSED — both halves shipped** (2026-08-02) | — | **This item has existed at `ROADMAP.md:158` since before the register was rebuilt (2026-07-27) and was never carried across — that omission is itself part of the finding**, because it is an open item *about work not getting done* that then went missing from the page that decides what gets done. Two separable parts, per R-29's own analysis: **(a)** the `docker_run_volume_path_gate` finding is benign and the fix is a 3-line ALLOWLIST addition with its why — **not** a rewrite of the flagged call — and it gets its own reviewed diff, never bundled into a feature commit; **(b)** the systemic half, the real item: decide where gates run (pre-push hook, `build.sh` step, or CI) and make a red gate block the train the way the Go green gate does. **Two further orphans confirmed 2026-07-29** by repo-wide grep across all file types + sibling repos + `~/.claude` settings/skills/hooks + `.git/hooks` (none non-sample) + Makefile/justfile/Taskfile find (only `hub/Makefile`, zero `gate` occurrences) + CI-directory find (**this repo has no CI at all**) — every one of the 19 hits is a docstring, a code comment or prose, and **not one is an invocation**: `scripts/hostinstall_gates.py` — **RED today** (`hub Setup-tab hostInstallVersion=1.19.0 != SCRIPT_VERSION=1.22.0`, exit 1), the same finding as **R-94 leg (b)** — and `scripts/hub_confirm_gate.py`. Of the four gates in `scripts/`, only `site_gates.py` is mandated anywhere (`CLAUDE.md:153`) and `manifest_bearer_gate.py` is named in `runbooks/secrets.md:76`. **In R-29's own words, carried forward deliberately: do not mint a new ID for a new instance** — the 2026-07-18 rehearsal independently re-raised this item and no second ID was minted then either **UPDATE 2026-08-02 — leg (a) CLOSED** (`felhom-controller` `c432f70`, its own reviewed diff as specified): `appexport/estimate.go`'s `-v` is a NAMED VOLUME mounted read-only into a throwaway container, no host path, structurally identical to the allowlisted `backup/backup.go` entry — allowlisted with its why; `realVolumeSize` untouched. **Leg (b) HALF-SHIPPED:** the 'decide where gates run' ruling is now made and half-implemented — **every repo has ONE entry point** (`felhom.eu/scripts/repo_gates.py`, `felhom-controller/controller/scripts/controller_gates.py`, `felhom-agent/scripts/agent_gates.py`, `app-catalog-felhom.eu/scripts/catalog_gates.py`), each mandated in its `CLAUDE.md` and each wired to `.githooks/pre-push` via `--fast`. **THE CENSUS, which is the finding:** thirteen gate scripts across four repos; **every gate a `CLAUDE.md` names was GREEN, and two of the four nobody names were RED** — `hostinstall_gates.py` (red since 2026-07-14) and `reuse_refs_check.py` (red on all four repos); a third, `docker_run_volume_path_gate.py`, was named only in `REUSE.md:284` and was also red. Correlation with 'named in a CLAUDE.md' was exact. **STAYS OPEN for the automatic half** — a hook is per-clone and `--no-verify` skips it; the unbypassable half is CI → **R-168** **CLOSED 2026-08-02, on the demonstrated ALARM and not on a green run.** Leg (b)'s automatic half is now live: a Gitea Actions runner re-runs every repo's entry point on every push, independent of who pushed and of what they typed (→ R-168). The class this row opened — *a gate that exists, asserts something true, is red, and is invoked by nothing* — is answered at both ends: the pre-push hook refuses locally, and CI catches a `--no-verify` bypass and **emails the operator**, proven with a real red run and a provider accepted-id. What remains is not this row's finding but a working-style choice — CI reports rather than blocks because there is no merge to gate (→ R-169) | CC |
|
||||
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
|
||||
| **R-86** | Restore-tests are interval-scheduled, not backup-aligned | **READY (M) — UNBLOCKED and re-ranked 2026-08-03** | — | Trigger a tier ~24 h after **its own** newest archive. **The R-90 dependency is discharged:** that row informed the cadence because ep0 had 3.8 GB and a 14.46 GB restore read had OOMed it, so a more frequent restore-test risked knocking the offsite endpoint over. ep0 is now a **CX33 with 8 GB RAM plus a 4 GiB swapfile** (measured 2026-08-03), so headroom is no longer what sets the cadence and this can be designed on its own merits. **Do not read that as "the ceiling is gone":** the OOM was a 14.46 GB restore against 3.8 GB, and 8 GB is comfortable rather than unbounded — the restore-test cadence should still be paced, just not by fear of the endpoint | CC |
|
||||
| **R-87** | The restic tier is never restore-tested | **READY** | — | Design a controller-side test (no scratch-guest analogue transfers) | CC |
|
||||
| **R-86** | Restore-tests are interval-scheduled, not backup-aligned | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.121.0**, hub **v0.91.0**) | — | **The rule that shipped:** *let A be the newest archive on a tier that has settled ≥24 h; the tier is DUE when A exists and A has not already been proven.* Daily tier → proved daily on yesterday's archive; weekly tier → weekly on its own; newborn → UNKNOWN. The daemon-start ticker survives only as the **evaluation interval**. **THE TRAP, recorded because it is the version a reasonable person writes:** the row's own wording implemented literally — *"due when the newest archive is ≥24 h old"* — is NEVER true on a **daily** tier, because a new archive resets the newest-archive age to zero long before it reaches the lag; it would have silently switched restore-testing OFF for the tier that matters most. Red-proved at **0 runs over 5 simulated days**. **The state now records WHICH archive was proven**, not when a tier last passed — a time cannot answer *have we proven this archive*. A pre-R-86 state file keeps its time (rotation ordering survives) and yields no proven archive, so each tier is due exactly once after the upgrade: the safe direction. **Two knobs replace one and the old one is not silently repurposed:** `restore_test_eval_interval_seconds` (6 h) and `restore_test_settle_seconds` (24 h); the deprecated `restore_test_cadence_seconds` keeps its DISABLE meaning verbatim, now seeds the settle lag, and the daemon WARNs once at start-up naming both. **6 h is bounded from both ends, not picked:** MEASURED cost of one evaluation on demo-felhom — local dir storage **18 ms**, PBS tier over the WAN to ep0 **392 ms**, both **430 ms** — so cost is irrelevant; the CEILING is that a FAILING tier stays due, making the evaluation interval its retry interval for a multi-GB restore. **Part 2 shipped WITH it and was not optional** — see the hub half in this row's sibling text and `07-backup-architecture.md` §3: `restoreProvenStaleAfter` was a flat 7 days derived from the very cadence this removed, and a healthy weekly tier's proof age reaches **exactly** 168 h against a 168 h window — it sat ON the line, so any ordinary delay tipped it into a nightly alarm about a working system. The window is now per tier from that tier's observed archive interval, ×4 generations, floored at the old 7 days and capped at 12 days (strictly inside the two-week offsite retention), falling back to the tier's DECLARED rhythm (26 h host / 8 d offsite — the backup-freshness checker's own thresholds) when history is too short to observe one. **A hollow test caught by its own red-proof:** the first Scenario-G fixture had no jitter and PASSED under the flat-window mutation, because a perfectly regular weekly tier sits exactly ON the line rather than over it. The jitter is what makes it a test. **Also fixed in passing:** the candidate picker now skips archives failing `archivePlausiblyComplete` (under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof and leave the tier due at EVERY evaluation), and the due-check runs BEFORE the heavy-operation gate is taken (a frequent poll must not be able to make a starting backup record a failure — F-A1). **Live proof:** see `felhom-agent/REPORT.md` | CC |
|
||||
| **R-87** | The restic tier is never restore-tested | **READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | — | Design a controller-side test (no scratch-guest analogue transfers). **Most of what this row needed now exists.** R-86 built the piece that was missing: a tier is proved **per archive generation**, on its own rhythm, with the proof recorded as *which archive* — which is exactly the shape a weekly-ish restic tier needs, and the reason this row could not simply reuse the whole-guest scheduler before. What remains is genuinely restic-specific and is NOT a scheduling problem: there is no scratch-guest analogue, so the test has to be a controller-side restore of a bounded sample into a throwaway path, with its own definition of "proved". **Two things to carry over rather than re-derive:** the proof must record the SNAPSHOT it proved (not a timestamp), and the hub's staleness window must learn this tier's rhythm the way `restoreProvenWindow` now does — a restic tier on a weekly cadence lands on the same false-alarm line the flat 7 days did. **And R-95 still applies:** that credential can delete, so a restic restore-test must never be able to write to the repo | CC |
|
||||
| **R-185** | **The agent cannot see the host backup tier's archives on demo-felhom — the PVE token has no ACL on `/storage/felhom-backup`, so the content listing returns EMPTY where root sees three archives.** Found 2026-08-03 while live-validating R-86. `pveum acl list` grants `FelhomAgentStore` on `/storage/{local,local-lvm,felhom-pbs}` and **not** on `felhom-backup`, which is the box's actual `local_backup_target`. Verified three ways: `pvesh` as root lists 3 archives (6.1–6.3 GB, 08-01/02/03); the same endpoint with the agent's token returns `{"data":[]}`; and `local` — which HAS a grant — returns its archives through the same token | **OPEN — filed, not fixed** | — | **Pre-existing and independent of R-86** (it is a property of the ACL, and the R-85 rotation had the same blindness). **Consequences:** the host tier has never been restore-testable on that box, and R-85's *"an empty tier is skipped, not failed"* rule made that silent. **The part worth fixing is the silence, not only the grant:** a permission-blinded tier is today INDISTINGUISHABLE from a newborn one — both report *"no settled archive yet"* — which is this project's own absence-is-not-evidence rule failing in a new place. The agent already knows better: it RECORDS successful backups to that target, so *"I wrote archives here and the tier lists none"* is a contradiction it can detect and should say loudly. **Do not fix by widening the token blind:** decide whether the host-install ACL set should follow `local_backup_target` (it currently hardcodes `local`), which is where the drift began | CC |
|
||||
| **R-186** | **A released agent binary's sha256 cannot be reproduced from its tag.** `release-agent.sh` builds at step 3 and tags at step 4, so Go's VCS stamp records a PSEUDO-version (`v0.120.1-0.20260803130452-4d825910…`) in the published bytes, while any rebuild after the tag exists stamps `v0.121.0` — a different binary. Measured 2026-08-03 on v0.121.0: published `b2128f3c…` (14 081 336 B) vs rebuild-at-tag `8302e396…` (14 077 240 B), identical source, identical toolchain, 4 096 bytes apart | **OPEN** | — | **Why it matters:** the sha the operator vouches is the one thing tying a machine to a binary, and today nobody can independently rebuild it to check. **The build order is deliberate** (the script's own comment: a tag with no package is caught by `check-published-versions.py`, a package with no tag is invisible to it), so the fix is not to swap the steps blind. Candidates: `-buildvcs=false` or `-trimpath` for a version-stable stamp, or tag-then-build with the tag deleted on a failed publish. **Mitigation used this session:** the DEPLOYED binary is the PUBLISHED artifact, downloaded from Gitea — not a local rebuild — so the running bytes are the vouchable ones | CC |
|
||||
| **R-187** | **R-115's one-command release had never actually run its publish leg — the first real use died there.** `scripts/publish-agent.sh` has been mode `0644` since it was created (2026-06-28), because every earlier caller invoked it as `bash scripts/publish-agent.sh`; `release-agent.sh` (written 2026-08-03) called it directly and got `Permission denied` on v0.121.0's release | **CLOSED — SHIPPED 2026-08-03** (`felhom-agent`) | — | **Fixed both ways in one commit:** the executable bit restored, and the caller changed to `bash "$REPO_ROOT/scripts/publish-agent.sh"` so the release no longer depends on a file mode — the kind of thing a checkout, an archive or a copy silently loses again. **The lesson is R-115's own, one level up:** the mechanism written to make a step unforgettable was itself never exercised end-to-end, so it failed the first time it mattered. A mechanism that has not been RUN is a note with better formatting | CC |
|
||||
| — | Storage Box **snapshots** on `storage-box-pool-1` — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm `size_snapshots > 0` tomorrow; until then the mitigation is armed, not proven | CC |
|
||||
| — | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | WAITING-ON-OPERATOR | operator console | Delete the box | operator |
|
||||
| **R-90** | ~~ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged~~ | **CLOSED — the operator rescaled ep0 to a CX33 on 2026-08-03** | — | **MEASURED ON THE BOX, not read from an invoice:** `felhom-hetzner` reports `Mem: 7757` MB total (**8 GB**, was 3.8) and `nproc` **4**. **The interim lever survived and was checked rather than assumed** — a resize is a stop/start, so "the swapfile is still there" was an assumption until measured: `/swapfile`, 4 GiB, dated `Jul 27 14:40`, **active** (`swapon --show` → `/swapfile file 4G 0B -2`), 0 B in use on an idle box. **THE 40 GB LOCAL DISK DID NOT CHANGE** and must not be "corrected" alongside the RAM: `/` is 38 G, 58% used. This was a CPU/RAM resize only, so every disk figure in the runbooks still stands — the separate 98 G volume at `/mnt/pbs-datastore` (R-82 P0.3) is unaffected. **Why this was BLOCKED and no longer is:** the row recorded CX33 as *"confirmed unavailable even powered OFF"* — the Cost-Optimized line's limited availability, not a power-state problem. It became available and the operator took it. **Documentation corrected** (`RUNBOOK-ep0-datastore-volume`, `RUNBOOK-pbs-prune-serverside` ×2, `runbooks/offsite-endpoint.md` ×2, `runbooks/target-selection.md`) and **audit/evidence documents ANNOTATED, not revised** (`SPIKE-connectivity-wireguard-2026-07-03`, campaign-10 `phaseA-journal`) — they record what was true when written and that is their value. **Still open and still the operator's, deliberately untouched:** `target-selection.md`'s *"D-d did not name ep0 either way. Confirm it explicitly."* | — |
|
||||
@@ -129,9 +132,15 @@ there is one ranking to maintain rather than two.
|
||||
already fetches 1.22.0. What remains is a wrong label plus two pieces of dead safety equipment —
|
||||
a drift gate nobody runs and a test that compares a constant to itself. Cheap and worth doing;
|
||||
not high-consequence, and it blocks nothing.
|
||||
3. **R-86** — an operator ruling already exists; it only waits on knowing what load ep0 can take.
|
||||
4. **R-87** — real and unbuilt, but needs its own design, so it should not jump work that is specified.
|
||||
5. **R-110** — last **because it is not a READY row**: the ruling is the operator's, not CC's, and
|
||||
3. ~~**R-86**~~ — **CLOSED 2026-08-03**, agent v0.121.0 + hub v0.91.0, proven live on demo-felhom.
|
||||
4. **R-87** — **re-ranked UP**: R-86 built most of what it was waiting for (per-archive due-ness, a
|
||||
proof that names its archive, and a staleness window that learns a tier's rhythm). What is left is
|
||||
restic-specific — there is no scratch-guest analogue — so it still needs its own design, but it is
|
||||
no longer waiting on a scheduling model that did not exist.
|
||||
5. **R-185** — the agent is blind to demo-felhom's host backup tier (a missing storage ACL), and the
|
||||
blindness reads exactly like a newborn tier. Small to fix, and the *silence* is the part worth
|
||||
fixing, not just the grant.
|
||||
6. **R-110** — last **because it is not a READY row**: the ruling is the operator's, not CC's, and
|
||||
there is nothing for CC to build until it lands. Ranked here rather than omitted because it is the
|
||||
only item on this page about the *publish channel* of the most privileged artifact Felhom ships,
|
||||
and today's exposure is zero — which makes now the cheapest moment it will ever be to decide.
|
||||
|
||||
Reference in New Issue
Block a user