docs(TASK-D): R-51/R-52 SHIPPED + new R-54 row; capability-map row; seam-discipline rider

R-51's roadmap diagnosis is corrected at the source: aggregation returned StateRunning
("partial") for a running/stopped mix, so the stack read RUNNING and IsDownState was never
consulted about  at all — the constraint that row protects was never in tension
with the fix.

New R-54 row closes the INCIDENT-guest-dhclient-killed-2026-07-20 §5 OPEN RISK, and records
the design fact that makes it work: liveness of the DHCP client is itself a probe, because
the damage is timed and the address outlives its cause by 1-2 hours. The static-guest leg is
deliberately deferred to R-50.

New capability-map row is IMPLEMENTED, not PROVEN-LIVE: one leg is live (the watchdog's
healthy cycle on felhom-pve), the three that matter are destructive and operator-present and
have not run.

PROMPT-TEMPLATE §10 gains the seam-discipline row, including that a strings.Contains source
assertion is NOT sufficient — a commented-out call still contains the string.
This commit is contained in:
2026-07-21 12:40:00 +02:00
parent 50a7ffacd2
commit 907e5ce65c
4 changed files with 78 additions and 92 deletions
+65 -90
View File
@@ -2,111 +2,86 @@
> **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md).
# TASK-B — R-39 fleet fix + R-50b(a) · hub **v0.67.0 → v0.68.0**
---
**Date:** 2026-07-21 · **Baseline:** `c35da9d` (clean, == `origin/main`) → `54a4644`.
Companion: agent **v0.91.2** (`felhom-agent`, see that repo's `REPORT.md` for the agent half,
STOP-1 evidence and the four red-proofs).
# TASK-D — unattended resilience: docs, roadmap, capability map, riders (2026-07-21)
## Status
**No hub code was touched** (§12 of the brief). This repo's share is documentation, the two Part-5
riders, and the honest status of the three claims the task makes.
| Leg | Status |
|---|---|
| Hub v0.68.0 | **SHIPPED + DEPLOYED** — GitOps manifest bump `0.67.0→0.68.0`, ArgoCD `Synced/Healthy`, rollout complete, pod 1/1 |
| Agent v0.91.2 | shipped + published + deployed to felhom-pve |
| Hub v0.68.1 | **SHIPPED + DEPLOYED** — fixes the Configuration layout the v0.68.0 field broke |
| STOP-1 | done + verified |
| **STOP-3 — manifest save** | **DONE by the operator 2026-07-21 08:31:12Z.** All six fields persisted, including `artifact_wrapper_sha256 = 104db0a4…` (matches the agent's reported hash → drift gauge reads **ok**) and `artifact_min_agent = 0.91.2`. |
| **STOP-2 — live re-issue** | **DONE + PROVEN 2026-07-21.** Chain closed in **13 seconds**; see below. |
Implementation reports live in the repos that shipped: `felhom-controller/REPORT.md` (R-51, R-52,
v0.156.0) and `felhom-agent/REPORT.md` (R-54, v0.92.1).
## STOP-2 — the proof (the whole task's reason to exist)
## 1. ROADMAP
The operator pressed **Re-issue PBS credentials**. The identical click on 2026-07-18 did nothing.
- **R-51 → SHIPPED (controller v0.156.0)**, and **the row's diagnosis is corrected in place**. It
claimed aggregation classified a dead-primary stack `unhealthy`, making the deliberate `unhealthy`
exclusion the suppressor. The source says otherwise: `aggregateState`'s final branch returned
`StateRunning` for any running/stopped mix ("report as running (partial)"), so the stack read
**running** and `IsDownState` was never consulted about `unhealthy` at all. The constraint the row
protects was therefore never in tension with the fix, and `downstate_test.go` ships untouched.
- **R-52 → SHIPPED (controller v0.156.0)**, `internal/bootrecon`. Note recorded that **P1 is not a
blocker**: the reconciliation is correct whether or not Docker recorded those two containers as
user-stopped, and P1's experiment falls out of STOP-1's reboot leg for free.
- **R-54 → NEW ROW, SHIPPED (agent v0.92.1).** Origin: `INCIDENT-guest-dhclient-killed-2026-07-20`
§5, whose OPEN RISK it closes. The row carries the load-bearing design fact — *liveness of the
DHCP client is itself a probe, because the damage is timed and the address survives the cause by
12 hours* — the live sudoers finding, and the explicit deferral of the **static-guest** leg to
**R-50**, which is where the incident's own option 2 ("give the guest a static address") belongs.
```
hub 08:39:31Z fresh mint, generation 0 -> 1
descriptor gains "secret_generation": 1
(token_id + fingerprint BYTE-IDENTICAL — the re-key shape that was invisible)
agent 10:39:34 felhom-pbs-apply read felhom-pbs <- leg (b): the read that was impossible
agent 10:39:38 ERROR "REJECTED ... applied and DEAD" previous_state=applied
<- leg (c): the R-39 state, loud at last
hub 08:39:45Z consumed_at stamped
agent 10:39:45 "one-time token secret consumed" secret_len=36
<- leg (a): NO short-circuit
agent 10:39:45 felhom-pbs-apply reconcile (set-only, no --server)
agent 10:39:47 "pbsdr: converged" state=applied
```
## 2. Capability map
(host CEST = UTC+2; hub timestamps UTC.)
New row in section: **"Box survives an unattended app or guest-network failure"** — status
**IMPLEMENTED**, deliberately **not** PROVEN-LIVE.
**Corroboration:** the agent marker hash moved to `afbb3b41…` — in the failure it was *byte-identical*
to the pre-reissue marker, which was the single-line proof of the defect. `consumed_at` stamped. The
on-disk secret's mtime moved `2026-07-18 20:28:52``2026-07-21 10:39:45`. A live probe with the NEW
credential returns **200**. Three consecutive hub reports trace the entire state machine
`applied → auth_failed → applied`. **Zero** `pbsdr_selfheal` escalations fired, with exactly ONE mint,
ONE consume and no `consumed-failed.json` — the box healed through the descriptor path before the
damper was ever needed.
One leg genuinely is live and is cited as such: the guest-network watchdog's healthy cycle on
felhom-pve (`ok=68 total=68 degraded=0` and `DEBUG guestnet: guest network healthy vmid=9201
mode=dhcp has_route=true dhclient_alive=true`, 12:34:15 CEST). The three legs that would earn
PROVEN-LIVE are all destructive and operator-present, and none has run: killing immich's primary,
rebooting 9201 to strand and recover boot orphans, and replaying the dhclient kill. The strict enum
says a row without that evidence is not PROVEN-LIVE, so it is not.
**First attempt, worth recording:** the operator initially pressed the **offsite** re-issue — there are
two distinct Re-issue actions and my instruction said only "press Re-issue". Harmless to PBS-DR, but it
rotated the restic password and correctly marked the escrow **stale**, so the recovery-code ceremony
had to be re-run (done). Name the surface explicitly in future runbook steps.
The existing **"Box survives a site/network change"** row (PARTIAL) already pointed at R-51/R-52 as
related work; it stays PARTIAL — R-50 is still its durable fix.
## What shipped hub-side
## 3. Part-5 riders
- **`host_pbs_secrets.generation`** — a monotonic per-host counter advanced by every fresh MINT and
by nothing else, stamped into the descriptor as `secret_generation`. Since the agent re-applies on
the descriptor's CONTENT HASH and a re-key returns byte-identical `token_id` / `fingerprint` /
`datastore` / `namespace`, this is the only field that moves — and therefore the thing that
re-arms a converged agent.
- A **re-stage** deliberately does not advance it (same secret, unchanged descriptor content).
- `omitempty` is load-bearing: emitting a zero would shift every pre-existing descriptor's hash at
once — a fleet-wide spurious re-apply.
- **Deviation from spec, deliberate:** the brief said to reuse "the new row's id … no schema
change". There is no row id — the table is `host_id PRIMARY KEY`, UPSERTed last-write-wins — and
`created_at` collides for two mints in one second. An additive counter column (existing
idempotent `ALTER TABLE` idiom) is the only monotonic source. **Verified applied on the live DB
after deploy.**
- **`pbsdrheal` gains an `auth_failed` trigger** — a new trigger in the existing machine, escalating
to a fresh mint (never a re-stage, which would re-feed the secret PBS just rejected) through the
**existing damper**, so a 401 flap cannot become a secret-minting chain.
- **`consumed_at` honesty gauge** — an unconsumed secret past a 15-minute grace under a box reporting
`applied` is the exact 2026-07-18 fingerprint and a disagreement **no single tier can detect
alone**. Surfaced with its own event, deliberately as a SURFACE not a heal: auto-re-issuing would
mint a second secret on top of an unconsumed one, which is the mint/consume race R-39(a) recorded.
- **Corrected a comment that stated a falsehood** — `ReissuePBSDR` claimed it refreshed the
descriptor "with the NEW token_id/fingerprint". False for a re-key, and believing it is why nobody
expected the descriptor to come back identical.
- **R-50b(a)** — `ArtifactManifest.WrapperSHA256` + operator field + host-page drift surface. **An
unknown on either side reads as quiet, never as drift.**
- **Hub `build.sh` deploy hint → GitOps wording.** The script printed `kubectl set image …` and
`kubectl apply -f manifests/hub.yaml` as the deploy instructions — both are reverted by the next
ArgoCD sync, which is the worst failure shape: it appears to work, then silently disappears. Now
prints the manifest-bump + hard-refresh + deliberate-sync sequence, and names the trap explicitly.
**Caveat worth knowing: `/mnt/5_hdd/felhom.eu/build/felhom-hub/build.sh` is NOT in any git repo**
it is a DooPlex-local build-dir script. So this rider is a fix to the operative file only, and it
is not versioned anywhere. Worth adopting into the repo as its own change.
- **`documentation/PROMPT-TEMPLATE.md` §10 gains the seam-discipline row.** Text generalised from
the brief's §9 rule 6, plus what this session added to it: three shipped inert-seam defects in
three days (controller v0.154.0, agent v0.91.0, agent v0.92.0's missing sudoers grant), all fully
green; and the finding that a `strings.Contains` source assertion is not sufficient, because a
commented-out call still contains the string — walk the AST.
## Method notes worth keeping
## 4. What was verified, and how
- **P2 confirmed GitOps-only, and the trap is real:** `build.sh` itself prints
`kubectl set image …` as its deploy hint, contradicting `CLAUDE.md`. Not used. Worth fixing in the
script — it will mislead exactly the session that trusts tool output over the runbook.
- **P4 read the fleet from a temporary copy of the hub DB**, which carries live credentials
(`api_key`, `host_pbs_secrets.value`). Copy shredded immediately after each read. Result: **one
enrolled host**, so the MinAgent raise strands nobody.
- Tests include a **flow-level** `ReissuePBSDR` test against a fake that models a real re-key
(identical token/fingerprint, rotated secret only). Its red-proof fails on the assertion with both
byte-identical blocks printed — the July-18 defect reproduced in a unit test.
Endpoint/host-level, no browser (none on DooPlex). Everything claimed live in the two implementation
reports was read out of journald or the agent's own capability self-check on felhom-pve. No claim in
any of the three reports rests on a test that only proves a seam.
## For the operator
## 5. An ordering question for the operator (STOP-1 vs STOP-3)
**STOP-2 (one click):** press **Re-issue PBS credentials** for the demo customer. Expected:
fresh secret row → `secret_generation` **0 → 1** (the live descriptor has no such key today) → poke →
agent re-applies with **no short-circuit** → fresh secret consumed → reconcile rc-0 → probe 200 → tier
`active`. The July-18 negative — the same click doing nothing — is the historical red-proof.
The brief's Phase A says **do not hand-deploy** the controller — the floor save at **STOP-3** is what
deploys 0.156.0, banking another single-fire self-update datapoint. But **STOP-1 exercises the
controller legs**, which need 0.156.0 to be live. As written the two are in the wrong order.
**STOP-3 (manifest save):** Agent `0.91.2` / sha256 `34d309be429473f3f0ab34e3185e17b22463a341b30bf46e162306bff4aec22a` /
**PBS wrapper sha256** `104db0a4401f65bbc476e82bfb1796433bcb36f8f8cce69efb3bb5c40fcb16b3` /
MinAgent `0.91.2`. Superseded, do not vouch: 0.91.0 (inert probe leg), 0.91.1. `0.90.1` correctly
stays 404.
The resolution that keeps both intentions is: **do the STOP-3 controller floor save first** (floor →
`0.156.0`), let the box self-update — that IS the R-23 datapoint — and then run STOP-1 against the
new version. The agent half of STOP-3 (manifest → 0.92.1) is independent and can happen at any
point. Flagged rather than assumed, because reordering an operator's STOP sequence is not CC's call.
## Residual
## 6. Follow-ups this session surfaced (none actioned here)
R-50b **(b)/(c) remain open** — the wrapper is still fetched unversioned from `raw/branch/main`; this
release makes drift visible, it does not fix the channel. The 0440 sudoers file is not agent-readable,
so its drift stays invisible. The DR-tier capability-map row is deliberately **not** upgraded to
PROVEN-LIVE until STOP-2 supplies the evidence.
1. `felhom-controller/controller/.gitignore`'s bare `controller` entry also matches the directory
`cmd/controller/`, so ripgrep silently skips `main.go` and new files there need `git add -f`.
Both directions produce inert-seam mistakes. XS fix: anchor it as `/controller`.
2. The hub `build.sh` above is unversioned.
3. `TestGenerateRecoveryCode_EntropyAndFormat` (agent) flakes on hyphenated wordlist entries
(`drop-down` → 11 tokens). Observed 3/8 this session — worse than the documented ~1/5, and it is
a one-line fix in the generator or the assertion, not a mystery.