diff --git a/REPORT.md b/REPORT.md index 939c9af..2a107b1 100644 --- a/REPORT.md +++ b/REPORT.md @@ -2,111 +2,86 @@ > **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md). -# TASK-B — R-39 fleet fix + R-50b(a) · hub **v0.67.0 → v0.68.0** +--- -**Date:** 2026-07-21 · **Baseline:** `c35da9d` (clean, == `origin/main`) → `54a4644`. -Companion: agent **v0.91.2** (`felhom-agent`, see that repo's `REPORT.md` for the agent half, -STOP-1 evidence and the four red-proofs). +# TASK-D — unattended resilience: docs, roadmap, capability map, riders (2026-07-21) -## Status +**No hub code was touched** (§12 of the brief). This repo's share is documentation, the two Part-5 +riders, and the honest status of the three claims the task makes. -| Leg | Status | -|---|---| -| Hub v0.68.0 | **SHIPPED + DEPLOYED** — GitOps manifest bump `0.67.0→0.68.0`, ArgoCD `Synced/Healthy`, rollout complete, pod 1/1 | -| Agent v0.91.2 | shipped + published + deployed to felhom-pve | -| Hub v0.68.1 | **SHIPPED + DEPLOYED** — fixes the Configuration layout the v0.68.0 field broke | -| STOP-1 | done + verified | -| **STOP-3 — manifest save** | **DONE by the operator 2026-07-21 08:31:12Z.** All six fields persisted, including `artifact_wrapper_sha256 = 104db0a4…` (matches the agent's reported hash → drift gauge reads **ok**) and `artifact_min_agent = 0.91.2`. | -| **STOP-2 — live re-issue** | **DONE + PROVEN 2026-07-21.** Chain closed in **13 seconds**; see below. | +Implementation reports live in the repos that shipped: `felhom-controller/REPORT.md` (R-51, R-52, +v0.156.0) and `felhom-agent/REPORT.md` (R-54, v0.92.1). -## STOP-2 — the proof (the whole task's reason to exist) +## 1. ROADMAP -The operator pressed **Re-issue PBS credentials**. The identical click on 2026-07-18 did nothing. +- **R-51 → SHIPPED (controller v0.156.0)**, and **the row's diagnosis is corrected in place**. It + claimed aggregation classified a dead-primary stack `unhealthy`, making the deliberate `unhealthy` + exclusion the suppressor. The source says otherwise: `aggregateState`'s final branch returned + `StateRunning` for any running/stopped mix ("report as running (partial)"), so the stack read + **running** and `IsDownState` was never consulted about `unhealthy` at all. The constraint the row + protects was therefore never in tension with the fix, and `downstate_test.go` ships untouched. +- **R-52 → SHIPPED (controller v0.156.0)**, `internal/bootrecon`. Note recorded that **P1 is not a + blocker**: the reconciliation is correct whether or not Docker recorded those two containers as + user-stopped, and P1's experiment falls out of STOP-1's reboot leg for free. +- **R-54 → NEW ROW, SHIPPED (agent v0.92.1).** Origin: `INCIDENT-guest-dhclient-killed-2026-07-20` + §5, whose OPEN RISK it closes. The row carries the load-bearing design fact — *liveness of the + DHCP client is itself a probe, because the damage is timed and the address survives the cause by + 1–2 hours* — the live sudoers finding, and the explicit deferral of the **static-guest** leg to + **R-50**, which is where the incident's own option 2 ("give the guest a static address") belongs. -``` -hub 08:39:31Z fresh mint, generation 0 -> 1 - descriptor gains "secret_generation": 1 - (token_id + fingerprint BYTE-IDENTICAL — the re-key shape that was invisible) -agent 10:39:34 felhom-pbs-apply read felhom-pbs <- leg (b): the read that was impossible -agent 10:39:38 ERROR "REJECTED ... applied and DEAD" previous_state=applied - <- leg (c): the R-39 state, loud at last -hub 08:39:45Z consumed_at stamped -agent 10:39:45 "one-time token secret consumed" secret_len=36 - <- leg (a): NO short-circuit -agent 10:39:45 felhom-pbs-apply reconcile (set-only, no --server) -agent 10:39:47 "pbsdr: converged" state=applied -``` +## 2. Capability map -(host CEST = UTC+2; hub timestamps UTC.) +New row in section: **"Box survives an unattended app or guest-network failure"** — status +**IMPLEMENTED**, deliberately **not** PROVEN-LIVE. -**Corroboration:** the agent marker hash moved to `afbb3b41…` — in the failure it was *byte-identical* -to the pre-reissue marker, which was the single-line proof of the defect. `consumed_at` stamped. The -on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`. A live probe with the NEW -credential returns **200**. Three consecutive hub reports trace the entire state machine -`applied → auth_failed → applied`. **Zero** `pbsdr_selfheal` escalations fired, with exactly ONE mint, -ONE consume and no `consumed-failed.json` — the box healed through the descriptor path before the -damper was ever needed. +One leg genuinely is live and is cited as such: the guest-network watchdog's healthy cycle on +felhom-pve (`ok=68 total=68 degraded=0` and `DEBUG guestnet: guest network healthy vmid=9201 +mode=dhcp has_route=true dhclient_alive=true`, 12:34:15 CEST). The three legs that would earn +PROVEN-LIVE are all destructive and operator-present, and none has run: killing immich's primary, +rebooting 9201 to strand and recover boot orphans, and replaying the dhclient kill. The strict enum +says a row without that evidence is not PROVEN-LIVE, so it is not. -**First attempt, worth recording:** the operator initially pressed the **offsite** re-issue — there are -two distinct Re-issue actions and my instruction said only "press Re-issue". Harmless to PBS-DR, but it -rotated the restic password and correctly marked the escrow **stale**, so the recovery-code ceremony -had to be re-run (done). Name the surface explicitly in future runbook steps. +The existing **"Box survives a site/network change"** row (PARTIAL) already pointed at R-51/R-52 as +related work; it stays PARTIAL — R-50 is still its durable fix. -## What shipped hub-side +## 3. Part-5 riders -- **`host_pbs_secrets.generation`** — a monotonic per-host counter advanced by every fresh MINT and - by nothing else, stamped into the descriptor as `secret_generation`. Since the agent re-applies on - the descriptor's CONTENT HASH and a re-key returns byte-identical `token_id` / `fingerprint` / - `datastore` / `namespace`, this is the only field that moves — and therefore the thing that - re-arms a converged agent. - - A **re-stage** deliberately does not advance it (same secret, unchanged descriptor content). - - `omitempty` is load-bearing: emitting a zero would shift every pre-existing descriptor's hash at - once — a fleet-wide spurious re-apply. - - **Deviation from spec, deliberate:** the brief said to reuse "the new row's id … no schema - change". There is no row id — the table is `host_id PRIMARY KEY`, UPSERTed last-write-wins — and - `created_at` collides for two mints in one second. An additive counter column (existing - idempotent `ALTER TABLE` idiom) is the only monotonic source. **Verified applied on the live DB - after deploy.** -- **`pbsdrheal` gains an `auth_failed` trigger** — a new trigger in the existing machine, escalating - to a fresh mint (never a re-stage, which would re-feed the secret PBS just rejected) through the - **existing damper**, so a 401 flap cannot become a secret-minting chain. -- **`consumed_at` honesty gauge** — an unconsumed secret past a 15-minute grace under a box reporting - `applied` is the exact 2026-07-18 fingerprint and a disagreement **no single tier can detect - alone**. Surfaced with its own event, deliberately as a SURFACE not a heal: auto-re-issuing would - mint a second secret on top of an unconsumed one, which is the mint/consume race R-39(a) recorded. -- **Corrected a comment that stated a falsehood** — `ReissuePBSDR` claimed it refreshed the - descriptor "with the NEW token_id/fingerprint". False for a re-key, and believing it is why nobody - expected the descriptor to come back identical. -- **R-50b(a)** — `ArtifactManifest.WrapperSHA256` + operator field + host-page drift surface. **An - unknown on either side reads as quiet, never as drift.** +- **Hub `build.sh` deploy hint → GitOps wording.** The script printed `kubectl set image …` and + `kubectl apply -f manifests/hub.yaml` as the deploy instructions — both are reverted by the next + ArgoCD sync, which is the worst failure shape: it appears to work, then silently disappears. Now + prints the manifest-bump + hard-refresh + deliberate-sync sequence, and names the trap explicitly. + **Caveat worth knowing: `/mnt/5_hdd/felhom.eu/build/felhom-hub/build.sh` is NOT in any git repo** — + it is a DooPlex-local build-dir script. So this rider is a fix to the operative file only, and it + is not versioned anywhere. Worth adopting into the repo as its own change. +- **`documentation/PROMPT-TEMPLATE.md` §10 gains the seam-discipline row.** Text generalised from + the brief's §9 rule 6, plus what this session added to it: three shipped inert-seam defects in + three days (controller v0.154.0, agent v0.91.0, agent v0.92.0's missing sudoers grant), all fully + green; and the finding that a `strings.Contains` source assertion is not sufficient, because a + commented-out call still contains the string — walk the AST. -## Method notes worth keeping +## 4. What was verified, and how -- **P2 confirmed GitOps-only, and the trap is real:** `build.sh` itself prints - `kubectl set image …` as its deploy hint, contradicting `CLAUDE.md`. Not used. Worth fixing in the - script — it will mislead exactly the session that trusts tool output over the runbook. -- **P4 read the fleet from a temporary copy of the hub DB**, which carries live credentials - (`api_key`, `host_pbs_secrets.value`). Copy shredded immediately after each read. Result: **one - enrolled host**, so the MinAgent raise strands nobody. -- Tests include a **flow-level** `ReissuePBSDR` test against a fake that models a real re-key - (identical token/fingerprint, rotated secret only). Its red-proof fails on the assertion with both - byte-identical blocks printed — the July-18 defect reproduced in a unit test. +Endpoint/host-level, no browser (none on DooPlex). Everything claimed live in the two implementation +reports was read out of journald or the agent's own capability self-check on felhom-pve. No claim in +any of the three reports rests on a test that only proves a seam. -## For the operator +## 5. An ordering question for the operator (STOP-1 vs STOP-3) -**STOP-2 (one click):** press **Re-issue PBS credentials** for the demo customer. Expected: -fresh secret row → `secret_generation` **0 → 1** (the live descriptor has no such key today) → poke → -agent re-applies with **no short-circuit** → fresh secret consumed → reconcile rc-0 → probe 200 → tier -`active`. The July-18 negative — the same click doing nothing — is the historical red-proof. +The brief's Phase A says **do not hand-deploy** the controller — the floor save at **STOP-3** is what +deploys 0.156.0, banking another single-fire self-update datapoint. But **STOP-1 exercises the +controller legs**, which need 0.156.0 to be live. As written the two are in the wrong order. -**STOP-3 (manifest save):** Agent `0.91.2` / sha256 `34d309be429473f3f0ab34e3185e17b22463a341b30bf46e162306bff4aec22a` / -**PBS wrapper sha256** `104db0a4401f65bbc476e82bfb1796433bcb36f8f8cce69efb3bb5c40fcb16b3` / -MinAgent `0.91.2`. Superseded, do not vouch: 0.91.0 (inert probe leg), 0.91.1. `0.90.1` correctly -stays 404. +The resolution that keeps both intentions is: **do the STOP-3 controller floor save first** (floor → +`0.156.0`), let the box self-update — that IS the R-23 datapoint — and then run STOP-1 against the +new version. The agent half of STOP-3 (manifest → 0.92.1) is independent and can happen at any +point. Flagged rather than assumed, because reordering an operator's STOP sequence is not CC's call. -## Residual +## 6. Follow-ups this session surfaced (none actioned here) -R-50b **(b)/(c) remain open** — the wrapper is still fetched unversioned from `raw/branch/main`; this -release makes drift visible, it does not fix the channel. The 0440 sudoers file is not agent-readable, -so its drift stays invisible. The DR-tier capability-map row is deliberately **not** upgraded to -PROVEN-LIVE until STOP-2 supplies the evidence. +1. `felhom-controller/controller/.gitignore`'s bare `controller` entry also matches the directory + `cmd/controller/`, so ripgrep silently skips `main.go` and new files there need `git add -f`. + Both directions produce inert-seam mistakes. XS fix: anchor it as `/controller`. +2. The hub `build.sh` above is unversioned. +3. `TestGenerateRecoveryCode_EntropyAndFormat` (agent) flakes on hyphenated wordlist entries + (`drop-down` → 11 tokens). Observed 3/8 this session — worse than the documented ~1/5, and it is + a one-line fix in the generator or the assertion, not a mystery. diff --git a/documentation/PROMPT-TEMPLATE.md b/documentation/PROMPT-TEMPLATE.md index 6a68c7f..8508e8b 100644 --- a/documentation/PROMPT-TEMPLATE.md +++ b/documentation/PROMPT-TEMPLATE.md @@ -226,6 +226,15 @@ Then: [exact refusal — HTTP status, error, and the proven non-effect, e.g. "m `copier` interface so tests don't shell out to docker/rsync). - **Generation/idempotency:** assert the negative (e.g. "fetch count does NOT increment on an unchanged heartbeat"; "a re-run finds its own prior `(N)` and adds no `(1)(1)`"). +- **Seam discipline — every seam added gets ONE test through the PRODUCTION wiring path.** An + injected-seam test proves the component, never the caller. Three shipped defects in three days + make this non-negotiable: controller v0.154.0 (the wizard read a flag the handler never sets), + agent v0.91.0 (`main.go` never called `SetAuthSink`, so the whole auth-honesty leg was inert), and + agent v0.92.0 (the watchdog had no sudoers grant for three of its four probes). **Every one of + them was fully green.** Where the caller is `func main()` and cannot be invoked from a test, walk + its AST for the call — and note that a `strings.Contains` on the source is NOT sufficient: a + commented-out call still contains the string, which is how the controller's first version of that + test passed its own red-proof (2026-07-21). --- diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 7744af5..0d4ca54 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -62,6 +62,7 @@ | Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D | | Guest-loss DR: PBS restore with full-fidelity layout from archive, restore-test verification | agent v0.75/0.76, PBS | **PROVEN-LIVE** | `CAMPAIGN-2` T-P9-DESTROY-RESTORE (whole-guest `pct restore` of 9201 → running+healthy) + T-PBS-VERIFY (`verify_state: ok`, 13 snapshots); `DRILL-GL6-2026-07-08` Phase 0d (restore-test `mount_parity: ok`) | (Cited `VALIDATION-newbox-restore` is offbox **restic** file-restore, wrong tier — corrected.) Real **offsite** guest-loss round-trip still R1-blocked → S5 DR drill | | PBS-DR secret self-heal on reused-peer re-provision | hub v0.56 | **IMPLEMENTED** | hub v0.56.0 (`pbsdrheal/reconciler.go`, `RestageHostPBSSecret`, all §10 red-proofs); `SPIKE-pbsdr-selfheal-2026-07-15` (root cause) | Reconciler is **scoped to one host** (`PBSDRHEAL_ONLY_HOST`), not fleet-wide; already fired live hands-free on drill qm300 (07-15) — real-customer firing + fleet-wide widening pending | +| Box survives an **unattended app or guest-network failure** (a dead app member, a boot-orphaned app, a dead DHCP client) — it is noticed, and where safe it is repaired | controller v0.156.0, agent v0.92.1 | **IMPLEMENTED** | Unit + red-proofed both repos (`felhom-controller/REPORT.md`, `felhom-agent/REPORT.md`, 2026-07-21). One leg IS live: the guest-network watchdog's healthy cycle on felhom-pve — caps `68/68 ok, degraded=0` and `DEBUG guestnet: guest network healthy vmid=9201 mode=dhcp has_route=true dhclient_alive=true` (2026-07-21 12:34:15 CEST) | Deliberately **not** PROVEN-LIVE: the three legs that matter are all destructive and operator-present, and none has run yet — killing immich's primary (R-51), rebooting 9201 to strand and then recover boot orphans (R-52), and a deliberate replay of the 2026-07-20 dhclient kill (R-54). Origin of all three: `audits/AUDIT-vacation-remote-ops-2026-07-20.md` F4/F5 + `audits/INCIDENT-guest-dhclient-killed-2026-07-20.md` §5. The **static-guest** half of the network leg is deliberately out of scope and belongs to **R-50** — a statically-configured guest missing its address is reported loudly and never healed. Upgrade this row only with STOP-1/STOP-2 evidence | | Crash/power-loss mid-backup/mid-migration → self-heal on next run | controller, agent | **PROVEN-LIVE** | `CAMPAIGN-6D` P5-REST (SIGKILL mid-offbox → auto-restart ~15s, run marked failed not false-success, no stale lock); `CAMPAIGN-6E` B1-B3 | (Cited `CAMPAIGN-2` T-RBT-* legs were empty / auth-hollow — corrected.) Live mid-**migration** crash→self-heal is the weakest sub-claim (P5-REST is mid-backup) | | Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent, controller, bootstrap | **PARTIAL** | `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — a real relocation of the demo box: guest + hub telemetry + WG/PBS + Cloudflare tunnel all survived untouched, but the **controller↔agent control plane did not** (agent binds a LAN literal → `bind: cannot assign requested address` → storage/PBS-backup/quiesce/restore-test/DR down until fixed). Mitigated for the window by pinning `vmbr0` static | **R-50** (island-bridge control plane, spike-first) is the durable fix. Related: **R-51** (dead-primary alerting) and **R-52** (boot desired-state reconciliation) — the same event left two apps `Exited` with no alarm and no recovery | | Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | | diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index 07f0a41..ff0aa78 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -60,8 +60,9 @@ | R-27c | **Customer self-bind, slice 2 — console-passphrase bind.** Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. | M | idea | **Security constraints from the session ruling, all load-bearing:** passphrase **issued at customer creation**; the global-lookup endpoint must be **spray-hardened** — per-appliance **and** per-IP caps, constant-time comparison, a **single generic failure** (no oracle), alerting on abuse; an **accent-free wordlist** (console keymaps are not Hungarian); the **web capability-link path is RETAINED**; **claim-by-email is RETAINED** as the delivery-channel proof. **Also under this item:** the self-bind email gains the **public universal-ISO download link + two-line instructions** (the DIY case). **Secret-bearing per-customer ISOs are ruled OUT.** Sibling of R-27b (second-box flow) — different axis, both build on the same `/bind/` page | | R-50 | **[P2-HIGH] Island-bridge control plane — make controller↔agent independent of the LAN.** The agent's `localapi` binds a **LAN literal** (`listen_addr`) and the guest dials that same literal from `bootstrap.json`. Move both onto a **host-internal bridge with a fixed, private address** that no router, DHCP lease, or site move can invalidate, so the control plane is location-independent and golden-image-stable. | M | **idea — SPIKE-FIRST** | Origin: `audits/AUDIT-vacation-remote-ops-2026-07-20.md` (F1), where this failed for real. The demo box moved to a remote site, DHCP handed it `.147` instead of `.162`, and the agent then **could not start at all** — `bind: cannot assign requested address`, systemd gave up after 4 retries — taking storage, PBS backup, quiesce, restore-test and DR down for as long as nobody noticed. Mitigated for that window by pinning `vmbr0` static back to `.162`; that is a **window mitigation, not the fix** — it still depends on the site's subnet being `192.168.0.0/24` and free at that address. **Spike-first is mandatory:** validate end-to-end on the drill environment (agent bind + guest dial + TLS SAN/pin + reinstall/golden survival + the bootstrap-config migration for already-deployed boxes) BEFORE any production spec. **Pin fact (verified 2026-07-20, `agentapi/client.go` L105-129 — supersedes the earlier "the SAN set must cover the new address" note in this entry, which was wrong):** the controller-to-agent leg sets `InsecureSkipVerify: true` and replaces chain verification with a custom `VerifyPeerCertificate` that does a raw **SHA-256 match on the leaf DER** against the bootstrap fingerprint. Hostname/SAN therefore never enters verification on this leg, so moving the agent listen address most likely needs **no cert re-issuance** - only the endpoint the guest dials. The spike must still confirm this empirically rather than trust the read. Flips a future "box survives a site/network change" map row | | R-50b | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\|reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | S–M | **(a) SHIPPED 2026-07-21; (b)/(c) open** | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. | -| R-51 | **Dead-primary alerting — a multi-container app whose MAIN container is dead must alert.** Aggregation currently classifies such a stack `unhealthy`, and `IsDownState` deliberately excludes `unhealthy`, so nothing fires. | S | idea | Origin: `AUDIT-vacation-remote-ops-2026-07-20.md` (F4). Observed live: `immich-server` was `Exited` for **18 h** with the app 100 % unreachable, and the box produced **no** dead-app banner and **no** `app_start_failed` hub event — while single-container Calibre-Web, down for the same reason, alerted correctly within 90 s. **Constraint (load-bearing): do NOT simply fold `unhealthy` into down.** That exclusion is deliberate (`stacks/manager.go` fix-3, `downstate_test.go`) and reverting it reintroduces the flapping it was added to stop. Direction: distinguish *member-container-exited* from *healthcheck-failing* in the aggregation, and treat a dead primary as down | -| R-52 | **Boot desired-state reconciliation — a `deployed: true` app should be running after boot.** The controller *reports* deployed-but-stopped apps (30 s `deadapp-check`) but never starts them, so an app that misses its boot start stays down until a human notices. | M | idea | Origin: `AUDIT-vacation-remote-ops-2026-07-20.md` (F5). Observed live: the pre-transport shutdown left `immich-server` and `calibre-web` `Exited`; **10 sibling containers came back and those two did not**, and they were still down ~18 h later. **Includes root-causing why `restart: unless-stopped` did not resurrect them** — both were stopped ~25 s before power-off, so Docker most likely recorded them as user-stopped; that hypothesis is untested because the guest journal is volatile and the controller's own logs were rotated by the container recreate. Direction: a bounded start-once reconciliation (N attempts, reusing the existing boot grace), never a restart loop. Pairs with R-51 — that one is the *alarm*, this one is the *recovery* | +| R-51 | **Dead-primary alerting — a multi-container app whose MAIN container is dead must alert.** Aggregation currently classifies such a stack `unhealthy`, and `IsDownState` deliberately excludes `unhealthy`, so nothing fires. | S | **SHIPPED 2026-07-21 — controller v0.156.0** | Origin: `AUDIT-vacation-remote-ops-2026-07-20.md` (F4). Observed live: `immich-server` was `Exited` for **18 h** with the app 100 % unreachable, and the box produced **no** dead-app banner and **no** `app_start_failed` hub event — while single-container Calibre-Web, down for the same reason, alerted correctly within 90 s. **Constraint (load-bearing): do NOT simply fold `unhealthy` into down.** That exclusion is deliberate (`stacks/manager.go` fix-3, `downstate_test.go`) and reverting it reintroduces the flapping it was added to stop. Direction: distinguish *member-container-exited* from *healthcheck-failing* in the aggregation, and treat a dead primary as down | **SHIPPED 2026-07-21 (controller v0.156.0).** **The diagnosis in this row was WRONG at the source and is corrected here:** aggregation did NOT classify the stack `unhealthy`. `aggregateState`'s final branch returned `StateRunning` for any mix of running and stopped members — the comment said "report as running (partial)" — so the stack read as RUNNING and `IsDownState` had nothing to fire on. The `unhealthy` exclusion was never involved, and the constraint it protects was therefore never in tension with the fix. New `StateDegraded`: a DOWN member whose docker restart policy is `always`/`unless-stopped` means docker was supposed to be keeping it up, so the stack is degraded (a down state, alerting through the EXISTING banner + `app_start_failed` path, unchanged); `no`/`on-failure` is a finished one-shot init/migrate container and stays benign. An UNREADABLE policy counts as supervised — fail-CLOSED, deliberately the opposite of `IsDownState`'s fail-open, because there the *state* is ambiguous while here a member is known dead and only the excuse is missing (P2 census 2026-07-21: all 53 catalog templates / 78 services are `unless-stopped`, zero one-shot containers exist today). `IsDownState` gained `degraded` and NOTHING else — the `unhealthy`/`restarting`/`paused`/`unknown` exclusions are byte-identical and `downstate_test.go` is untouched and green. Red-proof: the mix branch reverted to `return StateRunning` makes the immich fixture and both production-path tests fail with `"running"`. Live leg (STOP-1) pending. Evidence: `felhom-controller/REPORT.md` (2026-07-21) | +| R-52 | **Boot desired-state reconciliation — a `deployed: true` app should be running after boot.** The controller *reports* deployed-but-stopped apps (30 s `deadapp-check`) but never starts them, so an app that misses its boot start stays down until a human notices. | M | **SHIPPED 2026-07-21 — controller v0.156.0** | Origin: `AUDIT-vacation-remote-ops-2026-07-20.md` (F5). Observed live: the pre-transport shutdown left `immich-server` and `calibre-web` `Exited`; **10 sibling containers came back and those two did not**, and they were still down ~18 h later. **Includes root-causing why `restart: unless-stopped` did not resurrect them** — both were stopped ~25 s before power-off, so Docker most likely recorded them as user-stopped; that hypothesis is untested because the guest journal is volatile and the controller's own logs were rotated by the container recreate. Direction: a bounded start-once reconciliation (N attempts, reusing the existing boot grace), never a restart loop. Pairs with R-51 — that one is the *alarm*, this one is the *recovery* | **SHIPPED 2026-07-21 (controller v0.156.0), `internal/bootrecon`.** One bounded start-once sweep at controller startup: every deployed, non-protected, not-mid-deploy stack that still HAS containers and is down gets `StartStack`, at most **2 attempts 30 s apart**, then it stops and the alarm owns the problem. Never a restart loop. The whole sweep (5 s settle + one 30 s gap) fits inside the existing 90 s `deadAppBootGrace`, so a successful recovery never alerts and a failed one alerts honestly — asserted by a test rather than left to a comment. **The safety argument is the container gate:** the UI's Stop is `compose down`, which REMOVES the containers, while an interrupted boot leaves them behind as `Exited` — so "deployed, has containers, and they are down" is exactly the boot-orphan signature, and a zero-container stack is never touched. Red-proof: dropping that gate makes the user-stopped app get started, which is the one thing this must never do. **P1 (the `unless-stopped` root-cause probe) is deliberately NOT what this shipped on** — the reconciliation is correct whether or not Docker recorded those two containers as user-stopped, and the probe is recorded as an open question rather than a blocker. Live leg (STOP-1) pending. Evidence: `felhom-controller/REPORT.md` (2026-07-21) | +| R-54 | **[P2-HIGH] The guest's DHCP client is unsupervised — its death takes the box off the internet 1-2 hours later, invisibly.** ifupdown starts `dhclient` once at guest boot and nothing restarts it. | S-M | **SHIPPED 2026-07-21 — agent v0.92.1** | Origin: `audits/INCIDENT-guest-dhclient-killed-2026-07-20.md` §5 "OPEN RISK" — this row closes it. On 2026-07-20 a cleanup step killed guest 9201's dhclient (visible in the HOST's pid namespace; §4's `/proc//cgroup` rule exists because of it). **The guest then kept working for another ~80 minutes on its unexpired lease**; only at expiry did the address and default route vanish, taking the Cloudflare tunnel, hub reports, catalog sync and the controller→agent channel with them — 1h15m outage, and every observable signal said healthy for the first 80 minutes. **The design consequence: liveness of the DHCP client is itself a probe.** `internal/guestnet` flags a DHCP guest unhealthy on `pgrep -x dhclient` alone, while the lease is still live — waiting for the IP to disappear is waiting out precisely that silent window (red-proof: reverting to IP-presence-only makes the July-20 fixture report **healthy** with zero heals). Four fixed-shape `pct exec` probes (address / default route / `/etc/network/interfaces` mode / client liveness, parsers pinned to output captured live from 9201), the incident's restored invocation as the heal, verbatim, and dampers throughout: two CONSECUTIVE bad probes, ≥10 min between heals per guest, ≤3/hour, observe-only while guest or agent uptime < 3 min. Refuses to act on a static guest (dhclient must never fight a static config — reported loudly and left to **R-50**, which is where option 2 of the incident's three choices belongs), on an unknown mode, on an unprobeable guest, or on an ownership-unproven guest list (source is `ListLXC` ∩ the felhom pool, audit A1). Host-tier by necessity: a guest with no default route cannot repair its own default route. **A live finding during deployment:** the first sweep on felhom-pve logged `dhclient liveness probe failed: sudo: a password is required` and reported `state=unknown` — fail-safe, but blind. TASK-D assumed no sudoers change was needed; three of the four probes had no grant. `FELHOM_GUESTNET` + four `guestnet-*` capability rows shipped in v0.92.1 (v0.92.0 superseded, do not vouch). **Healthy cycle PROVEN LIVE 2026-07-21** on felhom-pve: caps `68/68 ok, degraded=0` and `level=DEBUG guestnet: guest network healthy vmid=9201 mode=dhcp has_route=true dhclient_alive=true`. The HEAL leg (STOP-2, a deliberate replay of the incident) is operator-present and pending. Note the guest and the host still differ (host static since the F1 mitigation, guest DHCP) — choosing one for both remains **R-50**'s call, not this row's. Evidence: `felhom-agent/REPORT.md` (2026-07-21) | | R-53 | **`app_export.html` substituted the CSRF token where the customer domain belongs** - the open-in-browser link was wrong for every app with a subdomain, and a session CSRF token landed in a URL. | XS | **SHIPPED (controller v0.150.0, 2026-07-20)** | One template token (`{{$.CSRFToken}}` -> `{{$.Domain}}`) plus the `Domain` key in `exportPageHandler`'s data map - that handler does not go through `baseData`, which is where every other page gets it, so the template had no domain to read. Render tests assert the joined `.` and that the token appears nowhere in that line; red-proofed against the pre-fix template. Origin: `audits/AUDIT-vacation-remote-ops-2026-07-20.md` (F7) | ## P3 — post-alpha