From 515e0c3cc70684552e8276b26efa81e88da4215d Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 25 Jul 2026 12:19:09 +0200 Subject: [PATCH] R-50 island-bridge: SPIKED -> GO (probes P1-P8 pass live on t740 drill) Provisioned nested-PVE drill 'drill-r50' (qm300 on demo-hp) via the v1.25.0 nested-vm ISO through the real day-0, then ran the R-50 empirical spike: - vmbr9 portless island bridge + guest island NIC hot-add (LAN undisturbed) - F1 replay money shot: LAN move survives on the island; LAN-literal bind reproduces the 2026-07-20 daemon-exit bug verbatim - dnsmasq trap confirmed live + lan_resolver.host_ip fix proven - pin address-independent (leaf SHA-256 unchanged, HTTP 200 over island) - survival matrix: agent/guest/host-cold-reboot all return on the island Docs: SPIKE verdict BLOCKED->GO, ROADMAP R-50 SPIKED->GO, nodes.md drill VM, REPORT overwrite. --- REPORT.md | 97 ++++++++++--------- .../audits/SPIKE-island-bridge-2026-07-25.md | 57 ++++++++++- documentation/backlog/ROADMAP.md | 2 +- documentation/operations/nodes.md | 22 +++-- 4 files changed, 117 insertions(+), 61 deletions(-) diff --git a/REPORT.md b/REPORT.md index 80f157b..28a1dd5 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,54 +1,55 @@ -# REPORT — R-50 island-bridge SPIKE, empirical attempt 2 (t740): STILL BLOCKED (2026-07-25) +# REPORT — R-50 island-bridge drill + empirical spike (RUNBOOK, 2026-07-25 PM) -**Overwritten** per the standing rule. This session re-attempted the empirical half of -`SPIKE-island-bridge-2026-07-25.md` on the t740 (per the operator's 2026-07-25 drill-host ruling). -Docs-only; no code, no version bumps. +## What ran +Provisioned the first **nested-PVE drill appliance on the t740 (`demo-hp`)** and ran the **R-50 +island-bridge empirical spike** end-to-end. Verdict: **GO.** -## t740 access method -Discovery was read-only (no LAN scan): the SSH config's `demo-hp` entry (t740 = Tailscale -`100.76.96.79` / LAN `192.168.0.87`, host_id `demo-hp-bb76ea`) + the hub registry. The t740 has **no -baked SSH key**, so access used the **hub-vaulted G1 break-glass root credential** — -`host_recovery/demo-hp-bb76ea` (read from the hub SQLite via `kubectl cp` + `sqlite3`), `sshpass -e` -(secret redacted, never printed). Connected: `felhom-host`, PVE 9.2.2. +## Part A — drill provisioned through the REAL day-0 +- **VM 300 `drill-r50`** on `demo-hp`: 8 GiB / 4 vCPU (cpu=host) / 32 GiB local-lvm / OVMF (SB off) / one + NIC on `vmbr0` DHCP. Nested-virt already enabled on the host (A2 no-op; no QEMU VM was running at A1). +- Installed from **`felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso`** (current release; the + RUNBOOK's stated *v1.22.0* was both stale **and** the wrong profile — v1.22.0-nested-**canary** is a + deliberate match-nothing installer that aborts touching no disk; corrected to the v1.25.0 **nested-vm** + profile. Friction finding, not a day-0 defect). +- Day-0 drove itself: install → first-boot → **self-register** (appliance #9, pairing `9MB-4QX`) → + **operator bind** to a scratch customer `drill-r50` (minimal config, no DNS/offsite/PBS side-effects) → + **deliver** → nested guest **9201** provisioned, controller stack healthy (agent 0.93.0, golden 0.161.0). +- **R-50 starting condition reached & confirmed:** agent `listen_addr = 192.168.0.176:8443` **and** guest + `bootstrap.json local_api.endpoint = 192.168.0.176:8443` — the F1 LAN literal baked in both places. +- Snapshot **`r50pre`** taken (clean day-0). Break-glass vaulted `host_recovery/drill-r50-0a4f9a`. +- **Day-0 tail not driven:** *claim* is customer-email-gated (scratch customer has no inbox; forcing it + needs a real email + a production Resend send) — left PENDING; *escrow* is N/A (no DR/PBS tier on the + minimal drill). Neither gates the control-plane probes. **Operator decision if the full customer-side + tail should be exercised (needs an email address + DR-tier choice).** -## P1 inventory — the finding: there is NO drill VM to probe -- **t740 (`demo-hp`):** `qm list` → **empty**; `/etc/pve/qemu-server/` → **empty**; standalone (no - cluster); only its own LXC guest 9201 + the guest-9201 LVM volumes. **No nested drill PVE VM.** -- **Hub registry:** only `demo-hp-bb76ea` (t740) + `demo-felhom-8363b5` (N100) — **no drill appliance** - (no `2f4b00`/`demo-vm`/`drill` host). -- **felhom-pve:** `qm list` empty (confirmed in attempt 1). -- **DooPlex:** the historical drill appliance `drill.qcow2` exists (`/mnt/5_hdd/felhom.eu/drill/`, 18G) - but is **powered off** — a **golden-bake** VM (bake logs to 0.153.0, last Jul 18–20) whose nested - guest is purged after each bake, so it has **no island-bridge topology** (agent + nested guest + - bootstrap) to probe. DooPlex is the production k3s node and is loaded (~17G of 62G free). +## Part B — probes P1–P8, ALL PASS (verdict GO) +- **P2** `vmbr9` portless island bridge `169.254.253.1/30` — vmbr0 untouched, LAN gw reachable. +- **P3** guest island NIC `eth1 169.254.253.2/30` hot-added — eth0/LAN undisturbed, bidirectional island ping. +- **P4 (dnsmasq trap + fix)** — with `host_ip` unset, moving `listen_addr` to the island rebound dnsmasq + to `169.254.253.1:53` and **killed LAN DNS**; `lan_resolver.host_ip=192.168.0.176` restored it while the + API bind stayed on the island. Finding 1 confirmed + fixed, live. +- **P5 (pin)** — served leaf SHA-256 over the island **identical** (`4ef1d953…`), authenticated + `GET /storage` from the guest over the island → **HTTP 200**. No cert re-issue (Finding 6 live). +- **P6 (F1 replay — money shot)** — LAN `192.168.0.176→.200`, `listen_addr` on the island → agent stays + active + serves (200). LAN-literal contrast reproduced the 2026-07-20 bug verbatim + (`localapi: bind 192.168.0.176:8443` → daemon exit 1). +- **P7 (survival)** — agent restart / guest reboot / **host cold reboot** all return the control plane on + the island with zero intervention (vmbr9, island bind, dnsmasq on LAN IP, guest autostart, controller). +- **Method caveat:** probes drove the runtime chain via manual config edits; the *provisioning* path + (host-install + golden-bake bootstrap template) is the implementation task — now fully de-risked. -**Snapshot: N/A** (no drill VM existed to snapshot). +## Docs updated +- `documentation/audits/SPIKE-island-bridge-2026-07-25.md` — verdict flipped BLOCKED → **GO**, empirical + P1–P8 results added, old BLOCKED/NOT-RUN sections marked superseded. +- `documentation/backlog/ROADMAP.md` — R-50 → **SPIKED → GO**; next = Phase A/B/C production spec. +- `documentation/operations/nodes.md` — the drill VM 300 now documented on the t740 (access, snapshot, + teardown). -## Per-probe outcomes -P2–P7 (bridge create / NIC hot-add / island bind / pin / **F1 replay** / survival) — **NOT RUN.** The -operator's ruling ("drill+build VMs on the HP from now on") is **forward-looking and not yet realized**: -no probeable drill appliance exists on the t740, and the only artifact is a stale bake VM on the -production node. **Per the hard rule ("do not improvise on production"), felhom-pve, the t740 host -networking, and guest 9201 were left UNTOUCHED.** No config was changed anywhere → **the rollback table -is empty**, and there is no island end-state to leave in place. +## State left behind +Drill VM 300 left **running in the working island configuration** (snapshot `r50pre` preserves clean +day-0). Scratch hub records `drill-r50` (customer + appliance #9) remain — throwaway; teardown noted in +nodes.md. `demo-hp` host networking, guest 9201, and felhom-pve were **untouched**. -## GO/NO-GO: PENDING (unchanged) — the design half still stands -The source-grounded half of the parent doc (address plan `169.254.253.1/30`↔`.2/30`; the F1 two-place -literal; the **dnsmasq trap** — `LANResolverConfig.WithDefaults` derives the DNS listen-addr from -`listen_addr`, so the spec MUST set `lan_resolver.host_ip = LAN IP`; leaf-DER pin → no cert re-issue -expected; provisioning inventory; cluster parity) is unchanged and ready. Only the **empirical** -validation remains blocked. - -## Operator decision + docs -**Operator chose (2026-07-25): they will provision a nested-PVE drill VM on the t740 (agent + a nested -guest); this spike re-runs then.** Commits (docs-only): the spike doc amended (attempt-2 section + -verdict), ROADMAP R-50 note appended, `documentation/operations/nodes.md` gains a "designated drill+build -VM host" subsection (t740 ruling + the not-yet-realized state + break-glass recipe), and the -controller/agent `CLAUDE.md` env tables gain a `demo-hp` row/note. **The build-PIPELINE relocation to the -t740 is explicitly NOT ruled or implemented — only the VM-hosting ruling is recorded.** - -## Observations -- The t740 is a genuinely better drill host than DooPlex (dedicated demo node, ~30G RAM / ~49G free - local-lvm, not the production k3s node) — the ruling is sound; it just needs the VM created. -- The t740's agent is **0.93.0**, behind demo-felhom's 0.95.0 — a publish-when-convenient gap (noted in - nodes.md), unrelated to this spike. +## Not done here (by design / veto) +- **Part C** (deploy agent 0.95.0 → demo-hp) — veto-able; pending after this evidence. +- **Forensic** (qm300 disappearance on felhom-pve) — optional, read-only; pending. diff --git a/documentation/audits/SPIKE-island-bridge-2026-07-25.md b/documentation/audits/SPIKE-island-bridge-2026-07-25.md index bc14658..1093b95 100644 --- a/documentation/audits/SPIKE-island-bridge-2026-07-25.md +++ b/documentation/audits/SPIKE-island-bridge-2026-07-25.md @@ -8,11 +8,60 @@ fixed private addresses, so the control plane survives any router/lease/site mov **validate the whole chain end-to-end on the DRILL environment (qm300 / `demo-vm-felhom-2f4b00`)** before any production spec exists. -## VERDICT UP FRONT: SPIKE STILL BLOCKED (empirically) — no probeable drill environment exists. +## VERDICT UP FRONT: **GO** — validated end-to-end on a live drill (2026-07-25, evening). -**Two empirical attempts, both blocked (2026-07-25). Design half below stands; GO/NO-GO PENDING.** +**A nested-PVE drill appliance was provisioned on the t740 (`demo-hp`) through the REAL day-0 pipeline, +and probes P1–P8 ALL PASS.** The island-bridge control plane works end-to-end, survives the exact F1 +failure that motivated the row, the pin is address-independent as predicted, the dnsmasq trap is +confirmed live AND its fix proven, and the whole topology survives a cold host reboot untouched. +**Recommendation: GO — write the production implementation spec (Phase A/B/C skeleton at the bottom).** +See **“## Empirical validation (2026-07-25 PM) — probes P1–P8”** below. The two earlier BLOCKED attempts +(no drill existed) are retained further down for provenance. -### Empirical attempt 2 (2026-07-25, PM) — the t740 was checked; there is no drill VM there. +### Method caveat (honest scope) +The probes drove the mechanism via **manual, surgical config edits** on the drill (append a `vmbr9` +stanza + `ifreload`, `pct set -net1`, hand-edit `agent.json` `listen_addr`/`lan_resolver.host_ip` and +the host-side `bootstrap.json` endpoint, restart). They prove the **runtime chain** is sound and +reboot-durable. They do **not** exercise the *provisioning* path that a real rollout needs +(`felhom-host-install.sh` writing these at install/golden-bake time) — that is Phase A of the impl spec, +still a future task. What is now de-risked: every assumption the spec rests on is empirically true. + +## Empirical validation (2026-07-25 PM) — probes P1–P8, ALL PASS + +**Drill:** nested-PVE appliance `drill-r50` (QEMU VM **300** on `demo-hp`/t740), installed from the +current release ISO `felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` through the **real day-0** +(self-register → operator bind → deliver → guest provision). Reached the exact R-50 starting condition: +agent `local_api.listen_addr = 192.168.0.176:8443` (the LAN literal) **and** the guest's +`bootstrap.json local_api.endpoint = 192.168.0.176:8443` — the F1 literal baked in both places. Nested +guest **9201** runs the controller (agent 0.93.0, golden 0.161.0). Access to the drill-PVE: +break-glass root via hub `host_recovery/drill-r50-0a4f9a`, reached through `demo-hp`. +(Day-0 tail: **claim** left PENDING — customer-email-gated, no inbox on a scratch customer; **escrow** +N/A — no DR/PBS tier on the minimal drill. Neither gates the control-plane probes.) + +| Probe | What it validates | Result | +|---|---|---| +| **P1** snapshot | rollback point before mutation | **PASS** — `qm snapshot 300 r50pre` (retained) | +| **P2 (3)** `vmbr9` create | portless host bridge, non-disruptive | **PASS** — `vmbr9 169.254.253.1/30` up via `ifreload -a`; **vmbr0 untouched** (`.176/24`), LAN gw still reachable, guest kept running | +| **P3 (4)** guest NIC hot-add | live hotplug, LAN leg undisturbed | **PASS** — `pct set 9201 -net1 …ip=169.254.253.2/30`; `eth1` up on the island, **`eth0`/LAN undisturbed** (`192.168.0.15/24`), host↔guest island ping both ways | +| **P4 (7)** dnsmasq trap **+ fix** | the Finding-1 coupling, live | **PASS (both)** — with `host_ip` **unset**, moving `listen_addr` to the island rebound dnsmasq to **`169.254.253.1:53`** and **LAN DNS `192.168.0.176:53` died**; setting `lan_resolver.host_ip=192.168.0.176` put dnsmasq **back on `192.168.0.176:53`** while the API bind stayed on the island | +| **P5 (6)** pin | pin survives the address move | **PASS** — served leaf SHA-256 over the island **identical** (`4ef1d953fe…f219bab`); authenticated `GET https://169.254.253.1:8443/storage` from the guest → **HTTP 200**. No cert re-issue. Finding 6 confirmed live | +| **P6 (5)** **F1 replay** | control plane survives a LAN move | **PASS — the money shot.** LAN moved `192.168.0.176 → .200`, `listen_addr` left on the island → agent **stays `active`**, still bound `169.254.253.1:8443`, control plane **HTTP 200**. **Contrast (original bug reproduced):** set `listen_addr` back to the now-absent `.176` → `level=ERROR "daemon: exited with error" err="localapi: bind 192.168.0.176:8443: listen…"`, systemd `status=1/FAILURE`, nothing bound — the 2026-07-20 incident, verbatim | +| **P7 (8)** survival matrix | topology is reboot-durable | **PASS** — **agent restart**: island bind + 200; **guest reboot**: `eth1` island NIC + controller + 200 returned; **host (qm300) COLD reboot**: `vmbr9`, agent island bind, dnsmasq on the LAN IP, guest autostart (`onboot=1`) with `eth1`, controller healthy, control plane **HTTP 200** — **all with zero intervention** | + +**All source-grounded findings are now empirically confirmed:** the F1 double-bake (agent bind + guest +dial) moves atomically and works; the pin is DER-based and address-transparent (Finding 6); the dnsmasq +trap is real and its `lan_resolver.host_ip = LAN IP` fix works (Finding 1); link-local `169.254.253.x/30` +is a clean host-internal fabric that no LAN move perturbs. **The drill was left in the working island +configuration** (not rolled back); snapshot `r50pre` preserves the clean LAN-literal day-0 if a re-run +is wanted. + +### GO/NO-GO: **GO.** Proceed to the implementation spec (Phase A/B/C below). No blocker remains; the one +residual is that the *provisioning* path (host-install writing the island config + golden-bake carrying +the new bootstrap template) is unbuilt — that IS the impl task, now de-risked. + +--- + +### Empirical attempt 2 (2026-07-25, PM) — the t740 was checked; there is no drill VM there. — **SUPERSEDED by the validation above (a drill was then provisioned and probed).** The operator ruled (2026-07-25) that **drill + build VMs are hosted on the HP t740 (`demo-hp`) from now on**, so this half was re-attempted there. Discovery (read-only, no LAN scan): the t740 = `demo-hp` (Tailscale `100.76.96.79` / LAN `192.168.0.87`), accessed via the hub-vaulted G1 break-glass root @@ -56,7 +105,7 @@ next drill run) start from a real base rather than zero. --- -## Probes 3–8 — NOT RUN (drill env absent) +## Probes 3–8 — NOT RUN (drill env absent) — **SUPERSEDED: all now PASS, see “Empirical validation (2026-07-25 PM)” above.** | Probe | Would validate | Status | |---|---|---| diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index 82629fb..5408764 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -58,7 +58,7 @@ | R-35 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. | S | idea | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | | R-36 | **Post-RESET re-enroll leaves offsite "enabled but unprovisioned" — silently.** The hub knows the state and says nothing on the customer page. | S | **SHIPPED (hub v0.67.0, 2026-07-18)** | Both halves delivered. **(1) The warning:** the customer page now names the state and the fix — enabled-but-unprovisioned raises an amber banner saying provisioning is *Save*-triggered (press Save once, then verify), reusing the exact `enabled && type == ""` predicate the offsite re-issue handler already refuses on. **(2) The related sub-item, also done:** the self-bind link is now **auto-minted at customer creation AND at RESET completion**, so the console banner's „e-mailben kapott link" is already true instead of true-once-the-operator-remembers. Extracting the shared `mintAndSendSelfBindLink` core keeps the button and the auto-mint callers on the same F1/F2 honesty rules, and the auto-mint never fails the operation it rides on. **Gap found and closed while wiring it:** `PurgeCustomerResetDBState` does NOT clear `selfbind_tokens`, so a link minted BEFORE a reset would have stayed live across it — the skip paths now clear stale tokens, giving the invariant "after auto-mint, the only live link is one we just issued, or none". Tests assert the banner is ABSENT in all three nominal cases too; red-proofed. — Original analysis: Source-cited behaviour, confirmed live in the rehearsal: **provisioning is Save-triggered** (`configs.go` `applyOffsite`) — which also answers S6's open question — and the re-enroll auto-re-issue **correctly** skips unprovisioned targets (`handler.go`). So nothing is broken; the gap is that nobody is told. Direction: flash it on the customer page. **Interim: an R-3 step.** **Related sub-item:** auto-mint the **self-bind link on customer create/RESET**, so the console banner's „e-mailben kapott link" is always already true instead of true-once-the-operator-remembers | | R-27c | **Customer self-bind, slice 2 — console-passphrase bind.** Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. | M | idea | **Security constraints from the session ruling, all load-bearing:** passphrase **issued at customer creation**; the global-lookup endpoint must be **spray-hardened** — per-appliance **and** per-IP caps, constant-time comparison, a **single generic failure** (no oracle), alerting on abuse; an **accent-free wordlist** (console keymaps are not Hungarian); the **web capability-link path is RETAINED**; **claim-by-email is RETAINED** as the delivery-channel proof. **Also under this item:** the self-bind email gains the **public universal-ISO download link + two-line instructions** (the DIY case). **Secret-bearing per-customer ISOs are ruled OUT.** Sibling of R-27b (second-box flow) — different axis, both build on the same `/bind/` page | -| R-50 | **[P2-HIGH] Island-bridge control plane — make controller↔agent independent of the LAN.** The agent's `localapi` binds a **LAN literal** (`listen_addr`) and the guest dials that same literal from `bootstrap.json`. Move both onto a **host-internal bridge with a fixed, private address** that no router, DHCP lease, or site move can invalidate, so the control plane is location-independent and golden-image-stable. | M | **SPIKED (partial) 2026-07-25 — empirical BLOCKED, design half source-grounded** | **2026-07-25 spike (`audits/SPIKE-island-bridge-2026-07-25.md`): the drill environment (qm300 / `demo-vm-felhom-2f4b00`) is GONE** (`qm list` empty; only guest 9201 remains), so per the hard drill-only rule the empirical probes (bridge create / NIC hot-add / island bind / **F1 replay** / survival) were NOT run — production was left untouched, **no live GO/NO-GO**. The **source-grounded** half IS done: recommend **link-local `169.254.253.1/30`↔`.2/30`** (structurally uncollidable vs LAN); the F1 literal is baked in TWO places to move atomically (`config.go:229` + `felhom-host-install.sh:2226` for the bind, `provision/backhalf.go:129`/`bootstrap.json` for the guest dial); **NEW dnsmasq trap CONFIRMED in source** — `LANResolverConfig.WithDefaults` (`config.go:208–210`) derives the DNS listen-addr from `listen_addr`, so moving the bind to the island silently kills LAN DNS → the spec MUST set `lan_resolver.host_ip = LAN IP` explicitly; pin is leaf-DER-based (address-independent → no cert re-issue expected); provisioning inventory + cluster-parity (SDN on Peti's 2 nodes) + an implementation skeleton recorded. **Empirical attempt 2 (2026-07-25 PM, on the t740 per the operator's drill-host ruling): STILL BLOCKED — there is NO drill VM on the t740** (`qm list` empty; it is a real demo node running only its own guest 9201). The only drill artifact is a stale golden-bake `drill.qcow2` on the production DooPlex node (off, nested guest purged). The ruling ("drill+build VMs on the HP from now on") is forward-looking and not yet realized. **Operator decision: they will provision a nested-PVE drill VM on the t740 (agent + nested guest); this spike re-runs then.** GO/NO-GO PENDING. **Remaining: a provisioned t740 drill VM to validate probes P2–P7 before the production spec.** Origin: `audits/AUDIT-vacation-remote-ops-2026-07-20.md` (F1), where this failed for real. The demo box moved to a remote site, DHCP handed it `.147` instead of `.162`, and the agent then **could not start at all** — `bind: cannot assign requested address`, systemd gave up after 4 retries — taking storage, PBS backup, quiesce, restore-test and DR down for as long as nobody noticed. Mitigated for that window by pinning `vmbr0` static back to `.162`; that is a **window mitigation, not the fix** — it still depends on the site's subnet being `192.168.0.0/24` and free at that address. **Spike-first is mandatory:** validate end-to-end on the drill environment (agent bind + guest dial + TLS SAN/pin + reinstall/golden survival + the bootstrap-config migration for already-deployed boxes) BEFORE any production spec. **Pin fact (verified 2026-07-20, `agentapi/client.go` L105-129 — supersedes the earlier "the SAN set must cover the new address" note in this entry, which was wrong):** the controller-to-agent leg sets `InsecureSkipVerify: true` and replaces chain verification with a custom `VerifyPeerCertificate` that does a raw **SHA-256 match on the leaf DER** against the bootstrap fingerprint. Hostname/SAN therefore never enters verification on this leg, so moving the agent listen address most likely needs **no cert re-issuance** - only the endpoint the guest dials. The spike must still confirm this empirically rather than trust the read. Flips a future "box survives a site/network change" map row | +| R-50 | **[P2-HIGH] Island-bridge control plane — make controller↔agent independent of the LAN.** The agent's `localapi` binds a **LAN literal** (`listen_addr`) and the guest dials that same literal from `bootstrap.json`. Move both onto a **host-internal bridge with a fixed, private address** that no router, DHCP lease, or site move can invalidate, so the control plane is location-independent and golden-image-stable. | M | **SPIKED → GO 2026-07-25 (probes P1–P8 PASS live)** | **GO (2026-07-25 PM, `audits/SPIKE-island-bridge-2026-07-25.md`): validated end-to-end on a real nested-PVE drill (`drill-r50`, qm300 on the t740, installed via the v1.25.0 ISO through the actual day-0).** All probes PASS: `vmbr9` portless island bridge (vmbr0 untouched); guest island NIC `eth1 169.254.253.2/30` hot-added (LAN leg undisturbed); **F1 replay = the money shot** — LAN moved `.176→.200` with `listen_addr` on the island → agent stays `active`, control plane HTTP 200; the LAN-literal contrast reproduced the 2026-07-20 bug verbatim (`localapi: bind 192.168.0.176:8443` → daemon exit 1); **dnsmasq trap CONFIRMED LIVE and its `lan_resolver.host_ip=LAN IP` fix PROVEN**; **pin address-independent** (served leaf SHA-256 unchanged, HTTP 200 over the island — no cert re-issue); **survival matrix** (agent restart / guest reboot / host COLD reboot) all return the control plane on the island with zero intervention. **Method caveat:** probes drove the runtime chain via manual config edits — the *provisioning* path (host-install writing the island config + golden-bake bootstrap template) is the impl task, now de-risked. **Next: write the Phase A/B/C production spec (drill-proven → demo → Peti SDN).** — Prior (superseded): **2026-07-25 spike: the drill environment (qm300 / `demo-vm-felhom-2f4b00`) is GONE** (`qm list` empty; only guest 9201 remains), so per the hard drill-only rule the empirical probes (bridge create / NIC hot-add / island bind / **F1 replay** / survival) were NOT run — production was left untouched, **no live GO/NO-GO**. The **source-grounded** half IS done: recommend **link-local `169.254.253.1/30`↔`.2/30`** (structurally uncollidable vs LAN); the F1 literal is baked in TWO places to move atomically (`config.go:229` + `felhom-host-install.sh:2226` for the bind, `provision/backhalf.go:129`/`bootstrap.json` for the guest dial); **NEW dnsmasq trap CONFIRMED in source** — `LANResolverConfig.WithDefaults` (`config.go:208–210`) derives the DNS listen-addr from `listen_addr`, so moving the bind to the island silently kills LAN DNS → the spec MUST set `lan_resolver.host_ip = LAN IP` explicitly; pin is leaf-DER-based (address-independent → no cert re-issue expected); provisioning inventory + cluster-parity (SDN on Peti's 2 nodes) + an implementation skeleton recorded. **Empirical attempt 2 (2026-07-25 PM, on the t740 per the operator's drill-host ruling): STILL BLOCKED — there is NO drill VM on the t740** (`qm list` empty; it is a real demo node running only its own guest 9201). The only drill artifact is a stale golden-bake `drill.qcow2` on the production DooPlex node (off, nested guest purged). The ruling ("drill+build VMs on the HP from now on") is forward-looking and not yet realized. **Operator decision: they will provision a nested-PVE drill VM on the t740 (agent + nested guest); this spike re-runs then.** GO/NO-GO PENDING. **Remaining: a provisioned t740 drill VM to validate probes P2–P7 before the production spec.** Origin: `audits/AUDIT-vacation-remote-ops-2026-07-20.md` (F1), where this failed for real. The demo box moved to a remote site, DHCP handed it `.147` instead of `.162`, and the agent then **could not start at all** — `bind: cannot assign requested address`, systemd gave up after 4 retries — taking storage, PBS backup, quiesce, restore-test and DR down for as long as nobody noticed. Mitigated for that window by pinning `vmbr0` static back to `.162`; that is a **window mitigation, not the fix** — it still depends on the site's subnet being `192.168.0.0/24` and free at that address. **Spike-first is mandatory:** validate end-to-end on the drill environment (agent bind + guest dial + TLS SAN/pin + reinstall/golden survival + the bootstrap-config migration for already-deployed boxes) BEFORE any production spec. **Pin fact (verified 2026-07-20, `agentapi/client.go` L105-129 — supersedes the earlier "the SAN set must cover the new address" note in this entry, which was wrong):** the controller-to-agent leg sets `InsecureSkipVerify: true` and replaces chain verification with a custom `VerifyPeerCertificate` that does a raw **SHA-256 match on the leaf DER** against the bootstrap fingerprint. Hostname/SAN therefore never enters verification on this leg, so moving the agent listen address most likely needs **no cert re-issuance** - only the endpoint the guest dials. The spike must still confirm this empirically rather than trust the read. Flips a future "box survives a site/network change" map row | | R-50b | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\|reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | S–M | **(a) SHIPPED 2026-07-21; (b)/(c) open** | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. | | R-51 | **Dead-primary alerting — a multi-container app whose MAIN container is dead must alert.** Aggregation currently classifies such a stack `unhealthy`, and `IsDownState` deliberately excludes `unhealthy`, so nothing fires. | S | **SHIPPED 2026-07-21 — controller v0.156.0** | Origin: `AUDIT-vacation-remote-ops-2026-07-20.md` (F4). Observed live: `immich-server` was `Exited` for **18 h** with the app 100 % unreachable, and the box produced **no** dead-app banner and **no** `app_start_failed` hub event — while single-container Calibre-Web, down for the same reason, alerted correctly within 90 s. **Constraint (load-bearing): do NOT simply fold `unhealthy` into down.** That exclusion is deliberate (`stacks/manager.go` fix-3, `downstate_test.go`) and reverting it reintroduces the flapping it was added to stop. Direction: distinguish *member-container-exited* from *healthcheck-failing* in the aggregation, and treat a dead primary as down | **SHIPPED 2026-07-21 (controller v0.156.0).** **The diagnosis in this row was WRONG at the source and is corrected here:** aggregation did NOT classify the stack `unhealthy`. `aggregateState`'s final branch returned `StateRunning` for any mix of running and stopped members — the comment said "report as running (partial)" — so the stack read as RUNNING and `IsDownState` had nothing to fire on. The `unhealthy` exclusion was never involved, and the constraint it protects was therefore never in tension with the fix. New `StateDegraded`: a DOWN member whose docker restart policy is `always`/`unless-stopped` means docker was supposed to be keeping it up, so the stack is degraded (a down state, alerting through the EXISTING banner + `app_start_failed` path, unchanged); `no`/`on-failure` is a finished one-shot init/migrate container and stays benign. An UNREADABLE policy counts as supervised — fail-CLOSED, deliberately the opposite of `IsDownState`'s fail-open, because there the *state* is ambiguous while here a member is known dead and only the excuse is missing (P2 census 2026-07-21: all 53 catalog templates / 78 services are `unless-stopped`, zero one-shot containers exist today). `IsDownState` gained `degraded` and NOTHING else — the `unhealthy`/`restarting`/`paused`/`unknown` exclusions are byte-identical and `downstate_test.go` is untouched and green. Red-proof: the mix branch reverted to `return StateRunning` makes the immich fixture and both production-path tests fail with `"running"`. Live leg (STOP-1) pending. Evidence: `felhom-controller/REPORT.md` (2026-07-21) **PROVEN LIVE 2026-07-21 (STOP-1, operator-present).** On guest 9201, `docker stop immich-server` at **12:50:40 CEST** (policy `unless-stopped`, three helpers left running — the exact F4 shape). **12:50:53 — 13 seconds later — the stack read `degraded`** where it read `running` for 18 hours on 2026-07-20. `10:51:11Z` **exactly ONE** `app_start_failed (warn) — Telepített alkalmazás nem fut: Immich` (single-fire verified by count, not by eye). Dashboard rendered the banner *Telepített alkalmazás nem fut: Immich (degraded)* plus the new „Részlegesen leállt" state label. `docker start` at 12:51:37 → `running` by 12:51:52 and the banner **self-cleared** (state-based, as designed). Note the banner appends the raw state in English — `(degraded)` — which is pre-existing behaviour, not introduced here, but now more visible | | R-52 | **Boot desired-state reconciliation — a `deployed: true` app should be running after boot.** The controller *reports* deployed-but-stopped apps (30 s `deadapp-check`) but never starts them, so an app that misses its boot start stays down until a human notices. | M | **SHIPPED 2026-07-21 — controller v0.156.0** | Origin: `AUDIT-vacation-remote-ops-2026-07-20.md` (F5). Observed live: the pre-transport shutdown left `immich-server` and `calibre-web` `Exited`; **10 sibling containers came back and those two did not**, and they were still down ~18 h later. **Includes root-causing why `restart: unless-stopped` did not resurrect them** — both were stopped ~25 s before power-off, so Docker most likely recorded them as user-stopped; that hypothesis is untested because the guest journal is volatile and the controller's own logs were rotated by the container recreate. Direction: a bounded start-once reconciliation (N attempts, reusing the existing boot grace), never a restart loop. Pairs with R-51 — that one is the *alarm*, this one is the *recovery* | **SHIPPED 2026-07-21 (controller v0.156.0), `internal/bootrecon`.** One bounded start-once sweep at controller startup: every deployed, non-protected, not-mid-deploy stack that still HAS containers and is down gets `StartStack`, at most **2 attempts 30 s apart**, then it stops and the alarm owns the problem. Never a restart loop. The whole sweep (5 s settle + one 30 s gap) fits inside the existing 90 s `deadAppBootGrace`, so a successful recovery never alerts and a failed one alerts honestly — asserted by a test rather than left to a comment. **The safety argument is the container gate:** the UI's Stop is `compose down`, which REMOVES the containers, while an interrupted boot leaves them behind as `Exited` — so "deployed, has containers, and they are down" is exactly the boot-orphan signature, and a zero-container stack is never touched. Red-proof: dropping that gate makes the user-stopped app get started, which is the one thing this must never do. **P1 (the `unless-stopped` root-cause probe) is deliberately NOT what this shipped on** — the reconciliation is correct whether or not Docker recorded those two containers as user-stopped, and the probe is recorded as an open question rather than a blocker. Live leg (STOP-1) pending. Evidence: `felhom-controller/REPORT.md` (2026-07-21) **PROVEN LIVE 2026-07-21 (STOP-1, operator-present), and it answered P1 for free.** Fixture on 9201: `docker stop` on bookstack (2 containers, left in place) + calibre-web, and a UI Stop on immich (`compose down` → 0 containers); then `pct reboot 9201`. Result: `10:53:22Z [bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [bookstack] — up to 2 attempt(s)` → `10:53:28Z attempt 1/2: started "bookstack" (took 6.0s)` → `complete: 1 app(s) recovered in 1 attempt(s)`, and **ZERO `app_start_failed`** — a successful recovery inside the boot grace is silent, exactly as designed. **P1 IS NOW ANSWERED, and the F5 hypothesis is CONFIRMED:** bookstack carries `restart=unless-stopped`, the Docker daemon came up at ~10:53:15Z, and the container's `StartedAt` is **`10:53:28.05Z` — the exact moment `bootrecon`'s `StartStack` returned**. Docker's own restart policy did NOT resurrect it; a container stopped before shutdown is recorded user-stopped and stays down. Only R-52 brought it back. **BUT the end-to-end "a deliberate Stop survives a reboot" property is NOT true today, for a reason outside R-52 — see R-55.** R-52's own gate is correct and was observed to be: immich (0 containers) was never a candidate | diff --git a/documentation/operations/nodes.md b/documentation/operations/nodes.md index cb792d6..8fc3442 100644 --- a/documentation/operations/nodes.md +++ b/documentation/operations/nodes.md @@ -65,14 +65,20 @@ No trace of `192.168.100.2` remains. `wg-felhom` `10.77.0.3/32` up to the hub. G them off DooPlex (the production k3s node). **This is a VM-HOSTING ruling only; the build-PIPELINE relocation to the t740 is NOT ruled or implemented here.** -**Current state (verified 2026-07-25):** the ruling is **forward-looking and not yet realized.** The -t740 hosts **no drill/build VM yet** — `qm list` is empty, `/etc/pve/qemu-server/` is empty, it is a -standalone PVE node running only its own LXC guest 9201. The historical drill appliance is still -`drill.qcow2` on **DooPlex** (`/mnt/5_hdd/felhom.eu/drill/`, ~18G, powered off — a golden-bake VM whose -nested guest is purged after each bake). felhom-pve likewise has no drill VM. **A nested-PVE drill VM -(agent + a nested guest) must be provisioned on the t740** before the R-50 island-bridge empirical spike -(`audits/SPIKE-island-bridge-2026-07-25.md`) can run its P2–P7 probes — that spike is blocked on exactly -this. The t740's ~30 GB RAM / ~49 GB free local-lvm suit it as the drill host. +**Current state (updated 2026-07-25 PM):** the ruling is **realized** — the t740 now hosts the first +drill appliance. **QEMU VM `300` = `drill-r50`**, a nested PVE-in-a-VM (8 GiB RAM, 4 vCPU cpu=host, 32 GiB +local-lvm disk, OVMF/SB-off, one NIC on `vmbr0` DHCP). Installed from +`felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` through the **real day-0** (self-register → +operator bind to scratch customer `drill-r50` → deliver → nested guest **9201** provisioned, controller +0.161.0 / agent 0.93.0 healthy). Its own break-glass root is vaulted in the hub +`host_recovery/drill-r50-0a4f9a`; reach it as `root@192.168.0.176` **through `demo-hp`** (it has no key +and no tailnet — it is a peer on demo-hp's LAN). It has snapshot **`r50pre`** (clean LAN-literal day-0). +This VM was provisioned to unblock and run the **R-50 island-bridge empirical spike** +(`audits/SPIKE-island-bridge-2026-07-25.md`) — **probes P1–P8 PASSED 2026-07-25, verdict GO**; the drill +was left in its working island configuration. It is a throwaway: destroy with `qm stop 300 && qm destroy +300 --purge 1` (and delete the `drill-r50` customer + appliance #9 hub-side) when no longer needed. The +historical golden-bake `drill.qcow2` still lives on **DooPlex** (`/mnt/5_hdd/felhom.eu/drill/`, ~18G, +powered off) and is unrelated. The build-PIPELINE relocation to the t740 remains unbuilt. ### Access — there is no baked SSH key