R-50 SHIPPED: fleet migrated to the island; Phase 5 docs
- B2 demo-hp + B3 demo-felhom migrated to the island (agent 0.96.0), apps served throughout (0 container restarts), island /storage 200, LAN DNS pinned to the LAN IP, hub reports 0.96.0. No rollback. - capability-map 'site/network change' row PARTIAL -> PROVEN-LIVE - ROADMAP R-50 -> SHIPPED (fleet-migrated); add R-74 (island on Peti's cluster) - nodes.md: both boxes island-bound, agent 0.96.0
This commit is contained in:
@@ -38,6 +38,14 @@ plane. This is the one hard ordering constraint.
|
||||
manual edits**) → the provisioning path is proven live (spike method caveat CLOSED). F1 replay + cold
|
||||
reboot were proven on the same drill (spike P6/P7). *(A literal ISO wipe-reinstall was not run — the
|
||||
surgical provision proves the R-50-specific behavior; the ISO/first-boot/day-0 glue is unchanged.)*
|
||||
- **B2 demo-hp / B3 demo-felhom:** await your go/no-go (each its own STOP). The runbook + rollback table
|
||||
are ready; apps stay up throughout.
|
||||
- **Phase C (Peti cluster): parked** — its own supervised runbook + a new ROADMAP row (to add at Phase 5).
|
||||
- **B2 demo-hp — DONE (2026-07-25).** Agent 0.95.0→0.96.0, vmbr9, guest 9201 net1, island bind, DNS pinned
|
||||
to `.87`. Verify: island `/storage` **200**, LAN DNS OK, **7/7 apps unchanged** (cloudflared/traefik uptime
|
||||
42h = never restarted), public URL **302**, hub reports 0.96.0. No rollback.
|
||||
- **B3 demo-felhom — DONE (2026-07-25).** Same procedure; DNS pinned to `.162`. Verify: island `/storage`
|
||||
**200**, LAN DNS OK, **15/15 apps unchanged**, public URL **302**, hub reports 0.96.0. No rollback.
|
||||
- Both boxes' `.bak-0.95.0` agent + `/etc/network/interfaces.pre-island.bak` retained for rollback.
|
||||
- **Phase 5 (closing docs): DONE.** Capability-map "site/network change" row → **PROVEN-LIVE**; ROADMAP R-50 →
|
||||
**SHIPPED (fleet-migrated)**; **R-74** added (island on Peti's cluster); nodes.md updated (both boxes island,
|
||||
agent 0.96.0). **Firewall (runbook §8):** on both boxes the island bind already moved the listener off the
|
||||
LAN (nothing on `<LAN>:8443`) — the LAN surface is closed by the bind; the portless-bridge narrowing is moot.
|
||||
- **Phase C (Peti cluster): parked as R-74** — its own supervised runbook (SDN/bridge parity), coordinated with Peti.
|
||||
|
||||
@@ -68,7 +68,7 @@
|
||||
| Box survives an **unattended app or guest-network failure** (a dead app member, a boot-orphaned app, a dead DHCP client) — it is noticed, and where safe it is repaired | controller v0.156.0, agent v0.92.1 | **PROVEN-LIVE** (2026-07-21) | All three legs exercised on the live demo box, operator-present, in one session — `felhom-controller/REPORT.md` + `felhom-agent/REPORT.md` (2026-07-21). **Dead primary:** `docker stop immich-server` 12:50:40 CEST → `degraded` 13 s later → **exactly one** `app_start_failed` + dashboard banner → restart → banner self-cleared (the 2026-07-20 shape that was silent for 18 h). **Boot orphan:** `pct reboot 9201` → `[bootrecon] 1 boot-orphaned app(s) found: [bookstack]` → started in 1 attempt of 2, **zero alerts** (success inside the boot grace is silent); `StartedAt` proves Docker's `unless-stopped` did NOT resurrect it — only the sweep did, which also answers P1 and confirms the F5 hypothesis. **Dead DHCP client:** deliberate replay of the incident — `kill -9` 12:43:18 → detected on process liveness 57 s later while the lease was still live → healed 12:45:18 with the incident's verbatim invocation; **the tunnel never dropped (`cloudflared Up 29 hours`)**, i.e. the outage was prevented rather than merely observed | **The gap this validation surfaced → R-55, now FIXED (controller v0.157.0, 2026-07-21).** For a *drive-backed* app a customer's deliberate Stop did NOT survive a reboot — the boot bind gate recreated and started every deployed drive-backed app unconditionally. Pre-existing, not introduced by R-52 (whose own gate was observed correct). The gate now also requires the app to still HAVE containers, which is R-52's own `existing-Exited vs absent` predicate: a UI Stop is `compose down` and removes them. So **"a stopped app stays stopped" now holds for drive-backed apps too — PROVEN LIVE 2026-07-21** (TASK-F Part 3, operator-present). immich was stopped through the real UI endpoint (`compose down` → 0 containers), calibre-web and bookstack left running, then `pct reboot 9201`: the gate recreated calibre-web and logged `1 drive-backed app(s) left stopped — zero containers means the customer stopped them on purpose`; immich came back **stopped**, where the identical fixture had brought it back running hours earlier. Zero alerts, ~15 s to steady state. The **static-guest** half of the network leg stays deliberately out of scope → **R-50** |
|
||||
| Crash/power-loss mid-backup/mid-migration → self-heal on next run | controller, agent | **PROVEN-LIVE** | `CAMPAIGN-6D` P5-REST (SIGKILL mid-offbox → auto-restart ~15s, run marked failed not false-success, no stale lock); `CAMPAIGN-6E` B1-B3 | (Cited `CAMPAIGN-2` T-RBT-* legs were empty / auth-hollow — corrected.) Live mid-**migration** crash→self-heal is the weakest sub-claim (P5-REST is mid-backup) |
|
||||
| An app can be **withdrawn from the catalog without orphaning the customers running it** (available / hidden / abandoned) | controller v0.158.1, catalog metadata | **PROVEN-LIVE** (2026-07-21) | TASK-F Part 1. Verified on 9201 through the real endpoints: `lifecycle: abandoned` arrived via the normal catalog sync; plant-it renders 0 times on the Alkalmazások page (control app renders 10); a direct `POST /api/stacks/plant-it/deploy` → **HTTP 409 "Ez az alkalmazás jelenleg nem telepíthető."**; the app page carries the permanent notice and offers no Telepítés button. `felhom-controller/REPORT.md` (2026-07-21) | Deployed instances keep FULL function in every state — lifecycle governs what is offered, never what runs. Orphan detection deliberately never sees the field (red-proofed): a withdrawn template stays in the catalog tree, or every deployed instance would read `Elavult` and be offered deletion. Unknown values fail OPEN; the deploy gate fails CLOSED. R-57 |
|
||||
| Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent, controller, bootstrap | **PARTIAL** | `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — a real relocation of the demo box: guest + hub telemetry + WG/PBS + Cloudflare tunnel all survived untouched, but the **controller↔agent control plane did not** (agent binds a LAN literal → `bind: cannot assign requested address` → storage/PBS-backup/quiesce/restore-test/DR down until fixed). Mitigated for the window by pinning `vmbr0` static | **R-50** (island-bridge control plane, spike-first) is the durable fix. Related: **R-51** (dead-primary alerting) and **R-52** (boot desired-state reconciliation) — the same event left two apps `Exited` with no alarm and no recovery |
|
||||
| Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap | **PROVEN-LIVE (2026-07-25)** | **R-50 SHIPPED and deployed to the whole fleet.** The control plane now rides a host-internal, portless island bridge (`vmbr9`, `169.254.253.1/30`↔`.2/30`) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original `bind: cannot assign requested address` daemon-death) + cold-reboot survival (`SPIKE-island-bridge-2026-07-25.md`), the migration runbook run verbatim (`RUNBOOK-island-migration.md`), a fresh provision auto-attaching the island `net1` (A4), and the live migration of **both demo boxes** (demo-hp + demo-felhom, 2026-07-25) — island `/storage` HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), **apps served throughout (0 container restarts)**, hub reporting 0.96.0. **Origin:** `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box | **Fleet: DONE.** Remaining: **R-74** — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped |
|
||||
| Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | |
|
||||
| **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. **2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped** (`audits/DIAG-immich-restore-round2-2026-07-19.md`): a deliberate run from snapshot `49e7cb46` did recover all 11 assets (`status=active`, files resolve), but the operation **reported failure** and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. **2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED.** *(The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.)* The C6 drill ran end-to-end **through the UI**: photos deleted, **trash emptied**, the full files+database restore pressed on `/backups/restore`, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: `felhom-controller/REPORT.md` 4e. **Residual: the run was performed by the OPERATOR, not by a customer** — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count |
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -13,7 +13,8 @@
|
||||
| Firmware | AMI AN3PLUS-class | **AMI M42 v01.10 (11/11/2020)** |
|
||||
| PVE node name | `demo-felhom` | `felhom-host` |
|
||||
| Customer | `demo-felhom` | `demo-hp` |
|
||||
| Agent | 0.95.0 | 0.95.0 |
|
||||
| Agent | 0.96.0 | 0.96.0 |
|
||||
| Control plane | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) |
|
||||
| SSH alias | `felhom-pve` | **`demo-hp`** |
|
||||
| Tailnet | `100.70.170.35` | **`100.76.96.79`** |
|
||||
| Loader used to install | `mkimage` (unsigned, SB **off** — firmware workaround) | **`shim`, Secure Boot ENABLED** |
|
||||
@@ -57,7 +58,8 @@ that looked installed and could never call home. Repaired on the console by brid
|
||||
|
||||
Current, post-repair: `vmbr0` static `192.168.0.87/24`, gw `192.168.0.1`, bridge-port `enp2s0f0`.
|
||||
No trace of `192.168.100.2` remains. `wg-felhom` `10.77.0.3/32` up to the hub. Guest **9201
|
||||
`demo-hp`** running. Agent **0.95.0** (deployed 2026-07-25 via break-glass; `.bak-0.93.0` retained; caps 68/68 ok).
|
||||
`demo-hp`** running. Agent **0.96.0** (R-50 island: `local_api` on `169.254.253.1:8443`/`vmbr9`,
|
||||
guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to `192.168.0.87`; `.bak-0.95.0` retained; caps 68/68 ok).
|
||||
|
||||
### Designated drill + build VM host (operator ruling, 2026-07-25)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user