docs: ISO train v1.25.0 — REPORT (belt/apt/R-63/gate/vault live), F9 resolved, R-63 shipped, nodes belt+vault, F8 checklist; critical golden<floor finding
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NKSN3gSg4TKVBBqkwW2djR
This commit is contained in:
@@ -107,7 +107,7 @@ repo consistent, 14 snapshots, latest 04:16 CEST (pre-outage), **zero locks** (r
|
||||
| # | Finding | Severity | Evidence | Disposition |
|
||||
|---|---|---|---|---|
|
||||
| **F8** | **Site power loss; no auto-power-on.** Power itself returned within minutes (router rebooted at ~14:42 and stayed up); both miniPCs remained off ~4.5 h until manual power-on. For a paying customer this converts a power blip into a half-day outage ending only when someone is physically present. | HIGH | p1-\*; router uptime operator-attested | **mitigated-on-site** — Viktor set BIOS AC-power-on on both boxes (attested; not OS-verifiable). **roadmap-candidate:** make BIOS "restore on AC power" a provisioning-checklist item + a host-install doc requirement for every fleet box. |
|
||||
| **F9** | **H1 OOB belt only partially installed on the current fleet.** `felhom-mgmt-watchdog.timer` is live on both boxes (heal marker absent = no heal needed), but `felhom-sshd.service` and `felhom-oob-nft.service` (TASK H1, present in `felhom-agent/configs/`) are installed on **neither** box; operator access rides stock sshd :22 + tailscale + G1 break-glass. Pre-existing (both boxes provisioned via the universal ISO), surfaced by this audit's access preflight. | MEDIUM | p2-felhom-pve-ssh-belt.txt, p2-\*-host.txt | **needs-ruling** — was H1 intentionally dropped from the universal-ISO provisioning path, or should host-install grow the belt? |
|
||||
| **F9** | **H1 OOB belt only partially installed on the current fleet.** `felhom-mgmt-watchdog.timer` is live on both boxes (heal marker absent = no heal needed), but `felhom-sshd.service` and `felhom-oob-nft.service` (TASK H1, present in `felhom-agent/configs/`) are installed on **neither** box; operator access rides stock sshd :22 + tailscale + G1 break-glass. Pre-existing (both boxes provisioned via the universal ISO), surfaced by this audit's access preflight. | MEDIUM | p2-felhom-pve-ssh-belt.txt, p2-\*-host.txt | **RESOLVED 2026-07-23 (ISO train v1.25.0).** Ruling: install everywhere. host-install grows the belt as a DEFAULT appliance leg (`--no-oob` opts out; byo still refuses — deliberate). Phase-0 confirmed the omission was NOT a coded exclusion, just `--enable-oob` never passed by the universal ISO. Belt installed + `oob.enabled` on BOTH live boxes; **login PROVEN end-to-end on felhom-pve** (`felhom-op@demo-felhom` over wg-felhom → belt); also re-anchored the orphaned operator identity to the operator's real machine. See `REPORT.md` (2026-07-23) + `operations/nodes.md`. |
|
||||
| **F10** | **demo-hp backup tiers incomplete.** Tier-2 dump tree exists (paperless-ngx, fresh), but **no offbox target is configured** and the PBS DR datastore holds **0 snapshots** for it. A power event with disk damage would have had no off-box recovery path for that guest. Pre-existing (box added 07-21), not outage-caused. | MEDIUM | p6-demo-hp-offbox-locks.txt (`NO-OFFBOX`), p6-demo-hp-backup.txt (`snapshots=0`), p6-demo-hp-tier2.txt | **OFFSITE LEG RESOLVED 2026-07-23** — root cause was NOT "provisioning unfinished": the day-0 managed update killed the apply-bridge after password-consume, burning the one-shot credential (full diagnosis, designed-path repair via operator Re-issue, escrow ceremony, and a byte-identical offsite restore round-trip: `DIAG-f10-demo-hp-offsite-2026-07-23.md`; product rows minted **R-70** visibility + **R-71** the race). **The PBS-DR-snapshot half STAYS OPEN** pending F13 (cadence ruling) and the deliberate DR ceremony R-moment on demo-hp. |
|
||||
| **F11** | **Recovery is silent.** `host_recovered`/`node_recovered` fired correctly but carry severity `info`, which the dispatcher deliberately does not email — the operator/customer only learns of recovery by looking. During a real customer outage the "it's back" signal is arguably the second-most valuable email. Working-as-coded, so a product question, not a defect. | LOW | p5-hub-events.txt; `hub/internal/notify/dispatcher.go` severityNotifies | **needs-ruling** — opt-in recovery notifications (operator at least)? |
|
||||
| **F12** | **demo-hp has no customer notification prefs** (`customer_notifications` has no row for it) → its "customer" received no node_down email and never would. Only demo-felhom is wired (doodoo21@freemail.hu). Pre-existing demo-box config gap; on a real onboarding this must not be skippable. | LOW | p5-customer-notification-prefs.txt | **roadmap-candidate** — make notification-prefs setup a claim/onboarding step, not an optional settings page. |
|
||||
|
||||
@@ -76,6 +76,7 @@
|
||||
| R-60 | **[P2] First-boot NIC sweep self-heal: if the hub is unreachable, try DHCP across every carrier-bearing NIC before settling.** | S | **SHIPPED v1.24.0 (2026-07-22) — spike + nested drill proven (SPIKE-firstboot-nic-sweep-2026-07-22.md): cable move → sweep → heal + hub registration unaided in <1 min; sweep is structurally first-boot-only (state.json gate + the unit's done-flag condition); drill also surfaced and fixed the baked-fallback-default-route trap (flush before the bounded dhclient)** | `felhom-bootstrap` currently accepts whatever addressing the installer left behind and, if the hub cannot be reached, simply stays broken. On demo-hp the fix was a human moving one cable from the 4-port card to the onboard port — **a sweep would have healed it unaided**: enumerate NICs with `carrier=1`, DHCP each in turn, and keep the first that reaches the hub. Cheap because the box has nothing to lose at first boot (no customer data, no running guests) and the failure it repairs is total. Deliberately scoped to FIRST BOOT and to the hub-unreachable condition only — a running box must never re-shuffle its own networking. Complements **R-59**: that one refuses to produce an unreachable box, this one repairs the case where the truth changed after the install (cable moved, switch port died, the installer guessed the wrong port) |
|
||||
| R-61 | **[P1] The baked root password must be knowable by the operator — the recurring console lockout.** | S | **slice 1 SHIPPED v1.24.0 (2026-07-22): the build emits the plaintext into a 0600 sibling `<iso>.rootpw.txt` (single record of truth — never logged/manifested/committed); drill-verified against the installed box's shadow hash. Follow-up (appliance-grade record-keeping) stays open** | The ISO mints a **fresh throwaway crypt hash per build** and the plaintext is discarded, so nobody — including the person holding the machine — can log into the console of a box they just installed. Today that meant reaching demo-hp only through the G1 break-glass credential vaulted in the hub, which is the right mechanism for a *lost* password and the wrong one for a *never-known* password: it requires a working hub, a working network, and operator tooling, at exactly the moment the likely reason you need the console is that one of those is broken. **Slice 1 (do this):** the ISO build emits the baked root password into the build REPORT and the operator cheat-sheet alongside the sha256 — it is already a per-build value, so surfacing it costs nothing and closes the lockout. **Follow-up (appliance-grade):** keep it per-build random and treat the build output as the record of truth. **A fixed well-known password is explicitly REJECTED (operator ruling 2026-07-21)** — a pre-pairing box sits on a stranger's LAN with a predictable root credential, which is a far worse exposure than the lockout it would fix. Relates to G1 break-glass (the vault stays; this is about the window before/without it) |
|
||||
| R-62 | **[P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape.** | XS | **idea (operator, 2026-07-22)** | Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; **wording only** — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). |
|
||||
| R-63 | **The install console learns ő/ű** — the kernel default console font lacks the Hungarian double-acute glyphs, so the R-59 network screen and pairing banner rendered ő as blanks. | XS | **SHIPPED (scripts v1.25.0, 2026-07-23)** | `felhom-bootstrap.sh` loads a Latin-2 console font (`Lat2-Terminus16` → `Lat2-Fixed16` → `Lat2-Terminus14`) ONCE before the first paint — idempotent, best-effort (a missing font/ioctl never blocks boot). Lat2 ships in the trixie/PVE base (console-setup), so no copy rewording needed. Font names verified against the package. Nested-console capture proof rides the v1.25.0 drill. Also in the same train: **F9 belt-everywhere RESOLVED** (host-install default appliance leg + live on both boxes + login proven) and the **R-71 build-gate** (`build-felhom-iso` asserts golden ≥ managed floor, `publish-train-rules.md` rule 5). **Live finding: golden 0.153.0 < floor 0.156.0 in production NOW** — the gate catches it; the fix is the golden republish at 0.161.0 (Part 4, pending; the managed floor stays 0.156.0). See `REPORT.md` (2026-07-23). |
|
||||
| R-64 | **„Felhom↔Felhom media pairing blessed" — the two-box SMB pairing (one box shares, the other mounts it as NAS storage) becomes a supported, documented flow.** | XS–S | idea (2026-07-22) | Origin: the operator ran the pairing drill on the live demo pair and it WORKS — the drill itself is the pending evidence leg (a written run-through with the R-66 surfaces in play). R-66 shipped the enabling visibility: the serving box's address is now on its own Beállítások → Rendszer „Hálózat" card, and the add form names the NetBIOS trap. Blessing = a short customer-facing recipe (`documentation/controller/network-storage-nas.md` naming-caveat paragraph is the seed) + one supported-path sentence in the capability map. Flips: would add a "Felhom↔Felhom media pairing" capability row (currently unlisted). Pairs with R-65 (same two-box topology, entirely different transport + guarantees) |
|
||||
| R-66 | **The box's own address becomes visible — „Hálózat" card, Debug network dump, NetBIOS hint.** | XS | **SHIPPED (controller v0.159.0, 2026-07-22)** | Origin: the pairing drill — the serving box's IP was findable only as a hint buried on the OTHER box's Megosztás page, and the add form's failure for „FELHOM" taught nothing. Three legs: (A) „Hálózat" card on Beállítások → Rendszer (Helyi cím / Hálózati név only-while-sharing / Átjáró; live per render, stored nowhere — S-5; „—" when unavailable); (B) `network` section in the Debug dump (interfaces/route/DNS/lan_address, best-effort per item); (C) the NetBIOS trap named (Szerver helper text + a purely lexical hint on `unreachable` for single-label non-IP names). **Design decision recorded:** the controller is bridge-netns'd, so ALL guest-net reads go through the one netns door (docker exec into host-networked felhom-samba, `stacks/guestnet.go`) — with Megosztás off the card honestly shows „—" rather than the plausible-wrong 172.x answer. Deployed demo-felhom + demo-hp 2026-07-22; demo-hp live-shows the closed-door path (sharing off → dashes + in-place dump errors), demo-felhom the open one (real .104/.1/\\FELHOM values). Flips no capability-map row (diagnosability/UX polish); enables R-64 |
|
||||
| R-67 | **The NAS share appears in FileBrowser — browse what you mounted.** | S | **SHIPPED (controller v0.160.0, 2026-07-22)** | Origin: the R-64 pairing drill — the share said „Elérhető" and the customer had no way to BROWSE it (FileBrowser synced drives only). **Couples to R-64: browsing was its missing UX half.** A registered network storage now binds its share ROOT into FileBrowser (`/mnt/felhom-drives/<name>:/srv/<name>:rslave`) with its display label as the sidebar source; NAS add/remove trigger the same debounced sync. Two classes, two gates: drives keep the drive-absent gate byte-identically (proven live: the drives-only box logged a no-op sync); network shares gate on the STUB classifier instead — idle autofs is HEALTHY and included (Phase-0 probe on demo-hp: an in-container access through an rslave bind WAKES the idle trigger), while a stub verdict excludes the share from mounts AND sources with a WARN (an exposed stub swallows uploads the real mount later shadows). Nothing is ever written toward the NAS (no skeleton — red-proven). Live leg: cross-box upload round-trip demo-hp → demo-felhom + dead-NAS check (`Host is down` in seconds, unaided recovery after samba restart). Operator residual: the FileBrowser UI click-through (its admin credential is customer-held by design). Evidence: `felhom-controller/REPORT.md` (2026-07-22) |
|
||||
|
||||
@@ -83,6 +83,36 @@ may be trying to fix. Slice 1 is to emit the baked password into the build repor
|
||||
|
||||
`demo-hp-lan` (`192.168.0.87` via `ProxyJump felhom-pve`) is the fallback while the box is away.
|
||||
|
||||
## OOB belt (H1) — both boxes, since 2026-07-23 (ISO train v1.25.0)
|
||||
|
||||
The dedicated OOB sshd belt (TASK H1: `felhom-sshd` + the static `inet felhom_oob` table + `felhom-op`)
|
||||
is installed and **active on BOTH fleet boxes** — the F9 gap (belt on neither) is closed. From
|
||||
v1.25.0 host-install installs it by default on every appliance install (`--no-oob` opts out; byo still
|
||||
refuses).
|
||||
|
||||
- **Claimed port: `8822` on both** (first-free from `[8822,2222,8022,62222]`; persisted per box).
|
||||
- **Reachability: the wg-felhom offsite tunnel ONLY** — the belt admits the operator `/32`
|
||||
(`10.77.0.250`) over `wg-felhom` to 8822 and drops everything else; `:22` and every other interface
|
||||
are untouched. **tailscale does NOT reach the belt** (wrong fabric, dropped by design).
|
||||
- **Operator login** (from the machine holding the wg-felhom operator tunnel + the registered
|
||||
`oob_operator_ssh_pubkey`): `ssh -p 8822 felhom-op@10.77.0.2` (felhom-pve) / `@10.77.0.3` (demo-hp).
|
||||
**PROVEN live 2026-07-23** on felhom-pve (`felhom-op@demo-felhom`).
|
||||
- **Operator tunnel**: the Mac/Windows operator peer dials `ep0.felhom.eu:443` (WireGuard), address
|
||||
`10.77.0.250/32`, AllowedIPs `10.77.0.0/24`, server pubkey `f3d1ZI7…`. ep0's `forward` chain
|
||||
(persisted in its `/etc/nftables.conf`) allows `10.77.0.250 → 10.77.0.2/.3`. If a work-network blocks
|
||||
UDP/443, the RheinMetall-style firewalls pass UDP/51820 — a home/hotspot network works on 443.
|
||||
- Register/rotate the operator identity hub-side: `PUT /api/v1/admin/wg/operator-peer` (global key)
|
||||
with `{pubkey, assigned_ip:"10.77.0.250", ssh_pubkey}`; wgsync pushes it to ep0 and the SSH key flows
|
||||
to both boxes' `felhom-op` authorized_keys within a tick.
|
||||
|
||||
## felhom-pve (the N100) — vault parity + access
|
||||
|
||||
felhom-pve has **operator SSH-key access** (over tailscale `100.70.170.35`) AND, since 2026-07-23,
|
||||
**G1 break-glass vault parity with demo-hp**: its root@pam password is freshly rotated and vaulted in
|
||||
the hub `host_recovery` row **`demo-felhom-8363b5`** (same PUT `…/recovery-credential` mechanism day-0
|
||||
uses; verified retrievable + authenticating over `:22`). Retrieval + shred-the-copy recipe is identical
|
||||
to demo-hp's below (swap the host_id). So a lost N100 key is recoverable the same way as the key-less HP.
|
||||
|
||||
## Tailscale on demo-hp is an OPERATOR-LAB EXCEPTION
|
||||
|
||||
> **Read this before any product-shape audit.** `demo-hp` is **customer-shaped** — it is a normal
|
||||
|
||||
@@ -201,6 +201,11 @@ Recommended order: **GL-1 and GL-3 immediately** (operator-heavy, unblock everyt
|
||||
|
||||
## 6. Open questions & operator actions
|
||||
|
||||
**Provisioning-checklist items (every fleet box, bench prep):**
|
||||
- **BIOS "restore on AC power" = ON** (power-outage audit F8, 2026-07-22): a site power blip must not
|
||||
become a half-day outage waiting for someone to press the button. Set + attest per box (not
|
||||
OS-verifiable). Both current demo boxes done (Viktor, 07-22).
|
||||
|
||||
**Operator actions (Viktor):**
|
||||
- **Hub manifest bump — DONE** (confirmed at GL-6 Phase 0: the manifest vouches agent `0.76.0` +
|
||||
golden `0.103.0`). No action.
|
||||
|
||||
Reference in New Issue
Block a user