bcdd5b2058
gates / gates (push) Successful in 28s
R-601 said demo-hp was unreachable. The operator looked at the hub and said it was online. It was, and had been up four and a half weeks, reporting every few minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan` at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub had carried the right address in every host report, and `ip neigh` on felhom-pve had .104 four lines above the .87 I quoted — I searched that output for the address I expected instead of reading it for the address that was there. Both ssh entries repointed and verified; nodes.md corrected, including that the tailnet route for this box does not exist. The hunt then found R-604, which is the real defect: demo-hp carried a per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. `managed floor SERVED` fires once per change by design, so a box behind a static override is silent for ever and its silence is indistinguishable from a box that already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes and its claim page now answers "Wrong or expired code" in English. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
213 lines
15 KiB
Markdown
213 lines
15 KiB
Markdown
# Fleet node inventory — the physical demo/lab hosts
|
||
|
||
> Added 2026-07-21, when the fleet stopped being one box. Two Proxmox hosts now check in to the hub.
|
||
> This is the operator-facing inventory: what the hardware is, how to reach it, and what is
|
||
> deliberately NOT enrolled on it.
|
||
|
||
## The fleet
|
||
|
||
| | `demo-felhom-8363b5` | `demo-hp-bb76ea` |
|
||
|---|---|---|
|
||
| Hardware | N100 mini-PC | **HP t740 Thin Client** (s/n `8CN944035T`) |
|
||
| CPU / RAM | Intel N100 | **AMD Ryzen Embedded V1756B**, 8 threads / **30 GiB** |
|
||
| Firmware | AMI AN3PLUS-class | **AMI M42 v01.10 (11/11/2020)** |
|
||
| PVE node name | `demo-felhom` | `felhom-host` |
|
||
| Customer | `demo-felhom` | `demo-hp` |
|
||
| Agent / controller | *not recorded here* — see the note below | *not recorded here* |
|
||
| Control plane | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) |
|
||
| SSH alias | `felhom-pve` | **`demo-hp`** |
|
||
| Tailnet | `100.70.170.35` | **NONE — tailscale is NOT INSTALLED on demo-hp** (checked on the box 2026-09-21: no `tailscaled`, no `tailscale` binary). The peer `100.76.96.79` still listed by `tailscale status` is a STALE entry from an earlier build and **can never answer**; it showed `offline, last seen 30d ago` while the box was up and reporting. Use the LAN address. |
|
||
| Loader used to install | `mkimage` (unsigned, SB **off** — firmware workaround) | **`shim`, Secure Boot ENABLED** |
|
||
|
||
> **No component versions are recorded on this page — deliberately.** Agent, controller, hub and
|
||
> host-install versions change several times a day, so any number written here is wrong within hours and
|
||
> is then read as fact. **Ask the fleet instead:** the hub host list (`/hosts`) and customer list
|
||
> (`/configs`) carry the live agent and controller versions per box; `felhom-agent --version` on the
|
||
> node and `pct exec <vmid> -- docker ps` in the guest are the authorities. Versions that must be pinned
|
||
> in writing belong in the per-repo `CHANGELOG.md` and the hub's Day-0 artifact manifest — not in an
|
||
> inventory. **The fleet is not uniform**: on 2026-07-30 the two boxes ran different agent *and*
|
||
> different controller versions, so a single number for "the fleet" would have been wrong regardless.
|
||
|
||
**Both are at the VACATION site** and travel home with the rest of the kit **~2026-08-02**. While
|
||
away, their LAN addresses are on that site's `192.168.0.0/24`: N100 `.147`/`.162`, HP `.87` — **re-check
|
||
rather than trusting these** (`ip -br addr show vmbr0`; the N100 read `.162` on 2026-07-30). The
|
||
tailnet addresses are the stable ones — use those. Direct LAN literals are **not** reachable from
|
||
DooPlex while the boxes are away (`felhom-pve-lan` → `No route to host`, 2026-07-30).
|
||
|
||
**Which box is safe to break, and what may be done to each:
|
||
[`../runbooks/target-selection.md`](../runbooks/target-selection.md).** This page is *what the hardware
|
||
is*; that page is *what you may do to it*.
|
||
|
||
## demo-hp — the HP t740, in detail
|
||
|
||
### Disks
|
||
|
||
| device | model | serial | role |
|
||
|---|---|---|---|
|
||
| `sda` | SanDisk X600 M.2 2280 SATA 128GB | `182195804614` | **system disk** — PVE, LVM (`pve-root` 39.6G, `pve-data` thin pool, guest 9201's three volumes) |
|
||
| `nvme0n1` | KXG50PNV1T02 NVMe TOSHIBA 1024GB | `58BS11AFT8MQ` | **ENROLLED** — mounted `/mnt/nvme-1tb`, the enrolled user-data drive **and** the `felhom-backup` whole-guest backup target |
|
||
|
||
> **The NVMe joined the product on 2026-07-22, through the normal Tárhely flow, as intended.** Enrolled
|
||
> to guest 9201 (PUBLISH TRAIN, `pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md`); the vzdump target was
|
||
> moved onto it by E-2a. Verified live 2026-07-30: `nvme0n1` → `/mnt/nvme-1tb`, and
|
||
> `dir: felhom-backup / path /mnt/nvme-1tb / is_mountpoint 1` in `storage.cfg`.
|
||
>
|
||
> **The "PRESENT AND UNENROLLED — do not touch" fence that stood here is RETRACTED**, and its reason is
|
||
> recorded so it is not mistaken for a live rule: the NVMe was outside everything because the install
|
||
> ISO's exact-serial filter pinned `sda` only, and it was to join *through the Tárhely flow rather than
|
||
> the installer or by hand*. **That condition was satisfied; the prohibition expired with it.** It stood
|
||
> for eight days after enrolment and contradicted the task specs that (correctly) sent drill-VM disks to
|
||
> `/mnt/nvme-1tb`.
|
||
>
|
||
> Still true, and now the operative caution: a dir storage there **must sit at the mountpoint root** (a
|
||
> subdirectory fails the agent's `exactMount` check → the storage reads `disconnected` forever), and it
|
||
> shares the device with the box's own backups — so remove scratch storages when done.
|
||
|
||
### NIC map — and the trap
|
||
|
||
This board has **five wired interfaces**, and the obvious one is the wrong one:
|
||
|
||
| interface | MAC | driver | what it is | state at this site |
|
||
|---|---|---|---|---|
|
||
| `enp1s0f0` | `a0:36:9f:5d:07:20` | `igb` | 4-port expansion card | no carrier, **no DHCP** |
|
||
| `enp1s0f1` | `a0:36:9f:5d:07:21` | `igb` | ″ | no carrier |
|
||
| `enp1s0f2` | `a0:36:9f:5d:07:22` | `igb` | ″ | no carrier |
|
||
| `enp1s0f3` | `a0:36:9f:5d:07:23` | `igb` | ″ | no carrier |
|
||
| **`enp2s0f0`** | `7c:d3:0a:77:d9:76` | `r8169` | **onboard port — the one that works** | carrier up, 1000 Mb, **this is `vmbr0`'s port** |
|
||
| `wlo1` | `24:ee:9a:e5:05:b0` | `iwlwifi` | wifi | unused |
|
||
|
||
**This trap cost the first install.** The 4-port card got no lease, and instead of aborting the
|
||
installer baked its `192.168.100.2` fallback as a **static** `vmbr0` address and completed — a box
|
||
that looked installed and could never call home. Repaired on the console by bridging `vmbr0` to
|
||
`enp2s0f0`. Filed as **R-59** (must hard-abort) and **R-60** (first-boot NIC sweep self-heal).
|
||
|
||
Current, read off the box 2026-09-21: `vmbr0` **static `192.168.0.104/24`**, gw `192.168.0.1`, bridge-port **`nic0`**. (It was `192.168.0.87/24` on `enp2s0f0` before a reprovision; both were stale here for long enough to cost a session a false "the box is down" — R-601.) **The hub always knows the truth:** every host report carries `addresses: [{iface, cidr}, …]`, so read it there rather than from this page.
|
||
No trace of `192.168.100.2` remains. `wg-felhom` `10.77.0.3/32` up to the hub. Guest **9201
|
||
`demo-hp`** running. Agent config shape (R-50 island): `local_api` on `169.254.253.1:8443`/`vmbr9`,
|
||
guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to the host's LAN address.
|
||
|
||
### Designated drill + build VM host (operator ruling, 2026-07-25)
|
||
|
||
**Ruling:** drill and build VMs are hosted on the **t740 from now on** — NOT on felhom-pve, and moving
|
||
them off DooPlex (the production k3s node). **This is a VM-HOSTING ruling only; the build-PIPELINE
|
||
relocation to the t740 is NOT ruled or implemented here.**
|
||
|
||
**Current state (updated 2026-07-25 PM):** the ruling is **realized** — the t740 now hosts the first
|
||
drill appliance. **QEMU VM `300` = `drill-r50`**, a nested PVE-in-a-VM (8 GiB RAM, 4 vCPU cpu=host, 32 GiB
|
||
local-lvm disk, OVMF/SB-off, one NIC on `vmbr0` DHCP). Installed from
|
||
`felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` through the **real day-0** (self-register →
|
||
operator bind to scratch customer `drill-r50` → deliver → nested guest **9201** provisioned and healthy
|
||
on the then-current controller + agent). Its own break-glass root is vaulted in the hub
|
||
`host_recovery/drill-r50-0a4f9a`; reach it as `root@192.168.0.176` **through `demo-hp`** (it has no key
|
||
and no tailnet — it is a peer on demo-hp's LAN). It has snapshot **`r50pre`** (clean LAN-literal day-0).
|
||
This VM was provisioned to unblock and run the **R-50 island-bridge empirical spike**
|
||
(`audits/SPIKE-island-bridge-2026-07-25.md`) — **probes P1–P8 PASSED 2026-07-25, verdict GO**; the drill
|
||
was left in its working island configuration. It is a throwaway: destroy with `qm stop 300 && qm destroy
|
||
300 --purge 1` (and delete the `drill-r50` customer + appliance #9 hub-side) when no longer needed. The
|
||
historical golden-bake `drill.qcow2` still lives on **DooPlex** (`/mnt/5_hdd/felhom.eu/drill/`, ~18G,
|
||
powered off) and is unrelated. The build-PIPELINE relocation to the t740 remains unbuilt.
|
||
|
||
### Access — there is no baked SSH key
|
||
|
||
`ssh demo-hp` resolves to the tailnet address, but **no operator public key is on this box** — the
|
||
HP profile deliberately left `FELHOM_ROOT_SSH_KEY` blank. Authentication is the **G1 break-glass
|
||
root password vaulted in the hub**, `host_recovery` row `demo-hp-bb76ea` (set at day-0,
|
||
2026-07-21 16:24 UTC).
|
||
|
||
Retrieval (operator-side, and **shred the copy** — that DB holds every host's secret):
|
||
|
||
> **WAL-AWARE SINCE HUB v0.88.0 — copying `hub.db` ALONE is no longer safe.** The hub runs SQLite in
|
||
> **WAL** mode (R-172), so a committed transaction may still live in `hub.db-wal` and not yet be in
|
||
> the main file. A bare `cat /data/hub.db` therefore yields a copy that is **valid but stale** — it
|
||
> opens cleanly and silently lacks the most recent writes, which is the worst failure shape for a
|
||
> credential lookup. Copy the `-wal` beside it and let SQLite replay it on open.
|
||
|
||
```bash
|
||
sudo kubectl -n felhom-system exec <hub-pod> -- cat /data/hub.db > /tmp/x.db
|
||
sudo kubectl -n felhom-system exec <hub-pod> -- cat /data/hub.db-wal > /tmp/x.db-wal 2>/dev/null || true
|
||
python3 -c "import sqlite3;print(sqlite3.connect('/tmp/x.db').execute(
|
||
\"SELECT secret FROM host_recovery WHERE host_id='demo-hp-bb76ea'\").fetchone()[0])"
|
||
shred -u /tmp/x.db /tmp/x.db-wal
|
||
```
|
||
|
||
The `|| true` is deliberate: an absent `-wal` is legitimate (a freshly checkpointed database), and
|
||
must not fail the retrieval. **Shred both files** — the WAL holds the same secrets as the DB.
|
||
|
||
Then `sshpass -e ssh root@demo-hp` (sshpass is on DooPlex, not on the nodes).
|
||
|
||
**This is the lockout filed as R-61**: the ISO mints a throwaway root password per build and discards
|
||
the plaintext, so the console is unreachable without a working hub and network — precisely what you
|
||
may be trying to fix. Slice 1 is to emit the baked password into the build report.
|
||
|
||
`demo-hp-lan` (**`192.168.0.104`** via `ProxyJump felhom-pve`) is the fallback when DooPlex is not on the home LAN. **`ssh demo-hp` now goes DIRECT to `192.168.0.104`** — DooPlex is on the same LAN, and the tailnet route for this box does not exist (see the table above). Both were verified 2026-09-21.
|
||
|
||
## OOB belt (H1) — both boxes, since 2026-07-23 (ISO train v1.25.0)
|
||
|
||
The dedicated OOB sshd belt (TASK H1: `felhom-sshd` + the static `inet felhom_oob` table + `felhom-op`)
|
||
is installed and **active on BOTH fleet boxes** — the F9 gap (belt on neither) is closed. From
|
||
v1.25.0 host-install installs it by default on every appliance install (`--no-oob` opts out; byo still
|
||
refuses).
|
||
|
||
- **Claimed port: `8822` on both** (first-free from `[8822,2222,8022,62222]`; persisted per box).
|
||
- **Reachability: the wg-felhom offsite tunnel ONLY** — the belt admits the operator `/32`
|
||
(`10.77.0.250`) over `wg-felhom` to 8822 and drops everything else; `:22` and every other interface
|
||
are untouched. **tailscale does NOT reach the belt** (wrong fabric, dropped by design).
|
||
- **Operator login** (from the machine holding the wg-felhom operator tunnel + the registered
|
||
`oob_operator_ssh_pubkey`): `ssh -p 8822 felhom-op@10.77.0.2` (felhom-pve) / `@10.77.0.3` (demo-hp).
|
||
**PROVEN live 2026-07-23** on felhom-pve (`felhom-op@demo-felhom`).
|
||
- **Operator tunnel**: the Mac/Windows operator peer dials `ep0.felhom.eu:443` (WireGuard), address
|
||
`10.77.0.250/32`, AllowedIPs `10.77.0.0/24`, server pubkey `f3d1ZI7…`. ep0's `forward` chain
|
||
(persisted in its `/etc/nftables.conf`) allows `10.77.0.250 → 10.77.0.2/.3`. If a work-network blocks
|
||
UDP/443, the RheinMetall-style firewalls pass UDP/51820 — a home/hotspot network works on 443.
|
||
- Register/rotate the operator identity hub-side: `PUT /api/v1/admin/wg/operator-peer` (global key)
|
||
with `{pubkey, assigned_ip:"10.77.0.250", ssh_pubkey}`; wgsync pushes it to ep0 and the SSH key flows
|
||
to both boxes' `felhom-op` authorized_keys within a tick.
|
||
|
||
## felhom-pve (the N100) — vault parity + access
|
||
|
||
felhom-pve has **operator SSH-key access** (over tailscale `100.70.170.35`) AND, since 2026-07-23,
|
||
**G1 break-glass vault parity with demo-hp**: its root@pam password is freshly rotated and vaulted in
|
||
the hub `host_recovery` row **`demo-felhom-8363b5`** (same PUT `…/recovery-credential` mechanism day-0
|
||
uses; verified retrievable + authenticating over `:22`). Retrieval + shred-the-copy recipe is identical
|
||
to demo-hp's below (swap the host_id). So a lost N100 key is recoverable the same way as the key-less HP.
|
||
|
||
## Tailscale on demo-hp is an OPERATOR-LAB EXCEPTION
|
||
|
||
> **Read this before any product-shape audit.** `demo-hp` is **customer-shaped** — it is a normal
|
||
> appliance install with a real customer record (`demo-hp`), a real guest, and a real day-0.
|
||
> **Tailscale is not part of that shape.** It was installed by hand on 2026-07-21 purely so the
|
||
> operator can reach a lab box that lives on someone else's LAN.
|
||
>
|
||
> **Real customer boxes never get tailscale.** Their operator access is the WireGuard tunnel plus
|
||
> the H1 OOB path, and nothing else. If a future audit finds tailscale on `demo-hp` and concludes
|
||
> the product ships it — that conclusion is wrong, and this paragraph is the reason. The same
|
||
> exception already applies to `felhom-pve`.
|
||
|
||
Details, and the two hard rules that apply to any host node, in `operations/tailscale.md`.
|
||
|
||
### demo-hp scratch guest — LXC 9202 `demo-hp-scratch` (R-481, built 2026-09-13)
|
||
|
||
- **Disposition:** the nightly rotation's throwaway host. Under the **demo-hp customer** (same domain,
|
||
same dashboard password), **hub reporting OFF, Cloudflare tunnel OFF, agent local API OFF, off-site
|
||
OFF, self-update OFF** (the image is set by hand in `/etc/felhom-controller-image`). **Persists across
|
||
nights on purpose.** Apps on it are throwaways; nothing on it is a customer promise. Not in the felhom
|
||
pool, so the hub never sees it. Tags `scratch,r481`; the same text sits in the guest at
|
||
`/etc/felhom-scratch-disposition`, in the host's `pct` description and in its bootstrap file.
|
||
- **Reach:** `https://192.168.0.114` with the demo-hp Host names (`felhom.enkisfelhom.hu`,
|
||
`<sub>.enkisfelhom.hu`) — LAN only. Helper: the scratch twin of `ctl.sh`.
|
||
- **Shape:** 7 cores, 25 898 MB cap, rootfs 32 G + data 70 G, both on the `nvme-scratch` dir storage
|
||
at `/mnt/hdd_1` (the runbook's NVMe path; a `dir` storage was re-added there). Unprivileged.
|
||
"Second drive" for Tier 2: host `/mnt/hdd_1/scratch-drives/scratch_hdd` mounted at
|
||
`/mnt/felhom-drives/scratch_hdd` (mp8), registered as the default storage path.
|
||
- **Rebuild from nothing:** restore the golden (`local:backup/felhom-golden-<ver>.tar.zst`) with
|
||
`pct restore … --storage nvme-scratch`, add mp0/mp8/mp9 as above, fix ownership from the host with
|
||
the unprivileged mapping (100000), seed `controller.yaml` (hub/tunnel/agent/off-site off) and a
|
||
claimed `settings.json`, then start `felhom-controller-bootstrap.service`.
|
||
|
||
## What is NOT enrolled here (deliberately)
|
||
|
||
- **No PBS datastore, no offsite target** on demo-hp; the DR tier is the N100's.
|
||
- **No second customer guest** beyond 9201.
|
||
|
||
*(The 1TB NVMe was on this list until 2026-07-22. It is enrolled — see the disks table above.)*
|