Files
felhom.eu/documentation/operations/nodes.md
T
admin bcdd5b2058
gates / gates (push) Successful in 28s
floor 0.259.0 raised; R-601 withdrawn as FALSE; R-604 filed
R-601 said demo-hp was unreachable. The operator looked at the hub and said it
was online. It was, and had been up four and a half weeks, reporting every few
minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a
tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan`
at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub
had carried the right address in every host report, and `ip neigh` on felhom-pve
had .104 four lines above the .87 I quoted — I searched that output for the
address I expected instead of reading it for the address that was there.

Both ssh entries repointed and verified; nodes.md corrected, including that the
tailnet route for this box does not exist.

The hunt then found R-604, which is the real defect: demo-hp carried a
per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so
it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0.
`managed floor SERVED` fires once per change by design, so a box behind a static
override is silent for ever and its silence is indistinguishable from a box that
already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes
and its claim page now answers "Wrong or expired code" in English.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 09:13:57 +02:00

213 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Fleet node inventory — the physical demo/lab hosts
> Added 2026-07-21, when the fleet stopped being one box. Two Proxmox hosts now check in to the hub.
> This is the operator-facing inventory: what the hardware is, how to reach it, and what is
> deliberately NOT enrolled on it.
## The fleet
| | `demo-felhom-8363b5` | `demo-hp-bb76ea` |
|---|---|---|
| Hardware | N100 mini-PC | **HP t740 Thin Client** (s/n `8CN944035T`) |
| CPU / RAM | Intel N100 | **AMD Ryzen Embedded V1756B**, 8 threads / **30 GiB** |
| Firmware | AMI AN3PLUS-class | **AMI M42 v01.10 (11/11/2020)** |
| PVE node name | `demo-felhom` | `felhom-host` |
| Customer | `demo-felhom` | `demo-hp` |
| Agent / controller | *not recorded here* — see the note below | *not recorded here* |
| Control plane | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) |
| SSH alias | `felhom-pve` | **`demo-hp`** |
| Tailnet | `100.70.170.35` | **NONE — tailscale is NOT INSTALLED on demo-hp** (checked on the box 2026-09-21: no `tailscaled`, no `tailscale` binary). The peer `100.76.96.79` still listed by `tailscale status` is a STALE entry from an earlier build and **can never answer**; it showed `offline, last seen 30d ago` while the box was up and reporting. Use the LAN address. |
| Loader used to install | `mkimage` (unsigned, SB **off** — firmware workaround) | **`shim`, Secure Boot ENABLED** |
> **No component versions are recorded on this page — deliberately.** Agent, controller, hub and
> host-install versions change several times a day, so any number written here is wrong within hours and
> is then read as fact. **Ask the fleet instead:** the hub host list (`/hosts`) and customer list
> (`/configs`) carry the live agent and controller versions per box; `felhom-agent --version` on the
> node and `pct exec <vmid> -- docker ps` in the guest are the authorities. Versions that must be pinned
> in writing belong in the per-repo `CHANGELOG.md` and the hub's Day-0 artifact manifest — not in an
> inventory. **The fleet is not uniform**: on 2026-07-30 the two boxes ran different agent *and*
> different controller versions, so a single number for "the fleet" would have been wrong regardless.
**Both are at the VACATION site** and travel home with the rest of the kit **~2026-08-02**. While
away, their LAN addresses are on that site's `192.168.0.0/24`: N100 `.147`/`.162`, HP `.87` — **re-check
rather than trusting these** (`ip -br addr show vmbr0`; the N100 read `.162` on 2026-07-30). The
tailnet addresses are the stable ones — use those. Direct LAN literals are **not** reachable from
DooPlex while the boxes are away (`felhom-pve-lan` → `No route to host`, 2026-07-30).
**Which box is safe to break, and what may be done to each:
[`../runbooks/target-selection.md`](../runbooks/target-selection.md).** This page is *what the hardware
is*; that page is *what you may do to it*.
## demo-hp — the HP t740, in detail
### Disks
| device | model | serial | role |
|---|---|---|---|
| `sda` | SanDisk X600 M.2 2280 SATA 128GB | `182195804614` | **system disk** — PVE, LVM (`pve-root` 39.6G, `pve-data` thin pool, guest 9201's three volumes) |
| `nvme0n1` | KXG50PNV1T02 NVMe TOSHIBA 1024GB | `58BS11AFT8MQ` | **ENROLLED** — mounted `/mnt/nvme-1tb`, the enrolled user-data drive **and** the `felhom-backup` whole-guest backup target |
> **The NVMe joined the product on 2026-07-22, through the normal Tárhely flow, as intended.** Enrolled
> to guest 9201 (PUBLISH TRAIN, `pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md`); the vzdump target was
> moved onto it by E-2a. Verified live 2026-07-30: `nvme0n1` → `/mnt/nvme-1tb`, and
> `dir: felhom-backup / path /mnt/nvme-1tb / is_mountpoint 1` in `storage.cfg`.
>
> **The "PRESENT AND UNENROLLED — do not touch" fence that stood here is RETRACTED**, and its reason is
> recorded so it is not mistaken for a live rule: the NVMe was outside everything because the install
> ISO's exact-serial filter pinned `sda` only, and it was to join *through the Tárhely flow rather than
> the installer or by hand*. **That condition was satisfied; the prohibition expired with it.** It stood
> for eight days after enrolment and contradicted the task specs that (correctly) sent drill-VM disks to
> `/mnt/nvme-1tb`.
>
> Still true, and now the operative caution: a dir storage there **must sit at the mountpoint root** (a
> subdirectory fails the agent's `exactMount` check → the storage reads `disconnected` forever), and it
> shares the device with the box's own backups — so remove scratch storages when done.
### NIC map — and the trap
This board has **five wired interfaces**, and the obvious one is the wrong one:
| interface | MAC | driver | what it is | state at this site |
|---|---|---|---|---|
| `enp1s0f0` | `a0:36:9f:5d:07:20` | `igb` | 4-port expansion card | no carrier, **no DHCP** |
| `enp1s0f1` | `a0:36:9f:5d:07:21` | `igb` | ″ | no carrier |
| `enp1s0f2` | `a0:36:9f:5d:07:22` | `igb` | ″ | no carrier |
| `enp1s0f3` | `a0:36:9f:5d:07:23` | `igb` | ″ | no carrier |
| **`enp2s0f0`** | `7c:d3:0a:77:d9:76` | `r8169` | **onboard port — the one that works** | carrier up, 1000 Mb, **this is `vmbr0`'s port** |
| `wlo1` | `24:ee:9a:e5:05:b0` | `iwlwifi` | wifi | unused |
**This trap cost the first install.** The 4-port card got no lease, and instead of aborting the
installer baked its `192.168.100.2` fallback as a **static** `vmbr0` address and completed — a box
that looked installed and could never call home. Repaired on the console by bridging `vmbr0` to
`enp2s0f0`. Filed as **R-59** (must hard-abort) and **R-60** (first-boot NIC sweep self-heal).
Current, read off the box 2026-09-21: `vmbr0` **static `192.168.0.104/24`**, gw `192.168.0.1`, bridge-port **`nic0`**. (It was `192.168.0.87/24` on `enp2s0f0` before a reprovision; both were stale here for long enough to cost a session a false "the box is down" — R-601.) **The hub always knows the truth:** every host report carries `addresses: [{iface, cidr}, …]`, so read it there rather than from this page.
No trace of `192.168.100.2` remains. `wg-felhom` `10.77.0.3/32` up to the hub. Guest **9201
`demo-hp`** running. Agent config shape (R-50 island): `local_api` on `169.254.253.1:8443`/`vmbr9`,
guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to the host's LAN address.
### Designated drill + build VM host (operator ruling, 2026-07-25)
**Ruling:** drill and build VMs are hosted on the **t740 from now on** — NOT on felhom-pve, and moving
them off DooPlex (the production k3s node). **This is a VM-HOSTING ruling only; the build-PIPELINE
relocation to the t740 is NOT ruled or implemented here.**
**Current state (updated 2026-07-25 PM):** the ruling is **realized** — the t740 now hosts the first
drill appliance. **QEMU VM `300` = `drill-r50`**, a nested PVE-in-a-VM (8 GiB RAM, 4 vCPU cpu=host, 32 GiB
local-lvm disk, OVMF/SB-off, one NIC on `vmbr0` DHCP). Installed from
`felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` through the **real day-0** (self-register →
operator bind to scratch customer `drill-r50` → deliver → nested guest **9201** provisioned and healthy
on the then-current controller + agent). Its own break-glass root is vaulted in the hub
`host_recovery/drill-r50-0a4f9a`; reach it as `root@192.168.0.176` **through `demo-hp`** (it has no key
and no tailnet — it is a peer on demo-hp's LAN). It has snapshot **`r50pre`** (clean LAN-literal day-0).
This VM was provisioned to unblock and run the **R-50 island-bridge empirical spike**
(`audits/SPIKE-island-bridge-2026-07-25.md`) — **probes P1–P8 PASSED 2026-07-25, verdict GO**; the drill
was left in its working island configuration. It is a throwaway: destroy with `qm stop 300 && qm destroy
300 --purge 1` (and delete the `drill-r50` customer + appliance #9 hub-side) when no longer needed. The
historical golden-bake `drill.qcow2` still lives on **DooPlex** (`/mnt/5_hdd/felhom.eu/drill/`, ~18G,
powered off) and is unrelated. The build-PIPELINE relocation to the t740 remains unbuilt.
### Access — there is no baked SSH key
`ssh demo-hp` resolves to the tailnet address, but **no operator public key is on this box** — the
HP profile deliberately left `FELHOM_ROOT_SSH_KEY` blank. Authentication is the **G1 break-glass
root password vaulted in the hub**, `host_recovery` row `demo-hp-bb76ea` (set at day-0,
2026-07-21 16:24 UTC).
Retrieval (operator-side, and **shred the copy** — that DB holds every host's secret):
> **WAL-AWARE SINCE HUB v0.88.0 — copying `hub.db` ALONE is no longer safe.** The hub runs SQLite in
> **WAL** mode (R-172), so a committed transaction may still live in `hub.db-wal` and not yet be in
> the main file. A bare `cat /data/hub.db` therefore yields a copy that is **valid but stale** — it
> opens cleanly and silently lacks the most recent writes, which is the worst failure shape for a
> credential lookup. Copy the `-wal` beside it and let SQLite replay it on open.
```bash
sudo kubectl -n felhom-system exec <hub-pod> -- cat /data/hub.db > /tmp/x.db
sudo kubectl -n felhom-system exec <hub-pod> -- cat /data/hub.db-wal > /tmp/x.db-wal 2>/dev/null || true
python3 -c "import sqlite3;print(sqlite3.connect('/tmp/x.db').execute(
\"SELECT secret FROM host_recovery WHERE host_id='demo-hp-bb76ea'\").fetchone()[0])"
shred -u /tmp/x.db /tmp/x.db-wal
```
The `|| true` is deliberate: an absent `-wal` is legitimate (a freshly checkpointed database), and
must not fail the retrieval. **Shred both files** — the WAL holds the same secrets as the DB.
Then `sshpass -e ssh root@demo-hp` (sshpass is on DooPlex, not on the nodes).
**This is the lockout filed as R-61**: the ISO mints a throwaway root password per build and discards
the plaintext, so the console is unreachable without a working hub and network — precisely what you
may be trying to fix. Slice 1 is to emit the baked password into the build report.
`demo-hp-lan` (**`192.168.0.104`** via `ProxyJump felhom-pve`) is the fallback when DooPlex is not on the home LAN. **`ssh demo-hp` now goes DIRECT to `192.168.0.104`** — DooPlex is on the same LAN, and the tailnet route for this box does not exist (see the table above). Both were verified 2026-09-21.
## OOB belt (H1) — both boxes, since 2026-07-23 (ISO train v1.25.0)
The dedicated OOB sshd belt (TASK H1: `felhom-sshd` + the static `inet felhom_oob` table + `felhom-op`)
is installed and **active on BOTH fleet boxes** — the F9 gap (belt on neither) is closed. From
v1.25.0 host-install installs it by default on every appliance install (`--no-oob` opts out; byo still
refuses).
- **Claimed port: `8822` on both** (first-free from `[8822,2222,8022,62222]`; persisted per box).
- **Reachability: the wg-felhom offsite tunnel ONLY** — the belt admits the operator `/32`
(`10.77.0.250`) over `wg-felhom` to 8822 and drops everything else; `:22` and every other interface
are untouched. **tailscale does NOT reach the belt** (wrong fabric, dropped by design).
- **Operator login** (from the machine holding the wg-felhom operator tunnel + the registered
`oob_operator_ssh_pubkey`): `ssh -p 8822 felhom-op@10.77.0.2` (felhom-pve) / `@10.77.0.3` (demo-hp).
**PROVEN live 2026-07-23** on felhom-pve (`felhom-op@demo-felhom`).
- **Operator tunnel**: the Mac/Windows operator peer dials `ep0.felhom.eu:443` (WireGuard), address
`10.77.0.250/32`, AllowedIPs `10.77.0.0/24`, server pubkey `f3d1ZI7…`. ep0's `forward` chain
(persisted in its `/etc/nftables.conf`) allows `10.77.0.250 → 10.77.0.2/.3`. If a work-network blocks
UDP/443, the RheinMetall-style firewalls pass UDP/51820 — a home/hotspot network works on 443.
- Register/rotate the operator identity hub-side: `PUT /api/v1/admin/wg/operator-peer` (global key)
with `{pubkey, assigned_ip:"10.77.0.250", ssh_pubkey}`; wgsync pushes it to ep0 and the SSH key flows
to both boxes' `felhom-op` authorized_keys within a tick.
## felhom-pve (the N100) — vault parity + access
felhom-pve has **operator SSH-key access** (over tailscale `100.70.170.35`) AND, since 2026-07-23,
**G1 break-glass vault parity with demo-hp**: its root@pam password is freshly rotated and vaulted in
the hub `host_recovery` row **`demo-felhom-8363b5`** (same PUT `…/recovery-credential` mechanism day-0
uses; verified retrievable + authenticating over `:22`). Retrieval + shred-the-copy recipe is identical
to demo-hp's below (swap the host_id). So a lost N100 key is recoverable the same way as the key-less HP.
## Tailscale on demo-hp is an OPERATOR-LAB EXCEPTION
> **Read this before any product-shape audit.** `demo-hp` is **customer-shaped** — it is a normal
> appliance install with a real customer record (`demo-hp`), a real guest, and a real day-0.
> **Tailscale is not part of that shape.** It was installed by hand on 2026-07-21 purely so the
> operator can reach a lab box that lives on someone else's LAN.
>
> **Real customer boxes never get tailscale.** Their operator access is the WireGuard tunnel plus
> the H1 OOB path, and nothing else. If a future audit finds tailscale on `demo-hp` and concludes
> the product ships it — that conclusion is wrong, and this paragraph is the reason. The same
> exception already applies to `felhom-pve`.
Details, and the two hard rules that apply to any host node, in `operations/tailscale.md`.
### demo-hp scratch guest — LXC 9202 `demo-hp-scratch` (R-481, built 2026-09-13)
- **Disposition:** the nightly rotation's throwaway host. Under the **demo-hp customer** (same domain,
same dashboard password), **hub reporting OFF, Cloudflare tunnel OFF, agent local API OFF, off-site
OFF, self-update OFF** (the image is set by hand in `/etc/felhom-controller-image`). **Persists across
nights on purpose.** Apps on it are throwaways; nothing on it is a customer promise. Not in the felhom
pool, so the hub never sees it. Tags `scratch,r481`; the same text sits in the guest at
`/etc/felhom-scratch-disposition`, in the host's `pct` description and in its bootstrap file.
- **Reach:** `https://192.168.0.114` with the demo-hp Host names (`felhom.enkisfelhom.hu`,
`<sub>.enkisfelhom.hu`) — LAN only. Helper: the scratch twin of `ctl.sh`.
- **Shape:** 7 cores, 25 898 MB cap, rootfs 32 G + data 70 G, both on the `nvme-scratch` dir storage
at `/mnt/hdd_1` (the runbook's NVMe path; a `dir` storage was re-added there). Unprivileged.
"Second drive" for Tier 2: host `/mnt/hdd_1/scratch-drives/scratch_hdd` mounted at
`/mnt/felhom-drives/scratch_hdd` (mp8), registered as the default storage path.
- **Rebuild from nothing:** restore the golden (`local:backup/felhom-golden-<ver>.tar.zst`) with
`pct restore … --storage nvme-scratch`, add mp0/mp8/mp9 as above, fix ownership from the host with
the unprivileged mapping (100000), seed `controller.yaml` (hub/tunnel/agent/off-site off) and a
claimed `settings.json`, then start `felhom-controller-bootstrap.service`.
## What is NOT enrolled here (deliberately)
- **No PBS datastore, no offsite target** on demo-hp; the DR tier is the N100's.
- **No second customer guest** beyond 9201.
*(The 1TB NVMe was on this list until 2026-07-22. It is enrolled — see the disks table above.)*