R-601 said demo-hp was unreachable. The operator looked at the hub and said it was online. It was, and had been up four and a half weeks, reporting every few minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan` at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub had carried the right address in every host report, and `ip neigh` on felhom-pve had .104 four lines above the .87 I quoted — I searched that output for the address I expected instead of reading it for the address that was there. Both ssh entries repointed and verified; nodes.md corrected, including that the tailnet route for this box does not exist. The hunt then found R-604, which is the real defect: demo-hp carried a per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. `managed floor SERVED` fires once per change by design, so a box behind a static override is silent for ever and its silence is indistinguishable from a box that already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes and its claim page now answers "Wrong or expired code" in English. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
15 KiB
Fleet node inventory — the physical demo/lab hosts
Added 2026-07-21, when the fleet stopped being one box. Two Proxmox hosts now check in to the hub. This is the operator-facing inventory: what the hardware is, how to reach it, and what is deliberately NOT enrolled on it.
The fleet
demo-felhom-8363b5 |
demo-hp-bb76ea |
|
|---|---|---|
| Hardware | N100 mini-PC | HP t740 Thin Client (s/n 8CN944035T) |
| CPU / RAM | Intel N100 | AMD Ryzen Embedded V1756B, 8 threads / 30 GiB |
| Firmware | AMI AN3PLUS-class | AMI M42 v01.10 (11/11/2020) |
| PVE node name | demo-felhom |
felhom-host |
| Customer | demo-felhom |
demo-hp |
| Agent / controller | not recorded here — see the note below | not recorded here |
| Control plane | island 169.254.253.1:8443 on vmbr9 (R-50, migrated 2026-07-25) |
island 169.254.253.1:8443 on vmbr9 (R-50, migrated 2026-07-25) |
| SSH alias | felhom-pve |
demo-hp |
| Tailnet | 100.70.170.35 |
NONE — tailscale is NOT INSTALLED on demo-hp (checked on the box 2026-09-21: no tailscaled, no tailscale binary). The peer 100.76.96.79 still listed by tailscale status is a STALE entry from an earlier build and can never answer; it showed offline, last seen 30d ago while the box was up and reporting. Use the LAN address. |
| Loader used to install | mkimage (unsigned, SB off — firmware workaround) |
shim, Secure Boot ENABLED |
No component versions are recorded on this page — deliberately. Agent, controller, hub and host-install versions change several times a day, so any number written here is wrong within hours and is then read as fact. Ask the fleet instead: the hub host list (
/hosts) and customer list (/configs) carry the live agent and controller versions per box;felhom-agent --versionon the node andpct exec <vmid> -- docker psin the guest are the authorities. Versions that must be pinned in writing belong in the per-repoCHANGELOG.mdand the hub's Day-0 artifact manifest — not in an inventory. The fleet is not uniform: on 2026-07-30 the two boxes ran different agent and different controller versions, so a single number for "the fleet" would have been wrong regardless.
Both are at the VACATION site and travel home with the rest of the kit ~2026-08-02. While
away, their LAN addresses are on that site's 192.168.0.0/24: N100 .147/.162, HP .87 — re-check
rather than trusting these (ip -br addr show vmbr0; the N100 read .162 on 2026-07-30). The
tailnet addresses are the stable ones — use those. Direct LAN literals are not reachable from
DooPlex while the boxes are away (felhom-pve-lan → No route to host, 2026-07-30).
Which box is safe to break, and what may be done to each:
../runbooks/target-selection.md. This page is what the hardware
is; that page is what you may do to it.
demo-hp — the HP t740, in detail
Disks
| device | model | serial | role |
|---|---|---|---|
sda |
SanDisk X600 M.2 2280 SATA 128GB | 182195804614 |
system disk — PVE, LVM (pve-root 39.6G, pve-data thin pool, guest 9201's three volumes) |
nvme0n1 |
KXG50PNV1T02 NVMe TOSHIBA 1024GB | 58BS11AFT8MQ |
ENROLLED — mounted /mnt/nvme-1tb, the enrolled user-data drive and the felhom-backup whole-guest backup target |
The NVMe joined the product on 2026-07-22, through the normal Tárhely flow, as intended. Enrolled to guest 9201 (PUBLISH TRAIN,
pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md); the vzdump target was moved onto it by E-2a. Verified live 2026-07-30:nvme0n1→/mnt/nvme-1tb, anddir: felhom-backup / path /mnt/nvme-1tb / is_mountpoint 1instorage.cfg.The "PRESENT AND UNENROLLED — do not touch" fence that stood here is RETRACTED, and its reason is recorded so it is not mistaken for a live rule: the NVMe was outside everything because the install ISO's exact-serial filter pinned
sdaonly, and it was to join through the Tárhely flow rather than the installer or by hand. That condition was satisfied; the prohibition expired with it. It stood for eight days after enrolment and contradicted the task specs that (correctly) sent drill-VM disks to/mnt/nvme-1tb.Still true, and now the operative caution: a dir storage there must sit at the mountpoint root (a subdirectory fails the agent's
exactMountcheck → the storage readsdisconnectedforever), and it shares the device with the box's own backups — so remove scratch storages when done.
NIC map — and the trap
This board has five wired interfaces, and the obvious one is the wrong one:
| interface | MAC | driver | what it is | state at this site |
|---|---|---|---|---|
enp1s0f0 |
a0:36:9f:5d:07:20 |
igb |
4-port expansion card | no carrier, no DHCP |
enp1s0f1 |
a0:36:9f:5d:07:21 |
igb |
″ | no carrier |
enp1s0f2 |
a0:36:9f:5d:07:22 |
igb |
″ | no carrier |
enp1s0f3 |
a0:36:9f:5d:07:23 |
igb |
″ | no carrier |
enp2s0f0 |
7c:d3:0a:77:d9:76 |
r8169 |
onboard port — the one that works | carrier up, 1000 Mb, this is vmbr0's port |
wlo1 |
24:ee:9a:e5:05:b0 |
iwlwifi |
wifi | unused |
This trap cost the first install. The 4-port card got no lease, and instead of aborting the
installer baked its 192.168.100.2 fallback as a static vmbr0 address and completed — a box
that looked installed and could never call home. Repaired on the console by bridging vmbr0 to
enp2s0f0. Filed as R-59 (must hard-abort) and R-60 (first-boot NIC sweep self-heal).
Current, read off the box 2026-09-21: vmbr0 static 192.168.0.104/24, gw 192.168.0.1, bridge-port nic0. (It was 192.168.0.87/24 on enp2s0f0 before a reprovision; both were stale here for long enough to cost a session a false "the box is down" — R-601.) The hub always knows the truth: every host report carries addresses: [{iface, cidr}, …], so read it there rather than from this page.
No trace of 192.168.100.2 remains. wg-felhom 10.77.0.3/32 up to the hub. Guest 9201
demo-hp running. Agent config shape (R-50 island): local_api on 169.254.253.1:8443/vmbr9,
guest eth1 169.254.253.2/30, lan_resolver.host_ip pinned to the host's LAN address.
Designated drill + build VM host (operator ruling, 2026-07-25)
Ruling: drill and build VMs are hosted on the t740 from now on — NOT on felhom-pve, and moving them off DooPlex (the production k3s node). This is a VM-HOSTING ruling only; the build-PIPELINE relocation to the t740 is NOT ruled or implemented here.
Current state (updated 2026-07-25 PM): the ruling is realized — the t740 now hosts the first
drill appliance. QEMU VM 300 = drill-r50, a nested PVE-in-a-VM (8 GiB RAM, 4 vCPU cpu=host, 32 GiB
local-lvm disk, OVMF/SB-off, one NIC on vmbr0 DHCP). Installed from
felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso through the real day-0 (self-register →
operator bind to scratch customer drill-r50 → deliver → nested guest 9201 provisioned and healthy
on the then-current controller + agent). Its own break-glass root is vaulted in the hub
host_recovery/drill-r50-0a4f9a; reach it as root@192.168.0.176 through demo-hp (it has no key
and no tailnet — it is a peer on demo-hp's LAN). It has snapshot r50pre (clean LAN-literal day-0).
This VM was provisioned to unblock and run the R-50 island-bridge empirical spike
(audits/SPIKE-island-bridge-2026-07-25.md) — probes P1–P8 PASSED 2026-07-25, verdict GO; the drill
was left in its working island configuration. It is a throwaway: destroy with qm stop 300 && qm destroy 300 --purge 1 (and delete the drill-r50 customer + appliance #9 hub-side) when no longer needed. The
historical golden-bake drill.qcow2 still lives on DooPlex (/mnt/5_hdd/felhom.eu/drill/, ~18G,
powered off) and is unrelated. The build-PIPELINE relocation to the t740 remains unbuilt.
Access — there is no baked SSH key
ssh demo-hp resolves to the tailnet address, but no operator public key is on this box — the
HP profile deliberately left FELHOM_ROOT_SSH_KEY blank. Authentication is the G1 break-glass
root password vaulted in the hub, host_recovery row demo-hp-bb76ea (set at day-0,
2026-07-21 16:24 UTC).
Retrieval (operator-side, and shred the copy — that DB holds every host's secret):
WAL-AWARE SINCE HUB v0.88.0 — copying
hub.dbALONE is no longer safe. The hub runs SQLite in WAL mode (R-172), so a committed transaction may still live inhub.db-waland not yet be in the main file. A barecat /data/hub.dbtherefore yields a copy that is valid but stale — it opens cleanly and silently lacks the most recent writes, which is the worst failure shape for a credential lookup. Copy the-walbeside it and let SQLite replay it on open.
sudo kubectl -n felhom-system exec <hub-pod> -- cat /data/hub.db > /tmp/x.db
sudo kubectl -n felhom-system exec <hub-pod> -- cat /data/hub.db-wal > /tmp/x.db-wal 2>/dev/null || true
python3 -c "import sqlite3;print(sqlite3.connect('/tmp/x.db').execute(
\"SELECT secret FROM host_recovery WHERE host_id='demo-hp-bb76ea'\").fetchone()[0])"
shred -u /tmp/x.db /tmp/x.db-wal
The || true is deliberate: an absent -wal is legitimate (a freshly checkpointed database), and
must not fail the retrieval. Shred both files — the WAL holds the same secrets as the DB.
Then sshpass -e ssh root@demo-hp (sshpass is on DooPlex, not on the nodes).
This is the lockout filed as R-61: the ISO mints a throwaway root password per build and discards the plaintext, so the console is unreachable without a working hub and network — precisely what you may be trying to fix. Slice 1 is to emit the baked password into the build report.
demo-hp-lan (192.168.0.104 via ProxyJump felhom-pve) is the fallback when DooPlex is not on the home LAN. ssh demo-hp now goes DIRECT to 192.168.0.104 — DooPlex is on the same LAN, and the tailnet route for this box does not exist (see the table above). Both were verified 2026-09-21.
OOB belt (H1) — both boxes, since 2026-07-23 (ISO train v1.25.0)
The dedicated OOB sshd belt (TASK H1: felhom-sshd + the static inet felhom_oob table + felhom-op)
is installed and active on BOTH fleet boxes — the F9 gap (belt on neither) is closed. From
v1.25.0 host-install installs it by default on every appliance install (--no-oob opts out; byo still
refuses).
- Claimed port:
8822on both (first-free from[8822,2222,8022,62222]; persisted per box). - Reachability: the wg-felhom offsite tunnel ONLY — the belt admits the operator
/32(10.77.0.250) overwg-felhomto 8822 and drops everything else;:22and every other interface are untouched. tailscale does NOT reach the belt (wrong fabric, dropped by design). - Operator login (from the machine holding the wg-felhom operator tunnel + the registered
oob_operator_ssh_pubkey):ssh -p 8822 felhom-op@10.77.0.2(felhom-pve) /@10.77.0.3(demo-hp). PROVEN live 2026-07-23 on felhom-pve (felhom-op@demo-felhom). - Operator tunnel: the Mac/Windows operator peer dials
ep0.felhom.eu:443(WireGuard), address10.77.0.250/32, AllowedIPs10.77.0.0/24, server pubkeyf3d1ZI7…. ep0'sforwardchain (persisted in its/etc/nftables.conf) allows10.77.0.250 → 10.77.0.2/.3. If a work-network blocks UDP/443, the RheinMetall-style firewalls pass UDP/51820 — a home/hotspot network works on 443. - Register/rotate the operator identity hub-side:
PUT /api/v1/admin/wg/operator-peer(global key) with{pubkey, assigned_ip:"10.77.0.250", ssh_pubkey}; wgsync pushes it to ep0 and the SSH key flows to both boxes'felhom-opauthorized_keys within a tick.
felhom-pve (the N100) — vault parity + access
felhom-pve has operator SSH-key access (over tailscale 100.70.170.35) AND, since 2026-07-23,
G1 break-glass vault parity with demo-hp: its root@pam password is freshly rotated and vaulted in
the hub host_recovery row demo-felhom-8363b5 (same PUT …/recovery-credential mechanism day-0
uses; verified retrievable + authenticating over :22). Retrieval + shred-the-copy recipe is identical
to demo-hp's below (swap the host_id). So a lost N100 key is recoverable the same way as the key-less HP.
Tailscale on demo-hp is an OPERATOR-LAB EXCEPTION
Read this before any product-shape audit.
demo-hpis customer-shaped — it is a normal appliance install with a real customer record (demo-hp), a real guest, and a real day-0. Tailscale is not part of that shape. It was installed by hand on 2026-07-21 purely so the operator can reach a lab box that lives on someone else's LAN.Real customer boxes never get tailscale. Their operator access is the WireGuard tunnel plus the H1 OOB path, and nothing else. If a future audit finds tailscale on
demo-hpand concludes the product ships it — that conclusion is wrong, and this paragraph is the reason. The same exception already applies tofelhom-pve.
Details, and the two hard rules that apply to any host node, in operations/tailscale.md.
demo-hp scratch guest — LXC 9202 demo-hp-scratch (R-481, built 2026-09-13)
- Disposition: the nightly rotation's throwaway host. Under the demo-hp customer (same domain,
same dashboard password), hub reporting OFF, Cloudflare tunnel OFF, agent local API OFF, off-site
OFF, self-update OFF (the image is set by hand in
/etc/felhom-controller-image). Persists across nights on purpose. Apps on it are throwaways; nothing on it is a customer promise. Not in the felhom pool, so the hub never sees it. Tagsscratch,r481; the same text sits in the guest at/etc/felhom-scratch-disposition, in the host'spctdescription and in its bootstrap file. - Reach:
https://192.168.0.114with the demo-hp Host names (felhom.enkisfelhom.hu,<sub>.enkisfelhom.hu) — LAN only. Helper: the scratch twin ofctl.sh. - Shape: 7 cores, 25 898 MB cap, rootfs 32 G + data 70 G, both on the
nvme-scratchdir storage at/mnt/hdd_1(the runbook's NVMe path; adirstorage was re-added there). Unprivileged. "Second drive" for Tier 2: host/mnt/hdd_1/scratch-drives/scratch_hddmounted at/mnt/felhom-drives/scratch_hdd(mp8), registered as the default storage path. - Rebuild from nothing: restore the golden (
local:backup/felhom-golden-<ver>.tar.zst) withpct restore … --storage nvme-scratch, add mp0/mp8/mp9 as above, fix ownership from the host with the unprivileged mapping (100000), seedcontroller.yaml(hub/tunnel/agent/off-site off) and a claimedsettings.json, then startfelhom-controller-bootstrap.service.
What is NOT enrolled here (deliberately)
- No PBS datastore, no offsite target on demo-hp; the DR tier is the N100's.
- No second customer guest beyond 9201.
(The 1TB NVMe was on this list until 2026-07-22. It is enrolled — see the disks table above.)