Files
felhom.eu/REPORT.md
T

8.1 KiB
Raw Blame History

REPORT — TASK-H: HP node access + docs + today's findings (2026-07-21, evening)

Baseline 59226ed, clean tree. No product code touched; the N100 and guest 9201 were not modified.

Part Status
1 — tailscale on demo-hp DONE, one item open (key expiry — needs an admin-console toggle)
2 — inventory + docs DONEdocumentation/operations/nodes.md
3 — findings filed DONER-59, R-60, R-61 + positive evidence recorded

Part 1 — tailscale on demo-hp

tailscale 1.98.9 (Debian trixie apt repo, host package — same shape as felhom-pve).

demo-hp   100.76.96.79   online   direct path 37.191.56.193:45127

Verified from DooPlex over the tailnet, no jumphost: ping ~40 ms, ssh root@100.76.96.79 works, and the path is direct, not a DERP relay. ~/.ssh/config gains demo-hp (tailnet) and demo-hp-lan (192.168.0.87 via ProxyJump felhom-pve) as the away-fallback.

The pre-auth key was passed as --authkey=file:<path> and shredded immediately, so it never entered the box's process list.

OPEN — key expiry is NOT disabled

demo-hp expires 2027-01-17; felhom-pve and dooplex both have expiry disabled, so this is the one place it diverges from the fleet convention. I could not do it from here: disabling per-device key expiry is an admin-console toggle or an API call, and the supplied tskey-auth-… pre-auth key cannot drive the API (verified — GET /api/v2/tailnet/-/devices returns 401).

One click when convenient: Tailscale admin → Machines → demo-hpDisable key expiry. Or hand me a tskey-api-… token and I will do it. Verify with: ssh demo-hp 'tailscale status --json' | python3 -c "import sys,json;print(json.load(sys.stdin)['Self'].get('KeyExpiry'))" → should print None.

A rule I broke, then fixed

The join omitted --accept-dns=false, and MagicDNS immediately rewrote /etc/resolv.conf to nameserver 100.100.100.100 — exactly what operations/tailscale.md forbids for a host node.

Nothing broke at the vacation site: there is no pi-hole there, and both gitea.dooplex.hu and hub.felhom.eu resolve publicly anyway. The damage would have surfaced silently when the box comes home, where split-horizon is what makes gitea.dooplex.hu resolve to 192.168.0.180 locally — a class of failure that looks like "the network is slow" rather than "DNS is wrong". tailscale set --accept-dns=false restored nameserver 192.168.0.1; tailnet and agent unaffected. The doc now carries the evidence and the instruction to pass the flag at join time.

Part 2 — the node inventory

documentation/operations/nodes.md (new). The fleet is now two hosts, both agent 0.92.1: demo-felhom-8363b5 (N100) and demo-hp-bb76ea (HP t740). Both are at the vacation site and travel home ~2026-08-02.

demo-hp: HP t740 Thin Client, s/n 8CN944035T, AMI M42 v01.10 (11/11/2020), Ryzen Embedded V1756B (8 threads), 30 GiB RAM, PVE 9.2.2 as node felhom-host, guest 9201 demo-hp running, WireGuard 10.77.0.3/32 up.

device serial role
sda SanDisk X600 128GB 182195804614 system disk — PVE + LVM + guest volumes
nvme0n1 Toshiba KXG50 1024GB 58BS11AFT8MQ PRESENT AND UNENROLLED — do not touch

The NVMe still holds its previous NTFS partition, is unmounted, and is in no LVM PV and no ZFS pool — the exact-serial filter did its job on hardware it had never seen. It joins later through the Tárhely flow, never the installer.

NIC map (documented in both nodes.md and scripts/iso/README.md): enp1s0f0f3 are the 4-port igb card with no carrier and no DHCP at this site; enp2s0f0 (r8169, MAC 7c:d3:0a:77:d9:76) is the onboard port that works and is now vmbr0's bridge-port. MACs for all five are in the doc.

Loader/firmware finding: the box installed with the shim loader and Secure Boot ENABLED (mokutil --sb-stateSecureBoot enabled). That retires an assumption — SB-off was an N100-firmware workaround, not a Felhom requirement.

Access, and why it is awkward: no operator SSH key is baked (the HP profile deliberately left FELHOM_ROOT_SSH_KEY blank), so authentication is the G1 break-glass root password vaulted in the hub (host_recovery row demo-hp-bb76ea, set 2026-07-21 16:24 UTC). I retrieved it by streaming /data/hub.db out of the hub pod, extracting one field to a 0600 file, and shredding the copy immediately — that DB holds every host's secret. Recipe is in the doc. This lockout is R-61.

The operator-lab exception is documented prominently, in both nodes.md and tailscale.md: demo-hp is customer-shaped but tailscale is not part of that shape. Real customer boxes get the WireGuard tunnel and the H1 OOB path and nothing else. A future product-shape audit finding tailscale here must not conclude the product ships it.

Part 3 — findings filed

TASK-G Part 3 verified as already filed — R-58 (assisted disk-picker) exists and includes the abort-screen candidate-table slice. Nothing to complete.

New rows:

  • R-59 [P1] — a no-DHCP install must HARD-ABORT. Instead it baked 192.168.100.2 as a static vmbr0 address and completed: the install "succeeded", the box looked finished, and it could never call home. The worst silent onboarding failure shape there is. The philosophy already exists one layer over — the disk filter refuses loudly and touches nothing when it cannot identify its target; networking should do the same, with the candidate NIC table on screen (same grammar as R-58 slice 1).
  • R-60 [P2] — first-boot NIC sweep self-heal. Hub unreachable ⇒ DHCP across every carrier-bearing NIC before settling. Today's repair was a human moving one cable; a sweep would have healed it unaided. Scoped to first boot and the hub-unreachable condition only — a running box must never re-shuffle its own networking.
  • R-61 [P1] — the baked root password must be knowable. The ISO mints a throwaway hash per build and discards the plaintext, so nobody can reach the console of a box they just installed. Slice 1: emit it into the build REPORT + operator cheat-sheet alongside the sha256. A fixed well-known password is explicitly rejected (operator ruling) — a pre-pairing box sits on a stranger's LAN.

Positive evidence, same session:

  • R-21 slice C is no longer a one-board result. The row and the capability-map ISO row now read PROVEN-LIVE on TWO different boards (N100 2026-07-18; HP t740 2026-07-21) — virgin hardware, one pass, self-registration → operator bind → day-0 → running guest + agent check-in. The Secure Boot finding is recorded there too.
  • Fresh-box floor lift: the new box came up on golden 0.153.0 and self-updated to the fleet floor 0.156.0 during day-0, unattended. Cited on the publish-train capability-map row — the floor mechanism works on first contact, not only on boxes with history.
  • The t740 five-NIC trap is recorded as a board gotcha in scripts/iso/README.md.

Observations

  1. The break-glass path works, and it is also the argument for R-61. Reaching a box whose root password was never known required a working hub, a working network, and operator tooling — at exactly the moment the reason you want the console is usually that one of those is broken.
  2. apt update fails on this box against the PVE enterprise repos (401, no subscription) — pre-existing from the install, unrelated to today. I worked around it with a list-scoped apt-get update rather than editing the box's repo config. Worth deciding whether host-install should switch fresh boxes to the no-subscription repo; left alone deliberately.
  3. A heredoc silently ate a piped secret. printf … | ssh host 'bash -s' <<'EOF' sends the heredoc as stdin, so the piped key never arrives and read consumes script text instead. The working shape is the command as an argument, with stdin free for the secret — then stage it as a file and use --authkey=file:. Worth remembering next time a credential has to cross an SSH hop.