Files
felhom.eu/REPORT.md
T

133 lines
8.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — TASK-H: HP node access + docs + today's findings (2026-07-21, evening)
Baseline `59226ed`, clean tree. No product code touched; the N100 and guest 9201 were not modified.
| Part | Status |
|---|---|
| 1 — tailscale on demo-hp | **DONE**, one item open (key expiry — needs an admin-console toggle) |
| 2 — inventory + docs | **DONE**`documentation/operations/nodes.md` |
| 3 — findings filed | **DONE****R-59, R-60, R-61** + positive evidence recorded |
## Part 1 — tailscale on demo-hp
`tailscale 1.98.9` (Debian trixie apt repo, host package — same shape as `felhom-pve`).
```
demo-hp 100.76.96.79 online direct path 37.191.56.193:45127
```
Verified **from DooPlex over the tailnet, no jumphost**: ping ~40 ms, `ssh root@100.76.96.79` works,
and the path is **direct**, not a DERP relay. `~/.ssh/config` gains `demo-hp` (tailnet) and
`demo-hp-lan` (`192.168.0.87` via `ProxyJump felhom-pve`) as the away-fallback.
The pre-auth key was passed as `--authkey=file:<path>` and shredded immediately, so it never entered
the box's process list.
### OPEN — key expiry is NOT disabled
`demo-hp` expires **2027-01-17**; `felhom-pve` and `dooplex` both have expiry disabled, so this is
the one place it diverges from the fleet convention. **I could not do it from here:** disabling
per-device key expiry is an admin-console toggle or an API call, and the supplied `tskey-auth-…`
pre-auth key cannot drive the API (verified — `GET /api/v2/tailnet/-/devices` returns **401**).
**One click when convenient:** Tailscale admin → Machines → `demo-hp`*Disable key expiry*. Or
hand me a `tskey-api-…` token and I will do it. Verify with:
`ssh demo-hp 'tailscale status --json' | python3 -c "import sys,json;print(json.load(sys.stdin)['Self'].get('KeyExpiry'))"` → should print `None`.
### A rule I broke, then fixed
The join omitted `--accept-dns=false`, and MagicDNS immediately rewrote `/etc/resolv.conf` to
`nameserver 100.100.100.100` — exactly what `operations/tailscale.md` forbids for a host node.
Nothing broke *at the vacation site*: there is no pi-hole there, and both `gitea.dooplex.hu` and
`hub.felhom.eu` resolve publicly anyway. **The damage would have surfaced silently when the box comes
home**, where split-horizon is what makes `gitea.dooplex.hu` resolve to `192.168.0.180` locally — a
class of failure that looks like "the network is slow" rather than "DNS is wrong".
`tailscale set --accept-dns=false` restored `nameserver 192.168.0.1`; tailnet and agent unaffected.
The doc now carries the evidence and the instruction to pass the flag **at join time**.
## Part 2 — the node inventory
**`documentation/operations/nodes.md`** (new). The fleet is now two hosts, both agent 0.92.1:
`demo-felhom-8363b5` (N100) and `demo-hp-bb76ea` (HP t740). Both are at the vacation site and travel
home ~2026-08-02.
**demo-hp:** HP t740 Thin Client, s/n `8CN944035T`, AMI M42 v01.10 (11/11/2020), Ryzen Embedded
V1756B (8 threads), **30 GiB RAM**, PVE 9.2.2 as node `felhom-host`, guest 9201 `demo-hp` running,
WireGuard `10.77.0.3/32` up.
| device | serial | role |
|---|---|---|
| `sda` SanDisk X600 128GB | `182195804614` | system disk — PVE + LVM + guest volumes |
| `nvme0n1` Toshiba KXG50 1024GB | `58BS11AFT8MQ` | **PRESENT AND UNENROLLED — do not touch** |
The NVMe still holds its previous **NTFS** partition, is unmounted, and is in no LVM PV and no ZFS
pool — the exact-serial filter did its job on hardware it had never seen. It joins later through the
**Tárhely flow, never the installer**.
**NIC map** (documented in both `nodes.md` and `scripts/iso/README.md`): `enp1s0f0f3` are the
4-port `igb` card with **no carrier and no DHCP** at this site; **`enp2s0f0`** (`r8169`, MAC
`7c:d3:0a:77:d9:76`) is the onboard port that works and is now `vmbr0`'s bridge-port. MACs for all
five are in the doc.
**Loader/firmware finding:** the box installed with the **shim** loader and **Secure Boot ENABLED**
(`mokutil --sb-state``SecureBoot enabled`). That retires an assumption — SB-off was an
N100-firmware workaround, not a Felhom requirement.
**Access, and why it is awkward:** no operator SSH key is baked (the HP profile deliberately left
`FELHOM_ROOT_SSH_KEY` blank), so authentication is the **G1 break-glass root password vaulted in the
hub** (`host_recovery` row `demo-hp-bb76ea`, set 2026-07-21 16:24 UTC). I retrieved it by streaming
`/data/hub.db` out of the hub pod, extracting one field to a 0600 file, and **shredding the copy
immediately** — that DB holds every host's secret. Recipe is in the doc. This lockout is R-61.
**The operator-lab exception is documented prominently**, in both `nodes.md` and `tailscale.md`:
demo-hp is *customer-shaped* but tailscale is **not** part of that shape. Real customer boxes get the
WireGuard tunnel and the H1 OOB path and nothing else. A future product-shape audit finding tailscale
here must not conclude the product ships it.
## Part 3 — findings filed
**TASK-G Part 3 verified as already filed** — R-58 (assisted disk-picker) exists and includes the
abort-screen candidate-table slice. Nothing to complete.
New rows:
- **R-59 [P1] — a no-DHCP install must HARD-ABORT.** Instead it baked `192.168.100.2` as a *static*
`vmbr0` address and completed: the install "succeeded", the box looked finished, and it could never
call home. The worst silent onboarding failure shape there is. The philosophy already exists one
layer over — the disk filter refuses loudly and touches nothing when it cannot identify its target;
networking should do the same, with the candidate NIC table on screen (same grammar as R-58 slice 1).
- **R-60 [P2] — first-boot NIC sweep self-heal.** Hub unreachable ⇒ DHCP across every carrier-bearing
NIC before settling. Today's repair was a human moving one cable; a sweep would have healed it
unaided. Scoped to first boot and the hub-unreachable condition only — a running box must never
re-shuffle its own networking.
- **R-61 [P1] — the baked root password must be knowable.** The ISO mints a throwaway hash per build
and discards the plaintext, so nobody can reach the console of a box they just installed. Slice 1:
emit it into the build REPORT + operator cheat-sheet alongside the sha256. **A fixed well-known
password is explicitly rejected** (operator ruling) — a pre-pairing box sits on a stranger's LAN.
Positive evidence, same session:
- **R-21 slice C is no longer a one-board result.** The row and the capability-map ISO row now read
**PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) — virgin hardware,
one pass, self-registration → operator bind → day-0 → running guest + agent check-in. The Secure
Boot finding is recorded there too.
- **Fresh-box floor lift**: the new box came up on golden **0.153.0** and self-updated to the fleet
floor **0.156.0** during day-0, unattended. Cited on the publish-train capability-map row — the
floor mechanism works on first contact, not only on boxes with history.
- **The t740 five-NIC trap** is recorded as a board gotcha in `scripts/iso/README.md`.
## Observations
1. **The break-glass path works, and it is also the argument for R-61.** Reaching a box whose root
password was never known required a working hub, a working network, and operator tooling — at
exactly the moment the reason you want the console is usually that one of those is broken.
2. **`apt update` fails on this box** against the PVE **enterprise** repos (401, no subscription) —
pre-existing from the install, unrelated to today. I worked around it with a list-scoped
`apt-get update` rather than editing the box's repo config. Worth deciding whether `host-install`
should switch fresh boxes to the no-subscription repo; left alone deliberately.
3. **A heredoc silently ate a piped secret.** `printf … | ssh host 'bash -s' <<'EOF'` sends the
*heredoc* as stdin, so the piped key never arrives and `read` consumes script text instead. The
working shape is the command as an argument, with stdin free for the secret — then stage it as a
file and use `--authkey=file:`. Worth remembering next time a credential has to cross an SSH hop.