Files
felhom.eu/documentation/runbooks/target-selection.md
T
admin 376365bb12 docs(target-selection): a fixture may prove a mechanism; only a fresh box may prove a path
The page said which machine is safe to break but not when reusing a test box is
legitimate. That distinction is exactly what surfaced R-120: R-116's closing run
deliberately did a real day-0 from the ISO instead of reusing the standing fixture,
and the fresh box installed the golden's controller -- a release behind -- and showed
the customer the wrong absent-target message. A fixture would have shown a controller
nobody installs.

Adds to the Tier 1 section: a reusable snapshot-reset fixture is the right default
for MECHANISM work (payload capture, fix cycles, claims about code behaviour), while
a fresh day-0 from the ISO is REQUIRED for any claim about the install path, the
golden image, agent publish/vouch or first-boot state -- naming the drift family it
exists to catch (R-111, R-115, R-120).

Also: a fixture must record its provenance (which golden, agent and controller, and
when), because a fixture whose versions drift silently is R-120's mechanism turned
into a permanent installation -- worse than no fixture, since it produces confident
wrong results quickly.

Part 1 of the R-120 task, committed alone and before the bake. Docs only.
2026-07-30 10:05:52 +02:00

8.2 KiB

Which box do I break? — target selection by blast radius

Read before picking a machine for a drill, a destructive test, or a throwaway VM. It answers one question: what is safe to lose. Reachability is in CLAUDE.md; hardware in operations/nodes.md. Added 2026-07-30.

The rule

Start at Tier 0 and work down only if Tier 0 genuinely cannot host the work. A Tier 2 machine is used only when a task says so explicitly — never by inference from what was not forbidden. An absent fence is not permission. If no tier fits, stop and ask.

Fences name acts, not machines. "Do not re-target demo-hp's backup target" and "do not use demo-hp to host a throwaway VM" are unrelated; only the first has ever been meant. Read a per-machine prohibition as covering the act it names and nothing more.

Tier Meaning Machines
0 — disposable. Reach here first. Exists to be broken; reinstalling is a routine afternoon, not an incident. A drill that needs a victim uses one of these. demo-hp (t740), demo-felhom (N100)
1 — create and destroy freely Throwaway VMs, guests, scratch customers — hosted on a Tier 0 machine drill VMs, scratch guests
2 — precious. Never a drill target. Losing it costs the recovery chain or a real relationship DooPlex, Peti's cluster, ep0 + the Hetzner Storage Boxes

DooPlex is Tier 2 because it is the recovery chain — hub, Gitea, registry, PBS, k3s + Longhorn. Everything else rebuilds from it; it rebuilds from nothing. A bad moment in a DR drill there costs the thing under test, the source of truth for it, and the backups, at once.

Standing ruling, 2026-07-25 (operations/nodes.md): drill and build VMs live on the t740 — not felhom-pve, and moved off DooPlex. This page exists because that ruling sat where no session reads.

Per machine — permitted / needs care / forbidden

demo-hp — HP t740 · Tier 0 · the designated drill + build VM host

  • Freely: host nested drill VMs; create/destroy guests and scratch customers; reinstall the box. This is the default answer to "where do I run this".
  • Care: local-lvm is a thin pool, over-subscribed (~144 GiB allocated over ~54 GiB) backing live guest 9201 — filling it corrupts every guest. Put VM disks on a dir storage at /mnt/nvme-1tb, at its root (a subdirectory fails the agent's exactMount check → storage reads disconnected forever). That warning is about one storage, not the box. /mnt/nvme-1tb is also the felhom-backup target and the enrolled user-data drive, so remove scratch storages when done.
  • Forbidden: do not re-target its backup configuration (below). Do not destroy or unblock drill-r50 (VM 300) — the only drift fixture (R-93). No baked SSH key; access is the hub-vaulted G1 break-glass password (R-61).

demo-felhom — N100 · Tier 0

  • Freely: create/destroy guests and scratch customers; reinstall the box.
  • Care: it carries the PBS-DR / offsite tier (demo-hp has none), so it is the Tier 0 box whose backup chain a test can actually disturb. Prefer demo-hp, per the 2026-07-25 ruling.
  • Forbidden: do not re-target its backup configuration.

Both Tier 0 boxes, one shared prohibition and its reason: do not re-point either backup target. They are the only two correctly configured boxes, so they are the regression path new installer logic is measured against — losing that reference costs more than the box does.

DooPlex — 192.168.0.180 · Tier 2

  • Freely: build, push images, sudo kubectl, read anything. Work inside /mnt/5_hdd/felhom.eu/.
  • Care: the golden-bake nested VM drill/drill.qcow2 lives here and is an accepted exception for bakes (last used 0.185.1, 2026-07-29). Using it as a drill victim is not covered by that exception and contradicts the 2026-07-25 ruling. If touched, restore qemu-img snapshot -a virgin.
  • Forbidden: no global Docker cleanup; do not touch k3s data dirs, Longhorn mounts, PBS datastores or Gitea storage; no destructive disk/guest ops; never run CC here with prompts disabled; abort large builds if df -h /mnt/5_hdd / shows either over 90 %. Because it is the recovery chain and a live k3s node.

Peti's cluster (peti-felhom) — Tier 2

Do not touch, at all. A real external pilot with a real person behind it; its whole-guest backup still shares a device with its guest, so a drive failure is offsite-only recovery. Deliberately not migrated, parked until the tester reinstalls (PETI in backlog/OPEN-ITEMS.md). Currently DOWN, no enrolled host. No access route from DooPlex, and nothing here needs one.

ep0 (felhom-hetzner, ep0.felhom.eu) + the Hetzner Storage Boxes — Tier 2

Reads are fine. It is the offsite of last resort (PBS-DR datastore, WireGuard hub, operator OOB path) and RAM-constrained (3.8 GB, R-90) so a large restore can OOM it. Do not delete datastores, prune jobs, tunnel config or nftables rules; never a drill target. The Storage Boxes hold the restic copy — customer documents and photos, on a credential that can still delete (R-95). Access: ssh root@167.233.158.164 from DooPlexnot felhom-pve → 10.77.0.1, the route that produced a false "unreachable" verdict (standing rule 2).

Tier 1 — drill VMs, scratch guests, scratch customers

Create and destroy freely on a Tier 0 host. Two exceptions: drill-r50 is a fixture, not scratch; and scratch customers outlive their VMs in the hub — delete those too, or they accumulate (sess-c and sess-d were both left behind before anyone noticed).

A fixture may prove a mechanism. Only a fresh box may prove a path.

Rebuilding from scratch every time is waste; reusing a box is legitimate — but not for every claim.

  • Reusable fixture — a snapshot-reset VM on a Tier 0 host. Use it for mechanism work: payload capture, fix cycles, anything whose claim is about code behaviour. Reset to virgin between runs. Fast, repeatable, and the right default for iteration.
  • Fresh day-0 from the ISO — required for any claim about the install path, the golden image, agent publish/vouch, or first-boot state. Slow, and the only thing that catches the drift family: R-111 (the golden's agent 17 releases behind), R-115 (an agent built and deployed but never published), R-120 (the golden a controller release behind).

A fixture must record its provenance — which golden, agent and controller it was built from, and when — alongside the VM. A fixture whose versions drift silently is R-120's mechanism turned into a permanent installation, and it is worse than no fixture, because it produces confident wrong results quickly.

Worked example. R-116's closing run deliberately did a real day-0 from the v1.25.0 ISO on demo-hp instead of reusing the standing fixture. That is how R-120 surfaced — the fresh box installed the golden's controller, which is a release behind, and showed the customer the wrong absent-target message. The fixture would have shown a controller nobody installs.

Not established — unknown, not guessed

  • Hetzner Storage Boxes (storage-box-pool-1, PBS-storage-1 / u629193 / box 611421) — state taken from backlog/OPEN-ITEMS.md; no direct access attempted or confirmed. Tier 2 by what they hold.
  • felhotest (legacy, ssh -p 33022 kisfenyo@router.abonet.hu) — Connection refused 2026-07-30; that was the only route tried. Untiered; assume nothing.
  • Peti's cluster hardware/storage layout — unverified. Tier 2 rests on the relationship, which needs no verification.

The gap this page closes

During the R-116 diagnosis session (2026-07-30) a nested-Proxmox drill ran on DooPlex, a Tier 2 machine. Not a newly provisioned VM — the standing golden-bake fixture, booted from virgin and restored to it, so nothing was lost. But it was chosen because DooPlex was the only machine no spec had fenced, while the ruling naming the t740 sat in operations/nodes.md with nothing pointing at it. The designation existed and was unreachable — the same class as an absent signal read as a positive one.