Files
felhom.eu/documentation/runbooks/target-selection.md
T

16 KiB

Which box do I break? — target selection by blast radius

Read before picking a machine for a drill, a destructive test, or a throwaway VM. It answers one question: what is safe to lose. Reachability is in CLAUDE.md; hardware in operations/nodes.md. Added 2026-07-30. Revised 2026-08-02 by operator decision D-d (CONTEXT.md S-5).

The rule

Two machines are protected: DooPlex and ep0. Everything else is disposable. Operator decision D-d, 2026-08-02, and the ep0 ruling of 2026-08-03. DooPlex because it holds Gitea, the hub, the backups and the registry — everything else rebuilds from it, and it rebuilds from nothing. ep0 for the reason below. Every other box, both demo boxes included, may be broken or reinstalled freely. (Peti's box was the second protected machine until it was RETIRED on 2026-09-25 — operator ruling; audits/RETIRE-peti-2026-09-25.md.)

This is a correction, not a relaxation. The earlier posture was costing whole sessions to caution and pushing drills onto DooPlex — the one machine that should never host them. If you are weighing whether a demo box can take a destructive test: it can. Reach for a Tier 0 box first and do not ask.

Still true: start at Tier 0 and work down only if Tier 0 genuinely cannot host the work, and a protected machine is used only when a task says so explicitly — never by inference from what was not forbidden. An absent fence is not permission. If no tier fits, stop and ask.

Fences name acts, not machines. "Do not destroy demo-hp's drill-r50 fixture" and "do not use demo-hp to host a throwaway VM" are unrelated; only the first has ever been meant. Read a per-machine prohibition as covering the act it names and nothing more.

A drill night that forbids baking a golden — what you will see, and why it is expected

This is written here because the instruction that caused it was mine and I left it out (R-417). An overnight drill runbook that says "no golden bake, no vouch, no floor change tonight — those are the operator's acts" is correct, and it also guarantees the golden-currency gate is RED for the whole night whenever a controller release is newer than the last bake. That is the gate telling the truth, not a fault to work around.

Since 2026-09-01 (R-404) it no longer refuses your pushes. A drill's own pushes — evidence, register rows, STATUS.md, REPORT.md — are documents-only, so the gate's conviction prints as a loud ADVISORY block and the push proceeds. Expect to see it, every push, all night. Do not silence it and do not bypass the hook for it.

Two things still hold:

  • A push that touches code is still refused — bake first or do not push code.
  • Every other gate still refuses every push. If a push is refused during a drill it is NOT the golden gate, and the message will name which gate it was.

If the drill genuinely must ship a release with no golden, the honest instrument is a waiver row in documentation/backlog/OPEN-ITEMS.md, never --no-verify.

Before you revert it — take the evidence off first

A phase's evidence is copied off the machine at the end of THAT phase, before any revert, snapshot restore or teardown. Not at the end of the session. (Standing rule 5, R-320.)

The intermediate revert is the one that gets forgotten. Both losses this project has recorded were the middle teardown, never the final one — the Phase A logs of the 2026-08-12 retained-key drill and the Part 1 logs of the 2026-08-13 R-316 run, both on drill-r50, both destroyed by a revert to virgin between phases, three days apart. Both times the conclusions survived only because the quotations had been read live and an independent reproduction happened to exist. That is luck.

The mechanism: make the pull the last act of the phase, not a step to remember later — pct pull / scp into documentation/audits/evidence-<topic>-<date>/ on DooPlex, then revert. Tier 0 machines are disposable, which is precisely why nothing you need may be left on one.

If you notice it is already gone: say so plainly in the report and reproduce the finding independently. That is the documented expectation, not an improvisation.

Tier Meaning Machines
0 — disposable. Reach here first. Exists to be broken; reinstalling is a routine afternoon, not an incident. A drill that needs a victim uses one of these. demo-hp (t740), demo-felhom (N100)
1 — create and destroy freely Throwaway VMs, guests, scratch customers — hosted on a Tier 0 machine drill VMs, scratch guests
2 — protected. Never a drill target. Losing it costs the recovery chain or a real relationship DooPlex, ep0 (operator ruling 2026-08-03) — and nothing else (Peti's cluster RETIRED 2026-09-25)

DooPlex is Tier 2 because it is the recovery chain — hub, Gitea, registry, PBS, k3s + Longhorn. Everything else rebuilds from it; it rebuilds from nothing. A bad moment in a DR drill there costs the thing under test, the source of truth for it, and the backups, at once.

ep0 is Tier 2 — PROTECTED. Operator ruling, 2026-08-03. D-d named two protected machines and did not name ep0 either way, so this page carried the question in writing for two days and read it the narrow way meanwhile (not protected, but not wipeable). The ruling settles it and extends D-d's protected list to three machines: DooPlex, Peti's cluster, ep0. Since 2026-09-25 (Peti's box retired) it is two: DooPlex and ep0.

The reason, as it stands on 2026-09-25: ep0 holds the PBS-DR datastore with the demo boxes' (and tester-1's) whole-guest copies, the WireGuard hub every box's off-site path runs through, and the operator's out-of-band path; the Hetzner Storage Boxes hold each box's restic copy. There is no paying customer yet, but ep0 is the only off-premises tier of the whole product, and a mistake there cannot be rebuilt from DooPlex. So deleting datastores, prune jobs, tunnel config or nftables rules was already forbidden by what it would destroy; the ruling makes the classification say so plainly instead of leaving each session to re-derive it. Reads are fine — including the ordinary off-site read a restore-test performs (R-86) — and it is never a drill target. The Hetzner Storage Boxes ride the same reasoning.

Standing ruling, 2026-07-25 (operations/nodes.md): drill and build VMs live on the t740 — not felhom-pve, and moved off DooPlex. This page exists because that ruling sat where no session reads.

Per machine — permitted / needs care / forbidden

demo-hp — HP t740 · Tier 0 · the designated drill + build VM host

  • Freely: host nested drill VMs; create/destroy guests and scratch customers; reinstall the box. This is the default answer to "where do I run this".
  • Care: local-lvm is a thin pool, over-subscribed (~144 GiB allocated over ~54 GiB) backing live guest 9201 — filling it corrupts every guest. Put VM disks on a dir storage at /mnt/hdd_1, at its root (R-461, measured 2026-09-13: there is NO /mnt/nvme-1tb on either box — demo-hp's 1 TB NVMe is nvme0n1 mounted at /mnt/hdd_1, the enrolled user-data drive and felhom-backup target, so "the NVMe" and "the data drive" are one disk) (a subdirectory fails the agent's exactMount check → storage reads disconnected forever). That warning is about one storage, not the box. /mnt/hdd_1 is also the felhom-backup target and the enrolled user-data drive, so remove scratch storages when done.
  • Forbidden: do not destroy or unblock drill-r50 (VM 300) — the only drift fixture (R-93). Measured 2026-09-13 (R-461): qm list is EMPTY on both demo-hp and demo-felhom — the VM does not exist anywhere, so the fence currently protects nothing and R-93's premise is gone. The fence stays as written for the day someone rebuilds it; do not read its presence here as evidence the fixture exists. (Access: the docs say no baked SSH key and G1 break-glass, but a key authenticated on 2026-07-31 — R-129, unresolved.)

demo-felhom — N100 · Tier 0

  • Freely: create/destroy guests and scratch customers; reinstall the box.
  • Care: it carries the PBS-DR / offsite tier — but so does demo-hp now, so this is no longer a reason to prefer one over the other. Prefer demo-hp per the 2026-07-25 ruling, on the ruling's own grounds (it is the designated drill host), not because its backup chain is safe to disturb.

Both Tier 0 boxes — the shared backup-target fence is DOWNGRADED to a cost, 2026-08-02 (D-d). It read "do not re-point either backup target", because these are the only two correctly configured boxes and therefore the regression path new installer logic is measured against. D-d makes both boxes freely breakable and reinstallable, which loses that reference just as thoroughly — so the fence was inconsistent with the decision and is not kept as a prohibition. What survives is the reason: re-pointing (or reinstalling) costs the reference configuration, so know that you are spending it and put the box back. If both are spent at once there is no correctly-configured box left to compare against.

DooPlex — 192.168.0.180 · Tier 2

  • Freely: build, push images, sudo kubectl, read anything. Work inside /mnt/5_hdd/felhom.eu/.
  • Care: the golden-bake nested VM drill/drill.qcow2 lives here and is an accepted exception for bakes (last used 0.185.1, 2026-07-29). Using it as a drill victim is not covered by that exception and contradicts the 2026-07-25 ruling. If touched, restore qemu-img snapshot -a virgin.
  • Forbidden: no global Docker cleanup; do not touch k3s data dirs, Longhorn mounts, PBS datastores or Gitea storage; no destructive disk/guest ops; never run CC here with prompts disabled; abort large builds if df -h /mnt/5_hdd / shows either over 90 %. Because it is the recovery chain and a live k3s node.

Peti's cluster (peti-felhom) — RETIRED 2026-09-25 (operator ruling)

No longer protected and no longer a machine this project knows: the tester wiped his server and it will not return. Its hub customer, Storage Box sub-account (u629488-sub2, which never held a backup) and records were removed through the hub's own customer delete; ep0 held nothing of it. The audit trail stays (events, the deletion and reset tombstones). Record: audits/RETIRE-peti-2026-09-25.md.

ep0 (felhom-hetzner, ep0.felhom.eu) + the Hetzner Storage Boxes — Tier 2, PROTECTED (operator ruling 2026-08-03)

Reads are fine. It is the offsite of last resort (PBS-DR datastore, WireGuard hub, operator OOB path) and — until 2026-08-03 — RAM-constrained (3.8 GB, R-90); it is now a CX33 with 8 GB RAM and a 4 GiB swapfile, which is what closed R-90. A very large restore is still worth watching — the 8 GB is comfortable, not unbounded, and the OOM that started R-90 was a 14.46 GB restore read against 3.8 GB. Do not delete datastores, prune jobs, tunnel config or nftables rules; never a drill target. The ordinary off-site READ a restore-test performs is permitted and unchanged by the ruling (R-86): the classification forbids destruction, not use. The Storage Boxes hold the restic copy — customer documents and photos, on a credential that can still delete (R-95). Access: ssh root@167.233.158.164 from DooPlex — not felhom-pve → 10.77.0.1, the route that produced a false "unreachable" verdict (standing rule 2).

Tier 1 — drill VMs, scratch guests, scratch customers

Create and destroy freely on a Tier 0 host. Two exceptions: drill-r50 is a fixture, not scratch; and scratch customers outlive their VMs in the hub — delete those too, or they accumulate (sess-c and sess-d were both left behind before anyone noticed).

A fixture may prove a mechanism. Only a fresh box may prove a path.

Rebuilding from scratch every time is waste; reusing a box is legitimate — but not for every claim.

  • Reusable fixture — a snapshot-reset VM on a Tier 0 host. Use it for mechanism work: payload capture, fix cycles, anything whose claim is about code behaviour. Reset to virgin between runs. Fast, repeatable, and the right default for iteration.
  • Fresh day-0 from the ISO — required for any claim about the install path, the golden image, agent publish/vouch, or first-boot state. Slow, and the only thing that catches the drift family: R-111 (the golden's agent 17 releases behind), R-115 (an agent built and deployed but never published), R-120 (the golden a controller release behind).

A fixture must record its provenance — which golden, agent and controller it was built from, and when — alongside the VM. A fixture whose versions drift silently is R-120's mechanism turned into a permanent installation, and it is worse than no fixture, because it produces confident wrong results quickly.

Worked example. R-116's closing run deliberately did a real day-0 from the v1.25.0 ISO on demo-hp instead of reusing the standing fixture. That is how R-120 surfaced — the fresh box installed the golden's controller, which is a release behind, and showed the customer the wrong absent-target message. The fixture would have shown a controller nobody installs.

Not established — unknown, not guessed

  • Hetzner Storage Boxes (storage-box-pool-1, PBS-storage-1 / u629193 / box 611421) — state taken from backlog/OPEN-ITEMS.md; no direct access attempted or confirmed. Tier 2 by what they hold.
  • felhotest (legacy, ssh -p 33022 kisfenyo@router.abonet.hu) — Connection refused 2026-07-30; that was the only route tried. Untiered; assume nothing.
  • Peti's cluster — RETIRED 2026-09-25 (operator ruling: the tester wiped his server; it will not return). Its hub records, its Storage Box sub-account and folder are gone; ep0 never held anything of it. audits/RETIRE-peti-2026-09-25.md.

The gap this page closes

During the R-116 diagnosis session (2026-07-30) a nested-Proxmox drill ran on DooPlex, a Tier 2 machine. Not a newly provisioned VM — the standing golden-bake fixture, booted from virgin and restored to it, so nothing was lost. But it was chosen because DooPlex was the only machine no spec had fenced, while the ruling naming the t740 sat in operations/nodes.md with nothing pointing at it. The designation existed and was unreachable — the same class as an absent signal read as a positive one.