Files
felhom.eu/documentation/runbooks/target-selection.md
T
admin 4d6ec7c7bb
gates / gates (push) Successful in 14s
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
Four paper debts and one fact given a reader. Hub-only — nothing to bake.

A4 — the entry about "the tester's machine" named a risk correctly and labelled it
in a way that invited deleting it. Established from the hub's own store: `peti-felhom`
is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and
the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record
with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47
this morning. The prompt's premise conflated the two; the register now says which is which.

A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out
of STATUS's "Waiting on you", which is now empty.

A3 — day0-install §C.1 said pushing the installer publishes it. It has not since
R-110. Corrected, with the two manifest pins named and an outside-verification command;
the one copy that repeated it (a dated audit, true when written) carries a superseded note.

A5 — standing rule 5: evidence comes off the machine at the end of the phase that
produced it, before any revert. Earned twice in three days on the same box at the same
point (R-320). Four homes, plus what to do when it is already gone.

R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll`
mail kind so the mail names the page a REBUILT box actually shows („A szerver
beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach.
Naming only; the acceptance pin proves the secret is untouched.

R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The
signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads
healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never
drawn as healthy — three absences, three sentences. No alarm, deliberately.
Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190.

B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now.
Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled`
reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed.

Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is
age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has
never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect).
2026-08-13 10:50:12 +02:00

204 lines
13 KiB
Markdown

# Which box do I break? — target selection by blast radius
> Read before picking a machine for a drill, a destructive test, or a throwaway VM. It answers one
> question: *what is safe to lose.* Reachability is in `CLAUDE.md`; hardware in `operations/nodes.md`.
> Added 2026-07-30. **Revised 2026-08-02 by operator decision D-d** (`CONTEXT.md` S-5).
## The rule
> **Two machines are protected: `DooPlex` and Peti's box. Everything else is disposable.**
> Operator decision **D-d**, 2026-08-02. DooPlex because it holds Gitea, the hub, the backups and the
> registry — everything else rebuilds from it, and it rebuilds from nothing. Peti's box because there
> is a real person behind it. **Every other box, both demo boxes included, may be broken or
> reinstalled freely.**
>
> **This is a correction, not a relaxation.** The earlier posture was costing whole sessions to
> caution and pushing drills onto DooPlex — the one machine that should never host them. If you are
> weighing whether a demo box can take a destructive test: it can. **Reach for a Tier 0 box first and
> do not ask.**
>
> Still true: **start at Tier 0 and work down only if Tier 0 genuinely cannot host the work**, and a
> protected machine is used only when a task says so explicitly — never by inference from what was
> not forbidden. An absent fence is not permission. If no tier fits, **stop and ask.**
Fences name **acts**, not machines. "Do not destroy demo-hp's `drill-r50` fixture" and "do not use
demo-hp to host a throwaway VM" are unrelated; only the first has ever been meant. Read a per-machine
prohibition as covering the act it names and nothing more.
## Before you revert it — take the evidence off first
> **A phase's evidence is copied off the machine at the end of THAT phase, before any revert, snapshot
> restore or teardown. Not at the end of the session.** (Standing rule 5, R-320.)
>
> **The intermediate revert is the one that gets forgotten.** Both losses this project has recorded
> were the *middle* teardown, never the final one — the Phase A logs of the 2026-08-12 retained-key
> drill and the Part 1 logs of the 2026-08-13 R-316 run, both on `drill-r50`, both destroyed by a
> revert to `virgin` between phases, three days apart. Both times the conclusions survived only
> because the quotations had been read live and an independent reproduction happened to exist. **That
> is luck.**
>
> **The mechanism:** make the pull the last act of the phase, not a step to remember later —
> `pct pull` / `scp` into `documentation/audits/evidence-<topic>-<date>/` on DooPlex, then revert.
> Tier 0 machines are *disposable*, which is precisely why nothing you need may be left on one.
>
> **If you notice it is already gone:** say so plainly in the report and **reproduce the finding
> independently**. That is the documented expectation, not an improvisation.
| Tier | Meaning | Machines |
|---|---|---|
| **0 — disposable. Reach here first.** | Exists to be broken; reinstalling is a routine afternoon, not an incident. **A drill that needs a victim uses one of these.** | `demo-hp` (t740), `demo-felhom` (N100) |
| **1 — create and destroy freely** | Throwaway VMs, guests, scratch customers — **hosted on a Tier 0 machine** | drill VMs, scratch guests |
| **2 — protected. Never a drill target.** | Losing it costs the recovery chain or a real relationship | **DooPlex**, **Peti's cluster**, **`ep0`** (operator ruling 2026-08-03) — and nothing else |
**DooPlex is Tier 2 because it *is* the recovery chain** — hub, Gitea, registry, PBS, k3s + Longhorn.
Everything else rebuilds from it; it rebuilds from nothing. A bad moment in a DR drill there costs the
thing under test, the source of truth for it, and the backups, at once.
**`ep0` is Tier 2 — PROTECTED. Operator ruling, 2026-08-03.** D-d named two protected machines and did
not name ep0 either way, so this page carried the question in writing for two days and read it the
narrow way meanwhile (not protected, but not wipeable). The ruling settles it and **extends D-d's
protected list to three machines**: DooPlex, Peti's cluster, ep0.
The reason it was never really in doubt: ep0 holds the **PBS-DR datastore and the restic copy of a
real customer's data**, which is the only off-premises copy that exists. So *deleting datastores,
prune jobs, tunnel config or nftables rules* was already forbidden by what it would destroy; the
ruling makes the classification say so plainly instead of leaving each session to re-derive it.
**Reads are fine** — including the ordinary off-site read a restore-test performs (R-86) — and it is
never a drill target. The Hetzner Storage Boxes ride the same reasoning.
**Standing ruling, 2026-07-25 (`operations/nodes.md`):** drill and build VMs live on the **t740** — not
felhom-pve, and **moved off DooPlex**. This page exists because that ruling sat where no session reads.
## Per machine — permitted / needs care / forbidden
### `demo-hp` — HP t740 · **Tier 0 · the designated drill + build VM host**
- **Freely:** host nested drill VMs; create/destroy guests and scratch customers; reinstall the box.
**This is the default answer to "where do I run this".**
- **Care:** `local-lvm` is a **thin pool, over-subscribed** (~144 GiB allocated over ~54 GiB) backing
live guest 9201 — filling it corrupts every guest. Put VM disks on a dir storage at
**`/mnt/nvme-1tb`, at its root** (a subdirectory fails the agent's `exactMount` check → storage reads
`disconnected` forever). **That warning is about one storage, not the box.** `/mnt/nvme-1tb` is also
the `felhom-backup` target and the enrolled user-data drive, so remove scratch storages when done.
- **Forbidden:** do not destroy or unblock **`drill-r50` (VM 300)** — the only drift fixture (R-93).
(Access: the docs say no baked SSH key and G1 break-glass, but a key authenticated on 2026-07-31 —
**R-129**, unresolved.)
### `demo-felhom` — N100 · **Tier 0**
- **Freely:** create/destroy guests and scratch customers; reinstall the box.
- **Care:** it carries the **PBS-DR / offsite tier** — but so does demo-hp now, so this is **no longer
a reason to prefer one over the other**. Prefer demo-hp per the 2026-07-25 ruling, on the ruling's
own grounds (it is the designated drill host), not because its backup chain is safe to disturb.
<!--
CORRECTED 2026-08-06. This line used to read "(demo-hp has none)" and was the runbook's stated reason
for steering backup-disturbing tests at demo-felhom. It was TRUE WHEN WRITTEN and went stale when F10
resolved on 2026-07-23; it was flagged as wrong for some time and carried anyway.
Measured before editing, on the box, from the tier's own evidence rather than a config entry:
demo-hp:/etc/pve/storage.cfg pbs: felhom-pbs -> datastore felhom-offsite, server 10.77.0.1,
namespace demo-hp, username felhom@pbs!demo-hp
demo-hp# LC_ALL=C pvesm list felhom-pbs
felhom-pbs:backup/ct/9201/2026-07-28T19:19:45Z pbs-ct 6264034048 9201
felhom-pbs:backup/ct/9201/2026-08-04T19:24:16Z pbs-ct 4637840512 9201
ep0:/mnt/pbs-datastore/ns/ c11 demo-felhom demo-hp rewalk
Two snapshots, the newest two days old, in demo-hp's OWN namespace on the shared off-site datastore.
The PBS-DR leg is what was measured here; the restic/offbox app-data leg is separately evidenced by
R-193, whose whole subject is demo-hp's off-site app-data tier being dropped by the 2026-08-03
rebuild and restaged.
Why this mattered more than a typo: this file's job is to say which machine may be destroyed, and the
sentence was load-bearing for that judgement — "its backup chain is not disturbable" is exactly the
kind of belief under which a drill lands somewhere it should not.
-->
**Both Tier 0 boxes — the shared backup-target fence is DOWNGRADED to a cost, 2026-08-02 (D-d).** It
read *"do not re-point either backup target"*, because these are the only two **correctly configured**
boxes and therefore the regression path new installer logic is measured against. D-d makes both boxes
freely breakable and reinstallable, which loses that reference just as thoroughly — so the fence was
inconsistent with the decision and is not kept as a prohibition. **What survives is the reason:**
re-pointing (or reinstalling) costs the reference configuration, so know that you are spending it and
put the box back. If both are spent at once there is no correctly-configured box left to compare
against.
### `DooPlex` — 192.168.0.180 · **Tier 2**
- **Freely:** build, push images, `sudo kubectl`, read anything. Work inside `/mnt/5_hdd/felhom.eu/`.
- **Care:** the golden-bake nested VM `drill/drill.qcow2` lives here and is an accepted exception **for
bakes** (last used 0.185.1, 2026-07-29). Using it as a *drill victim* is not covered by that exception
and contradicts the 2026-07-25 ruling. If touched, restore `qemu-img snapshot -a virgin`.
- **Forbidden:** no global Docker cleanup; do not touch k3s data dirs, Longhorn mounts, PBS datastores
or Gitea storage; no destructive disk/guest ops; never run CC here with prompts disabled; abort large
builds if `df -h /mnt/5_hdd /` shows either over 90 %. **Because it is the recovery chain and a live
k3s node.**
### `Peti's cluster` (`peti-felhom`) — **Tier 2**
**Do not touch, at all.** A real external pilot with a real person behind it; its whole-guest backup
still shares a device with its guest, so a drive failure is **offsite-only recovery**. Deliberately not
migrated, parked until the tester reinstalls (`PETI` in `backlog/OPEN-ITEMS.md`). Currently DOWN, no
enrolled host. No access route from DooPlex, and nothing here needs one.
### `ep0` (`felhom-hetzner`, `ep0.felhom.eu`) + the Hetzner Storage Boxes — **Tier 2, PROTECTED** (operator ruling 2026-08-03)
Reads are fine. It is the **offsite of last resort** (PBS-DR datastore, WireGuard hub, operator OOB
path) and — until 2026-08-03 — RAM-constrained (3.8 GB, R-90); it is now a **CX33 with 8 GB RAM and a
4 GiB swapfile**, which is what closed R-90. A very large restore is still worth watching — the 8 GB
is comfortable, not unbounded, and the OOM that started R-90 was a 14.46 GB restore read against
3.8 GB. Do not delete datastores, prune jobs, tunnel config or nftables rules; never a drill target.
**The ordinary off-site READ a restore-test performs is permitted and unchanged by the ruling**
(R-86): the classification forbids destruction, not use. The Storage Boxes hold the restic copy —
customer documents and photos, on a credential that can still delete (R-95).
**Access: `ssh root@167.233.158.164` from DooPlex***not* `felhom-pve → 10.77.0.1`, the route that
produced a false "unreachable" verdict (standing rule 2).
### Tier 1 — drill VMs, scratch guests, scratch customers
Create and destroy freely **on a Tier 0 host**. Two exceptions: **`drill-r50`** is a fixture, not
scratch; and scratch **customers** outlive their VMs in the hub — delete those too, or they accumulate
(`sess-c` and `sess-d` were both left behind before anyone noticed).
#### A fixture may prove a mechanism. Only a fresh box may prove a path.
Rebuilding from scratch every time is waste; reusing a box is legitimate — but not for every claim.
- **Reusable fixture** — a snapshot-reset VM on a Tier 0 host. Use it for **mechanism** work: payload
capture, fix cycles, anything whose claim is about *code behaviour*. Reset to `virgin` between runs.
Fast, repeatable, and the right default for iteration.
- **Fresh day-0 from the ISO** — required for any claim about the **install path, the golden image,
agent publish/vouch, or first-boot state**. Slow, and the only thing that catches the drift family:
**R-111** (the golden's agent 17 releases behind), **R-115** (an agent built and deployed but never
published), **R-120** (the golden a controller release behind).
**A fixture must record its provenance** — which golden, agent and controller it was built from, and
when — alongside the VM. A fixture whose versions drift silently is R-120's mechanism turned into a
permanent installation, and it is *worse* than no fixture, because it produces confident wrong results
quickly.
**Worked example.** R-116's closing run deliberately did a real day-0 from the v1.25.0 ISO on demo-hp
instead of reusing the standing fixture. **That is how R-120 surfaced** — the fresh box installed the
golden's controller, which is a release behind, and showed the customer the wrong absent-target message.
The fixture would have shown a controller nobody installs.
## Not established — unknown, not guessed
- **Hetzner Storage Boxes** (`storage-box-pool-1`, `PBS-storage-1` / u629193 / box 611421) — state taken
from `backlog/OPEN-ITEMS.md`; **no direct access attempted or confirmed.** Tier 2 by what they hold.
- **`felhotest`** (legacy, `ssh -p 33022 kisfenyo@router.abonet.hu`) — **`Connection refused`**
2026-07-30; that was the only route tried. Untiered; assume nothing.
- **Peti's cluster** hardware/storage layout — unverified. Tier 2 rests on the relationship, which
needs no verification.
## The gap this page closes
During the R-116 diagnosis session (2026-07-30) a nested-Proxmox drill ran **on DooPlex**, a Tier 2
machine. Not a newly provisioned VM — the standing golden-bake fixture, booted from `virgin` and
restored to it, so nothing was lost. But it was chosen because DooPlex was the only machine no spec had
fenced, while the ruling naming the t740 sat in `operations/nodes.md` with nothing pointing at it.
**The designation existed and was unreachable** — the same class as an absent signal read as a positive
one.