eb1c56a981
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
232 lines
16 KiB
Markdown
232 lines
16 KiB
Markdown
# Which box do I break? — target selection by blast radius
|
|
|
|
> Read before picking a machine for a drill, a destructive test, or a throwaway VM. It answers one
|
|
> question: *what is safe to lose.* Reachability is in `CLAUDE.md`; hardware in `operations/nodes.md`.
|
|
> Added 2026-07-30. **Revised 2026-08-02 by operator decision D-d** (`CONTEXT.md` S-5).
|
|
|
|
## The rule
|
|
|
|
> **Two machines are protected: `DooPlex` and `ep0`. Everything else is disposable.**
|
|
> Operator decision **D-d**, 2026-08-02, and the ep0 ruling of 2026-08-03. DooPlex because it holds
|
|
> Gitea, the hub, the backups and the registry — everything else rebuilds from it, and it rebuilds from
|
|
> nothing. ep0 for the reason below. **Every other box, both demo boxes included, may be broken or
|
|
> reinstalled freely.** *(Peti's box was the second protected machine until it was RETIRED on
|
|
> 2026-09-25 — operator ruling; `audits/RETIRE-peti-2026-09-25.md`.)*
|
|
>
|
|
> **This is a correction, not a relaxation.** The earlier posture was costing whole sessions to
|
|
> caution and pushing drills onto DooPlex — the one machine that should never host them. If you are
|
|
> weighing whether a demo box can take a destructive test: it can. **Reach for a Tier 0 box first and
|
|
> do not ask.**
|
|
>
|
|
> Still true: **start at Tier 0 and work down only if Tier 0 genuinely cannot host the work**, and a
|
|
> protected machine is used only when a task says so explicitly — never by inference from what was
|
|
> not forbidden. An absent fence is not permission. If no tier fits, **stop and ask.**
|
|
|
|
Fences name **acts**, not machines. "Do not destroy demo-hp's `drill-r50` fixture" and "do not use
|
|
demo-hp to host a throwaway VM" are unrelated; only the first has ever been meant. Read a per-machine
|
|
prohibition as covering the act it names and nothing more.
|
|
|
|
## A drill night that forbids baking a golden — what you will see, and why it is expected
|
|
|
|
**This is written here because the instruction that caused it was mine and I left it out** (R-417).
|
|
An overnight drill runbook that says *"no golden bake, no vouch, no floor change tonight — those are
|
|
the operator's acts"* is correct, and it also guarantees the `golden-currency` gate is RED for the
|
|
whole night whenever a controller release is newer than the last bake. **That is the gate telling the
|
|
truth**, not a fault to work around.
|
|
|
|
Since 2026-09-01 (R-404) it no longer refuses your pushes. A drill's own pushes — evidence, register
|
|
rows, `STATUS.md`, `REPORT.md` — are documents-only, so the gate's conviction prints as a loud
|
|
**ADVISORY** block and the push proceeds. **Expect to see it, every push, all night. Do not silence
|
|
it and do not bypass the hook for it.**
|
|
|
|
Two things still hold:
|
|
|
|
* A push that touches **code** is still refused — bake first or do not push code.
|
|
* **Every other gate still refuses every push.** If a push is refused during a drill it is NOT the
|
|
golden gate, and the message will name which gate it was.
|
|
|
|
If the drill genuinely must ship a release with no golden, the honest instrument is a **waiver row**
|
|
in `documentation/backlog/OPEN-ITEMS.md`, never `--no-verify`.
|
|
|
|
## Before you revert it — take the evidence off first
|
|
|
|
> **A phase's evidence is copied off the machine at the end of THAT phase, before any revert, snapshot
|
|
> restore or teardown. Not at the end of the session.** (Standing rule 5, R-320.)
|
|
>
|
|
> **The intermediate revert is the one that gets forgotten.** Both losses this project has recorded
|
|
> were the *middle* teardown, never the final one — the Phase A logs of the 2026-08-12 retained-key
|
|
> drill and the Part 1 logs of the 2026-08-13 R-316 run, both on `drill-r50`, both destroyed by a
|
|
> revert to `virgin` between phases, three days apart. Both times the conclusions survived only
|
|
> because the quotations had been read live and an independent reproduction happened to exist. **That
|
|
> is luck.**
|
|
>
|
|
> **The mechanism:** make the pull the last act of the phase, not a step to remember later —
|
|
> `pct pull` / `scp` into `documentation/audits/evidence-<topic>-<date>/` on DooPlex, then revert.
|
|
> Tier 0 machines are *disposable*, which is precisely why nothing you need may be left on one.
|
|
>
|
|
> **If you notice it is already gone:** say so plainly in the report and **reproduce the finding
|
|
> independently**. That is the documented expectation, not an improvisation.
|
|
|
|
| Tier | Meaning | Machines |
|
|
|---|---|---|
|
|
| **0 — disposable. Reach here first.** | Exists to be broken; reinstalling is a routine afternoon, not an incident. **A drill that needs a victim uses one of these.** | `demo-hp` (t740), `demo-felhom` (N100) |
|
|
| **1 — create and destroy freely** | Throwaway VMs, guests, scratch customers — **hosted on a Tier 0 machine** | drill VMs, scratch guests |
|
|
| **2 — protected. Never a drill target.** | Losing it costs the recovery chain or a real relationship | **DooPlex**, **`ep0`** (operator ruling 2026-08-03) — and nothing else (Peti's cluster RETIRED 2026-09-25) |
|
|
|
|
**DooPlex is Tier 2 because it *is* the recovery chain** — hub, Gitea, registry, PBS, k3s + Longhorn.
|
|
Everything else rebuilds from it; it rebuilds from nothing. A bad moment in a DR drill there costs the
|
|
thing under test, the source of truth for it, and the backups, at once.
|
|
|
|
**`ep0` is Tier 2 — PROTECTED. Operator ruling, 2026-08-03.** D-d named two protected machines and did
|
|
not name ep0 either way, so this page carried the question in writing for two days and read it the
|
|
narrow way meanwhile (not protected, but not wipeable). The ruling settles it and **extends D-d's
|
|
protected list to three machines**: DooPlex, Peti's cluster, ep0. **Since 2026-09-25 (Peti's box
|
|
retired) it is two: DooPlex and ep0.**
|
|
|
|
The reason, as it stands on 2026-09-25: ep0 holds the **PBS-DR datastore with the demo boxes' (and
|
|
tester-1's) whole-guest copies, the WireGuard hub every box's off-site path runs through, and the
|
|
operator's out-of-band path**; the Hetzner Storage Boxes hold each box's restic copy. There is no paying
|
|
customer yet, but ep0 is the only off-premises tier of the whole product, and a mistake there cannot be
|
|
rebuilt from DooPlex. So *deleting datastores,
|
|
prune jobs, tunnel config or nftables rules* was already forbidden by what it would destroy; the
|
|
ruling makes the classification say so plainly instead of leaving each session to re-derive it.
|
|
**Reads are fine** — including the ordinary off-site read a restore-test performs (R-86) — and it is
|
|
never a drill target. The Hetzner Storage Boxes ride the same reasoning.
|
|
|
|
**Standing ruling, 2026-07-25 (`operations/nodes.md`):** drill and build VMs live on the **t740** — not
|
|
felhom-pve, and **moved off DooPlex**. This page exists because that ruling sat where no session reads.
|
|
|
|
## Per machine — permitted / needs care / forbidden
|
|
|
|
### `demo-hp` — HP t740 · **Tier 0 · the designated drill + build VM host**
|
|
|
|
- **Freely:** host nested drill VMs; create/destroy guests and scratch customers; reinstall the box.
|
|
**This is the default answer to "where do I run this".**
|
|
- **Care:** `local-lvm` is a **thin pool, over-subscribed** (~144 GiB allocated over ~54 GiB) backing
|
|
live guest 9201 — filling it corrupts every guest. Put VM disks on a dir storage at
|
|
**`/mnt/hdd_1`, at its root** (R-461, measured 2026-09-13: there is NO `/mnt/nvme-1tb` on either box — demo-hp's 1 TB NVMe is `nvme0n1` mounted at `/mnt/hdd_1`, the enrolled user-data drive and `felhom-backup` target, so "the NVMe" and "the data drive" are one disk) (a subdirectory fails the agent's `exactMount` check → storage reads
|
|
`disconnected` forever). **That warning is about one storage, not the box.** `/mnt/hdd_1` is also
|
|
the `felhom-backup` target and the enrolled user-data drive, so remove scratch storages when done.
|
|
- **Forbidden:** do not destroy or unblock **`drill-r50` (VM 300)** — the only drift fixture (R-93). **Measured 2026-09-13 (R-461): `qm list` is EMPTY on both demo-hp and demo-felhom — the VM does not exist anywhere, so the fence currently protects nothing and R-93's premise is gone. The fence stays as written for the day someone rebuilds it; do not read its presence here as evidence the fixture exists.**
|
|
(Access: the docs say no baked SSH key and G1 break-glass, but a key authenticated on 2026-07-31 —
|
|
**R-129**, unresolved.)
|
|
|
|
### `demo-felhom` — N100 · **Tier 0**
|
|
|
|
- **Freely:** create/destroy guests and scratch customers; reinstall the box.
|
|
- **Care:** it carries the **PBS-DR / offsite tier** — but so does demo-hp now, so this is **no longer
|
|
a reason to prefer one over the other**. Prefer demo-hp per the 2026-07-25 ruling, on the ruling's
|
|
own grounds (it is the designated drill host), not because its backup chain is safe to disturb.
|
|
|
|
<!--
|
|
CORRECTED 2026-08-06. This line used to read "(demo-hp has none)" and was the runbook's stated reason
|
|
for steering backup-disturbing tests at demo-felhom. It was TRUE WHEN WRITTEN and went stale when F10
|
|
resolved on 2026-07-23; it was flagged as wrong for some time and carried anyway.
|
|
|
|
Measured before editing, on the box, from the tier's own evidence rather than a config entry:
|
|
|
|
demo-hp:/etc/pve/storage.cfg pbs: felhom-pbs -> datastore felhom-offsite, server 10.77.0.1,
|
|
namespace demo-hp, username felhom@pbs!demo-hp
|
|
demo-hp# LC_ALL=C pvesm list felhom-pbs
|
|
felhom-pbs:backup/ct/9201/2026-07-28T19:19:45Z pbs-ct 6264034048 9201
|
|
felhom-pbs:backup/ct/9201/2026-08-04T19:24:16Z pbs-ct 4637840512 9201
|
|
ep0:/mnt/pbs-datastore/ns/ c11 demo-felhom demo-hp rewalk
|
|
|
|
Two snapshots, the newest two days old, in demo-hp's OWN namespace on the shared off-site datastore.
|
|
The PBS-DR leg is what was measured here; the restic/offbox app-data leg is separately evidenced by
|
|
R-193, whose whole subject is demo-hp's off-site app-data tier being dropped by the 2026-08-03
|
|
rebuild and restaged.
|
|
|
|
Why this mattered more than a typo: this file's job is to say which machine may be destroyed, and the
|
|
sentence was load-bearing for that judgement — "its backup chain is not disturbable" is exactly the
|
|
kind of belief under which a drill lands somewhere it should not.
|
|
-->
|
|
|
|
|
|
**Both Tier 0 boxes — the shared backup-target fence is DOWNGRADED to a cost, 2026-08-02 (D-d).** It
|
|
read *"do not re-point either backup target"*, because these are the only two **correctly configured**
|
|
boxes and therefore the regression path new installer logic is measured against. D-d makes both boxes
|
|
freely breakable and reinstallable, which loses that reference just as thoroughly — so the fence was
|
|
inconsistent with the decision and is not kept as a prohibition. **What survives is the reason:**
|
|
re-pointing (or reinstalling) costs the reference configuration, so know that you are spending it and
|
|
put the box back. If both are spent at once there is no correctly-configured box left to compare
|
|
against.
|
|
|
|
### `DooPlex` — 192.168.0.180 · **Tier 2**
|
|
|
|
- **Freely:** build, push images, `sudo kubectl`, read anything. Work inside `/mnt/5_hdd/felhom.eu/`.
|
|
- **Care:** the golden-bake nested VM `drill/drill.qcow2` lives here and is an accepted exception **for
|
|
bakes** (last used 0.185.1, 2026-07-29). Using it as a *drill victim* is not covered by that exception
|
|
and contradicts the 2026-07-25 ruling. If touched, restore `qemu-img snapshot -a virgin`.
|
|
- **Forbidden:** no global Docker cleanup; do not touch k3s data dirs, Longhorn mounts, PBS datastores
|
|
or Gitea storage; no destructive disk/guest ops; never run CC here with prompts disabled; abort large
|
|
builds if `df -h /mnt/5_hdd /` shows either over 90 %. **Because it is the recovery chain and a live
|
|
k3s node.**
|
|
|
|
### `Peti's cluster` (`peti-felhom`) — **RETIRED 2026-09-25** (operator ruling)
|
|
|
|
No longer protected and no longer a machine this project knows: the tester wiped his server and it will not
|
|
return. Its hub customer, Storage Box sub-account (`u629488-sub2`, which never held a backup) and records were
|
|
removed through the hub's own customer delete; ep0 held nothing of it. The audit trail stays (events, the
|
|
deletion and reset tombstones). Record: `audits/RETIRE-peti-2026-09-25.md`.
|
|
|
|
### `ep0` (`felhom-hetzner`, `ep0.felhom.eu`) + the Hetzner Storage Boxes — **Tier 2, PROTECTED** (operator ruling 2026-08-03)
|
|
|
|
Reads are fine. It is the **offsite of last resort** (PBS-DR datastore, WireGuard hub, operator OOB
|
|
path) and — until 2026-08-03 — RAM-constrained (3.8 GB, R-90); it is now a **CX33 with 8 GB RAM and a
|
|
4 GiB swapfile**, which is what closed R-90. A very large restore is still worth watching — the 8 GB
|
|
is comfortable, not unbounded, and the OOM that started R-90 was a 14.46 GB restore read against
|
|
3.8 GB. Do not delete datastores, prune jobs, tunnel config or nftables rules; never a drill target.
|
|
**The ordinary off-site READ a restore-test performs is permitted and unchanged by the ruling**
|
|
(R-86): the classification forbids destruction, not use. The Storage Boxes hold the restic copy —
|
|
customer documents and photos, on a credential that can still delete (R-95).
|
|
**Access: `ssh root@167.233.158.164` from DooPlex** — *not* `felhom-pve → 10.77.0.1`, the route that
|
|
produced a false "unreachable" verdict (standing rule 2).
|
|
|
|
### Tier 1 — drill VMs, scratch guests, scratch customers
|
|
|
|
Create and destroy freely **on a Tier 0 host**. Two exceptions: **`drill-r50`** is a fixture, not
|
|
scratch; and scratch **customers** outlive their VMs in the hub — delete those too, or they accumulate
|
|
(`sess-c` and `sess-d` were both left behind before anyone noticed).
|
|
|
|
#### A fixture may prove a mechanism. Only a fresh box may prove a path.
|
|
|
|
Rebuilding from scratch every time is waste; reusing a box is legitimate — but not for every claim.
|
|
|
|
- **Reusable fixture** — a snapshot-reset VM on a Tier 0 host. Use it for **mechanism** work: payload
|
|
capture, fix cycles, anything whose claim is about *code behaviour*. Reset to `virgin` between runs.
|
|
Fast, repeatable, and the right default for iteration.
|
|
- **Fresh day-0 from the ISO** — required for any claim about the **install path, the golden image,
|
|
agent publish/vouch, or first-boot state**. Slow, and the only thing that catches the drift family:
|
|
**R-111** (the golden's agent 17 releases behind), **R-115** (an agent built and deployed but never
|
|
published), **R-120** (the golden a controller release behind).
|
|
|
|
**A fixture must record its provenance** — which golden, agent and controller it was built from, and
|
|
when — alongside the VM. A fixture whose versions drift silently is R-120's mechanism turned into a
|
|
permanent installation, and it is *worse* than no fixture, because it produces confident wrong results
|
|
quickly.
|
|
|
|
**Worked example.** R-116's closing run deliberately did a real day-0 from the v1.25.0 ISO on demo-hp
|
|
instead of reusing the standing fixture. **That is how R-120 surfaced** — the fresh box installed the
|
|
golden's controller, which is a release behind, and showed the customer the wrong absent-target message.
|
|
The fixture would have shown a controller nobody installs.
|
|
|
|
## Not established — unknown, not guessed
|
|
|
|
- **Hetzner Storage Boxes** (`storage-box-pool-1`, `PBS-storage-1` / u629193 / box 611421) — state taken
|
|
from `backlog/OPEN-ITEMS.md`; **no direct access attempted or confirmed.** Tier 2 by what they hold.
|
|
- **`felhotest`** (legacy, `ssh -p 33022 kisfenyo@router.abonet.hu`) — **`Connection refused`**
|
|
2026-07-30; that was the only route tried. Untiered; assume nothing.
|
|
- **Peti's cluster** — RETIRED 2026-09-25 (operator ruling: the tester wiped his server; it will not
|
|
return). Its hub records, its Storage Box sub-account and folder are gone; ep0 never held anything of
|
|
it. `audits/RETIRE-peti-2026-09-25.md`.
|
|
|
|
## The gap this page closes
|
|
|
|
During the R-116 diagnosis session (2026-07-30) a nested-Proxmox drill ran **on DooPlex**, a Tier 2
|
|
machine. Not a newly provisioned VM — the standing golden-bake fixture, booted from `virgin` and
|
|
restored to it, so nothing was lost. But it was chosen because DooPlex was the only machine no spec had
|
|
fenced, while the ruling naming the t740 sat in `operations/nodes.md` with nothing pointing at it.
|
|
**The designation existed and was unreachable** — the same class as an absent signal read as a positive
|
|
one.
|