Peti's box retired (operator ruling 2026-09-25): hub customer + Storage Box sub-account removed, ep0 held nothing; protected list = DooPlex + ep0; decision 34 (no leg resume, R-686 closed); R-688
gates / gates (push) Successful in 25s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-25 11:04:52 +02:00
parent 6b2176e480
commit eb1c56a981
28 changed files with 398 additions and 28 deletions
@@ -45,7 +45,7 @@
| Scenario | Components | Status | Evidence | Gap / roadmap |
|---|---|---|---|---|
| Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box | installer, agent, hub, golden | **PROVEN-LIVE** (nested VM) | `DRILL-day0-vm-2026-07-12`, `DRILL-day0-take2-2026-07-12` | First firing on real customer hardware pending → R-1 |
| BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1 |
| BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | the first real BYO run of the current path is still owed (R-1) — the planned venue, Peti's clean-slate reinstall, is gone: Peti's box was RETIRED 2026-09-25 |
| **The installer is PUBLISHED, not pushed — the artifact that runs as root on a virgin box is served from a version-controlled ref, and rolling back is one act** | scripts **v1.23.0** + `manifests/webpage.yaml` (R-110, operator ruling option (b)) | **PROVEN-LIVE (2026-08-03)** | `scripts/CHANGELOG.md` v1.23.0 + `REPORT.md`. **Proven by HTTP against the real URL, not from a pod's filesystem.** *Scenario A:* a real push to `main` without moving the tag left the served script **byte-identical** (`sha256 2f859555…`), and a marker comment planted in that very commit was **absent** from the served bytes, while the website tree advanced to the new commit in the same observation — both halves of the split in one measurement. *Scenario B:* moving the tag published in **~40 s** (`sha → ea2b4aa9…`, marker present) and moving it back restored **exactly** the pre-publish sha. `https://felhom.eu/` returned 200 throughout. *P-A, measured BEFORE the manifest was touched because the model rests on it:* git-sync v4.4.0 follows a tag **and notices a moved one** (`update required … local:<old> remote:<new>` → `updated successfully`) | **Two syncs, deliberately: the WEBSITE still tracks `main`.** Pinning both would turn every copy edit into a release, which makes the release meaningless and the site slow to fix. **Publish** = cut `installer-v<SCRIPT_VERSION>` + bump the manifest `--ref` + sync; **roll back** = move the tag back, which needs **no ArgoCD sync and no deploy**. Continuity is structural rather than lucky: both trees are seeded by init containers so a fresh pod is not Ready until the tag is checked out, and `maxUnavailable` rounds to 0 on one replica, so a failed scripts-init leaves the OLD pod serving — the failure direction is *no update*, never *no `/scripts/`*. **The URL never carried a ref**, so the bootstrap script and the hub's day-0 command follow the tag with no edit and **no hub change**. The installer's own sixteen run-time fetches are a separate channel pinned to the AGENT's version (**R-183**), because they are the agent's configs and not this repo's — leaving them on `main` would have made the whole change cosmetic |
| Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row | **Second board, 2026-07-21 (demo-hp, HP t740 / Ryzen V1756B / AMI M42):** the whole chain ran on virgin hardware in one pass — armed install → self-registration as an unclaimed appliance → operator bind → day-0 → running guest 9201 + agent 0.92.1 as `demo-hp-bb76ea`. **The shim loader booted with Secure Boot ENABLED**, which retires the assumption that Felhom installs need SB off — that was an N100-firmware workaround. The exact-serial disk filter took the system SSD and left the box's 1TB NVMe untouched/unenrolled on hardware it had never seen. Two failures filed rather than smoothed over: **R-59** (no DHCP → the installer baked a static fallback instead of aborting) and **R-61** (baked root password unknowable → no console access).
| Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (`<iso>.rootpw.txt`) | scripts v1.24.0 (`scripts/iso/felhom-bootstrap.sh` `network_gate`/`sweep_nics`, `build-felhom-iso.sh` rootpw emission) | **PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install)** | `audits/SPIKE-firstboot-nic-sweep-2026-07-22.md` — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and **self-registered at the hub unaided in under a minute**; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash | R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (`state.json` gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched |
+1 -1
View File
@@ -334,7 +334,7 @@ per Part 1: **snapshot** (LVM-thin, transient, whole-guest rollback — not a ba
> pool has room.
>
> **Delivered 2026-09-24 night** to demo-felhom and demo-hp by CC-signed `agent_update` jobs; the restore test is
> back ON on both (the `-1` config kept as `agent.json.night-0925-off`). Peti's box has not received it.
> back ON on both (the `-1` config kept as `agent.json.night-0925-off`). Peti's box never received it — it was RETIRED 2026-09-25 and will not return.
> **A whole-box backup that cannot fit is skipped with a reason, before anything starts (R-685, agent v0.134.0).**
> Before a vzdump to a LOCAL (non-PBS) target: free ≥ the guest's newest archive on that target × 1.25 + 1 GiB,
@@ -314,7 +314,8 @@ caveats, recorded not papered over: **(1)** this SIM was handed a **public mobil
stays a *retest-on-a-CGNAT-SIM-when-available* follow-up (low risk: mapping-hold is NAT-tier-agnostic
by mechanism). **(2)** a dual-stack mobile uplink made `wg-quick` prefer the endpoint **AAAA and ride
un-NATed IPv6** until v4 was forced — functionally fine, but see §4.2. The deferred second-ISP
vantage (Peti VM 110) remains the thorough confirmation but no longer gates anything. Runbook:
vantage (Peti VM 110) is gone — Peti's box was RETIRED 2026-09-25 — so a second-ISP confirmation needs another
venue; it gates nothing. Runbook:
`RUNBOOK-s3-cgnat-smoke`.
**Open sub-decisions (deferred by design):**
@@ -394,6 +394,14 @@ R-636's louder repeated alarm.
each step then gets a night of the household using it before the next one, and a failure is one step wide.
Operator may widen it. `TestLeg_OneStepPerAppPerNight`.
### 2026-09-25 — an operator ruling
34. **A controller restart during the night's update leg does NOT resume the leg** (operator ruling 2026-09-25,
option B of R-686). The apps the leg had not reached wait for the next night. A step already pressed is
finished or put back by the guarded update's own journal, exactly as before. **Nothing is built:** the
behaviour shipped in v0.271.0 is now the rule. **Why:** one night's delay for an app is cheap, and a resume
would be one more mechanism acting with nobody watching.
---
## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed
@@ -0,0 +1,81 @@
# RETIRE — Peti's box (`peti-felhom`), 2026-09-25
Operator ruling 2026-09-25: Peti's box is retired (the tester wiped his server; it will not return); its off-site
backup holds no user data and is to be deleted; it leaves the protected list. Architecture read first:
`runbooks/target-selection.md`, `05-hub-architecture.md` §14 (what a customer delete deprovisions),
`06-offsite-connectivity.md`, `04-control-plane-authorization.md`. Evidence: `audits/retire-peti-2026-09-25/`.
## Not done, or changed
- **Nothing of Peti's was on ep0.** No PBS namespace (`ns`: demo-felhom, demo-hp, tester-1), no WireGuard peer in
the live `wg show`, no config naming him. His only off-site item was a Storage Box **sub-account**
(`u629488-sub2`, id 269130, label `felhom-customer=peti-felhom`, home `felhom-peti-felhom`) on the pool box.
- **The "backup" was never a backup.** The sub-account's folder held ONE file: `.ssh/authorized_keys`, 81 bytes. No
restic repository was ever created — all 482 of his reports (2026-07-10 … 07-15) show 0 off-site snapshots and
0 bytes; his escrow stayed pending. The operator's "no user data" is confirmed by the listing, not only by word.
- **To list and empty the folder I reset the sub-account's password through the Hetzner API** (the product's own
`reset_subaccount_password` action, scoped to id 269130 after re-reading its label). The key file was removed from
inside the folder, then the hub deleted the sub-account. Whether Hetzner deletes a sub-account's folder with it
is NOT measured here — so the folder was emptied first, and nothing of his could remain either way.
- **Cloudflare:** his customer config carried a Cloudflare tunnel token and an API token (for `sajatfelhom.hu`).
They were purged with the customer record. The hub's delete dialog says "tunnel/zone removed", but no leg of the
cascade calls Cloudflare — the tunnel and DNS records on Cloudflare's side, if any remain, were NOT inventoried or
removed (a row: R-688).
- **A second, older Storage Box** (`PBS-storage-1`, u629193, box 611421) once held a folder `felhom-peti-spike/`
(spike leftovers, per `RUNBOOK-ep0-datastore-volume-2026-07-27.md`). Its mount on ep0 is gone, and the box is
not visible to either API token (404). The register row "delete the box" is the operator's. Not touched.
- **Code comments that name Peti were left** (hub `rollup.go`, `appliances.go`, `notify/templates.go`; agent pbsdr
tests). They explain why code is shaped as it is; no behaviour depends on Peti.
## A1 — inventory (before anything was removed)
| where | item | how found |
|---|---|---|
| hub | customer `peti-felhom` („Peti Proxmox", domain `sajatfelhom.hu`, dr_tier 0) with 482 reports, 123 events, 1,036 telemetry rows, DR recipe, one-time secret, claim, 1 log tail | copy of the hub DB (with its -wal), every table's `customer_id` / `host_id` |
| hub | host `peti-felhom-86d37d` — already DELETED 2026-07-15 08:56:22 (`host_deletions`, escrow_acked 0); no host, guest, escrow, recovery or PBS-secret rows | same |
| hub | WireGuard peer — none in `wg_peers` | same |
| Storage Box | sub-account id 269130 `u629488-sub2`, home `felhom-peti-felhom`, label `felhom-customer=peti-felhom`, created 2026-07-10 | Hetzner API with the hub's own token, matched on the LABEL |
| ep0 | nothing (namespaces, peers, `/etc`, `/root`, `/srv` grep — one false hit, "competing" in `lvm.conf`) | ssh read-only |
| documents | 48 non-history files + workspace rules + 21 memory notes (whole-word pattern; positive control target-selection.md 7 hits, negative control catalog templates 0) | `A1-documents.txt` |
**Control, before:** pool box `size_data 3,120,562,176`, sub-accounts sub1 (demo-felhom), sub3 (demo-hp), sub4
(tester-1); ep0 namespaces demo-felhom 2 snapshots / 220K, demo-hp 2 / 540K, tester-1 2 / 368K; demo boxes'
off-site `last_status ok` (11 and 90 snapshots, runs 02:15:46Z / 02:18:41Z).
## A3 — removal
1. `POST …/subaccounts/269130/actions/reset_subaccount_password` → success (password kept 0600 in the scratchpad,
deleted afterwards).
2. As `u629488-sub2`: `rm .ssh/authorized_keys`, `rmdir .ssh` → `ls -la` empty, `du -s .` 1.
3. Hub preview `GET /configs/peti-felhom/delete` — 0 hosts, off-site `u629488-sub2`, residue 1,519. Then
`POST /configs/peti-felhom/delete` (ack_hosts, ack_reset, ack_purge, confirm_id, expect_hosts=0) → 303. The hub's
log: sub-account 269130 deprovisioned; PBS tenancy `existed=false`; claim reset; residue purged (reports 482,
telemetry 1,036, log tails 1); cascade COMPLETE (journal #20).
## A4 — nothing else moved
| item | before | after |
|---|---|---|
| sub-accounts | sub1, sub2 (Peti), sub3, sub4 | sub1, sub3, sub4 — labels and homes unchanged |
| pool box `size_data` | 3,120,562,176 | 3,120,562,176 |
| login as `u629488-sub2` | worked | "Permission denied" |
| ep0 namespaces (snapshots / du) | demo-felhom 2/220K, demo-hp 2/540K, tester-1 2/368K | identical |
| ep0 WireGuard peers | .2 .3 .4 .250 | identical |
| hub hosts | demo-felhom, demo-hp, drill-r50 | identical |
| hub customers | demo-felhom, demo-hp, drill-r50, peti-felhom, tester-1 | without peti-felhom |
| hub rows naming peti | — | events 124, notification_log 80, host_deletions 1, customer_resets 1 (audit, by design), app_log_issues 8 (R-244) |
## A5 — documents changed
`runbooks/target-selection.md` (Tier 2 = DooPlex + ep0, ep0's reason rewritten, Peti's section retired); the
unprompted-work rule (4 identical copies); `03`, `06`, `00-capability-map`; runbooks `TASK-identity-only-escrow`
(obsolete note), `RUNBOOK-vzdump-target-move` (row 9), `RUNBOOK-island-migration`; retired banners on
`pilot/PETI-tester-agreement.md`, `RUNBOOK-peti-return`, `RUNBOOK-peti-pbsdr`; register (PETI closed as retired,
R-530, R-244, R-600 annotated); CONTEXT; STATUS; memory notes. Historic audits, tests and CHANGELOGs keep their text.
## Claims in the brief
1. *Peti's backup is on ep0* — **WRONG.** ep0 held nothing of his; the item was a Storage Box sub-account.
2. *The hub's host delete removes the WireGuard peer* — **not testable here:** the host was deleted in July and
ep0 carries no peer for it now; whether that delete removed one is not recorded (R-600 annotated).
3. *The backup holds no user data* — **TRUE, and stronger:** it held no backup at all (one 81-byte key file).
@@ -0,0 +1,4 @@
# A1 control BEFORE — demo boxes' off-site status from their latest hub report
demo-felhom 2026-09-25 08:44:15 {'snapshot_count': 11, 'repo_size_bytes': 147483, 'last_run': '2026-09-25T02:15:46Z', 'last_status': 'ok', 'state': None, 'escrow_state': 'escrowed'}
demo-hp 2026-09-25 08:44:17 {'snapshot_count': 90, 'repo_size_bytes': 215007449, 'last_run': '2026-09-25T02:18:41Z', 'last_status': 'ok', 'state': None, 'escrow_state': 'escrowed'}
tester-1 2026-09-17 00:28:20 {'snapshot_count': 0, 'repo_size_bytes': 0, 'last_run': '2026-09-16T21:09:00Z', 'last_status': 'error', 'state': None, 'escrow_state': 'escrowed'}
@@ -0,0 +1,75 @@
# A1 documents — whole-word pattern: \bpeti\b|peti-felhom|\bPETI\b|Peti's|Petinek|Petié|petis — 2026-09-25T08:54:25+00:00
positive control (pattern on a known line: target-selection.md): 7
negative control (same pattern on app-catalog templates): 0
== non-history files (excluding audits/, tests/ evidence, CHANGELOGs, REPORT-*.md, pilot/ runbooks of the past):
felhom.eu/STATUS.md
felhom.eu/CONTEXT.md
felhom.eu/documentation/architecture/06-offsite-connectivity.md
felhom.eu/documentation/architecture/07-backup-architecture.md
felhom.eu/documentation/architecture/03-host-agent.md
felhom.eu/documentation/architecture/_recovery-inventory-2026-07-28.md
felhom.eu/documentation/architecture/10-localisation.md
felhom.eu/documentation/runbooks/RUNBOOK-byo-deployment.md
felhom.eu/documentation/runbooks/RUNBOOK-island-migration.md
felhom.eu/documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md
felhom.eu/documentation/runbooks/publish-train-rules.md
felhom.eu/documentation/runbooks/RUNBOOK-onboarding-draft-v4.md
felhom.eu/documentation/runbooks/RUNBOOK-manual-build.md
felhom.eu/documentation/runbooks/RUNBOOK-ep0-datastore-volume-2026-07-27.md
felhom.eu/documentation/architecture/00-capability-map.md
felhom.eu/documentation/runbooks/TASK-identity-only-escrow.md
felhom.eu/documentation/runbooks/target-selection.md
felhom.eu/documentation/pilot/RUNBOOK-peti-pbsdr-2026-07-11.md
felhom.eu/documentation/pilot/RUNBOOK-publish-0.85-0.120-2026-07-12.md
felhom.eu/documentation/pilot/RUNBOOK-publish-0.79-0.110-2026-07-10.md
felhom.eu/documentation/pilot/PETI-tester-agreement.md
felhom.eu/documentation/pilot/RUNBOOK-peti-return-2026-07-13.md
felhom.eu/documentation/pilot/RUNBOOK-publish-0.90-0.143-2026-07-18.md
felhom.eu/documentation/pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md
felhom.eu/documentation/pilot/RUNBOOK-publish-0.81-0.113-2026-07-11.md
felhom.eu/documentation/pilot/GO-LIVE-PACKAGE.md
felhom.eu/documentation/pilot/DRILL-GL6-2026-07-08.md
felhom.eu/.claude/rules/unprompted-work.md
felhom.eu/documentation/backlog/ROADMAP.md
felhom.eu/documentation/backlog/CLOSED-ITEMS.md
felhom.eu/hub/internal/notify/templates.go
felhom.eu/documentation/backlog/OPEN-ITEMS.md
felhom.eu/hub/internal/tenantsync/client_test.go
felhom.eu/hub/internal/monitor/deadline_unbound_test.go
felhom.eu/hub/internal/web/rollup.go
felhom.eu/hub/internal/web/rollup_test.go
felhom.eu/hub/internal/web/pbsdr_generation_test.go
felhom.eu/hub/internal/web/pbsdr_poke_test.go
felhom.eu/hub/internal/web/pbsdr_test.go
felhom.eu/hub/internal/web/render_test.go
felhom.eu/hub/internal/web/appliances.go
felhom-controller/.claude/rules/unprompted-work.md
felhom-controller/CONTEXT.md
felhom-agent/CONTEXT.md
felhom-agent/internal/pbsdr/seed_reassert_test.go
felhom-agent/internal/pbsdr/manager_test.go
felhom-agent/internal/pbsdr/rearm_test.go
app-catalog-felhom.eu/.claude/rules/unprompted-work.md
== workspace level:
.claude/rules/unprompted-work.md
.claude-memory/campaign3-nightrun-2026-07-11.md
.claude-memory/agent-update-needs-signed-job.md
.claude-memory/current-state-2026-07-11.md
.claude-memory/current-state-2026-07-21.md
.claude-memory/hub-closing-bundle-2026-07-13.md
.claude-memory/pbs-tier-provisioning-spike.md
.claude-memory/drtier-by-default-2026-07-12.md
.claude-memory/peti-return-p1-stop-2026-07-13.md
.claude-memory/spike-lan-discovery-r6-2026-07-18.md
.claude-memory/restore-test-fills-thin-pool.md
.claude-memory/gl2-byo-install-profile-shipped.md
.claude-memory/customer-claim-arc-2026-07-12.md
.claude-memory/polish-batch-2026-07-13.md
.claude-memory/observability-pass-shipped.md
.claude-memory/r50-island-bridge-go-2026-07-25.md
.claude-memory/fleet-identity-peti-vs-tester1.md
.claude-memory/power-outage-audit-2026-07-22.md
.claude-memory/hub-offsite-provisioning-e2e.md
.claude-memory/nas-network-storage.md
.claude-memory/remote-applog-diagnostics-shipped.md
.claude-memory/MEMORY.md
@@ -0,0 +1,35 @@
# A1 ep0 read-only — 2026-09-25T08:57:45+00:00
felhom-hetzner
== PBS namespaces (live store /mnt/pbs-datastore):
demo-felhom
demo-hp
tester-1
demo-felhom: 220K snapshots=2
demo-hp: 540K snapshots=2
tester-1: 368K snapshots=2
== grep peti anywhere in PBS config:
== wg show (peers):
peer: snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc=
allowed ips: 10.77.0.3/32
latest handshake: 57 seconds ago
peer: KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo=
allowed ips: 10.77.0.2/32
latest handshake: 1 minute, 39 seconds ago
peer: yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00=
allowed ips: 10.77.0.250/32
latest handshake: 18 hours, 59 minutes, 18 seconds ago
peer: kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M=
allowed ips: 10.77.0.4/32
latest handshake: 43 days, 2 hours, 41 minutes, 56 seconds ago
== wg config peers:
8:PublicKey = KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo=
9:AllowedIPs = 10.77.0.2/32
12:PublicKey = yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00=
13:AllowedIPs = 10.77.0.250/32
16:PublicKey = snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc=
17:AllowedIPs = 10.77.0.3/32
20:PublicKey = kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M=
21:AllowedIPs = 10.77.0.4/32
== grep peti in /etc and /root, /srv:
/etc/lvm/lvm.conf
1053:# When there are competing read-only and rea
@@ -0,0 +1,7 @@
# A1 Hetzner (the hub's own token, read-only) — 2026-09-25T08:57:06.131163+00:00
pool box id=611714 name=storage-box-pool-1 user=u629488 stats: size=3233808384 size_data=3120562176 size_snapshots=113246208
sub id=273581 user=u629488-sub1 home=felhom-demo-felhom labels={'felhom-customer': 'demo-felhom'} created=2026-07-18T16:53:43Z desc=felhom offsite demo-felhom
sub id=269130 user=u629488-sub2 home=felhom-peti-felhom labels={'felhom-customer': 'peti-felhom'} created=2026-07-10T07:29:44Z desc=felhom offsite peti-felhom
sub id=275124 user=u629488-sub3 home=felhom-demo-hp labels={'felhom-customer': 'demo-hp'} created=2026-07-21T16:01:40Z desc=felhom offsite demo-hp
sub id=311327 user=u629488-sub4 home=felhom-tester-1 labels={'felhom-customer': 'tester-1'} created=2026-09-16T15:03:56Z desc=felhom offsite tester-1
all boxes visible to this token: [(611714, 'storage-box-pool-1')]
@@ -0,0 +1,10 @@
# A1b ep0 second storage box mount — 2026-09-25T09:02:23+00:00
inactive
ls: cannot access '/mnt/pbs-storagebox': No such file or directory
== felhom-peti-spike: du: cannot access '/mnt/pbs-storagebox/felhom-peti-spike': No such file or directory
ls: cannot access '/mnt/pbs-storagebox/felhom-peti-spike': No such file or directory
== felhom-demo: du: cannot access '/mnt/pbs-storagebox/felhom-demo': No such file or directory
ls: cannot access '/mnt/pbs-storagebox/felhom-demo': No such file or directory
== spike-sub: du: cannot access '/mnt/pbs-storagebox/spike-sub': No such file or directory
ls: cannot access '/mnt/pbs-storagebox/spike-sub': No such file or directory
== is it a PBS datastore anywhere?
@@ -0,0 +1,34 @@
# A2 listing of Peti's sub-account home (u629488-sub2), read-only — 2026-09-25T08:59:31+00:00
$ pwd
Warning: Permanently added '[u629488-sub2.your-storagebox.de]:23' (ED25519) to the list of known hosts.
/home
$ df
Filesystem 1K-blocks Used Available Use% Mounted on
u629488-sub2 1073631360 3047552 1070583808 1% /home
$ du -s .
2 .
$ ls -la
total 3
drwxr-xr-x 3 u629488-sub2 1006 3 Jul 10 09:49 .
dr-x--x--x 7 root root 11 Jul 10 07:30 ..
drwx------ 2 u629488-sub2 1006 3 Jul 10 09:49 .ssh
$ ls -la felhom-repo
/usr/bin/ls: cannot access 'felhom-repo': No such file or directory
$ du -s felhom-repo
/usr/bin/du: cannot access 'felhom-repo': No such file or directory
$ tree -a -L 2 felhom-repo
felhom-repo [error opening dir]
0 directories, 0 files
$ ls -la felhom-repo/snapshots
/usr/bin/ls: cannot access 'felhom-repo/snapshots': No such file or directory
$ ls -la felhom-repo/keys
/usr/bin/ls: cannot access 'felhom-repo/keys': No such file or directory
$ ls -la .ssh
total 2
drwx------ 2 u629488-sub2 1006 3 Jul 10 09:49 .
drwxr-xr-x 3 u629488-sub2 1006 3 Jul 10 09:49 ..
-rw------- 1 u629488-sub2 1006 81 Jul 10 09:49 authorized_keys
$ stat -c %n .ssh/authorized_keys .ssh/nonexistent-control-zq
.ssh/authorized_keys
/usr/bin/stat: cannot statx '.ssh/nonexistent-control-zq': No such file or directory
@@ -0,0 +1,10 @@
# A2 Peti's LAST report to the hub: id 12564 received 2026-07-15 08:39:00 UTC; controller 0.115.0
deployed apps: ['rallly']
offsite status: {"enabled": true, "escrow_state": "pending", "snapshot_count": 0, "repo_size_bytes": 0, "quota_gb": 50}
backup section keys: ['enabled', 'last_db_dump', 'snapshot_count', 'repo_size_mb', 'integrity_ok']
storage: [('/', 'SSD')]
newest report WITH an offsite object: 12564 2026-07-15 08:39:00 {"enabled": true, "escrow_state": "pending", "snapshot_count": 0, "repo_size_bytes": 0, "quota_gb": 50}
== all 482 Peti reports (2026-07-10 09:49:09 .. 2026-07-15 08:39:00 UTC): 482 carry an offsite object; MAX offsite snapshot_count=0, MAX repo_size_bytes=0
apps ever deployed (report count): {'rallly': 481, 'calibre-web': 73}
offsite-related events for Peti: 1
@@ -0,0 +1,3 @@
# A3-1 2026-09-25T08:59:13.275582+00:00 TARGET CHECK: u629488-sub2 felhom-peti-felhom {'felhom-customer': 'peti-felhom'} access: {'samba_enabled': False, 'ssh_enabled': True, 'webdav_enabled': False, 'reachable_externally': True, 'readonly': False}
action 657507966278067 reset_subaccount_password running
action status: success (password stored out-of-band in the session scratchpad, 0600)
@@ -0,0 +1,9 @@
# A3-2 empty Peti's home (target: u629488-sub2 = sub 269130, label peti-felhom; its home is /home) — 2026-09-25T08:59:59+00:00
$ rm .ssh/authorized_keys
$ rmdir .ssh
$ ls -la
total 2
drwxr-xr-x 2 u629488-sub2 1006 2 Sep 25 09:00 .
dr-x--x--x 7 root root 11 Jul 10 07:30 ..
$ du -s .
1 .
@@ -0,0 +1,28 @@
# A3-3 hub delete PREVIEW for peti-felhom — 2026-09-25T09:00:09+00:00
{
"claim_present": true,
"customer_id": "peti-felhom",
"customer_name": "Peti Proxmox",
"dr_recipe_present": true,
"has_config": true,
"host_count": 0,
"hosts": [],
"offsite_enabled": true,
"offsite_identifier": "u629488-sub2",
"offsite_type": "shared",
"one_time_secret": true,
"online_host_present": false,
"pbs_tenancy_configured": true,
"pending_journal": null,
"residue": {
"app_log_tails": 1,
"app_telemetry": 1036,
"appliance_registrations": 0,
"log_tail_requests": 0,
"notification_prefs": 0,
"reports": 482,
"selfbind_tokens": 0
},
"residue_total": 1519,
"superseded_blobs": 0
}
@@ -0,0 +1,12 @@
# A3-4 hub customer DELETE cascade for peti-felhom — 2026-09-25T09:00:21Z
POST /configs/peti-felhom/delete -> 303 Location: /configs?flash=deleted
== the hub's own log:
2026/09/25 11:00:21 [INFO] customer DELETE cascade started for peti-felhom (journal #20, 0 host(s))
2026/09/25 11:00:29 [offsite] deprovisioned shared sub-account 269130 for peti-felhom (repo data destroyed)
2026/09/25 11:00:29 [INFO] reset peti-felhom: offsite deprovisioned (repo data destroyed)
2026/09/25 11:00:29 [INFO] tenantsync: deprovision ok for peti-felhom (ns=peti-felhom, existed=false)
2026/09/25 11:00:29 [INFO] reset peti-felhom: PBS tenancy deprovisioned
2026/09/25 11:00:29 [INFO] [claim] reset to unclaimed for peti-felhom (customer RESET) — next onboarding mints a fresh code
2026/09/25 11:00:33 [INFO] delete peti-felhom: residue purged (reports=482 app_telemetry=1036 app_log_tails=1 log_tail_requests=0 notif_prefs=0 selfbind_tokens=0 appliance_registrations=0)
2026/09/25 11:00:33 [INFO] customer DELETE cascade COMPLETE for peti-felhom (journal #20) — full teardown
@@ -0,0 +1,28 @@
# A4 AFTER — 2026-09-25T09:00:49.893123+00:00
pool box 611714 stats: size=3233808384 size_data=3120562176 size_snapshots=113246208
sub id=273581 user=u629488-sub1 home=felhom-demo-felhom labels={'felhom-customer': 'demo-felhom'}
sub id=275124 user=u629488-sub3 home=felhom-demo-hp labels={'felhom-customer': 'demo-hp'}
sub id=311327 user=u629488-sub4 home=felhom-tester-1 labels={'felhom-customer': 'tester-1'}
== login as the deleted sub-account must now FAIL:
Permission denied, please try again.
login_rc=0
== ep0 after:
demo-felhom
demo-hp
tester-1
demo-felhom snapshots=2 du=220K
demo-hp snapshots=2 du=540K
tester-1 snapshots=2 du=368K
KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo= 10.77.0.2/32
yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00= 10.77.0.250/32
kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M= 10.77.0.4/32
snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc= 10.77.0.3/32
== hub census AFTER: every table, every text column, rows matching 'peti'
notification_log: 80
events: 124
app_log_issues: 8
host_deletions: 1
customer_resets: 1
control — demo rows still present: {'reports': 15337, 'customer_configs': 2, 'events': 2262}
hosts: ['demo-felhom-8363b5', 'demo-hp-bb76ea', 'drill-r50-0a4f9a']
customer_configs: ['demo-felhom', 'demo-hp', 'drill-r50', 'tester-1']
+2
View File
@@ -392,3 +392,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-680** | **The box did not remember a failed update step (P2).** Controller v0.271.0: an undone or held step is recorded in app.yaml (`failed_update_step`, tied to the ladder's print); the automatic leg skips it until the catalog's ladder changes; a person can still press. Live on 9202: vikunja undone night 1, skipped `failed_before` night 2, re-tried after the catalog re-tested it. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/`; `B/redproofs/R680-*` |
| **R-678** | **After a step ended `done`, steps-left and the badge stayed stale (P3).** Controller v0.271.0: the update re-reads the app's catalog fields BEFORE it says done (and after an undo). Live on 9202: every automatic step's page read current at the leg's end. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `B/redproofs/R678-*` |
| **R-643** | **The ruled chain left the automatic update leg at most 15 minutes a night (P2).** Decision 20, built in controller v0.271.0: the full-system backup's gate defers while the leg runs, until W+5h (then only for a step in flight, cap W+5h30m — decision 31); the leg starts no step at or after W+5h; one shared constant. Unit + red-proof (`TestD20_GateWaitsForTheLeg`); live on the demo boxes: see the night record Part D. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/B/redproofs/D20-*` |
| **PETI** | **`peti-felhom` deliberately not migrated; parked until the tester reinstalls.** **RETIRED 2026-09-25 (operator ruling):** the tester wiped his server and the box will not return. Removed through the hub's customer delete (journal #20): the customer record and 1,519 residue rows, Storage Box sub-account `u629488-sub2` (id 269130) — which held only one 81-byte `authorized_keys`, never a repository (all 482 reports: 0 off-site snapshots, 0 bytes); ep0 held nothing of it (no PBS namespace, no WireGuard peer). The audit trail stays by design. | retired 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/RETIRE-peti-2026-09-25.md` |
| **R-686** | **The automatic update leg is not resumed after a controller restart during the night.** **RULED 2026-09-25 (operator, option B):** the apps the leg had not reached wait for the next night; nothing is built — `09` §3 decision 34. The page-line side effect (a resumed step's `last_auto_update` not written) stays as measured. | ruled 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/night3-kill/`, `night4-power/` |
File diff suppressed because one or more lines are too long
@@ -1,5 +1,8 @@
# Felhom pilot — tester agreement (Peti)
> **RETIRED 2026-09-25 (operator ruling).** Peti's box will not return — the tester wiped his server. Its hub customer and off-site sub-account are gone (`audits/RETIRE-peti-2026-09-25.md`). This document is kept as a record; do not act on it.
> The terms of the first external pilot: what Peti runs, what Felhom can and cannot do on his
> hardware, the honest limitations of the pilot, and his exit rights. Lives at
> `felhom.eu/documentation/pilot/PETI-tester-agreement.md`.
@@ -1,5 +1,8 @@
# RUNBOOK — Peti PBS DR tier enable (the epic's payoff) — 2026-07-11
> **RETIRED 2026-09-25 (operator ruling).** Peti's box will not return — the tester wiped his server. Its hub customer and off-site sub-account are gone (`audits/RETIRE-peti-2026-09-25.md`). This document is kept as a record; do not act on it.
> **State when written:** agent **0.80.0 published** (sha `f2ba62ca6aca6e24a8d08706ea0dc3ae43940a63e9bdf57d4606ad1756cf06d2`,
> anon round-trip verified; the demo runs the IDENTICAL bytes, adoption-proven). Hub v0.44.0 live;
> ep0 tenantsync surface live. Peti's box: `peti-felhom-86d37d`, agent 0.79.0, **wg_tunnel DISABLED
@@ -1,5 +1,8 @@
# RUNBOOK — Peti's return: convergence, the parked train, the first real-customer onboarding, and the credential rotations
> **RETIRED 2026-09-25 (operator ruling).** Peti's box will not return — the tester wiped his server. Its hub customer and off-site sub-account are gone (`audits/RETIRE-peti-2026-09-25.md`). This document is kept as a record; do not act on it.
<!--
Class: operational runbook (GL pattern). Actors: Viktor (operator, signs ops, leads the call),
CC (read-only verification + supervised host steps over SSH), and — for the first time — PETI
@@ -7,7 +7,7 @@
> **v1.19.0**. Params are LAW (spike-validated): `vmbr9` portless, host `169.254.253.1/30`, guest
> `169.254.253.2/30` on `net1`, `listen_addr=169.254.253.1:8443`, `lan_resolver.host_ip=<LAN IP>`.
>
> **Scope:** ONE host with ONE customer guest (the current fleet shape). A cluster (Peti, 2 nodes) is
> **Scope:** ONE host with ONE customer guest (the current fleet shape). A cluster (Peti, 2 nodes — RETIRED 2026-09-25) is
> **out** — it needs bridge parity on both nodes / an SDN vnet; that is Phase C, its own runbook.
>
> **The apps never stop.** Only the management channel moves; `cloudflared` + the app containers keep
@@ -435,7 +435,7 @@ label. Filed under E-2.
| 6 | **Absent-target policy** per §6: decide fallback-vs-fail, and if fallback, alarm that protection is degraded rather than reporting a healthy tier. |
| 7 | **Retention and space accounting** on a drive the customer also uses — today `keep-last=3` competes with customer data with no reservation and no ceiling. |
| 8 | The honest **single-drive label**. |
| 9 | **Migrating the remaining fleet** — `peti-felhom` and any box not covered here. |
| 9 | **Migrating the remaining fleet** — any box not covered here (`peti-felhom` was RETIRED 2026-09-25 and needs nothing). |
| 10 | Naming guard: `backupIsPBS` matches the substring `pbs`, so a target so named would be mislabelled to the customer as offsite (§1.5). |
---
@@ -41,7 +41,7 @@ stays valid for old history only.
## 4. Deploy / live
Build v0.80.0 → felhom-pve (demo agent) → healthy, 56/56 caps. Publish 0.80.0 to Gitea + bump the hub
Day-0 agent manifest (the publish-train pattern; the operator-sign step is Viktor's 🛑 as per GL-1).
**Peti's agent update path:** the agent self-update is operator-signed + pinned — confirm from the go-live
**[OBSOLETE 2026-09-25 — Peti's box RETIRED; no BYO box remains.]** **Peti's agent update path:** the agent self-update is operator-signed + pinned — confirm from the go-live
record how a BYO agent updates (self-update channel armed at his install? operator pubkey file was NOT
passed on his install form) — if his box cannot self-update the agent, REPORT must say so and the ceremony
waits for the next Peti-touch window (he runs one update command). Do not improvise a new update path.
+22 -16
View File
@@ -6,11 +6,12 @@
## The rule
> **Two machines are protected: `DooPlex` and Peti's box. Everything else is disposable.**
> Operator decision **D-d**, 2026-08-02. DooPlex because it holds Gitea, the hub, the backups and the
> registry — everything else rebuilds from it, and it rebuilds from nothing. Peti's box because there
> is a real person behind it. **Every other box, both demo boxes included, may be broken or
> reinstalled freely.**
> **Two machines are protected: `DooPlex` and `ep0`. Everything else is disposable.**
> Operator decision **D-d**, 2026-08-02, and the ep0 ruling of 2026-08-03. DooPlex because it holds
> Gitea, the hub, the backups and the registry — everything else rebuilds from it, and it rebuilds from
> nothing. ep0 for the reason below. **Every other box, both demo boxes included, may be broken or
> reinstalled freely.** *(Peti's box was the second protected machine until it was RETIRED on
> 2026-09-25 — operator ruling; `audits/RETIRE-peti-2026-09-25.md`.)*
>
> **This is a correction, not a relaxation.** The earlier posture was costing whole sessions to
> caution and pushing drills onto DooPlex — the one machine that should never host them. If you are
@@ -70,7 +71,7 @@ in `documentation/backlog/OPEN-ITEMS.md`, never `--no-verify`.
|---|---|---|
| **0 — disposable. Reach here first.** | Exists to be broken; reinstalling is a routine afternoon, not an incident. **A drill that needs a victim uses one of these.** | `demo-hp` (t740), `demo-felhom` (N100) |
| **1 — create and destroy freely** | Throwaway VMs, guests, scratch customers — **hosted on a Tier 0 machine** | drill VMs, scratch guests |
| **2 — protected. Never a drill target.** | Losing it costs the recovery chain or a real relationship | **DooPlex**, **Peti's cluster**, **`ep0`** (operator ruling 2026-08-03) — and nothing else |
| **2 — protected. Never a drill target.** | Losing it costs the recovery chain or a real relationship | **DooPlex**, **`ep0`** (operator ruling 2026-08-03) — and nothing else (Peti's cluster RETIRED 2026-09-25) |
**DooPlex is Tier 2 because it *is* the recovery chain** — hub, Gitea, registry, PBS, k3s + Longhorn.
Everything else rebuilds from it; it rebuilds from nothing. A bad moment in a DR drill there costs the
@@ -79,10 +80,14 @@ thing under test, the source of truth for it, and the backups, at once.
**`ep0` is Tier 2 — PROTECTED. Operator ruling, 2026-08-03.** D-d named two protected machines and did
not name ep0 either way, so this page carried the question in writing for two days and read it the
narrow way meanwhile (not protected, but not wipeable). The ruling settles it and **extends D-d's
protected list to three machines**: DooPlex, Peti's cluster, ep0.
protected list to three machines**: DooPlex, Peti's cluster, ep0. **Since 2026-09-25 (Peti's box
retired) it is two: DooPlex and ep0.**
The reason it was never really in doubt: ep0 holds the **PBS-DR datastore and the restic copy of a
real customer's data**, which is the only off-premises copy that exists. So *deleting datastores,
The reason, as it stands on 2026-09-25: ep0 holds the **PBS-DR datastore with the demo boxes' (and
tester-1's) whole-guest copies, the WireGuard hub every box's off-site path runs through, and the
operator's out-of-band path**; the Hetzner Storage Boxes hold each box's restic copy. There is no paying
customer yet, but ep0 is the only off-premises tier of the whole product, and a mistake there cannot be
rebuilt from DooPlex. So *deleting datastores,
prune jobs, tunnel config or nftables rules* was already forbidden by what it would destroy; the
ruling makes the classification say so plainly instead of leaving each session to re-derive it.
**Reads are fine** — including the ordinary off-site read a restore-test performs (R-86) — and it is
@@ -158,12 +163,12 @@ against.
builds if `df -h /mnt/5_hdd /` shows either over 90 %. **Because it is the recovery chain and a live
k3s node.**
### `Peti's cluster` (`peti-felhom`) — **Tier 2**
### `Peti's cluster` (`peti-felhom`) — **RETIRED 2026-09-25** (operator ruling)
**Do not touch, at all.** A real external pilot with a real person behind it; its whole-guest backup
still shares a device with its guest, so a drive failure is **offsite-only recovery**. Deliberately not
migrated, parked until the tester reinstalls (`PETI` in `backlog/OPEN-ITEMS.md`). Currently DOWN, no
enrolled host. No access route from DooPlex, and nothing here needs one.
No longer protected and no longer a machine this project knows: the tester wiped his server and it will not
return. Its hub customer, Storage Box sub-account (`u629488-sub2`, which never held a backup) and records were
removed through the hub's own customer delete; ep0 held nothing of it. The audit trail stays (events, the
deletion and reset tombstones). Record: `audits/RETIRE-peti-2026-09-25.md`.
### `ep0` (`felhom-hetzner`, `ep0.felhom.eu`) + the Hetzner Storage Boxes — **Tier 2, PROTECTED** (operator ruling 2026-08-03)
@@ -212,8 +217,9 @@ The fixture would have shown a controller nobody installs.
from `backlog/OPEN-ITEMS.md`; **no direct access attempted or confirmed.** Tier 2 by what they hold.
- **`felhotest`** (legacy, `ssh -p 33022 kisfenyo@router.abonet.hu`) — **`Connection refused`**
2026-07-30; that was the only route tried. Untiered; assume nothing.
- **Peti's cluster** hardware/storage layout — unverified. Tier 2 rests on the relationship, which
needs no verification.
- **Peti's cluster** — RETIRED 2026-09-25 (operator ruling: the tester wiped his server; it will not
return). Its hub records, its Storage Box sub-account and folder are gone; ep0 never held anything of
it. `audits/RETIRE-peti-2026-09-25.md`.
## The gap this page closes