From eb1c56a981109a7fa1843dd2d72056ee6f705e9b Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 25 Sep 2026 11:04:52 +0200 Subject: [PATCH] Peti's box retired (operator ruling 2026-09-25): hub customer + Storage Box sub-account removed, ep0 held nothing; protected list = DooPlex + ep0; decision 34 (no leg resume, R-686 closed); R-688 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .claude/rules/unprompted-work.md | 2 +- CONTEXT.md | 9 +++ .../architecture/00-capability-map.md | 2 +- documentation/architecture/03-host-agent.md | 2 +- .../architecture/06-offsite-connectivity.md | 3 +- .../architecture/09-update-architecture.md | 8 ++ .../audits/RETIRE-peti-2026-09-25.md | 81 +++++++++++++++++++ .../A1-demo-control-before.txt | 4 + .../retire-peti-2026-09-25/A1-documents.txt | 75 +++++++++++++++++ .../retire-peti-2026-09-25/A1-ep0-before.txt | 35 ++++++++ .../A1-hetzner-before.txt | 7 ++ .../A1b-ep0-storagebox-mount.txt | 10 +++ .../retire-peti-2026-09-25/A2-listing.txt | 34 ++++++++ .../A2-peti-last-report.txt | 10 +++ .../A3-1-reset-password.txt | 3 + .../A3-2-empty-home.txt | 9 +++ .../A3-3-hub-preview.txt | 28 +++++++ .../A3-4-hub-delete.txt | 12 +++ .../retire-peti-2026-09-25/A4-after.txt | 28 +++++++ documentation/backlog/CLOSED-ITEMS.md | 2 + documentation/backlog/OPEN-ITEMS.md | 9 +-- documentation/pilot/PETI-tester-agreement.md | 3 + .../pilot/RUNBOOK-peti-pbsdr-2026-07-11.md | 3 + .../pilot/RUNBOOK-peti-return-2026-07-13.md | 3 + .../runbooks/RUNBOOK-island-migration.md | 2 +- .../RUNBOOK-vzdump-target-move-2026-07-29.md | 2 +- .../runbooks/TASK-identity-only-escrow.md | 2 +- documentation/runbooks/target-selection.md | 38 +++++---- 28 files changed, 398 insertions(+), 28 deletions(-) create mode 100644 documentation/audits/RETIRE-peti-2026-09-25.md create mode 100644 documentation/audits/retire-peti-2026-09-25/A1-demo-control-before.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A1-documents.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A1-ep0-before.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A1-hetzner-before.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A1b-ep0-storagebox-mount.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A2-listing.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A2-peti-last-report.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A3-1-reset-password.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A3-2-empty-home.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A3-3-hub-preview.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A3-4-hub-delete.txt create mode 100644 documentation/audits/retire-peti-2026-09-25/A4-after.txt diff --git a/.claude/rules/unprompted-work.md b/.claude/rules/unprompted-work.md index c28407e8..80157dd2 100644 --- a/.claude/rules/unprompted-work.md +++ b/.claude/rules/unprompted-work.md @@ -20,7 +20,7 @@ unconditional: true **Not yours, ever, without a task file or an operator word:** money; anything that changes risk to customer data; anything that changes a promise the product makes to a customer; anything that reverses a documented design decision (`documentation/architecture/` — a design decision is not a -defect, R-370); anything on DooPlex, Peti's box or ep0; baking or vouching a golden; promoting a +defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a catalog version; a new external dependency. ## 2. When you may decide instead of ask (operator grant, 2026-09-14) diff --git a/CONTEXT.md b/CONTEXT.md index a64b3dfb..d0568c2c 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,15 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## Rulings 2026-09-25 (operator) + +- **Peti's box (`peti-felhom`, host `peti-felhom-86d37d`) is RETIRED.** It will not return — the tester wiped his + server. Its off-site backup holds no user data (operator's words; measured: all 482 of its reports show 0 off-site + snapshots, 0 bytes) and is deleted. It leaves the protected list; DooPlex and ep0 stay Tier 2. Record: + `documentation/audits/RETIRE-peti-2026-09-25.md`. +- **A controller restart during the night's update leg does not resume the leg** (R-686, option B) — `09` §3 + decision 34. Nothing is built. + ## 2026-09-25 (night shift) — agents 0.133.0 + 0.134.0 delivered, demo-hp backs up, automatic updates shipped (controller v0.271.0) - **Agent v0.133.0 and v0.134.0 delivered to both demo boxes** by CC-signed `agent_update` jobs (ruling 1); restore diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 169a79c4..b887488c 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -45,7 +45,7 @@ | Scenario | Components | Status | Evidence | Gap / roadmap | |---|---|---|---|---| | Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box | installer, agent, hub, golden | **PROVEN-LIVE** (nested VM) | `DRILL-day0-vm-2026-07-12`, `DRILL-day0-take2-2026-07-12` | First firing on real customer hardware pending → R-1 | -| BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1 | +| BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | the first real BYO run of the current path is still owed (R-1) — the planned venue, Peti's clean-slate reinstall, is gone: Peti's box was RETIRED 2026-09-25 | | **The installer is PUBLISHED, not pushed — the artifact that runs as root on a virgin box is served from a version-controlled ref, and rolling back is one act** | scripts **v1.23.0** + `manifests/webpage.yaml` (R-110, operator ruling option (b)) | **PROVEN-LIVE (2026-08-03)** | `scripts/CHANGELOG.md` v1.23.0 + `REPORT.md`. **Proven by HTTP against the real URL, not from a pod's filesystem.** *Scenario A:* a real push to `main` without moving the tag left the served script **byte-identical** (`sha256 2f859555…`), and a marker comment planted in that very commit was **absent** from the served bytes, while the website tree advanced to the new commit in the same observation — both halves of the split in one measurement. *Scenario B:* moving the tag published in **~40 s** (`sha → ea2b4aa9…`, marker present) and moving it back restored **exactly** the pre-publish sha. `https://felhom.eu/` returned 200 throughout. *P-A, measured BEFORE the manifest was touched because the model rests on it:* git-sync v4.4.0 follows a tag **and notices a moved one** (`update required … local: remote:` → `updated successfully`) | **Two syncs, deliberately: the WEBSITE still tracks `main`.** Pinning both would turn every copy edit into a release, which makes the release meaningless and the site slow to fix. **Publish** = cut `installer-v` + bump the manifest `--ref` + sync; **roll back** = move the tag back, which needs **no ArgoCD sync and no deploy**. Continuity is structural rather than lucky: both trees are seeded by init containers so a fresh pod is not Ready until the tag is checked out, and `maxUnavailable` rounds to 0 on one replica, so a failed scripts-init leaves the OLD pod serving — the failure direction is *no update*, never *no `/scripts/`*. **The URL never carried a ref**, so the bootstrap script and the hub's day-0 command follow the tag with no edit and **no hub change**. The installer's own sixteen run-time fetches are a separate channel pinned to the AGENT's version (**R-183**), because they are the agent's configs and not this repo's — leaving them on `main` would have made the whole change cosmetic | | Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row | **Second board, 2026-07-21 (demo-hp, HP t740 / Ryzen V1756B / AMI M42):** the whole chain ran on virgin hardware in one pass — armed install → self-registration as an unclaimed appliance → operator bind → day-0 → running guest 9201 + agent 0.92.1 as `demo-hp-bb76ea`. **The shim loader booted with Secure Boot ENABLED**, which retires the assumption that Felhom installs need SB off — that was an N100-firmware workaround. The exact-serial disk filter took the system SSD and left the box's 1TB NVMe untouched/unenrolled on hardware it had never seen. Two failures filed rather than smoothed over: **R-59** (no DHCP → the installer baked a static fallback instead of aborting) and **R-61** (baked root password unknowable → no console access). | Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (`.rootpw.txt`) | scripts v1.24.0 (`scripts/iso/felhom-bootstrap.sh` `network_gate`/`sweep_nics`, `build-felhom-iso.sh` rootpw emission) | **PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install)** | `audits/SPIKE-firstboot-nic-sweep-2026-07-22.md` — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and **self-registered at the hub unaided in under a minute**; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash | R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (`state.json` gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched | diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index a649010f..6d6e8aa3 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -334,7 +334,7 @@ per Part 1: **snapshot** (LVM-thin, transient, whole-guest rollback — not a ba > pool has room. > > **Delivered 2026-09-24 night** to demo-felhom and demo-hp by CC-signed `agent_update` jobs; the restore test is -> back ON on both (the `-1` config kept as `agent.json.night-0925-off`). Peti's box has not received it. +> back ON on both (the `-1` config kept as `agent.json.night-0925-off`). Peti's box never received it — it was RETIRED 2026-09-25 and will not return. > **A whole-box backup that cannot fit is skipped with a reason, before anything starts (R-685, agent v0.134.0).** > Before a vzdump to a LOCAL (non-PBS) target: free ≥ the guest's newest archive on that target × 1.25 + 1 GiB, diff --git a/documentation/architecture/06-offsite-connectivity.md b/documentation/architecture/06-offsite-connectivity.md index 8e19a4d5..938d3640 100644 --- a/documentation/architecture/06-offsite-connectivity.md +++ b/documentation/architecture/06-offsite-connectivity.md @@ -314,7 +314,8 @@ caveats, recorded not papered over: **(1)** this SIM was handed a **public mobil stays a *retest-on-a-CGNAT-SIM-when-available* follow-up (low risk: mapping-hold is NAT-tier-agnostic by mechanism). **(2)** a dual-stack mobile uplink made `wg-quick` prefer the endpoint **AAAA and ride un-NATed IPv6** until v4 was forced — functionally fine, but see §4.2. The deferred second-ISP -vantage (Peti VM 110) remains the thorough confirmation but no longer gates anything. Runbook: +vantage (Peti VM 110) is gone — Peti's box was RETIRED 2026-09-25 — so a second-ISP confirmation needs another +venue; it gates nothing. Runbook: `RUNBOOK-s3-cgnat-smoke`. **Open sub-decisions (deferred by design):** diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index c873798b..cd332c60 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -394,6 +394,14 @@ R-636's louder repeated alarm. each step then gets a night of the household using it before the next one, and a failure is one step wide. Operator may widen it. `TestLeg_OneStepPerAppPerNight`. +### 2026-09-25 — an operator ruling + +34. **A controller restart during the night's update leg does NOT resume the leg** (operator ruling 2026-09-25, + option B of R-686). The apps the leg had not reached wait for the next night. A step already pressed is + finished or put back by the guarded update's own journal, exactly as before. **Nothing is built:** the + behaviour shipped in v0.271.0 is now the rule. **Why:** one night's delay for an app is cheap, and a resume + would be one more mechanism acting with nobody watching. + --- ## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed diff --git a/documentation/audits/RETIRE-peti-2026-09-25.md b/documentation/audits/RETIRE-peti-2026-09-25.md new file mode 100644 index 00000000..25455374 --- /dev/null +++ b/documentation/audits/RETIRE-peti-2026-09-25.md @@ -0,0 +1,81 @@ +# RETIRE — Peti's box (`peti-felhom`), 2026-09-25 + +Operator ruling 2026-09-25: Peti's box is retired (the tester wiped his server; it will not return); its off-site +backup holds no user data and is to be deleted; it leaves the protected list. Architecture read first: +`runbooks/target-selection.md`, `05-hub-architecture.md` §14 (what a customer delete deprovisions), +`06-offsite-connectivity.md`, `04-control-plane-authorization.md`. Evidence: `audits/retire-peti-2026-09-25/`. + +## Not done, or changed + +- **Nothing of Peti's was on ep0.** No PBS namespace (`ns`: demo-felhom, demo-hp, tester-1), no WireGuard peer in + the live `wg show`, no config naming him. His only off-site item was a Storage Box **sub-account** + (`u629488-sub2`, id 269130, label `felhom-customer=peti-felhom`, home `felhom-peti-felhom`) on the pool box. +- **The "backup" was never a backup.** The sub-account's folder held ONE file: `.ssh/authorized_keys`, 81 bytes. No + restic repository was ever created — all 482 of his reports (2026-07-10 … 07-15) show 0 off-site snapshots and + 0 bytes; his escrow stayed pending. The operator's "no user data" is confirmed by the listing, not only by word. +- **To list and empty the folder I reset the sub-account's password through the Hetzner API** (the product's own + `reset_subaccount_password` action, scoped to id 269130 after re-reading its label). The key file was removed from + inside the folder, then the hub deleted the sub-account. Whether Hetzner deletes a sub-account's folder with it + is NOT measured here — so the folder was emptied first, and nothing of his could remain either way. +- **Cloudflare:** his customer config carried a Cloudflare tunnel token and an API token (for `sajatfelhom.hu`). + They were purged with the customer record. The hub's delete dialog says "tunnel/zone removed", but no leg of the + cascade calls Cloudflare — the tunnel and DNS records on Cloudflare's side, if any remain, were NOT inventoried or + removed (a row: R-688). +- **A second, older Storage Box** (`PBS-storage-1`, u629193, box 611421) once held a folder `felhom-peti-spike/` + (spike leftovers, per `RUNBOOK-ep0-datastore-volume-2026-07-27.md`). Its mount on ep0 is gone, and the box is + not visible to either API token (404). The register row "delete the box" is the operator's. Not touched. +- **Code comments that name Peti were left** (hub `rollup.go`, `appliances.go`, `notify/templates.go`; agent pbsdr + tests). They explain why code is shaped as it is; no behaviour depends on Peti. + +## A1 — inventory (before anything was removed) + +| where | item | how found | +|---|---|---| +| hub | customer `peti-felhom` („Peti Proxmox", domain `sajatfelhom.hu`, dr_tier 0) with 482 reports, 123 events, 1,036 telemetry rows, DR recipe, one-time secret, claim, 1 log tail | copy of the hub DB (with its -wal), every table's `customer_id` / `host_id` | +| hub | host `peti-felhom-86d37d` — already DELETED 2026-07-15 08:56:22 (`host_deletions`, escrow_acked 0); no host, guest, escrow, recovery or PBS-secret rows | same | +| hub | WireGuard peer — none in `wg_peers` | same | +| Storage Box | sub-account id 269130 `u629488-sub2`, home `felhom-peti-felhom`, label `felhom-customer=peti-felhom`, created 2026-07-10 | Hetzner API with the hub's own token, matched on the LABEL | +| ep0 | nothing (namespaces, peers, `/etc`, `/root`, `/srv` grep — one false hit, "competing" in `lvm.conf`) | ssh read-only | +| documents | 48 non-history files + workspace rules + 21 memory notes (whole-word pattern; positive control target-selection.md 7 hits, negative control catalog templates 0) | `A1-documents.txt` | + +**Control, before:** pool box `size_data 3,120,562,176`, sub-accounts sub1 (demo-felhom), sub3 (demo-hp), sub4 +(tester-1); ep0 namespaces demo-felhom 2 snapshots / 220K, demo-hp 2 / 540K, tester-1 2 / 368K; demo boxes' +off-site `last_status ok` (11 and 90 snapshots, runs 02:15:46Z / 02:18:41Z). + +## A3 — removal + +1. `POST …/subaccounts/269130/actions/reset_subaccount_password` → success (password kept 0600 in the scratchpad, + deleted afterwards). +2. As `u629488-sub2`: `rm .ssh/authorized_keys`, `rmdir .ssh` → `ls -la` empty, `du -s .` 1. +3. Hub preview `GET /configs/peti-felhom/delete` — 0 hosts, off-site `u629488-sub2`, residue 1,519. Then + `POST /configs/peti-felhom/delete` (ack_hosts, ack_reset, ack_purge, confirm_id, expect_hosts=0) → 303. The hub's + log: sub-account 269130 deprovisioned; PBS tenancy `existed=false`; claim reset; residue purged (reports 482, + telemetry 1,036, log tails 1); cascade COMPLETE (journal #20). + +## A4 — nothing else moved + +| item | before | after | +|---|---|---| +| sub-accounts | sub1, sub2 (Peti), sub3, sub4 | sub1, sub3, sub4 — labels and homes unchanged | +| pool box `size_data` | 3,120,562,176 | 3,120,562,176 | +| login as `u629488-sub2` | worked | "Permission denied" | +| ep0 namespaces (snapshots / du) | demo-felhom 2/220K, demo-hp 2/540K, tester-1 2/368K | identical | +| ep0 WireGuard peers | .2 .3 .4 .250 | identical | +| hub hosts | demo-felhom, demo-hp, drill-r50 | identical | +| hub customers | demo-felhom, demo-hp, drill-r50, peti-felhom, tester-1 | without peti-felhom | +| hub rows naming peti | — | events 124, notification_log 80, host_deletions 1, customer_resets 1 (audit, by design), app_log_issues 8 (R-244) | + +## A5 — documents changed + +`runbooks/target-selection.md` (Tier 2 = DooPlex + ep0, ep0's reason rewritten, Peti's section retired); the +unprompted-work rule (4 identical copies); `03`, `06`, `00-capability-map`; runbooks `TASK-identity-only-escrow` +(obsolete note), `RUNBOOK-vzdump-target-move` (row 9), `RUNBOOK-island-migration`; retired banners on +`pilot/PETI-tester-agreement.md`, `RUNBOOK-peti-return`, `RUNBOOK-peti-pbsdr`; register (PETI closed as retired, +R-530, R-244, R-600 annotated); CONTEXT; STATUS; memory notes. Historic audits, tests and CHANGELOGs keep their text. + +## Claims in the brief + +1. *Peti's backup is on ep0* — **WRONG.** ep0 held nothing of his; the item was a Storage Box sub-account. +2. *The hub's host delete removes the WireGuard peer* — **not testable here:** the host was deleted in July and + ep0 carries no peer for it now; whether that delete removed one is not recorded (R-600 annotated). +3. *The backup holds no user data* — **TRUE, and stronger:** it held no backup at all (one 81-byte key file). diff --git a/documentation/audits/retire-peti-2026-09-25/A1-demo-control-before.txt b/documentation/audits/retire-peti-2026-09-25/A1-demo-control-before.txt new file mode 100644 index 00000000..3ab54923 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A1-demo-control-before.txt @@ -0,0 +1,4 @@ +# A1 control BEFORE — demo boxes' off-site status from their latest hub report +demo-felhom 2026-09-25 08:44:15 {'snapshot_count': 11, 'repo_size_bytes': 147483, 'last_run': '2026-09-25T02:15:46Z', 'last_status': 'ok', 'state': None, 'escrow_state': 'escrowed'} +demo-hp 2026-09-25 08:44:17 {'snapshot_count': 90, 'repo_size_bytes': 215007449, 'last_run': '2026-09-25T02:18:41Z', 'last_status': 'ok', 'state': None, 'escrow_state': 'escrowed'} +tester-1 2026-09-17 00:28:20 {'snapshot_count': 0, 'repo_size_bytes': 0, 'last_run': '2026-09-16T21:09:00Z', 'last_status': 'error', 'state': None, 'escrow_state': 'escrowed'} diff --git a/documentation/audits/retire-peti-2026-09-25/A1-documents.txt b/documentation/audits/retire-peti-2026-09-25/A1-documents.txt new file mode 100644 index 00000000..107a2712 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A1-documents.txt @@ -0,0 +1,75 @@ +# A1 documents — whole-word pattern: \bpeti\b|peti-felhom|\bPETI\b|Peti's|Petinek|Petié|petis — 2026-09-25T08:54:25+00:00 +positive control (pattern on a known line: target-selection.md): 7 +negative control (same pattern on app-catalog templates): 0 +== non-history files (excluding audits/, tests/ evidence, CHANGELOGs, REPORT-*.md, pilot/ runbooks of the past): + felhom.eu/STATUS.md + felhom.eu/CONTEXT.md + felhom.eu/documentation/architecture/06-offsite-connectivity.md + felhom.eu/documentation/architecture/07-backup-architecture.md + felhom.eu/documentation/architecture/03-host-agent.md + felhom.eu/documentation/architecture/_recovery-inventory-2026-07-28.md + felhom.eu/documentation/architecture/10-localisation.md + felhom.eu/documentation/runbooks/RUNBOOK-byo-deployment.md + felhom.eu/documentation/runbooks/RUNBOOK-island-migration.md + felhom.eu/documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md + felhom.eu/documentation/runbooks/publish-train-rules.md + felhom.eu/documentation/runbooks/RUNBOOK-onboarding-draft-v4.md + felhom.eu/documentation/runbooks/RUNBOOK-manual-build.md + felhom.eu/documentation/runbooks/RUNBOOK-ep0-datastore-volume-2026-07-27.md + felhom.eu/documentation/architecture/00-capability-map.md + felhom.eu/documentation/runbooks/TASK-identity-only-escrow.md + felhom.eu/documentation/runbooks/target-selection.md + felhom.eu/documentation/pilot/RUNBOOK-peti-pbsdr-2026-07-11.md + felhom.eu/documentation/pilot/RUNBOOK-publish-0.85-0.120-2026-07-12.md + felhom.eu/documentation/pilot/RUNBOOK-publish-0.79-0.110-2026-07-10.md + felhom.eu/documentation/pilot/PETI-tester-agreement.md + felhom.eu/documentation/pilot/RUNBOOK-peti-return-2026-07-13.md + felhom.eu/documentation/pilot/RUNBOOK-publish-0.90-0.143-2026-07-18.md + felhom.eu/documentation/pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md + felhom.eu/documentation/pilot/RUNBOOK-publish-0.81-0.113-2026-07-11.md + felhom.eu/documentation/pilot/GO-LIVE-PACKAGE.md + felhom.eu/documentation/pilot/DRILL-GL6-2026-07-08.md + felhom.eu/.claude/rules/unprompted-work.md + felhom.eu/documentation/backlog/ROADMAP.md + felhom.eu/documentation/backlog/CLOSED-ITEMS.md + felhom.eu/hub/internal/notify/templates.go + felhom.eu/documentation/backlog/OPEN-ITEMS.md + felhom.eu/hub/internal/tenantsync/client_test.go + felhom.eu/hub/internal/monitor/deadline_unbound_test.go + felhom.eu/hub/internal/web/rollup.go + felhom.eu/hub/internal/web/rollup_test.go + felhom.eu/hub/internal/web/pbsdr_generation_test.go + felhom.eu/hub/internal/web/pbsdr_poke_test.go + felhom.eu/hub/internal/web/pbsdr_test.go + felhom.eu/hub/internal/web/render_test.go + felhom.eu/hub/internal/web/appliances.go + felhom-controller/.claude/rules/unprompted-work.md + felhom-controller/CONTEXT.md + felhom-agent/CONTEXT.md + felhom-agent/internal/pbsdr/seed_reassert_test.go + felhom-agent/internal/pbsdr/manager_test.go + felhom-agent/internal/pbsdr/rearm_test.go + app-catalog-felhom.eu/.claude/rules/unprompted-work.md +== workspace level: + .claude/rules/unprompted-work.md + .claude-memory/campaign3-nightrun-2026-07-11.md + .claude-memory/agent-update-needs-signed-job.md + .claude-memory/current-state-2026-07-11.md + .claude-memory/current-state-2026-07-21.md + .claude-memory/hub-closing-bundle-2026-07-13.md + .claude-memory/pbs-tier-provisioning-spike.md + .claude-memory/drtier-by-default-2026-07-12.md + .claude-memory/peti-return-p1-stop-2026-07-13.md + .claude-memory/spike-lan-discovery-r6-2026-07-18.md + .claude-memory/restore-test-fills-thin-pool.md + .claude-memory/gl2-byo-install-profile-shipped.md + .claude-memory/customer-claim-arc-2026-07-12.md + .claude-memory/polish-batch-2026-07-13.md + .claude-memory/observability-pass-shipped.md + .claude-memory/r50-island-bridge-go-2026-07-25.md + .claude-memory/fleet-identity-peti-vs-tester1.md + .claude-memory/power-outage-audit-2026-07-22.md + .claude-memory/hub-offsite-provisioning-e2e.md + .claude-memory/nas-network-storage.md + .claude-memory/remote-applog-diagnostics-shipped.md + .claude-memory/MEMORY.md diff --git a/documentation/audits/retire-peti-2026-09-25/A1-ep0-before.txt b/documentation/audits/retire-peti-2026-09-25/A1-ep0-before.txt new file mode 100644 index 00000000..b1f248a7 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A1-ep0-before.txt @@ -0,0 +1,35 @@ +# A1 ep0 read-only — 2026-09-25T08:57:45+00:00 +felhom-hetzner +== PBS namespaces (live store /mnt/pbs-datastore): +demo-felhom +demo-hp +tester-1 + demo-felhom: 220K snapshots=2 + demo-hp: 540K snapshots=2 + tester-1: 368K snapshots=2 +== grep peti anywhere in PBS config: +== wg show (peers): +peer: snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc= + allowed ips: 10.77.0.3/32 + latest handshake: 57 seconds ago +peer: KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo= + allowed ips: 10.77.0.2/32 + latest handshake: 1 minute, 39 seconds ago +peer: yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00= + allowed ips: 10.77.0.250/32 + latest handshake: 18 hours, 59 minutes, 18 seconds ago +peer: kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M= + allowed ips: 10.77.0.4/32 + latest handshake: 43 days, 2 hours, 41 minutes, 56 seconds ago +== wg config peers: +8:PublicKey = KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo= +9:AllowedIPs = 10.77.0.2/32 +12:PublicKey = yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00= +13:AllowedIPs = 10.77.0.250/32 +16:PublicKey = snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc= +17:AllowedIPs = 10.77.0.3/32 +20:PublicKey = kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M= +21:AllowedIPs = 10.77.0.4/32 +== grep peti in /etc and /root, /srv: +/etc/lvm/lvm.conf +1053:# When there are competing read-only and rea diff --git a/documentation/audits/retire-peti-2026-09-25/A1-hetzner-before.txt b/documentation/audits/retire-peti-2026-09-25/A1-hetzner-before.txt new file mode 100644 index 00000000..d0d8f5fb --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A1-hetzner-before.txt @@ -0,0 +1,7 @@ +# A1 Hetzner (the hub's own token, read-only) — 2026-09-25T08:57:06.131163+00:00 +pool box id=611714 name=storage-box-pool-1 user=u629488 stats: size=3233808384 size_data=3120562176 size_snapshots=113246208 + sub id=273581 user=u629488-sub1 home=felhom-demo-felhom labels={'felhom-customer': 'demo-felhom'} created=2026-07-18T16:53:43Z desc=felhom offsite demo-felhom + sub id=269130 user=u629488-sub2 home=felhom-peti-felhom labels={'felhom-customer': 'peti-felhom'} created=2026-07-10T07:29:44Z desc=felhom offsite peti-felhom + sub id=275124 user=u629488-sub3 home=felhom-demo-hp labels={'felhom-customer': 'demo-hp'} created=2026-07-21T16:01:40Z desc=felhom offsite demo-hp + sub id=311327 user=u629488-sub4 home=felhom-tester-1 labels={'felhom-customer': 'tester-1'} created=2026-09-16T15:03:56Z desc=felhom offsite tester-1 +all boxes visible to this token: [(611714, 'storage-box-pool-1')] diff --git a/documentation/audits/retire-peti-2026-09-25/A1b-ep0-storagebox-mount.txt b/documentation/audits/retire-peti-2026-09-25/A1b-ep0-storagebox-mount.txt new file mode 100644 index 00000000..b66e3bb5 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A1b-ep0-storagebox-mount.txt @@ -0,0 +1,10 @@ +# A1b ep0 second storage box mount — 2026-09-25T09:02:23+00:00 +inactive +ls: cannot access '/mnt/pbs-storagebox': No such file or directory +== felhom-peti-spike: du: cannot access '/mnt/pbs-storagebox/felhom-peti-spike': No such file or directory +ls: cannot access '/mnt/pbs-storagebox/felhom-peti-spike': No such file or directory +== felhom-demo: du: cannot access '/mnt/pbs-storagebox/felhom-demo': No such file or directory +ls: cannot access '/mnt/pbs-storagebox/felhom-demo': No such file or directory +== spike-sub: du: cannot access '/mnt/pbs-storagebox/spike-sub': No such file or directory +ls: cannot access '/mnt/pbs-storagebox/spike-sub': No such file or directory +== is it a PBS datastore anywhere? diff --git a/documentation/audits/retire-peti-2026-09-25/A2-listing.txt b/documentation/audits/retire-peti-2026-09-25/A2-listing.txt new file mode 100644 index 00000000..a6553e68 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A2-listing.txt @@ -0,0 +1,34 @@ +# A2 listing of Peti's sub-account home (u629488-sub2), read-only — 2026-09-25T08:59:31+00:00 +$ pwd +Warning: Permanently added '[u629488-sub2.your-storagebox.de]:23' (ED25519) to the list of known hosts. +/home +$ df +Filesystem 1K-blocks Used Available Use% Mounted on +u629488-sub2 1073631360 3047552 1070583808 1% /home +$ du -s . +2 . +$ ls -la +total 3 +drwxr-xr-x 3 u629488-sub2 1006 3 Jul 10 09:49 . +dr-x--x--x 7 root root 11 Jul 10 07:30 .. +drwx------ 2 u629488-sub2 1006 3 Jul 10 09:49 .ssh +$ ls -la felhom-repo +/usr/bin/ls: cannot access 'felhom-repo': No such file or directory +$ du -s felhom-repo +/usr/bin/du: cannot access 'felhom-repo': No such file or directory +$ tree -a -L 2 felhom-repo +felhom-repo [error opening dir] + +0 directories, 0 files +$ ls -la felhom-repo/snapshots +/usr/bin/ls: cannot access 'felhom-repo/snapshots': No such file or directory +$ ls -la felhom-repo/keys +/usr/bin/ls: cannot access 'felhom-repo/keys': No such file or directory +$ ls -la .ssh +total 2 +drwx------ 2 u629488-sub2 1006 3 Jul 10 09:49 . +drwxr-xr-x 3 u629488-sub2 1006 3 Jul 10 09:49 .. +-rw------- 1 u629488-sub2 1006 81 Jul 10 09:49 authorized_keys +$ stat -c %n .ssh/authorized_keys .ssh/nonexistent-control-zq +.ssh/authorized_keys +/usr/bin/stat: cannot statx '.ssh/nonexistent-control-zq': No such file or directory diff --git a/documentation/audits/retire-peti-2026-09-25/A2-peti-last-report.txt b/documentation/audits/retire-peti-2026-09-25/A2-peti-last-report.txt new file mode 100644 index 00000000..4ebf478d --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A2-peti-last-report.txt @@ -0,0 +1,10 @@ +# A2 Peti's LAST report to the hub: id 12564 received 2026-07-15 08:39:00 UTC; controller 0.115.0 +deployed apps: ['rallly'] +offsite status: {"enabled": true, "escrow_state": "pending", "snapshot_count": 0, "repo_size_bytes": 0, "quota_gb": 50} +backup section keys: ['enabled', 'last_db_dump', 'snapshot_count', 'repo_size_mb', 'integrity_ok'] +storage: [('/', 'SSD')] +newest report WITH an offsite object: 12564 2026-07-15 08:39:00 {"enabled": true, "escrow_state": "pending", "snapshot_count": 0, "repo_size_bytes": 0, "quota_gb": 50} + +== all 482 Peti reports (2026-07-10 09:49:09 .. 2026-07-15 08:39:00 UTC): 482 carry an offsite object; MAX offsite snapshot_count=0, MAX repo_size_bytes=0 +apps ever deployed (report count): {'rallly': 481, 'calibre-web': 73} +offsite-related events for Peti: 1 diff --git a/documentation/audits/retire-peti-2026-09-25/A3-1-reset-password.txt b/documentation/audits/retire-peti-2026-09-25/A3-1-reset-password.txt new file mode 100644 index 00000000..44ca6c08 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A3-1-reset-password.txt @@ -0,0 +1,3 @@ +# A3-1 2026-09-25T08:59:13.275582+00:00 TARGET CHECK: u629488-sub2 felhom-peti-felhom {'felhom-customer': 'peti-felhom'} access: {'samba_enabled': False, 'ssh_enabled': True, 'webdav_enabled': False, 'reachable_externally': True, 'readonly': False} +action 657507966278067 reset_subaccount_password running +action status: success (password stored out-of-band in the session scratchpad, 0600) diff --git a/documentation/audits/retire-peti-2026-09-25/A3-2-empty-home.txt b/documentation/audits/retire-peti-2026-09-25/A3-2-empty-home.txt new file mode 100644 index 00000000..aec647c2 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A3-2-empty-home.txt @@ -0,0 +1,9 @@ +# A3-2 empty Peti's home (target: u629488-sub2 = sub 269130, label peti-felhom; its home is /home) — 2026-09-25T08:59:59+00:00 +$ rm .ssh/authorized_keys +$ rmdir .ssh +$ ls -la +total 2 +drwxr-xr-x 2 u629488-sub2 1006 2 Sep 25 09:00 . +dr-x--x--x 7 root root 11 Jul 10 07:30 .. +$ du -s . +1 . diff --git a/documentation/audits/retire-peti-2026-09-25/A3-3-hub-preview.txt b/documentation/audits/retire-peti-2026-09-25/A3-3-hub-preview.txt new file mode 100644 index 00000000..169484a5 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A3-3-hub-preview.txt @@ -0,0 +1,28 @@ +# A3-3 hub delete PREVIEW for peti-felhom — 2026-09-25T09:00:09+00:00 +{ + "claim_present": true, + "customer_id": "peti-felhom", + "customer_name": "Peti Proxmox", + "dr_recipe_present": true, + "has_config": true, + "host_count": 0, + "hosts": [], + "offsite_enabled": true, + "offsite_identifier": "u629488-sub2", + "offsite_type": "shared", + "one_time_secret": true, + "online_host_present": false, + "pbs_tenancy_configured": true, + "pending_journal": null, + "residue": { + "app_log_tails": 1, + "app_telemetry": 1036, + "appliance_registrations": 0, + "log_tail_requests": 0, + "notification_prefs": 0, + "reports": 482, + "selfbind_tokens": 0 + }, + "residue_total": 1519, + "superseded_blobs": 0 +} diff --git a/documentation/audits/retire-peti-2026-09-25/A3-4-hub-delete.txt b/documentation/audits/retire-peti-2026-09-25/A3-4-hub-delete.txt new file mode 100644 index 00000000..a3c1f779 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A3-4-hub-delete.txt @@ -0,0 +1,12 @@ +# A3-4 hub customer DELETE cascade for peti-felhom — 2026-09-25T09:00:21Z +POST /configs/peti-felhom/delete -> 303 Location: /configs?flash=deleted + +== the hub's own log: +2026/09/25 11:00:21 [INFO] customer DELETE cascade started for peti-felhom (journal #20, 0 host(s)) +2026/09/25 11:00:29 [offsite] deprovisioned shared sub-account 269130 for peti-felhom (repo data destroyed) +2026/09/25 11:00:29 [INFO] reset peti-felhom: offsite deprovisioned (repo data destroyed) +2026/09/25 11:00:29 [INFO] tenantsync: deprovision ok for peti-felhom (ns=peti-felhom, existed=false) +2026/09/25 11:00:29 [INFO] reset peti-felhom: PBS tenancy deprovisioned +2026/09/25 11:00:29 [INFO] [claim] reset to unclaimed for peti-felhom (customer RESET) — next onboarding mints a fresh code +2026/09/25 11:00:33 [INFO] delete peti-felhom: residue purged (reports=482 app_telemetry=1036 app_log_tails=1 log_tail_requests=0 notif_prefs=0 selfbind_tokens=0 appliance_registrations=0) +2026/09/25 11:00:33 [INFO] customer DELETE cascade COMPLETE for peti-felhom (journal #20) — full teardown diff --git a/documentation/audits/retire-peti-2026-09-25/A4-after.txt b/documentation/audits/retire-peti-2026-09-25/A4-after.txt new file mode 100644 index 00000000..66e20f39 --- /dev/null +++ b/documentation/audits/retire-peti-2026-09-25/A4-after.txt @@ -0,0 +1,28 @@ +# A4 AFTER — 2026-09-25T09:00:49.893123+00:00 +pool box 611714 stats: size=3233808384 size_data=3120562176 size_snapshots=113246208 + sub id=273581 user=u629488-sub1 home=felhom-demo-felhom labels={'felhom-customer': 'demo-felhom'} + sub id=275124 user=u629488-sub3 home=felhom-demo-hp labels={'felhom-customer': 'demo-hp'} + sub id=311327 user=u629488-sub4 home=felhom-tester-1 labels={'felhom-customer': 'tester-1'} +== login as the deleted sub-account must now FAIL: +Permission denied, please try again. +login_rc=0 +== ep0 after: +demo-felhom +demo-hp +tester-1 + demo-felhom snapshots=2 du=220K + demo-hp snapshots=2 du=540K + tester-1 snapshots=2 du=368K +KNaFXHiY9l08UhPO4gxUBFQz2WF3RYWyUGL8k+hzDVo= 10.77.0.2/32 +yNVpWt9Ekid9p459/VLHbmCDHIi8xTUYHbgazUp5S00= 10.77.0.250/32 +kdhOHMyADRpaTaFmh8ypwsWngvBpIyMyon80DJK+00M= 10.77.0.4/32 +snkgWxlcN7zXxjy1rG/jefe6uysGq/27E736CVPkbBc= 10.77.0.3/32 +== hub census AFTER: every table, every text column, rows matching 'peti' + notification_log: 80 + events: 124 + app_log_issues: 8 + host_deletions: 1 + customer_resets: 1 +control — demo rows still present: {'reports': 15337, 'customer_configs': 2, 'events': 2262} +hosts: ['demo-felhom-8363b5', 'demo-hp-bb76ea', 'drill-r50-0a4f9a'] +customer_configs: ['demo-felhom', 'demo-hp', 'drill-r50', 'tester-1'] diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 06b6d46c..f1eb7b6f 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -392,3 +392,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-680** | **The box did not remember a failed update step (P2).** Controller v0.271.0: an undone or held step is recorded in app.yaml (`failed_update_step`, tied to the ladder's print); the automatic leg skips it until the catalog's ladder changes; a person can still press. Live on 9202: vikunja undone night 1, skipped `failed_before` night 2, re-tried after the catalog re-tested it. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/`; `B/redproofs/R680-*` | | **R-678** | **After a step ended `done`, steps-left and the badge stayed stale (P3).** Controller v0.271.0: the update re-reads the app's catalog fields BEFORE it says done (and after an undo). Live on 9202: every automatic step's page read current at the leg's end. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `B/redproofs/R678-*` | | **R-643** | **The ruled chain left the automatic update leg at most 15 minutes a night (P2).** Decision 20, built in controller v0.271.0: the full-system backup's gate defers while the leg runs, until W+5h (then only for a step in flight, cap W+5h30m — decision 31); the leg starts no step at or after W+5h; one shared constant. Unit + red-proof (`TestD20_GateWaitsForTheLeg`); live on the demo boxes: see the night record Part D. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/B/redproofs/D20-*` | +| **PETI** | **`peti-felhom` deliberately not migrated; parked until the tester reinstalls.** **RETIRED 2026-09-25 (operator ruling):** the tester wiped his server and the box will not return. Removed through the hub's customer delete (journal #20): the customer record and 1,519 residue rows, Storage Box sub-account `u629488-sub2` (id 269130) — which held only one 81-byte `authorized_keys`, never a repository (all 482 reports: 0 off-site snapshots, 0 bytes); ep0 held nothing of it (no PBS namespace, no WireGuard peer). The audit trail stays by design. | retired 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/RETIRE-peti-2026-09-25.md` | +| **R-686** | **The automatic update leg is not resumed after a controller restart during the night.** **RULED 2026-09-25 (operator, option B):** the apps the leg had not reached wait for the next night; nothing is built — `09` §3 decision 34. The page-line side effect (a resumed step's `last_auto_update` not written) stays as measured. | ruled 2026-09-25 | `git show 6b2176e:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/night3-kill/`, `night4-power/` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index d4b4d2fd..08530548 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -214,7 +214,7 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | **R-240** | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | | **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | | **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | -| **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. | **READY** — owner Viktor | +| **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | | **R-245** | **Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT — and the reasoning against it is recorded with it so the decision can be revisited properly.** **The operator's proposal (2026-08-07):** a box that has been offered recovery for 30 days without the customer deciding is auto-abandoned, entering the 14-day grace, so an undecided box does not sit for ever holding history nobody has claimed. **What was built instead:** the escalating reminders (1/3/7/14 days) and the operator levers `--abandon-extend` / `--abandon-stop`. **The reasoning, as settled with the operator the same day:** (1) **nobody is absent** — a box does not reinstall itself, so whoever rebuilt it was standing there and met the recovery question; the "customer away for months" case does not arise from this situation, because **a reinstall implies a person**. (2) **A customer who cannot find their code will get in touch**, which is the moment to extend or disarm by hand — so the automation would be firing at people we are already talking to, which is why the levers were the thing worth building. (3) **The cost is theirs**: the old history sits in the customer's own storage allowance, and if they are paying to keep something they have not decided about, that is their call and they feel it before we do. (4) **The real harm, if it comes, is QUOTA** — old history blocking new backups — and **that is a condition, not a calendar**. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. **If this is ever built, build it that way.** **✅ RE-FILED 2026-08-08 AS A DECISION TAKEN, not a question pending.** It sat in the operator's queue as `WAITING-ON-OPERATOR` for a day, and **nothing was actually pending** — the operator and the reviewer settled it on 2026-08-07: it is **not built**, the levers were built instead, and the whole reasoning above is the record of why. A settled decision parked in a queue is a queue nobody trusts, and an audit of every `WAITING-ON-OPERATOR` row the same day found this was the ONLY one — so the drift was caught while it was still a single row. **THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups.** Not a calendar. If a customer's retained history ever refuses a new backup, revisit this with a dated warning that triggers on the refusal; until then it stays decided. | **DECIDED 2026-08-07 — not built; reopens on quota** | | **R-246** | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor | | **R-248** | **A flag that changes behaviour is visible to nobody who would look for it.** Q4 of the 2026-08-08 spike, answered plainly. **The customer** sees only a derived card stating a false reason (R-247). **The box** cannot see it at all (R-247's dropped field). **The operator** can see it on exactly ONE page — the **PBS-DR** view (`hub/internal/web/pbsdr.go:487`, `v.EscrowStale = escrow.StaleAt != ""`) — which is the wrong tier for this symptom: an operator investigating an OFF-SITE problem has no reason to open a PBS-DR page. **No alert, no report field, no off-site surface.** The one-shot `escrow_stale` event fired on 2026-08-04 and **was never notified** (a full census of `notification_log` for that customer that day returns 8 rows, none of them this one); it has fired twice ever, both on 4 August. **So the practical answer is: only a database read.** That is a finding in its own right — a flag that silently changes behaviour and cannot be seen is the shape this fortnight has been about (R-241's discarded comparison, R-228's unread field, R-243's unobserved state). **What it needs:** surface `stale_at` on the off-site operator surface and in the host report, or stop using a field nobody can observe to change what a customer is told. | **READY** — owner Viktor | @@ -307,7 +307,6 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2a** | ~~The target move needs a root-fenced wrapper — the agent cannot do it~~ | **SHIPPED + PROVEN-LIVE** (agent v0.113.0 + host-install v1.22.0, 2026-07-29) | — | `felhom-backup-target-apply` behind a literal `FELHOM_BACKUPTARGET` sudoers alias; the agent's PVE role was NOT widened. Enforces F-1 (`mountpoint -q`) and F-2 (`is_mountpoint 1` hardcoded), refuses a root-device target, has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint. All five laws proven live as root on demo-hp with 0 stray storages | — | | **E-2b** | ~~`NotifyStorageDisconnected`/`Reconnected` defined and called NOWHERE — a drive going absent emitted no event on any channel~~ | **SHIPPED + PROVEN-LIVE** (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29) | — | Seam wired in `ReconcileDriveGates`; a target drive raises the specific `backup_target_absent` instead. **A keying bug was caught before deploy:** `a.Path` is the registered GUEST path, not the agent's host `MountPath`, so the target branch was unreachable — every absent drive, target included, fell through to the generic event (v0.184.1). Tests observe the WIRE (httptest hub), not a mock | — | | **E-2c** | ~~E-1 put the whole-guest backups on a drive `POST /disks/eject` would eject~~ | **SHIPPED + PROVEN-LIVE** (agent v0.112.0, 2026-07-29) | — | Eject + decommission refuse 409 on the backup-target mount, naming the storage and the remedy. **Live on BOTH boxes:** demo-hp `/mnt/nvme-1tb` and demo-felhom `/mnt/hdd_1` both refused, drives unmoved. NOT a role reclassification — `RoleForStorage` untouched, because on both boxes that drive is ALSO the enrolled user-data drive; `TestEjectStillAllowedOnANonTargetDrive` pins the non-over-correction and `/var/lib/vz` is still refused by the PRE-EXISTING role gate, not this one | — | -| **PETI** | **`peti-felhom` deliberately NOT migrated — and the mitigation this row used to name DOES NOT EXIST.** This row said a drive failure there is *"offsite-only recovery"*. Re-read from the hub's own store on 2026-08-10 and again on 2026-08-12, with a control run first (the escrow query returns 1+1 rows for each demo box and 0+0 for `drill-r50`, so it distinguishes the states): **there is no off-site copy, no key, and no local backup either.** Three independent reasons, each a fact rather than an inference: **(1)** the host row was DELETED — `host_deletions` id=1, `peti-felhom-86d37d`, **2026-07-15 08:56:22**, `escrow_acked = 0` — long before host-delete-demotes-escrow-to-retained-custody existed, so nothing was carried over; **(2)** there is **no escrow row of any kind**, current or superseded (only 4 exist hub-wide, all belonging to the two demo boxes), and the off-site restic REPOSITORY password lives in `identity_blob` on that row (`hub/internal/store/store.go:381`) — its only other copy is `/offbox/repo_password` (`controller/internal/backup/offbox.go:395`) **on the very disk whose failure is the scenario**; **(3)** off-site backup **never ran once** — its last report carried `offsite: {escrow_state: "pending", snapshot_count: 0, repo_size_bytes: 0}`, and that is the fork-4 guard working exactly as designed (`controller/internal/settings/settings.go:317-320`: *"no offsite run proceeds until an operator confirms the escrow ceremony"*), not a fault. The local app-data restic repo was also empty (`snapshot_count: 0, integrity_ok: false`), and the whole-guest vzdump shares the failing device. **SO: if that drive fails today, everything on it is lost.** **Size, so this is not read as larger than it is:** one lightly-used test box — a single `/` mount, **3.6 GB used of 48.9 GB**, one catalogued app (`rallly`); the dashboard was never claimed (`customer_claims.claimed_at` NULL). **BOUNDARY:** every fact is as of the last report, **2026-07-15 08:39:00 UTC** (controller 0.115.0); the hub has heard nothing since and confirming today's state would mean contacting the machine, which is fenced. **Contact since deletion:** no inbound row of any kind after 2026-07-15 08:39 — the only later rows are the hub's OWN staleness alarms (`source = hub`: `node_stale` 09:09:32, `node_down` 09:39:32) — and no contact attempt, accepted or rejected, in the current hub pod's logs (since 2026-08-09 17:26 UTC; grep proven by 851 `demo-hp` hits against 0 for peti, 0 unauthorized). **The window 2026-07-15 → 2026-08-09 cannot be answered from records**: a report from a deleted host 401s and is not persisted, and those logs are gone. **THE RULING IS LEFT OPEN DELIBERATELY** — whether the machine stays parked is the operator's call and does not need restating here; this row records the FACT, which does not need his opinion to be true. **⚠ WHO THIS IS ABOUT — CORRECTED 2026-08-13, because the row was read as describing a record rather than a machine and that reading nearly deleted a true risk.** Every summary since 2026-08-09 has led with *"the tester's machine"*, and there are now **two different things** that word can mean, one of which carries no risk at all. **This row is about the machine: a physical 80-core Proxmox server belonging to a named person, running Felhom as a BYO guest** (`pilot/PETI-tester-agreement.md`, *"Operator: Viktor. Tester: Peti"*), which **reported to the hub 482 times between 2026-02-27 and 2026-07-15 08:39:00 UTC** (`reports`, counted). It is Tier 2 — protected — *"because there is a real person behind it"* (D-d). **The 3.6 GB and the missing recovery route are therefore real, and this row keeps its rank.** **The other thing is `tester-1`, and it is not this:** a customer record created **2026-08-13 07:56:47**, one minute after the record previously called `david` was torn down (`customer_resets` id 16, 07:55:49→07:55:50, every leg `ok`, `hetzner` skipped; `customer_deleted` event 2960). It has **no host row, no escrow row, no host-report and no controller report — measured, all zero** — and `david` before it had the same and was never anything else (its only four events were three hub-side `expected_dbdump_missed` false alarms and its own deletion; that is R-195's subject). **A record with no machine behind it can lose nothing.** So the operator's *"there is no actual tester yet — only a pre-created customer, now renamed"* is true of the **pilot programme and of `tester-1`**, and not of this row: the agreement was drafted, the onboarding runbook stopped at P1, and no pilot ever formally began — **while the hardware and the data have existed the whole time.** **Nothing about the risk changed; only the word that names it.** Wherever a document says *"the tester's machine"*, read `peti-felhom` | **PARKED — the recorded mitigation is void; ruling OPEN for the operator** | — | **First act of the visit: copy that ~3.6 GB off before anything is reinstalled — it is currently the only copy in existence.** Do not migrate, do not contact | operator | | **R-124** | **The recipe spells PBS's root namespace `"root"`, but the PBS API spells it `""`** and no namespace is literally named `root` — an operator pasting the field into `pct restore --ns root` gets a failure | READY (XS) | — | Pre-existing wire convention (`ToHub` has normalised empty→`"root"` since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at `hub.PBSRootNamespace`. Affects only a box with no `namespace` line — **no real customer today**, all three are per-customer. Fix = emit `""` + rely on `namespace_state`, or emit a `--ns`-ready form | CC | | **R-89** | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | | **R-92** | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC | @@ -758,7 +757,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-597** | **[P2-MEDIUM] The setup code is three Hungarian words, inside an otherwise fully English e-mail, sent to a household the hub knows is English.** FOUND 2026-09-20 by the slice-6 drill. The mail is English end to end (slice 3 working); the code it carries was **`képző-szkítia-ásatás`** — 20 characters, 3 words, **5 of them outside ASCII** (ő, í, á×2, é). An English speaker must copy three words they cannot read, spell or say aloud, and type them into a box on a keyboard that has no ő. They can paste — until the day they read the code to someone over the telephone, which is precisely what a three-word code is FOR. **The same generator feeds the recovery code (10 words) and the owner passphrase (5 words)**, so the fault is one wordlist wide, not one mail wide: this walk saw the passphrase too and it is Hungarian. **Fix shape:** an English wordlist chosen per `customer.language`, with the same word count and the same entropy, and a test that pins BOTH lists' entropy and that no word in either needs a character outside the reader's keyboard. **Not a rename of the existing words** — a second list. **CLOSED 2026-09-21, hub v0.119.0.** **One third of the row was wrong: the recovery code was never Hungarian.** `felhom-agent` mints it (`internal/escrow`) from the **EFF large wordlist** and always has — ten English words, ≈129 bits. The hub does not own that secret and no row was opened for it: a second definition here is the drift `backupTargetAbsentText` already demonstrates across two repos. The two the hub DOES mint now follow the household: setup code 3 hu words (44.6 bits) → **4 en words (51.7)**, owner passphrase 5 hu (74.3) → **6 en (77.5)**, list and count chosen together by `RandomPassphraseFor(lang, use)` so a caller cannot pair an English list with a Hungarian count. **The floor is computed from the embedded lists at test time, not compared with a constant** — red-proofed at 3 English words (38.77 vs 44.56). Hungarian is byte-unchanged, and the list length is pinned so a swap cannot move it quietly. **The task's proposed "read it over the phone" filter was MEASURED and NOT adopted** — it removes 5270 of 7772 words (68%, 12.92 → 11.29 bits/word) and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster; what it reached for is kept as an assertion (`TestEnglishListIsTranscribable`: 3-9 lower-case ASCII letters, no digit, no separator). Decision recorded in source, **operator may reverse**. Also: **no claim mail ever stated a word count** — the only count wording was the bind page's passphrase hint, whose English half is now count-free. | **CLOSED 2026-09-21 — hub v0.119.0** | | **R-598** | **[P2-MEDIUM] The Backup page's two protection warnings — the ones that say whether the household's files are safe — are Hungarian on an English dashboard.** FOUND 2026-09-20 by the slice-6 drill on a fresh box, and confirmed on the demo box. Of 73 lines on `/backups` exactly four are Hungarian: **„Csak egy másolat készül (nincs második meghajtó) — a 3-2-1 mentéshez csatlakoztasson egy második meghajtót vagy offsite tárolót"**, **„A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem"**, and the two backup-target names **„Helyi tároló (local)"** and **„Biztonsági szerver – külön hardver (PBS)"**. They come from `internal/web/backup_handlers.go` (12 Hungarian literals) and `internal/web/backup_target_offer.go` (10) — again composed sentences handed to the page, the R-573/R-596 shape. **It matters more than its line count:** those two warnings are the only place the product tells a household that one copy on one disk is not protection, and the volunteer guide's §9 sends every tester to exactly this page to read exactly these two sentences. **Fix shape:** keys + args for both files, with the retrieval-promise gate run over the English (these sentences are about what a backup does and does not protect). **CLOSED 2026-09-21, controller v0.259.0.** The row's count of `backup_handlers.go` was 12; **nine are code and three are Hungarian inside COMMENTS**. The offer file's ten is right. `degradedMessageFor` now returns a **KEY** — the decision stays language-free and in one place, the words are chosen by the caller that knows the reader — and `buildTierViews` / `backupTargetLabel` / `loadGuestBackup` take the language the way `buildDataPathCards` already did. **The English is asserted to carry the same NEGATION the Hungarian does** ("protects against corrupted files, **but not** against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian. **Proven LIVE on guest 9201** for the two tier names ("Local storage (felhom-backup)", "Backup server – separate hardware (PBS)"); **the two warnings themselves were NOT walked live** — that box is healthy and a healthy box renders nothing by design, and producing the state would mean un-assigning a live backup target. They are covered by render tests through the real handler in both states. **An apostrophe cost a render:** the first English absent-drive sentence never matched because `html/template` escapes `'` to `'` — caught by the test, not by review. | **CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live** | | **R-599** | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused **both** `POST /configs//delete` (409, *host … is ONLINE*) and the host delete (`deletable:false`) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a **report-staleness window**, and the window is **45 minutes** — `manifests/hub.yaml` sets `alerting.stale_threshold: "45m"`, which `hostStatus()` reads (`ok` under it, `stale` over, `down` at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. **Measured the boring way, and worth recording:** this row first said 30 minutes, because `monitor/host_staleness.go`'s literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. **The consequence is not theoretical:** a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. **Fix shape (smallest first):** the 409 body says *how long* it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and `runbooks/target-selection.md`'s drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. | **READY — rank P3-LOW; owner: CC (hub)** | -| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. | **READY - rank P2-MEDIUM; owner: CC (hub)** | +| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. | **READY - rank P2-MEDIUM; owner: CC (hub)** | | **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** | | **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | | **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** | @@ -778,7 +777,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** | | **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** | | **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. | -| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | +| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | | **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** | | **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** | | **R-615** | **[P3-LOW] Pointing a box at a different app catalog by `git.repo_url` alone is INERT — the box keeps fetching from the repository it first cloned.** FOUND 2026-09-21 by reading `sync.go` **before** running it, which is the only reason the update night's drill catalog worked at all. `Syncer.gitCloneOrPull` (`controller/internal/sync/sync.go:274-306`) clones **only when `/catalog-cache/.git` is absent**; on every later cycle it runs `git fetch --depth 1 origin ` + `git reset --hard origin/` **against the remote stored in the clone**, which `buildRepoURL` wrote at clone time. Changing `git.repo_url` in `controller.yaml` and restarting therefore changes **nothing**: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, `git -C /catalog-cache remote -v` still read `app-catalog-felhom.eu`; the box only followed the drill repo once the cache directory was removed. **Why it matters beyond a drill:** this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. **Nothing is wrong with the CACHING**, which is right; what is missing is that a changed `repo_url` must invalidate the clone. **Fix shape:** on start, compare `git.repo_url` with the clone's `origin` and re-clone when they differ (or `git remote set-url` + a full fetch); log which happened. A test that changes `repo_url` under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: `audits/update-night-2026-09-21/04-9202-config-pre.txt`, `05-9202-follows-drill.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | @@ -813,8 +812,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** | | **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | | **R-685** | **[P2-MEDIUM] A box whose whole-box backup cannot fit must say so BEFORE the night — an operator event and a line on the backup page — never only a nightly failure a log shows.** The product lesson of R-684 (operator ruling 2026-09-24 evening, option A for demo-hp): demo-hp's root-disk target held three 6–7 GB archives with ~4 GB free and failed `No space left on device` every night from 2026-09-23 while nothing but the vzdump log said why. Measured 2026-09-24 night: a 9201 archive is 8.18 GB (22.6 GB uncompressed); PVE prunes AFTER a successful backup, so a target must hold keep-last + 1 archives during the run — lowering the retention alone does not un-stick a full target. **Fix direction:** the agent predicts the archive size (last archive × margin) against the target's free space before a whole-box backup, skips with a reason reported to the hub (operator event), and the controller's backup page shows the sentence. `audits/night-2026-09-25/A/A4-hp-backup-space.txt` | **READY — P2; owner: CC (agent + controller)** | -| **R-686** | **[P3-LOW] The automatic update leg is not resumed after a controller restart during the night — the apps it had not reached wait a whole day.** MEASURED 2026-09-24 night on 9202 (v0.271.0, night 3): the controller was killed (kill -9) during romm's step; the step itself was put back correctly (pin back, the page says the update was interrupted), but vikunja — next in line, re-tested and ready — was not pressed that night. **Night 4 (power cut during romm's verify) showed a second effect:** the interrupted step was RESUMED after the boot and ended `done` (data read back), but the leg that pressed it was gone, so `last_auto_update` was never written and romm's page carries no „Automatikus frissítés … — sikeres" line for a step the box did take by itself. By design today: the leg is one call of the `offbox-backup` job, and a Daily job that already fired does not fire again. The controller's own self-update now waits for the leg (decision 32), so the common restart cause is excluded; a crash or a power cut still ends the night's leg. **Fix direction (needs a ruling):** persist "leg started, not finished, window W" and resume at boot while before W+5h — or accept one night's delay. `audits/night-2026-09-25/C/night3-kill/` | **OPEN — P3; owner: CC (operator ruling on resume)** | | **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** | +| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` | **READY — P3; owner: CC (hub) / operator (Peti's Cloudflare leftovers, if any)** |