docs(runbooks): onboarding draft v3 -> v4 — post-ship refresh (R-36/R-39 workarounds deleted, HP t740 second datapoint)

Per the 2026-07-21 refresh brief: R-39 interim blocks (B4/E1) and the R-36
manual-Save block (C4) deleted — both shipped and proven live; freemail.hu
gate proven (R-4 COMPLETE); golden/floor-lift note now cites two shapes
(rehearsal + virgin HP t740 day-0 lift 0.153.0->0.156.0); A3 loader table
per operations/nodes.md (N100=mkimage/SB-off per record, HP t740=shim/SB
ENABLED); B2 multi-NIC cabled-port gotcha (R-59/R-60 pending); new A5 gate
(agent >=0.93.0 deployed box-side before the first escrow ceremony); D
offboarding pointer to §G (R-25b). DRAFT status and the C7 graduation gate
unchanged. ROADMAP R-25b pointer follows the rename.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
This commit is contained in:
2026-07-22 08:14:50 +02:00
parent 9b3381be0a
commit 05dfaa1a10
3 changed files with 122 additions and 225 deletions
+47 -178
View File
@@ -1,187 +1,56 @@
# REPORT — TASK-I: R-25b, customer DELETE becomes the guided full-teardown cascade
# REPORT — Runbook refresh: RUNBOOK-onboarding-draft v3 → v4 (docs-only, XS)
**Two releases: v0.69.0 (the cascade) and v0.70.0 (the residue leg + ghost cleanup, found validating v0.69.0 against the live hub).**
**Date:** 2026-07-21 · **Repo:** `felhom.eu` · **Baseline:** `f59aa97` (clean, HEAD == origin/main)
**Scope:** hub only — **v0.68.1 → v0.69.0**. No agent / controller / catalog change.
## What shipped
`POST /configs/{id}/delete` (same route, new behaviour) is now the guided full-teardown cascade.
`GET` on the same path returns the dialog's live inventory. The shallow `handleConfigDelete` is
**gone**.
Three legs, fixed order:
1. **hosts** — every host row via `store.DeleteHost(hostID, true)`; escrow **DEMOTED** to retained
custody, never destroyed. Host-delete's own ONLINE rule is kept: an ONLINE host refuses the whole
cascade, checked for every host up front so it never half-runs.
2. **reset** — the committed RESET sequence verbatim (Hetzner → PBS → claim → descriptor → DB purge)
through the newly extracted `commitCustomerReset`, called with `purgeEscrow=false`.
3. **purge**`store.DeleteCustomerConfig`: the customer record **and all escrow ciphertext**.
Two invariants are asserted, not merely commented:
- **Ruling 3 by construction** — leg 2 can only run after leg 1, so the RESET sequence never sees a
host row. The standalone RESET handler's 409 gate is untouched.
- **Custody purged exactly ONCE, in leg 3** — leg 1 demotes; leg 2 runs with `purgeEscrow=false`;
leg 3 is the one true purge point (v0.60.1).
**Gates, all before any write** (a refusal has zero side effects): three acknowledgements
(`ack_hosts` / `ack_reset` / `ack_purge`, each exactly `1`), the typed customer-id, a **stale-preview**
check (the acknowledged host count must still match live → else 409), and the ONLINE-host refusal.
No force flag, no skip flag, no partial-run downgrade.
**Resume:** a failed leg retains the `customer_resets` journal row and the HTTP error names the leg.
The dialog renders the incomplete journal and offers **Resume**; a re-run is idempotent and must pass
every gate again (acknowledgements are not cached across attempts).
**UI:** Danger zone → **Delete customer…** → guided dialog (inventory panel: hosts by name + status,
offsite repository identifier, PBS namespace, custody state; three consequence checkboxes; typed
customer-id; one submit). Client-side checks are convenience only.
## Refactor — standalone RESET behaviour unchanged
`handleCustomerReset`'s committed half became
`commitCustomerReset(ctx, cfg, resetID int64, purgeEscrow bool) *resetLegError`. The standalone path
is byte-identical to v0.68.1: same leg order, same leg names, same operator-facing messages, same
status codes. Its existing suite is untouched and green.
## Files
| File | Change |
|---|---|
| `hub/internal/web/customer_delete.go` | NEW — the cascade + preview |
| `hub/internal/web/customer_delete_test.go` | NEW — scenarios AE |
| `hub/internal/web/customer_reset.go` | `commitCustomerReset` + `resetLegError` extracted |
| `hub/internal/web/configs.go` | `handleConfigDelete` removed (replaced by a do-not-reintroduce note) |
| `hub/internal/web/server.go` | route: GET → preview, POST → cascade |
| `hub/internal/web/templates/customer_unified.html` | guided dialog replaces the one-click Delete |
| `hub/internal/web/customer_edit_tab_test.go` | the delete redirect case now posts the full acks |
| `hub/CHANGELOG.md`, `REUSE.md`, `documentation/backlog/ROADMAP.md`, `documentation/architecture/00-capability-map.md`, `documentation/runbooks/RUNBOOK-onboarding-draft-v3.md` | docs |
## Tests + red-proofs
Green gate in `hub/`: `go build ./... && go vet ./... && go test ./...`**all green**, no flakes.
New coverage (`customer_delete_test.go`):
- **A — leg ORDER**, observed from *inside* leg 2 via a `tenancyProvisioner` fake whose `Deprovision`
snapshots store state: at that instant hosts = 0 (leg 1 done), the customer row is still present
(leg 3 not started), retained custody still present (leg 1 demoted, did not purge). Plus final
state, the four journal legs stamped `ok`, completion stamped, and the `customer_deleted` audit
event surviving the record.
- **B — 9 fail-closed gate cases** (each missing ack, an ack sent as `yes`, id mismatch, id absent,
host count moved, host count absent, ONLINE host). Each asserts the status code **and** that the
host, the customer row, the current escrow and the retained custody are untouched, **and** that
zero external calls fired, **and** that no journal row was opened.
- **C — resume**: an injected PBS failure → 502 naming the leg; hosts gone, customer + custody
SURVIVE, journal retained with `hosts=ok pbs=failed`; re-run converges and completes. Plus: a
resume without ack #3 is still refused.
- **E — custody**: `commitCustomerReset(..., purgeEscrow=false)` leaves the retained blobs; leg 3
purges them.
- **Preview**: names the real host/offsite/PBS/custody facts and leaks no secret (one-time password,
API keys, escrow blobs all asserted absent).
**Five red-proofs run, each failed red with the wrong value visible, then restored (`git diff` clean):**
| # | Pre-fix shape restored | Failure observed |
|---|---|---|
| 1 | ack gate disabled | `status = 303, want 400` + host deleted, customer deleted, custody destroyed, PBS deprovision fired, journal row opened |
| 2 | stale-preview gate weakened to always pass | `status = 303, want 409` + the same six non-effect assertions |
| 3 | ONLINE-host gate removed | `status = 303, want 409` + live host deleted |
| 4 | leg order inverted (RESET leg before the host leg) | `at the RESET leg the customer still had 1 host(s) — leg 1 must complete FIRST` |
| 5 | cascade's RESET leg called with `purgeEscrow=true` | `retained blobs after the RESET leg = 0, want 2` |
## Method / limits
Unit-land only so far — **the STOP-gated live leg has not been run** (see below). Validation method:
Go tests against a real SQLite store on `t.TempDir()` with fakes at the existing `tenancyProvisioner`
seam. No browser is available on DooPlex; the dialog's markup is covered by the existing
customer-page render tests (exactly one `/configs/{id}/delete` form on the page) — a strict
click-through remains a manual operator pass.
**Not covered by unit tests:** the Hetzner offsite `Deprovision` leg (`offsite.Provisioner` is a
concrete type, no interface seam — same as the standalone RESET suite). That leg is exactly what the
live run is for.
## STOP — operator-present live leg (NOT yet run)
Create a scratch customer on a throwaway domain, provision **offsite only** (no host — cheap), run
the cascade end-to-end, then verify from **outside** the hub that the Hetzner repository is gone and
the customer row is purged. **Never run against Demo Ügyfél, Demo HP, or Peti.** Failure paths are
unit-proven; the live leg proves the happy path + external teardown only.
## Deploy status
**LIVE.** Code `61dbd87``main`; image `gitea.dooplex.hu/admin/felhom-hub:0.69.0` built + pushed;
`manifests/hub.yaml` bumped in `5dcb72a`; ArgoCD hard-refresh + deliberate sync → app `felhom`
**Synced / Healthy**; `deploy/hub` rolled out on image `:0.69.0`; startup log
`[INFO] felhom-hub 0.69.0 starting``Listening on :8080`, no errors.
The STOP-gated live cascade run above is still **outstanding** — the code is deployed, the scratch
customer teardown has not been performed.
---
# Addendum — v0.70.0: a deleted customer actually disappears
## How it was found
The operator reported that `demo-vm-felhom` "was deleted but is still here". It had NOT failed:
config row gone, both hosts deleted (07-16, 07-18, escrow acked), `host_escrow` and
`host_escrow_superseded` empty, the 07-18 RESET journal complete with every leg `ok`.
The customer was still listed because **`store.GetCustomers()` derives the customer list purely from
the REPORT stream**, and no lifecycle tier has ever deleted a report. 502 report rows kept the ghost
alive.
**This was not cosmetic.** The staleness and offsite checkers iterate the same report-derived list,
so the hub kept raising `offsite_stale` for a customer that no longer exists — **10 events, the most
recent 2026-07-21 17:34, three days after deletion, with an operator email sent at 19:34** (after the
v0.69.0 deploy). Verified by streaming the live `hub.db` out read-only and querying it.
Two further residue rows are **credential-bearing**, not telemetry:
`appliance_registrations` (a `token_hash` with `status='delivered'`, still bound to the dead customer
— confirmed present for `demo-vm-felhom`) and `selfbind_tokens` (an unconsumed bind token would be a
working path to bind a box to a nonexistent customer).
**Date:** 2026-07-22 · **Repo:** `felhom.eu` · **Scope:** documentation only — no code, no version
bump, no deploy. The onboarding runbook was refreshed against what shipped 2026-07-18..21 and
against the second full onboarding on virgin hardware (HP t740, 2026-07-21).
## What changed
- **New leg 3, `residue`** (`store.PurgeCustomerResidue`, one transaction): `reports`,
`app_telemetry`, `app_log_tails`, `log_tail_requests`, `customer_notifications`,
`selfbind_tokens`, `appliance_registrations`. Runs BEFORE the record purge — `customer_configs` is
the identifying descriptor and goes last. The cascade is now `hosts → RESET → residue → purge`.
`events`, `notification_log`, `host_deletions`, `customer_resets` still survive.
The counter and the purge walk **one shared `residueQueries` list**, so a table cannot be
counted-but-not-purged.
- **Ghost customers are deletable.** Both the cascade and its preview used to 404 whenever the config
row was missing — meaning no operator surface could clear a customer deleted by any earlier path.
**404 now means "there is nothing here"** (no config, no host, no residue). With no config row the
offsite descriptor is unknowable, so `commitCustomerReset` records **`skipped_no_config`** for the
Hetzner and descriptor legs — never a bare `skipped`, which would read as "nothing to do". PBS is
id-keyed and idempotent, so it still runs. The dialog labels the ghost case and names the row count.
`documentation/runbooks/RUNBOOK-onboarding-draft-v3.md`**`RUNBOOK-onboarding-draft-v4.md`**
(git mv, one commit), per the project-Claude refresh brief of 2026-07-21:
## Tests + red-proofs (v0.70.0)
1. **R-39 interim blocks DELETED** (B4 descriptor caveat + E1 manual `pvesm status` check) —
replaced with one sentence each: the DR tier self-detects a dead credential (loud `auth_failed`)
and self-heals via damped re-issue; the hub gauge is trustworthy. Cited: `backlog/ROADMAP.md`
R-39 (CLOSED, proven live 2026-07-21, hub 0.68.1 + agent 0.91.2, 13 s chain
`applied → auth_failed → applied`).
2. **R-36 workaround block in C4 DELETED** (re-onboarding manual-Save) — replaced with the shipped
behaviour: enabled-but-unprovisioned raises an amber banner (hub v0.67.0). R-31 (Save-once/504)
and the R-32 orphan-card note KEPT. Consistency follow-through: A1 and the B3 note now say the
self-bind link is **auto-minted at create/RESET** (also R-36) — the manual mint is the re-mint
fallback.
3. **C1 freemail gate** — freemail.hu PROVEN 2026-07-21 (operator test-send received; R-4 COMPLETE
per `backlog/ROADMAP.md`); the check-spam habit line kept.
4. **A1 golden note** — golden 0.153.0 carries all four infra images; floor-lift proven on TWO
shapes (rehearsal 0.143→0.145 in 5 s; virgin HP t740 0.153.0→0.156.0 during day-0, `CONTEXT.md`).
"Rebuild before first tester" softened to a pointer at the standing rule in F.
5. **A3 loader note** — the single F1 line replaced with the loader table. **Deviation from the
brief, on the record's side:** the brief said "N100/AMI = shim"; the record
(`operations/nodes.md` + `tests/VALIDATION-n100-baremetal-2026-07-16.md` F1) says the N100
installed with **mkimage, SB off** — the AMI AN3PLUS firmware cannot relocate the signed GRUB,
which is exactly why shim was not usable there. Table written per the record, as the brief's
"(SB state as recorded)" directs. HP t740 = shim with Secure Boot ENABLED (2026-07-21).
6. **B2/B3 second datapoint** — the HP onboarding walked register → pairing banner → unclaimed
list → self-bind (code + passphrase) → day-0 → dashboard on never-seen hardware (cited
`CONTEXT.md` + `operations/nodes.md`). New B2 gotcha: on multi-NIC boards verify the CABLED
port got the lease during install (the t740 4-port trap; R-59/R-60 pending).
7. **New A5 checklist item** — agent train ≥ v0.93.0 (hyphen-free recovery wordlist,
`felhom-agent/CHANGELOG.md`) must be BUILT + DEPLOYED to the customer's box before the FIRST
real tester's escrow ceremony; ceremonies run box-side, source-only does not count.
8. **D offboarding pointer** — R-25b ruling noted: full teardown = DELETE cascade (TASK-I),
RESET = identity-preserving re-onboarding; pointer to §G.
9. **KEPT unchanged:** DRAFT status, the C7 graduation-gate paragraph (still the last unwalked
step), all reference wall-clocks, all F rules, all of §G.
- `TestDeleteCascade_PurgesResidueAndUnlistsCustomer` — residue zeroed, customer absent from
`GetCustomers()`, appliance registration and self-bind token gone **by name**, audit + F-14
provenance intact, journal `residue=ok customer_delete=ok`.
- `TestDeleteCascade_GhostCustomerIsDeletable` — the exact `demo-vm-felhom` shape (hosts deleted,
config dropped, residue alive): preview 200 with `has_config:false`, cascade completes, journal
records `skipped_no_config` for hetzner + descriptor.
- `TestDeleteCascade_404WhenNothingRemains`.
Also: `backlog/ROADMAP.md` R-25b's live pointer updated `…draft-v3.md``…draft-v4.md` so the
citation resolves (historical references in the rehearsal validation doc and past REPORTs left
as-is — they describe the file as it was named then).
| # | Pre-fix shape restored | Failure observed |
|---|---|---|
| 6 | residue leg removed (the v0.69.0 shape) | `residue after cascade = {Reports:1 AppTelemetry:1 NotificationPrefs:1 SelfBindTokens:1 ApplianceRegistrations:1}` + "customer is STILL on the Customers list" + appliance/self-bind rows outlived their customer |
| 7 | `cfg == nil` 404 restored | `ghost preview = 404, want 200` |
## Not done / unchanged
Full suite green (`go build ./... && go vet ./... && go test ./...`), hub confirm gate OK.
- C7 not flipped; the capability-map row untouched; no other runbook touched.
- No CHANGELOG entry: this repo's changelogs are per-area (`hub/`, `scripts/`, `website/`) and none
of those areas changed; prior docs-only commits (e.g. TASK-H nodes.md) followed the same pattern.
## Still outstanding
The STOP-gated live leg. `demo-vm-felhom` is now the natural subject — it is a real ghost, the
operator has authorised it, and clearing it proves the v0.70.0 path end to end on production data.
It cannot prove the **Hetzner** teardown (no config row → `skipped_no_config`), so a scratch customer
with offsite provisioned is still needed for that half.
*(Previous REPORT — TASK-I, hub v0.69.0/v0.70.0 DELETE cascade — preserved in git history at
`9b3381b` and summarized in `backlog/ROADMAP.md` R-25b.)*