Files
felhom.eu/REPORT.md
T
admin 61dbd870c3 feat(hub): v0.69.0 — customer DELETE is the guided full-teardown cascade (R-25b)
POST /configs/{id}/delete now runs hosts -> RESET -> purge behind three
acknowledgements, a typed customer-id, a stale-preview check and the
ONLINE-host refusal (every gate before any write, so a refusal has zero
side effects). The shallow handleConfigDelete is gone.

Two invariants are asserted, not just commented: ruling 3 is preserved by
construction (leg 2 never sees a host row) and retained escrow custody is
purged exactly once, in leg 3 (leg 2 runs with purgeEscrow=false).

handleCustomerReset's committed half was extracted as commitCustomerReset;
the standalone RESET path is byte-identical to v0.68.1 and its suite is
untouched. Five red-proofs run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J55BQE1gE2V4ffud5jweGS
2026-07-21 19:31:48 +02:00

6.9 KiB
Raw Blame History

REPORT — TASK-I: R-25b, customer DELETE becomes the guided full-teardown cascade

Date: 2026-07-21 · Repo: felhom.eu · Baseline: f59aa97 (clean, HEAD == origin/main) Scope: hub only — v0.68.1 → v0.69.0. No agent / controller / catalog change.

What shipped

POST /configs/{id}/delete (same route, new behaviour) is now the guided full-teardown cascade. GET on the same path returns the dialog's live inventory. The shallow handleConfigDelete is gone.

Three legs, fixed order:

  1. hosts — every host row via store.DeleteHost(hostID, true); escrow DEMOTED to retained custody, never destroyed. Host-delete's own ONLINE rule is kept: an ONLINE host refuses the whole cascade, checked for every host up front so it never half-runs.
  2. reset — the committed RESET sequence verbatim (Hetzner → PBS → claim → descriptor → DB purge) through the newly extracted commitCustomerReset, called with purgeEscrow=false.
  3. purgestore.DeleteCustomerConfig: the customer record and all escrow ciphertext.

Two invariants are asserted, not merely commented:

  • Ruling 3 by construction — leg 2 can only run after leg 1, so the RESET sequence never sees a host row. The standalone RESET handler's 409 gate is untouched.
  • Custody purged exactly ONCE, in leg 3 — leg 1 demotes; leg 2 runs with purgeEscrow=false; leg 3 is the one true purge point (v0.60.1).

Gates, all before any write (a refusal has zero side effects): three acknowledgements (ack_hosts / ack_reset / ack_purge, each exactly 1), the typed customer-id, a stale-preview check (the acknowledged host count must still match live → else 409), and the ONLINE-host refusal. No force flag, no skip flag, no partial-run downgrade.

Resume: a failed leg retains the customer_resets journal row and the HTTP error names the leg. The dialog renders the incomplete journal and offers Resume; a re-run is idempotent and must pass every gate again (acknowledgements are not cached across attempts).

UI: Danger zone → Delete customer… → guided dialog (inventory panel: hosts by name + status, offsite repository identifier, PBS namespace, custody state; three consequence checkboxes; typed customer-id; one submit). Client-side checks are convenience only.

Refactor — standalone RESET behaviour unchanged

handleCustomerReset's committed half became commitCustomerReset(ctx, cfg, resetID int64, purgeEscrow bool) *resetLegError. The standalone path is byte-identical to v0.68.1: same leg order, same leg names, same operator-facing messages, same status codes. Its existing suite is untouched and green.

Files

File Change
hub/internal/web/customer_delete.go NEW — the cascade + preview
hub/internal/web/customer_delete_test.go NEW — scenarios AE
hub/internal/web/customer_reset.go commitCustomerReset + resetLegError extracted
hub/internal/web/configs.go handleConfigDelete removed (replaced by a do-not-reintroduce note)
hub/internal/web/server.go route: GET → preview, POST → cascade
hub/internal/web/templates/customer_unified.html guided dialog replaces the one-click Delete
hub/internal/web/customer_edit_tab_test.go the delete redirect case now posts the full acks
hub/CHANGELOG.md, REUSE.md, documentation/backlog/ROADMAP.md, documentation/architecture/00-capability-map.md, documentation/runbooks/RUNBOOK-onboarding-draft-v3.md docs

Tests + red-proofs

Green gate in hub/: go build ./... && go vet ./... && go test ./...all green, no flakes.

New coverage (customer_delete_test.go):

  • A — leg ORDER, observed from inside leg 2 via a tenancyProvisioner fake whose Deprovision snapshots store state: at that instant hosts = 0 (leg 1 done), the customer row is still present (leg 3 not started), retained custody still present (leg 1 demoted, did not purge). Plus final state, the four journal legs stamped ok, completion stamped, and the customer_deleted audit event surviving the record.
  • B — 9 fail-closed gate cases (each missing ack, an ack sent as yes, id mismatch, id absent, host count moved, host count absent, ONLINE host). Each asserts the status code and that the host, the customer row, the current escrow and the retained custody are untouched, and that zero external calls fired, and that no journal row was opened.
  • C — resume: an injected PBS failure → 502 naming the leg; hosts gone, customer + custody SURVIVE, journal retained with hosts=ok pbs=failed; re-run converges and completes. Plus: a resume without ack #3 is still refused.
  • E — custody: commitCustomerReset(..., purgeEscrow=false) leaves the retained blobs; leg 3 purges them.
  • Preview: names the real host/offsite/PBS/custody facts and leaks no secret (one-time password, API keys, escrow blobs all asserted absent).

Five red-proofs run, each failed red with the wrong value visible, then restored (git diff clean):

# Pre-fix shape restored Failure observed
1 ack gate disabled status = 303, want 400 + host deleted, customer deleted, custody destroyed, PBS deprovision fired, journal row opened
2 stale-preview gate weakened to always pass status = 303, want 409 + the same six non-effect assertions
3 ONLINE-host gate removed status = 303, want 409 + live host deleted
4 leg order inverted (RESET leg before the host leg) at the RESET leg the customer still had 1 host(s) — leg 1 must complete FIRST
5 cascade's RESET leg called with purgeEscrow=true retained blobs after the RESET leg = 0, want 2

Method / limits

Unit-land only so far — the STOP-gated live leg has not been run (see below). Validation method: Go tests against a real SQLite store on t.TempDir() with fakes at the existing tenancyProvisioner seam. No browser is available on DooPlex; the dialog's markup is covered by the existing customer-page render tests (exactly one /configs/{id}/delete form on the page) — a strict click-through remains a manual operator pass.

Not covered by unit tests: the Hetzner offsite Deprovision leg (offsite.Provisioner is a concrete type, no interface seam — same as the standalone RESET suite). That leg is exactly what the live run is for.

STOP — operator-present live leg (NOT yet run)

Create a scratch customer on a throwaway domain, provision offsite only (no host — cheap), run the cascade end-to-end, then verify from outside the hub that the Hetzner repository is gone and the customer row is purged. Never run against Demo Ügyfél, Demo HP, or Peti. Failure paths are unit-proven; the live leg proves the happy path + external teardown only.

Deploy status

Code pushed to main; hub image 0.69.0 built and pushed; manifests/hub.yaml bumped and synced — see the verification lines below.