POST /configs/{id}/delete now runs hosts -> RESET -> purge behind three
acknowledgements, a typed customer-id, a stale-preview check and the
ONLINE-host refusal (every gate before any write, so a refusal has zero
side effects). The shallow handleConfigDelete is gone.
Two invariants are asserted, not just commented: ruling 3 is preserved by
construction (leg 2 never sees a host row) and retained escrow custody is
purged exactly once, in leg 3 (leg 2 runs with purgeEscrow=false).
handleCustomerReset's committed half was extracted as commitCustomerReset;
the standalone RESET path is byte-identical to v0.68.1 and its suite is
untouched. Five red-proofs run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J55BQE1gE2V4ffud5jweGS
6.9 KiB
REPORT — TASK-I: R-25b, customer DELETE becomes the guided full-teardown cascade
Date: 2026-07-21 · Repo: felhom.eu · Baseline: f59aa97 (clean, HEAD == origin/main)
Scope: hub only — v0.68.1 → v0.69.0. No agent / controller / catalog change.
What shipped
POST /configs/{id}/delete (same route, new behaviour) is now the guided full-teardown cascade.
GET on the same path returns the dialog's live inventory. The shallow handleConfigDelete is
gone.
Three legs, fixed order:
- hosts — every host row via
store.DeleteHost(hostID, true); escrow DEMOTED to retained custody, never destroyed. Host-delete's own ONLINE rule is kept: an ONLINE host refuses the whole cascade, checked for every host up front so it never half-runs. - reset — the committed RESET sequence verbatim (Hetzner → PBS → claim → descriptor → DB purge)
through the newly extracted
commitCustomerReset, called withpurgeEscrow=false. - purge —
store.DeleteCustomerConfig: the customer record and all escrow ciphertext.
Two invariants are asserted, not merely commented:
- Ruling 3 by construction — leg 2 can only run after leg 1, so the RESET sequence never sees a host row. The standalone RESET handler's 409 gate is untouched.
- Custody purged exactly ONCE, in leg 3 — leg 1 demotes; leg 2 runs with
purgeEscrow=false; leg 3 is the one true purge point (v0.60.1).
Gates, all before any write (a refusal has zero side effects): three acknowledgements
(ack_hosts / ack_reset / ack_purge, each exactly 1), the typed customer-id, a stale-preview
check (the acknowledged host count must still match live → else 409), and the ONLINE-host refusal.
No force flag, no skip flag, no partial-run downgrade.
Resume: a failed leg retains the customer_resets journal row and the HTTP error names the leg.
The dialog renders the incomplete journal and offers Resume; a re-run is idempotent and must pass
every gate again (acknowledgements are not cached across attempts).
UI: Danger zone → Delete customer… → guided dialog (inventory panel: hosts by name + status, offsite repository identifier, PBS namespace, custody state; three consequence checkboxes; typed customer-id; one submit). Client-side checks are convenience only.
Refactor — standalone RESET behaviour unchanged
handleCustomerReset's committed half became
commitCustomerReset(ctx, cfg, resetID int64, purgeEscrow bool) *resetLegError. The standalone path
is byte-identical to v0.68.1: same leg order, same leg names, same operator-facing messages, same
status codes. Its existing suite is untouched and green.
Files
| File | Change |
|---|---|
hub/internal/web/customer_delete.go |
NEW — the cascade + preview |
hub/internal/web/customer_delete_test.go |
NEW — scenarios A–E |
hub/internal/web/customer_reset.go |
commitCustomerReset + resetLegError extracted |
hub/internal/web/configs.go |
handleConfigDelete removed (replaced by a do-not-reintroduce note) |
hub/internal/web/server.go |
route: GET → preview, POST → cascade |
hub/internal/web/templates/customer_unified.html |
guided dialog replaces the one-click Delete |
hub/internal/web/customer_edit_tab_test.go |
the delete redirect case now posts the full acks |
hub/CHANGELOG.md, REUSE.md, documentation/backlog/ROADMAP.md, documentation/architecture/00-capability-map.md, documentation/runbooks/RUNBOOK-onboarding-draft-v3.md |
docs |
Tests + red-proofs
Green gate in hub/: go build ./... && go vet ./... && go test ./... — all green, no flakes.
New coverage (customer_delete_test.go):
- A — leg ORDER, observed from inside leg 2 via a
tenancyProvisionerfake whoseDeprovisionsnapshots store state: at that instant hosts = 0 (leg 1 done), the customer row is still present (leg 3 not started), retained custody still present (leg 1 demoted, did not purge). Plus final state, the four journal legs stampedok, completion stamped, and thecustomer_deletedaudit event surviving the record. - B — 9 fail-closed gate cases (each missing ack, an ack sent as
yes, id mismatch, id absent, host count moved, host count absent, ONLINE host). Each asserts the status code and that the host, the customer row, the current escrow and the retained custody are untouched, and that zero external calls fired, and that no journal row was opened. - C — resume: an injected PBS failure → 502 naming the leg; hosts gone, customer + custody
SURVIVE, journal retained with
hosts=ok pbs=failed; re-run converges and completes. Plus: a resume without ack #3 is still refused. - E — custody:
commitCustomerReset(..., purgeEscrow=false)leaves the retained blobs; leg 3 purges them. - Preview: names the real host/offsite/PBS/custody facts and leaks no secret (one-time password, API keys, escrow blobs all asserted absent).
Five red-proofs run, each failed red with the wrong value visible, then restored (git diff clean):
| # | Pre-fix shape restored | Failure observed |
|---|---|---|
| 1 | ack gate disabled | status = 303, want 400 + host deleted, customer deleted, custody destroyed, PBS deprovision fired, journal row opened |
| 2 | stale-preview gate weakened to always pass | status = 303, want 409 + the same six non-effect assertions |
| 3 | ONLINE-host gate removed | status = 303, want 409 + live host deleted |
| 4 | leg order inverted (RESET leg before the host leg) | at the RESET leg the customer still had 1 host(s) — leg 1 must complete FIRST |
| 5 | cascade's RESET leg called with purgeEscrow=true |
retained blobs after the RESET leg = 0, want 2 |
Method / limits
Unit-land only so far — the STOP-gated live leg has not been run (see below). Validation method:
Go tests against a real SQLite store on t.TempDir() with fakes at the existing tenancyProvisioner
seam. No browser is available on DooPlex; the dialog's markup is covered by the existing
customer-page render tests (exactly one /configs/{id}/delete form on the page) — a strict
click-through remains a manual operator pass.
Not covered by unit tests: the Hetzner offsite Deprovision leg (offsite.Provisioner is a
concrete type, no interface seam — same as the standalone RESET suite). That leg is exactly what the
live run is for.
STOP — operator-present live leg (NOT yet run)
Create a scratch customer on a throwaway domain, provision offsite only (no host — cheap), run the cascade end-to-end, then verify from outside the hub that the Hetzner repository is gone and the customer row is purged. Never run against Demo Ügyfél, Demo HP, or Peti. Failure paths are unit-proven; the live leg proves the happy path + external teardown only.
Deploy status
Code pushed to main; hub image 0.69.0 built and pushed; manifests/hub.yaml bumped and synced —
see the verification lines below.