R-534 and R-511 CLOSED — the ep0 grant proven end to end on a fresh box
gates / gates (push) Successful in 20s

The rebuilt-customer case reproduced by itself: the WG-registration hook refused
exactly as R-511 describes and named the Re-issue action. Pressing it then worked —
reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two
seconds later. No permission error.

This morning the identical action returned „missing Datastore.Modify … status 255"
and a 502. The only change in between is the narrow grant on ep0, and the narrowest
role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only,
and PBS has no custom roles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 19:26:54 +02:00
parent 63e2de9b9e
commit a73abf04db
3 changed files with 73 additions and 2 deletions
@@ -10,3 +10,33 @@
## correctly rather than a blocked step. The grant is proven end-to-end by the FRESH BOX in Part E:
## its WG registration must provision `felhom-pbs` with no hand — which is exactly the path that
## failed on 2026-09-16 with „missing Datastore.Modify". R-534 and R-511 stay open until that run.
## 2026-09-16T17:25:34Z PART C.2 — pressing „Re-issue PBS credentials" on the FRESH box (the hub itself asked for it)
submitting the customer form to the re-issue action, fields: _csrf, cf_api_token, cf_tunnel_token, customer_id, customer_name, domain, dr_tier, email, git_token, git_username, offsite_box_type, offsite_enabled, offsite_quota_gb, offsite_type, pbsdr_storage_id
POST pbsdr-reissue -> http=200
hub log right after:
2026/09/16 19:25:36 [INFO] tenantsync: reissue ok for tester-1 (ns=tester-1, token_id=felhom@pbs!tester-1; secret withheld from logs)
2026/09/16 19:25:36 [INFO] pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, ns tester-1, token_id felhom@pbs!tester-1, gen 2; fresh consume-once secret stored, withheld from logs)
2026/09/16 19:25:38 [INFO] host-report from tester-1-33b6a9 (1 guests, 2 storage targets, 1 backups, 0 restore-tests, 0 pbs-snapshots, 11186 bytes)
2026/09/16 19:25:38 [INFO] DR-recipe host-half stored for customer tester-1 (host tester-1-33b6a9, v1)
2026/09/16 19:25:38 [INFO] pbs token secret consumed by host tester-1-33b6a9 (single-use; value withheld from logs)
## R-511 REPRODUCED ON THE FRESH BOX, from the hub's own log (2026-09-16 18:01:19 CEST):
## „[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a
## PBS token for tester-1 but the hub has no descriptor — use the explicit „Re-issue PBS
## credentials" action — save the customer config to retry"
## That is R-511's exact shape: a customer whose box was rebuilt keeps the ep0 token, the automatic
## hook refuses, and the product names the one action that helps. It happened by itself on a box
## installed an hour after the ep0 grant, so the walk handed back the very test that could not be
## run this morning (the button renders only once a descriptor exists, and there was no host then).
## AND THE ADOPT SUCCEEDED — the ep0 grant proven END TO END (2026-09-16 19:25:36 CEST):
## „tenantsync: reissue ok for tester-1 (ns=tester-1, token_id=felhom@pbs!tester-1; secret withheld
## from logs)"
## „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, ns tester-1, token_id felhom@pbs!tester-1,
## gen 2; fresh consume-once secret stored, withheld from logs)"
## „pbs token secret consumed by host tester-1-33b6a9 (single-use; value withheld from logs)"
## Two seconds later the box reported back and the hub stored the DR-recipe host half.
## THE COMPARISON THAT MATTERS: this morning the same action ended with
## „Process exited with status 255 (stderr: Error: permission check failed — missing
## Datastore.Modify on /datastore/felhom-offsite)" -> HTTP 502, nothing written.
## The only thing that changed in between is the narrow grant on ep0 (DatastoreAdmin for the hub's
## felhom@pbs at the datastore root). R-534 and R-511 are therefore both closed by a live run, not
## by an argument — and the rebuilt-customer case they describe reproduced BY ITSELF on a fresh box.
@@ -53,3 +53,44 @@
## nothing asks the household to perform. `POST /backup/offbox/run` -> 302 and produced no snapshot;
## the controller log shows only `offsite-credential-retry`, no restic activity. So on day one the
## sentence above promises protection by a copy that does not exist yet.
## A TRAP THIS PROJECT ALREADY KNOWS, hit again and recorded: my `grep -oE` with an accented pattern
## („Helyre…") died with „exceeds complexity limits" and told me nothing, while two guessed paths
## (/backups/escrow, /escrow) 404'd. Parsing the page in Python with the ASCII fragment „Helyre"
## found the link at once: /backup/escrow („Helyreállítási kód létrehozása").
## The standing rule is „never let an accented pattern gate a conclusion" — the failure mode here
## was not a false 0 but a dead tool, and the cost was the same: two wrong guesses.
## THE ESCROW CEREMONY PAGE (/backup/escrow), in the household's own words:
## „1. Előfeltételek ellenőrzése — Ellenőrzés folyamatban…"
## „2. Fontos tudnivalók — A helyreállítási kód a mentései utolsó kulcsa. Pontosan egyszer jelenik
## meg — a rendszer sehol nem tárolja, és a Felhom sem ismeri. Ha a szerver megsemmisül, a távoli
## mentések CSAK ezzel a kóddal állíthatók vissza."
## So the ceremony is customer-facing, honest about what the code is, and shows it exactly once.
## That is also why this walk captures it into an out-of-band 0600 file at the moment it appears:
## the off-site restore later in this phase cannot happen without it, and nothing can re-issue it.
## 2026-09-16T17:2xZ THE ESCROW CEREMONY CANNOT START YET — preflight, verbatim:
## agent_supported: true · escrow_state: „pending"
## pbs_storage_id NOT OK — „escrow.pbs_storage_id not configured"
## dr_tier NOT OK — „DR tier not applied on this host"
## age_binary ok — /usr/bin/age
## hub_upload ok — hub upload target configured
## staged_secret ok — staged secret present
## sudo_grant ok — sudo grant listed (list-mode)
## So four of six prerequisites are already in place on a box that installed itself an hour ago;
## the two red ones are the PBS-DR cascade (host → WG peer → apply), which is precisely what the
## ep0 grant of this morning (R-534) exists to unblock. The hub had said at 17:05 that the
## descriptor „applies once the cascade is ready (host → WG peer → apply)" — a host now exists.
## Checked next on the hub side; whatever it says is the end-to-end verdict on that grant.
## THE CASCADE, in the hub's own words (customer page, DR tier widget):
## done „host enrolled (tester-1-33b6a9)"
## done „WG tunnel peer registered"
## waiting „provisions automatically when the WG peer registers"
## waiting „ceremony possible once the descriptor is applied on the box"
## and the agent's capability list on the SAME page:
## „escrow-ceremony critical inactive — customer recovery-code ceremony (controller-driven)
## — disabled by configuration"
## „pbsdr-create / pbsdr-grant / pbsdr-read inactive — disabled by configuration"
## The WG peer really is registered: ep0 shows FIVE wg peers with fresh handshakes, and the hub
## pushes them every five minutes („wgsync: pushed 5 peers").
## So the chain stops at the APPLY step, because the agent's PBS-DR capabilities are switched off by
## configuration — not because ep0 refused anything. This morning's grant (R-534) is therefore still
## unproven end-to-end: nothing has yet asked ep0 to do the thing the grant allows.