the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s

The capability map's journey row now carries the half it could never finish: five
photos in, deleted the way a child would, the old route refusing and touching
nothing, the off-site restore returning them, and them opening — sha256 identical,
5 of 5, with a negative control.

Stated with it, because both are true: the bind needed ZERO operator presses (the
box registered itself and used the mail the hub sent itself), but the PBS cascade
needed ONE — the Re-issue press R-511 documents, which then succeeded because of
this morning's ep0 grant.

R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on
day one — a fresh box waits at „Kulcsletétre vár" until the household creates its
recovery code, and nothing asks them to, while the tier-1 row already promises that
copy. R-544 records a log line that says „escrow deleted" where the effect is
demotion to retained custody.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 20:24:50 +02:00
parent c18efc0610
commit 3f7ac8ee6e
5 changed files with 177 additions and 69 deletions
+77 -68
View File
@@ -1,81 +1,90 @@
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
## Claims in the prompt that turned out wrong — settled so far (measured, not argued)
## Claims in the prompt that turned out wrong — first, as asked
1. **"A custom PBS role can carry just `Datastore.Modify`" — FALSE, and the prompt itself flagged it
as unverified.** Proxmox Backup Server has **no role-create command** and no custom roles: the CLI
describes `<role>` as "Enum representing roles via their [PRIVILEGES] combination", and
`proxmox-backup-manager` offers no `role` subcommand at all. I then measured the narrowest
BUILT-IN role by applying it and reading the effective permissions back: `DatastorePowerUser`
grants **Datastore.Backup + Datastore.Prune only** — it does not help. `DatastoreAdmin` grants
Audit, Backup, Modify, Prune, Read, Verify, and is therefore the narrowest role that works. It is
applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only; the per-customer
`DatastoreBackup` entries are untouched.
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
own bake, and the guest runs controller 0.243.0 with no self-update.
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
customer's namespace, empty before and after.
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
true.
2. **"The banner unit can learn the claimed state" — NOT REACHABLE, as the prompt suspected.** Two
measurements: the one-shot bind delivery emits `FELHOM_CUSTOMER_ID`, `FELHOM_RETRIEVAL_PASSPHRASE`,
`FELHOM_MODE` and `FELHOM_EXTRA_ARGS` — **no domain** — so the console cannot name the dashboard
URL without inventing it; and the unit hands over to the host install and exits, so the later CLAIM
happens when nothing is watching. What IS reachable, and is what shipped: the pairing code stops
being the last thing on the screen the moment the bind lands. The rest of R-535 is recorded as a
residue rather than implied away.
## What I exercised
3. **The baseline "ISO 1.27.1 published → target 1.28.0" needed care, and the care found a trap.**
`installer-v1.28.0` **already existed as a git tag** — from 2026-08-13 — because the install SCRIPT
and the ISO IMAGE are two separately numbered artifacts (`SCRIPT_VERSION` vs `ISO_VERSION`, and
`build-felhom-iso.sh` says so in a comment). The published ISO really was 1.27.1 (confirmed live:
the bucket serves 1.27.1 and 404s 1.28.0), so the target is right — but a session that read the tag
list as the ISO history would have concluded 1.28.0 was already published.
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
morning-after checks and a three-layer teardown.
4. **Still open at the time of writing:** "the off-site wizard's full restore brings back deleted files
for nextcloud" — proven for immich in July and read from the design for nextcloud; Part E walks it.
## What broke — product, and mine
5. **„The off-site wizard's full restore brings back deleted files for nextcloud" — TRUE, now measured.**
It was proven for immich in July and read from the design for nextcloud. Tonight it was walked:
five photos, deleted, returned byte-identical (sha256 5/5, negative control).
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
the app listing files it cannot open, after making the app's own wastebasket unreachable).
6. **My own wrong reading, recorded because I nearly filed it as a product fault.** I reported the data
drive as „formatted but not mounted, 42 minutes on" from `/api/disks/candidates`. The storage page
said the opposite and was right — the drive was mounted, registered and default. The endpoint reports
raw disks from the agent, not what the controller has registered (now **R-542**).
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
restoring caps after the memory test — the second time this exact mistake has happened, repaired
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
before reading it off the dashboard.
## What shipped
**controller v0.244.0** — the backup label is PER TIER and no longer claims files a Tier-1 unit cannot
hold (R-537); a unit restore REFUSES before touching anything when it cannot return the app's drive-side
files, and names the route that can (R-538); `app_deployed` moved from the deploy's acceptance to its
completion, with `app_deploy_started` / `app_deploy_failed` as the honest pair (R-536). Plus the pending
„0 B" tile fix.
**hub v0.116.0** — off-site backup is ON by default for a new customer (shared, 100 GB prefilled), and
the two new deploy event types are registered in both `allowedEventTypes` and `customerMessages`.
**ep0** — one narrow grant: `DatastoreAdmin` for the hub's `felhom@pbs` on `/datastore/felhom-offsite`
only. The narrowest role was MEASURED: `DatastorePowerUser` carries Backup+Prune only, and PBS has no
custom roles. Per-customer `DatastoreBackup` entries untouched.
**ISO 1.28.0** — built, gate-checked, installed and walked. **NOT published** — that is the operator's
call and the one STOP of this task.
**golden 0.244.0** — baked, published, vouched as a three-field change (agent and min_agent unchanged at
0.131.0), fleet floor raised 0.242.0 → 0.244.0 and already delivering (demo-felhom moved itself).
## Red-proofs
Each fix seen failing with its own sentence, then passing: the app-shaped label restored → the Tier-1
assertion fails; the guard disabled → „a restore that cannot return the files must refuse" fails; the
accept-time call put back → „the deploy handler announces an INSTALLED app at accept time" fails; the
success hook removed → „the deploy ended and nothing was told about it" fails; the hub default dropped →
„the new-customer form does not default the off-site copy ON" fails.
## The walk, end to end (evidence: `audits/evidence-backup-promise-2026-09-16/`)
Fresh VM from the BUILT image → Felhom's own first screen, no admin URL → registered itself → **bound
with zero operator presses** (the mail the hub sent itself after the morning's host delete) → claimed →
landed on agent 0.131.0 + controller 0.244.0 → data drive registered → Nextcloud deployed → five photos
in → tier-1 leg → **the PBS cascade stopped at R-511's refusal and needed ONE operator press**, which
then succeeded because of this morning's grant → escrow ceremony (re-auth required; code shown once,
captured out-of-band) → tier-3 „Sikeres" → photos deleted → **the old route REFUSED and touched nothing**
→ off-site restore (verification copy, then reconstitution: „5 fájl és 3 adatkötet és az adatbázis") →
**the photos open, byte-identical** → teardown in three layers → the automatic connect e-mail again, one
second after the host delete.
## Rows
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
## The automatic connect e-mail (R-509) — PASSED
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
requirement was two minutes.
## Verdict
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
things worse.
## Teardown
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
Opened: R-539, R-540, R-541, R-542, **R-543** (P1 — off-site on by default is not off-site working on day
one), R-544. Closed: R-511, R-534, R-536, R-537, R-538. Register 236 → 244 open.
## Checks
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
`head_sha`.
`repo_gates.py --fast` green at every push; controller `go build/vet/test` green; controller gates 15/15;
`unproven.py --summary` unchanged at 35 of 55 not walked. Secret-leak check on the evidence: six real
secret values as needles, planted control matched 6/6, committed evidence 0.