the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
The capability map's journey row now carries the half it could never finish: five photos in, deleted the way a child would, the old route refusing and touching nothing, the off-site restore returning them, and them opening — sha256 identical, 5 of 5, with a negative control. Stated with it, because both are true: the bind needed ZERO operator presses (the box registered itself and used the mail the hub sent itself), but the PBS cascade needed ONE — the Re-issue press R-511 documents, which then succeeded because of this morning's ep0 grant. R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on day one — a fresh box waits at „Kulcsletétre vár" until the household creates its recovery code, and nothing asks them to, while the tier-1 row already promises that copy. R-544 records a log line that says „escrow deleted" where the effect is demotion to retained custody. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,81 +1,90 @@
|
||||
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
|
||||
## Claims in the prompt that turned out wrong — settled so far (measured, not argued)
|
||||
|
||||
## Claims in the prompt that turned out wrong — first, as asked
|
||||
1. **"A custom PBS role can carry just `Datastore.Modify`" — FALSE, and the prompt itself flagged it
|
||||
as unverified.** Proxmox Backup Server has **no role-create command** and no custom roles: the CLI
|
||||
describes `<role>` as "Enum representing roles via their [PRIVILEGES] combination", and
|
||||
`proxmox-backup-manager` offers no `role` subcommand at all. I then measured the narrowest
|
||||
BUILT-IN role by applying it and reading the effective permissions back: `DatastorePowerUser`
|
||||
grants **Datastore.Backup + Datastore.Prune only** — it does not help. `DatastoreAdmin` grants
|
||||
Audit, Backup, Modify, Prune, Read, Verify, and is therefore the narrowest role that works. It is
|
||||
applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only; the per-customer
|
||||
`DatastoreBackup` entries are untouched.
|
||||
|
||||
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
|
||||
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
|
||||
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
|
||||
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
|
||||
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
|
||||
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
|
||||
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
|
||||
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
|
||||
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
|
||||
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
|
||||
own bake, and the guest runs controller 0.243.0 with no self-update.
|
||||
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
|
||||
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
|
||||
customer's namespace, empty before and after.
|
||||
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
|
||||
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
|
||||
true.
|
||||
2. **"The banner unit can learn the claimed state" — NOT REACHABLE, as the prompt suspected.** Two
|
||||
measurements: the one-shot bind delivery emits `FELHOM_CUSTOMER_ID`, `FELHOM_RETRIEVAL_PASSPHRASE`,
|
||||
`FELHOM_MODE` and `FELHOM_EXTRA_ARGS` — **no domain** — so the console cannot name the dashboard
|
||||
URL without inventing it; and the unit hands over to the host install and exits, so the later CLAIM
|
||||
happens when nothing is watching. What IS reachable, and is what shipped: the pairing code stops
|
||||
being the last thing on the screen the moment the bind lands. The rest of R-535 is recorded as a
|
||||
residue rather than implied away.
|
||||
|
||||
## What I exercised
|
||||
3. **The baseline "ISO 1.27.1 published → target 1.28.0" needed care, and the care found a trap.**
|
||||
`installer-v1.28.0` **already existed as a git tag** — from 2026-08-13 — because the install SCRIPT
|
||||
and the ISO IMAGE are two separately numbered artifacts (`SCRIPT_VERSION` vs `ISO_VERSION`, and
|
||||
`build-felhom-iso.sh` says so in a comment). The published ISO really was 1.27.1 (confirmed live:
|
||||
the bucket serves 1.27.1 and 404s 1.28.0), so the target is right — but a session that read the tag
|
||||
list as the ISO history would have concluded 1.28.0 was already published.
|
||||
|
||||
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
|
||||
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
|
||||
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
|
||||
morning-after checks and a three-layer teardown.
|
||||
4. **Still open at the time of writing:** "the off-site wizard's full restore brings back deleted files
|
||||
for nextcloud" — proven for immich in July and read from the design for nextcloud; Part E walks it.
|
||||
|
||||
## What broke — product, and mine
|
||||
5. **„The off-site wizard's full restore brings back deleted files for nextcloud" — TRUE, now measured.**
|
||||
It was proven for immich in July and read from the design for nextcloud. Tonight it was walked:
|
||||
five photos, deleted, returned byte-identical (sha256 5/5, negative control).
|
||||
|
||||
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
|
||||
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
|
||||
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
|
||||
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
|
||||
the app listing files it cannot open, after making the app's own wastebasket unreachable).
|
||||
6. **My own wrong reading, recorded because I nearly filed it as a product fault.** I reported the data
|
||||
drive as „formatted but not mounted, 42 minutes on" from `/api/disks/candidates`. The storage page
|
||||
said the opposite and was right — the drive was mounted, registered and default. The endpoint reports
|
||||
raw disks from the agent, not what the controller has registered (now **R-542**).
|
||||
|
||||
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
|
||||
restoring caps after the memory test — the second time this exact mistake has happened, repaired
|
||||
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
|
||||
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
|
||||
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
|
||||
before reading it off the dashboard.
|
||||
## What shipped
|
||||
|
||||
**controller v0.244.0** — the backup label is PER TIER and no longer claims files a Tier-1 unit cannot
|
||||
hold (R-537); a unit restore REFUSES before touching anything when it cannot return the app's drive-side
|
||||
files, and names the route that can (R-538); `app_deployed` moved from the deploy's acceptance to its
|
||||
completion, with `app_deploy_started` / `app_deploy_failed` as the honest pair (R-536). Plus the pending
|
||||
„0 B" tile fix.
|
||||
|
||||
**hub v0.116.0** — off-site backup is ON by default for a new customer (shared, 100 GB prefilled), and
|
||||
the two new deploy event types are registered in both `allowedEventTypes` and `customerMessages`.
|
||||
|
||||
**ep0** — one narrow grant: `DatastoreAdmin` for the hub's `felhom@pbs` on `/datastore/felhom-offsite`
|
||||
only. The narrowest role was MEASURED: `DatastorePowerUser` carries Backup+Prune only, and PBS has no
|
||||
custom roles. Per-customer `DatastoreBackup` entries untouched.
|
||||
|
||||
**ISO 1.28.0** — built, gate-checked, installed and walked. **NOT published** — that is the operator's
|
||||
call and the one STOP of this task.
|
||||
|
||||
**golden 0.244.0** — baked, published, vouched as a three-field change (agent and min_agent unchanged at
|
||||
0.131.0), fleet floor raised 0.242.0 → 0.244.0 and already delivering (demo-felhom moved itself).
|
||||
|
||||
## Red-proofs
|
||||
|
||||
Each fix seen failing with its own sentence, then passing: the app-shaped label restored → the Tier-1
|
||||
assertion fails; the guard disabled → „a restore that cannot return the files must refuse" fails; the
|
||||
accept-time call put back → „the deploy handler announces an INSTALLED app at accept time" fails; the
|
||||
success hook removed → „the deploy ended and nothing was told about it" fails; the hub default dropped →
|
||||
„the new-customer form does not default the off-site copy ON" fails.
|
||||
|
||||
## The walk, end to end (evidence: `audits/evidence-backup-promise-2026-09-16/`)
|
||||
|
||||
Fresh VM from the BUILT image → Felhom's own first screen, no admin URL → registered itself → **bound
|
||||
with zero operator presses** (the mail the hub sent itself after the morning's host delete) → claimed →
|
||||
landed on agent 0.131.0 + controller 0.244.0 → data drive registered → Nextcloud deployed → five photos
|
||||
in → tier-1 leg → **the PBS cascade stopped at R-511's refusal and needed ONE operator press**, which
|
||||
then succeeded because of this morning's grant → escrow ceremony (re-auth required; code shown once,
|
||||
captured out-of-band) → tier-3 „Sikeres" → photos deleted → **the old route REFUSED and touched nothing**
|
||||
→ off-site restore (verification copy, then reconstitution: „5 fájl és 3 adatkötet és az adatbázis") →
|
||||
**the photos open, byte-identical** → teardown in three layers → the automatic connect e-mail again, one
|
||||
second after the host delete.
|
||||
|
||||
## Rows
|
||||
|
||||
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
|
||||
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
|
||||
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
|
||||
|
||||
## The automatic connect e-mail (R-509) — PASSED
|
||||
|
||||
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
|
||||
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
|
||||
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
|
||||
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
|
||||
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
|
||||
requirement was two minutes.
|
||||
|
||||
## Verdict
|
||||
|
||||
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
|
||||
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
|
||||
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
|
||||
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
|
||||
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
|
||||
things worse.
|
||||
|
||||
## Teardown
|
||||
|
||||
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
|
||||
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
|
||||
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
|
||||
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
|
||||
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
|
||||
Opened: R-539, R-540, R-541, R-542, **R-543** (P1 — off-site on by default is not off-site working on day
|
||||
one), R-544. Closed: R-511, R-534, R-536, R-537, R-538. Register 236 → 244 open.
|
||||
|
||||
## Checks
|
||||
|
||||
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
|
||||
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
|
||||
`head_sha`.
|
||||
`repo_gates.py --fast` green at every push; controller `go build/vet/test` green; controller gates 15/15;
|
||||
`unproven.py --summary` unchanged at 35 of 55 not walked. Secret-leak check on the evidence: six real
|
||||
secret values as needles, planted control matched 6/6, committed evidence 0.
|
||||
|
||||
Reference in New Issue
Block a user