Files
felhom.eu/REPORT.md
T
admin 3f7ac8ee6e
gates / gates (push) Successful in 20s
the backup promise is kept: photos deleted and returned byte-identical
The capability map's journey row now carries the half it could never finish: five
photos in, deleted the way a child would, the old route refusing and touching
nothing, the off-site restore returning them, and them opening — sha256 identical,
5 of 5, with a negative control.

Stated with it, because both are true: the bind needed ZERO operator presses (the
box registered itself and used the mail the hub sent itself), but the PBS cascade
needed ONE — the Re-issue press R-511 documents, which then succeeded because of
this morning's ep0 grant.

R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on
day one — a fresh box waits at „Kulcsletétre vár" until the household creates its
recovery code, and nothing asks them to, while the tier-1 row already promises that
copy. R-544 records a log line that says „escrow deleted" where the effect is
demotion to retained custody.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 20:24:50 +02:00

91 lines
6.2 KiB
Markdown

## Claims in the prompt that turned out wrong — settled so far (measured, not argued)
1. **"A custom PBS role can carry just `Datastore.Modify`" — FALSE, and the prompt itself flagged it
as unverified.** Proxmox Backup Server has **no role-create command** and no custom roles: the CLI
describes `<role>` as "Enum representing roles via their [PRIVILEGES] combination", and
`proxmox-backup-manager` offers no `role` subcommand at all. I then measured the narrowest
BUILT-IN role by applying it and reading the effective permissions back: `DatastorePowerUser`
grants **Datastore.Backup + Datastore.Prune only** — it does not help. `DatastoreAdmin` grants
Audit, Backup, Modify, Prune, Read, Verify, and is therefore the narrowest role that works. It is
applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only; the per-customer
`DatastoreBackup` entries are untouched.
2. **"The banner unit can learn the claimed state" — NOT REACHABLE, as the prompt suspected.** Two
measurements: the one-shot bind delivery emits `FELHOM_CUSTOMER_ID`, `FELHOM_RETRIEVAL_PASSPHRASE`,
`FELHOM_MODE` and `FELHOM_EXTRA_ARGS` — **no domain** — so the console cannot name the dashboard
URL without inventing it; and the unit hands over to the host install and exits, so the later CLAIM
happens when nothing is watching. What IS reachable, and is what shipped: the pairing code stops
being the last thing on the screen the moment the bind lands. The rest of R-535 is recorded as a
residue rather than implied away.
3. **The baseline "ISO 1.27.1 published → target 1.28.0" needed care, and the care found a trap.**
`installer-v1.28.0` **already existed as a git tag** — from 2026-08-13 — because the install SCRIPT
and the ISO IMAGE are two separately numbered artifacts (`SCRIPT_VERSION` vs `ISO_VERSION`, and
`build-felhom-iso.sh` says so in a comment). The published ISO really was 1.27.1 (confirmed live:
the bucket serves 1.27.1 and 404s 1.28.0), so the target is right — but a session that read the tag
list as the ISO history would have concluded 1.28.0 was already published.
4. **Still open at the time of writing:** "the off-site wizard's full restore brings back deleted files
for nextcloud" — proven for immich in July and read from the design for nextcloud; Part E walks it.
5. **„The off-site wizard's full restore brings back deleted files for nextcloud" — TRUE, now measured.**
It was proven for immich in July and read from the design for nextcloud. Tonight it was walked:
five photos, deleted, returned byte-identical (sha256 5/5, negative control).
6. **My own wrong reading, recorded because I nearly filed it as a product fault.** I reported the data
drive as „formatted but not mounted, 42 minutes on" from `/api/disks/candidates`. The storage page
said the opposite and was right — the drive was mounted, registered and default. The endpoint reports
raw disks from the agent, not what the controller has registered (now **R-542**).
## What shipped
**controller v0.244.0** — the backup label is PER TIER and no longer claims files a Tier-1 unit cannot
hold (R-537); a unit restore REFUSES before touching anything when it cannot return the app's drive-side
files, and names the route that can (R-538); `app_deployed` moved from the deploy's acceptance to its
completion, with `app_deploy_started` / `app_deploy_failed` as the honest pair (R-536). Plus the pending
„0 B" tile fix.
**hub v0.116.0** — off-site backup is ON by default for a new customer (shared, 100 GB prefilled), and
the two new deploy event types are registered in both `allowedEventTypes` and `customerMessages`.
**ep0** — one narrow grant: `DatastoreAdmin` for the hub's `felhom@pbs` on `/datastore/felhom-offsite`
only. The narrowest role was MEASURED: `DatastorePowerUser` carries Backup+Prune only, and PBS has no
custom roles. Per-customer `DatastoreBackup` entries untouched.
**ISO 1.28.0** — built, gate-checked, installed and walked. **NOT published** — that is the operator's
call and the one STOP of this task.
**golden 0.244.0** — baked, published, vouched as a three-field change (agent and min_agent unchanged at
0.131.0), fleet floor raised 0.242.0 → 0.244.0 and already delivering (demo-felhom moved itself).
## Red-proofs
Each fix seen failing with its own sentence, then passing: the app-shaped label restored → the Tier-1
assertion fails; the guard disabled → „a restore that cannot return the files must refuse" fails; the
accept-time call put back → „the deploy handler announces an INSTALLED app at accept time" fails; the
success hook removed → „the deploy ended and nothing was told about it" fails; the hub default dropped →
„the new-customer form does not default the off-site copy ON" fails.
## The walk, end to end (evidence: `audits/evidence-backup-promise-2026-09-16/`)
Fresh VM from the BUILT image → Felhom's own first screen, no admin URL → registered itself → **bound
with zero operator presses** (the mail the hub sent itself after the morning's host delete) → claimed →
landed on agent 0.131.0 + controller 0.244.0 → data drive registered → Nextcloud deployed → five photos
in → tier-1 leg → **the PBS cascade stopped at R-511's refusal and needed ONE operator press**, which
then succeeded because of this morning's grant → escrow ceremony (re-auth required; code shown once,
captured out-of-band) → tier-3 „Sikeres" → photos deleted → **the old route REFUSED and touched nothing**
→ off-site restore (verification copy, then reconstitution: „5 fájl és 3 adatkötet és az adatbázis") →
**the photos open, byte-identical** → teardown in three layers → the automatic connect e-mail again, one
second after the host delete.
## Rows
Opened: R-539, R-540, R-541, R-542, **R-543** (P1 — off-site on by default is not off-site working on day
one), R-544. Closed: R-511, R-534, R-536, R-537, R-538. Register 236 → 244 open.
## Checks
`repo_gates.py --fast` green at every push; controller `go build/vet/test` green; controller gates 15/15;
`unproven.py --summary` unchanged at 35 of 55 not walked. Secret-leak check on the evidence: six real
secret values as needles, planted control matched 6/6, committed evidence 0.