the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
The capability map's journey row now carries the half it could never finish: five photos in, deleted the way a child would, the old route refusing and touching nothing, the off-site restore returning them, and them opening — sha256 identical, 5 of 5, with a negative control. Stated with it, because both are true: the bind needed ZERO operator presses (the box registered itself and used the mail the hub sent itself), but the PBS cascade needed ONE — the Re-issue press R-511 documents, which then succeeded because of this morning's ep0 grant. R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on day one — a fresh box waits at „Kulcsletétre vár" until the household creates its recovery code, and nothing asks them to, while the tier-1 row already promises that copy. R-544 records a log line that says „escrow deleted" where the effect is demotion to retained custody. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,81 +1,90 @@
|
||||
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
|
||||
## Claims in the prompt that turned out wrong — settled so far (measured, not argued)
|
||||
|
||||
## Claims in the prompt that turned out wrong — first, as asked
|
||||
1. **"A custom PBS role can carry just `Datastore.Modify`" — FALSE, and the prompt itself flagged it
|
||||
as unverified.** Proxmox Backup Server has **no role-create command** and no custom roles: the CLI
|
||||
describes `<role>` as "Enum representing roles via their [PRIVILEGES] combination", and
|
||||
`proxmox-backup-manager` offers no `role` subcommand at all. I then measured the narrowest
|
||||
BUILT-IN role by applying it and reading the effective permissions back: `DatastorePowerUser`
|
||||
grants **Datastore.Backup + Datastore.Prune only** — it does not help. `DatastoreAdmin` grants
|
||||
Audit, Backup, Modify, Prune, Read, Verify, and is therefore the narrowest role that works. It is
|
||||
applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only; the per-customer
|
||||
`DatastoreBackup` entries are untouched.
|
||||
|
||||
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
|
||||
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
|
||||
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
|
||||
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
|
||||
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
|
||||
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
|
||||
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
|
||||
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
|
||||
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
|
||||
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
|
||||
own bake, and the guest runs controller 0.243.0 with no self-update.
|
||||
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
|
||||
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
|
||||
customer's namespace, empty before and after.
|
||||
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
|
||||
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
|
||||
true.
|
||||
2. **"The banner unit can learn the claimed state" — NOT REACHABLE, as the prompt suspected.** Two
|
||||
measurements: the one-shot bind delivery emits `FELHOM_CUSTOMER_ID`, `FELHOM_RETRIEVAL_PASSPHRASE`,
|
||||
`FELHOM_MODE` and `FELHOM_EXTRA_ARGS` — **no domain** — so the console cannot name the dashboard
|
||||
URL without inventing it; and the unit hands over to the host install and exits, so the later CLAIM
|
||||
happens when nothing is watching. What IS reachable, and is what shipped: the pairing code stops
|
||||
being the last thing on the screen the moment the bind lands. The rest of R-535 is recorded as a
|
||||
residue rather than implied away.
|
||||
|
||||
## What I exercised
|
||||
3. **The baseline "ISO 1.27.1 published → target 1.28.0" needed care, and the care found a trap.**
|
||||
`installer-v1.28.0` **already existed as a git tag** — from 2026-08-13 — because the install SCRIPT
|
||||
and the ISO IMAGE are two separately numbered artifacts (`SCRIPT_VERSION` vs `ISO_VERSION`, and
|
||||
`build-felhom-iso.sh` says so in a comment). The published ISO really was 1.27.1 (confirmed live:
|
||||
the bucket serves 1.27.1 and 404s 1.28.0), so the target is right — but a session that read the tag
|
||||
list as the ISO history would have concluded 1.28.0 was already published.
|
||||
|
||||
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
|
||||
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
|
||||
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
|
||||
morning-after checks and a three-layer teardown.
|
||||
4. **Still open at the time of writing:** "the off-site wizard's full restore brings back deleted files
|
||||
for nextcloud" — proven for immich in July and read from the design for nextcloud; Part E walks it.
|
||||
|
||||
## What broke — product, and mine
|
||||
5. **„The off-site wizard's full restore brings back deleted files for nextcloud" — TRUE, now measured.**
|
||||
It was proven for immich in July and read from the design for nextcloud. Tonight it was walked:
|
||||
five photos, deleted, returned byte-identical (sha256 5/5, negative control).
|
||||
|
||||
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
|
||||
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
|
||||
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
|
||||
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
|
||||
the app listing files it cannot open, after making the app's own wastebasket unreachable).
|
||||
6. **My own wrong reading, recorded because I nearly filed it as a product fault.** I reported the data
|
||||
drive as „formatted but not mounted, 42 minutes on" from `/api/disks/candidates`. The storage page
|
||||
said the opposite and was right — the drive was mounted, registered and default. The endpoint reports
|
||||
raw disks from the agent, not what the controller has registered (now **R-542**).
|
||||
|
||||
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
|
||||
restoring caps after the memory test — the second time this exact mistake has happened, repaired
|
||||
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
|
||||
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
|
||||
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
|
||||
before reading it off the dashboard.
|
||||
## What shipped
|
||||
|
||||
**controller v0.244.0** — the backup label is PER TIER and no longer claims files a Tier-1 unit cannot
|
||||
hold (R-537); a unit restore REFUSES before touching anything when it cannot return the app's drive-side
|
||||
files, and names the route that can (R-538); `app_deployed` moved from the deploy's acceptance to its
|
||||
completion, with `app_deploy_started` / `app_deploy_failed` as the honest pair (R-536). Plus the pending
|
||||
„0 B" tile fix.
|
||||
|
||||
**hub v0.116.0** — off-site backup is ON by default for a new customer (shared, 100 GB prefilled), and
|
||||
the two new deploy event types are registered in both `allowedEventTypes` and `customerMessages`.
|
||||
|
||||
**ep0** — one narrow grant: `DatastoreAdmin` for the hub's `felhom@pbs` on `/datastore/felhom-offsite`
|
||||
only. The narrowest role was MEASURED: `DatastorePowerUser` carries Backup+Prune only, and PBS has no
|
||||
custom roles. Per-customer `DatastoreBackup` entries untouched.
|
||||
|
||||
**ISO 1.28.0** — built, gate-checked, installed and walked. **NOT published** — that is the operator's
|
||||
call and the one STOP of this task.
|
||||
|
||||
**golden 0.244.0** — baked, published, vouched as a three-field change (agent and min_agent unchanged at
|
||||
0.131.0), fleet floor raised 0.242.0 → 0.244.0 and already delivering (demo-felhom moved itself).
|
||||
|
||||
## Red-proofs
|
||||
|
||||
Each fix seen failing with its own sentence, then passing: the app-shaped label restored → the Tier-1
|
||||
assertion fails; the guard disabled → „a restore that cannot return the files must refuse" fails; the
|
||||
accept-time call put back → „the deploy handler announces an INSTALLED app at accept time" fails; the
|
||||
success hook removed → „the deploy ended and nothing was told about it" fails; the hub default dropped →
|
||||
„the new-customer form does not default the off-site copy ON" fails.
|
||||
|
||||
## The walk, end to end (evidence: `audits/evidence-backup-promise-2026-09-16/`)
|
||||
|
||||
Fresh VM from the BUILT image → Felhom's own first screen, no admin URL → registered itself → **bound
|
||||
with zero operator presses** (the mail the hub sent itself after the morning's host delete) → claimed →
|
||||
landed on agent 0.131.0 + controller 0.244.0 → data drive registered → Nextcloud deployed → five photos
|
||||
in → tier-1 leg → **the PBS cascade stopped at R-511's refusal and needed ONE operator press**, which
|
||||
then succeeded because of this morning's grant → escrow ceremony (re-auth required; code shown once,
|
||||
captured out-of-band) → tier-3 „Sikeres" → photos deleted → **the old route REFUSED and touched nothing**
|
||||
→ off-site restore (verification copy, then reconstitution: „5 fájl és 3 adatkötet és az adatbázis") →
|
||||
**the photos open, byte-identical** → teardown in three layers → the automatic connect e-mail again, one
|
||||
second after the host delete.
|
||||
|
||||
## Rows
|
||||
|
||||
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
|
||||
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
|
||||
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
|
||||
|
||||
## The automatic connect e-mail (R-509) — PASSED
|
||||
|
||||
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
|
||||
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
|
||||
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
|
||||
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
|
||||
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
|
||||
requirement was two minutes.
|
||||
|
||||
## Verdict
|
||||
|
||||
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
|
||||
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
|
||||
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
|
||||
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
|
||||
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
|
||||
things worse.
|
||||
|
||||
## Teardown
|
||||
|
||||
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
|
||||
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
|
||||
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
|
||||
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
|
||||
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
|
||||
Opened: R-539, R-540, R-541, R-542, **R-543** (P1 — off-site on by default is not off-site working on day
|
||||
one), R-544. Closed: R-511, R-534, R-536, R-537, R-538. Register 236 → 244 open.
|
||||
|
||||
## Checks
|
||||
|
||||
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
|
||||
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
|
||||
`head_sha`.
|
||||
`repo_gates.py --fast` green at every push; controller `go build/vet/test` green; controller gates 15/15;
|
||||
`unproven.py --summary` unchanged at 35 of 55 not walked. Secret-leak check on the evidence: six real
|
||||
secret values as needles, planted control matched 6/6, committed evidence 0.
|
||||
|
||||
@@ -1,5 +1,41 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-16 (evening) — the photos come back, and they open.**
|
||||
|
||||
> **Ready for a volunteer: almost — one thing stands in the way, and it is small.** A brand-new box now
|
||||
> protects the household's own files: I put five photos in, deleted them the way a child would, and got
|
||||
> them back byte-for-byte from the off-site copy. The old route that used to lie now refuses politely
|
||||
> and points at the one that works. What is missing: on a new box the off-site copy is switched on but
|
||||
> **paused** until the household creates their recovery code, and nothing asks them to do it.
|
||||
|
||||
**What changed today.** The backup page stops claiming it holds files it does not hold. A restore that
|
||||
cannot bring your files back now refuses instead of reporting success — and it no longer wipes the app's
|
||||
own wastebasket on the way. „Alkalmazás telepítve" now means installed, not merely started. Every new
|
||||
customer gets the off-site copy by default, 100 GB. The off-site server got the one permission it was
|
||||
missing, so a rebuilt customer's box can be set up again without hand-work.
|
||||
|
||||
**What I proved on a box that installed itself this evening.** It installed from the new image, showed
|
||||
Felhom's own screen with no Proxmox address, registered itself, and **bound with nothing pressed on your
|
||||
side** — the connect e-mail it used was the one the system sent itself. It landed on today's golden.
|
||||
Then: five photos in, the local backup, the recovery-code ceremony, the off-site copy, the deletion, the
|
||||
refusal, the restore, and five photos that open — identical to the originals.
|
||||
|
||||
**Decisions I took.** None under the unattended rule.
|
||||
|
||||
**Needs you.**
|
||||
1. **Say yes or no to publishing the new installer image (1.28.0).** It is built and passed every check,
|
||||
and it fixes the screen that kept showing the pairing code after the box was connected. Nothing is
|
||||
published without your word. If you do nothing: new volunteers keep getting the older image, which
|
||||
works but shows that stale screen.
|
||||
2. **One small fix before a volunteer: tell the household to create their recovery code.** Until they do,
|
||||
the off-site copy is paused — so „your files are protected" is a promise with a delay in it. I can
|
||||
add the prompt and make the sentence state the real state.
|
||||
3. **The slow-crash-loop counter** (yesterday's ruling) is still owed, and is a job for the nightly.
|
||||
|
||||
---
|
||||
|
||||
## Previous note
|
||||
|
||||
**Updated 2026-09-16 (drill on a fresh box) — the fixes hold; the backup promise does not.**
|
||||
|
||||
> **Ready for a volunteer: NO — one reason, and it is new.** On a brand-new box with one drive, the
|
||||
@@ -786,3 +822,4 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
||||
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|
||||
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -20,3 +20,64 @@ VMID Status Lock Name
|
||||
## against all 41 evidence files. Planted positive control matched 6/6, so the grep works; the
|
||||
## committed evidence matched 0 files. The earlier count of „22" was the WORD „password" in labels
|
||||
## like „Password set" — a word count, not a leak check.
|
||||
## 2026-09-16T18:16:43Z TEARDOWN, LAYER 3 — delete the HOST record, keep the customer (never RESET)
|
||||
pre-state: {"deletable":true,"escrow_present":true,"guests":1,"log_bundles":0,"pbs_secret_present":true,"recovery_present":true,"reports":9,"status":"stale","wg_peer_bound":true}
|
||||
delete POST at 2026-09-16T18:16:43Z
|
||||
http=409
|
||||
host record after: 200 (404 = gone)
|
||||
customer record after: 200 (200 = KEPT)
|
||||
hub log right after the delete:
|
||||
2026/09/16 20:16:30 [INFO] Host staleness: tester-1-33b6a9 ok → stale (host_stale)
|
||||
2026/09/16 20:16:31 [INFO] Operator email sent for tester-1/host_stale
|
||||
2026/09/16 20:16:43 [WARN] host delete refused: tester-1-33b6a9 has key escrow (acknowledgement missing)
|
||||
## 2026-09-16T18:17:45Z delete retried WITH the escrow acknowledgement (retained custody)
|
||||
delete POST at 2026-09-16T18:17:45Z
|
||||
http=303
|
||||
host record after: 404 (404 = gone)
|
||||
customer record after: 200 (200 = KEPT)
|
||||
hub log right after:
|
||||
2026/09/16 20:16:31 [INFO] Operator email sent for tester-1/host_stale
|
||||
2026/09/16 20:16:43 [WARN] host delete refused: tester-1-33b6a9 has key escrow (acknowledgement missing)
|
||||
2026/09/16 20:17:30 [INFO] Staleness: tester-1 ok → stale (node_stale)
|
||||
2026/09/16 20:17:31 [INFO] Operator email sent for tester-1/node_stale
|
||||
2026/09/16 20:17:46 [INFO] host deleted: tester-1-33b6a9 (escrow deleted: true)
|
||||
2026/09/16 20:17:46 [INFO] self-bind link emailed to the registered address of tester-1
|
||||
2026/09/16 20:17:46 [INFO] self-bind link (hash 69b8422e…, valid 7 days) emailed to the registered address of tester-1
|
||||
2026/09/16 20:17:46 [INFO] self-bind link auto-minted for tester-1 on host delete (the console banner's promised email now exists)
|
||||
## THE FIRST DELETE WAS REFUSED, and it was right to refuse (18:16:43Z):
|
||||
## „host delete refused: tester-1-33b6a9 has key escrow (acknowledgement missing)" -> HTTP 409,
|
||||
## host record still 200, customer still 200 — nothing was dropped.
|
||||
## This box carries a REAL key escrow because the household performed the ceremony an hour ago, and
|
||||
## the hub will not discard a customer's sealed package on an unacknowledged delete. The documented
|
||||
## action is to acknowledge it, which moves the escrow to RETAINED custody rather than destroying it.
|
||||
## AND THE STALENESS ALARM FIRED TRUTHFULLY: „Host staleness: tester-1-33b6a9 ok → stale (host_stale)"
|
||||
## at 20:16:30 CEST with an operator mail one second later — the box really was gone by then.
|
||||
## LAYER 3 DONE (18:17:45Z), with the acknowledgement given:
|
||||
## POST /hosts/tester-1-33b6a9/delete (confirm_host_id + delete_escrow=1) -> 303
|
||||
## host record -> 404 (gone) customer record -> 200 (KEPT; RESET was never used)
|
||||
## hub log, same second:
|
||||
## „host deleted: tester-1-33b6a9 (escrow deleted: true)"
|
||||
## „self-bind link emailed to the registered address of tester-1"
|
||||
## „self-bind link (hash 69b8422e…, valid 7 days) emailed to the registered address of tester-1"
|
||||
## „self-bind link auto-minted for tester-1 on host delete (the console banner's promised email
|
||||
## now exists)"
|
||||
## So R-509's automatic e-mail fired again, on a second box, at the second of the delete.
|
||||
## A WORDING MISMATCH WORTH CHECKING RATHER THAN REPEATING: the refusal text says the acknowledgement
|
||||
## moves the escrow „to retained custody", while the log line says „escrow deleted: true". Those are
|
||||
## two different statements about a household's last key, so the customer record is read below
|
||||
## instead of trusting either sentence.
|
||||
## THE AUTOMATIC CONNECT E-MAIL — PROVEN AGAIN, on a second box in the same session:
|
||||
## host delete at 2026-09-16T18:17:45Z -> mail in the customer's inbox at 18:17:46Z (ONE second),
|
||||
## „[Felhom] Kösd össze a Felhom dobozodat", to tester1@felhom.eu, body „Elkészült a Felhom dobozod,
|
||||
## és készen áll az összekötésre… https://hub.felhom.eu/bind/…", link valid 7 days.
|
||||
## Requirement was two minutes. R-509 now has two independent live proofs today (12:23:00Z and
|
||||
## 18:17:46Z), on two different boxes.
|
||||
## THE ESCROW'S FATE, answered from the hub's own text rather than from either log line:
|
||||
## „host deletion only demotes custody, never destroys it"
|
||||
## „1. The host(s) will be deleted — recovery-key custody is demoted to RETAINED custody, not
|
||||
## destroyed." …and the customer delete is named as „the one true purge point".
|
||||
## So the acknowledgement demoted the custody; it did not destroy the household's sealed package.
|
||||
## The log line „escrow deleted: true" is the FLAG's name, not the effect — recorded as a row.
|
||||
## CUSTOMER RECORD AFTER EVERYTHING: present (200), with e-mail, domain, tunnel and DR tier intact,
|
||||
## and „waiting — no host enrolled yet — the Day-0 install enrolls…" — exactly the state a customer
|
||||
## is in between boxes. RESET was never used.
|
||||
|
||||
@@ -726,6 +726,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-541** | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** |
|
||||
| **R-542** | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** |
|
||||
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. | **READY — rank P1-HIGH; owner: CC (controller copy + first-run prompt)** |
|
||||
| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
|
||||
|
||||
Reference in New Issue
Block a user