the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
The capability map's journey row now carries the half it could never finish: five photos in, deleted the way a child would, the old route refusing and touching nothing, the off-site restore returning them, and them opening — sha256 identical, 5 of 5, with a negative control. Stated with it, because both are true: the bind needed ZERO operator presses (the box registered itself and used the mail the hub sent itself), but the PBS cascade needed ONE — the Re-issue press R-511 documents, which then succeeded because of this morning's ep0 grant. R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on day one — a fresh box waits at „Kulcsletétre vár" until the household creates its recovery code, and nothing asks them to, while the tier-1 row already promises that copy. R-544 records a log line that says „escrow deleted" where the effect is demotion to retained custody. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
File diff suppressed because one or more lines are too long
@@ -20,3 +20,64 @@ VMID Status Lock Name
|
||||
## against all 41 evidence files. Planted positive control matched 6/6, so the grep works; the
|
||||
## committed evidence matched 0 files. The earlier count of „22" was the WORD „password" in labels
|
||||
## like „Password set" — a word count, not a leak check.
|
||||
## 2026-09-16T18:16:43Z TEARDOWN, LAYER 3 — delete the HOST record, keep the customer (never RESET)
|
||||
pre-state: {"deletable":true,"escrow_present":true,"guests":1,"log_bundles":0,"pbs_secret_present":true,"recovery_present":true,"reports":9,"status":"stale","wg_peer_bound":true}
|
||||
delete POST at 2026-09-16T18:16:43Z
|
||||
http=409
|
||||
host record after: 200 (404 = gone)
|
||||
customer record after: 200 (200 = KEPT)
|
||||
hub log right after the delete:
|
||||
2026/09/16 20:16:30 [INFO] Host staleness: tester-1-33b6a9 ok → stale (host_stale)
|
||||
2026/09/16 20:16:31 [INFO] Operator email sent for tester-1/host_stale
|
||||
2026/09/16 20:16:43 [WARN] host delete refused: tester-1-33b6a9 has key escrow (acknowledgement missing)
|
||||
## 2026-09-16T18:17:45Z delete retried WITH the escrow acknowledgement (retained custody)
|
||||
delete POST at 2026-09-16T18:17:45Z
|
||||
http=303
|
||||
host record after: 404 (404 = gone)
|
||||
customer record after: 200 (200 = KEPT)
|
||||
hub log right after:
|
||||
2026/09/16 20:16:31 [INFO] Operator email sent for tester-1/host_stale
|
||||
2026/09/16 20:16:43 [WARN] host delete refused: tester-1-33b6a9 has key escrow (acknowledgement missing)
|
||||
2026/09/16 20:17:30 [INFO] Staleness: tester-1 ok → stale (node_stale)
|
||||
2026/09/16 20:17:31 [INFO] Operator email sent for tester-1/node_stale
|
||||
2026/09/16 20:17:46 [INFO] host deleted: tester-1-33b6a9 (escrow deleted: true)
|
||||
2026/09/16 20:17:46 [INFO] self-bind link emailed to the registered address of tester-1
|
||||
2026/09/16 20:17:46 [INFO] self-bind link (hash 69b8422e…, valid 7 days) emailed to the registered address of tester-1
|
||||
2026/09/16 20:17:46 [INFO] self-bind link auto-minted for tester-1 on host delete (the console banner's promised email now exists)
|
||||
## THE FIRST DELETE WAS REFUSED, and it was right to refuse (18:16:43Z):
|
||||
## „host delete refused: tester-1-33b6a9 has key escrow (acknowledgement missing)" -> HTTP 409,
|
||||
## host record still 200, customer still 200 — nothing was dropped.
|
||||
## This box carries a REAL key escrow because the household performed the ceremony an hour ago, and
|
||||
## the hub will not discard a customer's sealed package on an unacknowledged delete. The documented
|
||||
## action is to acknowledge it, which moves the escrow to RETAINED custody rather than destroying it.
|
||||
## AND THE STALENESS ALARM FIRED TRUTHFULLY: „Host staleness: tester-1-33b6a9 ok → stale (host_stale)"
|
||||
## at 20:16:30 CEST with an operator mail one second later — the box really was gone by then.
|
||||
## LAYER 3 DONE (18:17:45Z), with the acknowledgement given:
|
||||
## POST /hosts/tester-1-33b6a9/delete (confirm_host_id + delete_escrow=1) -> 303
|
||||
## host record -> 404 (gone) customer record -> 200 (KEPT; RESET was never used)
|
||||
## hub log, same second:
|
||||
## „host deleted: tester-1-33b6a9 (escrow deleted: true)"
|
||||
## „self-bind link emailed to the registered address of tester-1"
|
||||
## „self-bind link (hash 69b8422e…, valid 7 days) emailed to the registered address of tester-1"
|
||||
## „self-bind link auto-minted for tester-1 on host delete (the console banner's promised email
|
||||
## now exists)"
|
||||
## So R-509's automatic e-mail fired again, on a second box, at the second of the delete.
|
||||
## A WORDING MISMATCH WORTH CHECKING RATHER THAN REPEATING: the refusal text says the acknowledgement
|
||||
## moves the escrow „to retained custody", while the log line says „escrow deleted: true". Those are
|
||||
## two different statements about a household's last key, so the customer record is read below
|
||||
## instead of trusting either sentence.
|
||||
## THE AUTOMATIC CONNECT E-MAIL — PROVEN AGAIN, on a second box in the same session:
|
||||
## host delete at 2026-09-16T18:17:45Z -> mail in the customer's inbox at 18:17:46Z (ONE second),
|
||||
## „[Felhom] Kösd össze a Felhom dobozodat", to tester1@felhom.eu, body „Elkészült a Felhom dobozod,
|
||||
## és készen áll az összekötésre… https://hub.felhom.eu/bind/…", link valid 7 days.
|
||||
## Requirement was two minutes. R-509 now has two independent live proofs today (12:23:00Z and
|
||||
## 18:17:46Z), on two different boxes.
|
||||
## THE ESCROW'S FATE, answered from the hub's own text rather than from either log line:
|
||||
## „host deletion only demotes custody, never destroys it"
|
||||
## „1. The host(s) will be deleted — recovery-key custody is demoted to RETAINED custody, not
|
||||
## destroyed." …and the customer delete is named as „the one true purge point".
|
||||
## So the acknowledgement demoted the custody; it did not destroy the household's sealed package.
|
||||
## The log line „escrow deleted: true" is the FLAG's name, not the effect — recorded as a row.
|
||||
## CUSTOMER RECORD AFTER EVERYTHING: present (200), with e-mail, domain, tunnel and DR tier intact,
|
||||
## and „waiting — no host enrolled yet — the Day-0 install enrolls…" — exactly the state a customer
|
||||
## is in between boxes. RESET was never used.
|
||||
|
||||
@@ -726,6 +726,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-541** | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** |
|
||||
| **R-542** | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** |
|
||||
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. | **READY — rank P1-HIGH; owner: CC (controller copy + first-run prompt)** |
|
||||
| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
|
||||
|
||||
Reference in New Issue
Block a user