R-543 closed: the household is asked for the recovery code (controller v0.245.0)
gates / gates (push) Successful in 21s

The tier-3 pause is the zero-knowledge escrow design and is untouched. What was
missing was the ASK, while the backup page promised the copy that had never run.

- VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and
  before the first app - what the code is, where, write it on PAPER, and that
  Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13.
- day0-install.md A.2b: the operator step for a REBUILT box, which was missing.
  Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of
  "Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go).
  This is the correction to last night's "zero presses" note.
- 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and
  2 records that the household is asked from first login.
- capability map: the first-hour row's last gap closed, with what it still does
  not claim (no volunteer has walked the ask from the written guide).
- register: R-543 CLOSED with the live measurements; R-545 filed (nothing
  un-configures an off-site target). R-511 was already closed yesterday.
- STATUS: the answered publish question removed (1.28.0 is live), readiness yes.
- evidence: red-proofs, the two-box live validation, teardown on three layers,
  and both of my own mistakes in this session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 21:20:51 +02:00
parent 1acd693854
commit d124c77e17
9 changed files with 309 additions and 33 deletions
File diff suppressed because one or more lines are too long
@@ -77,6 +77,13 @@ So the honest statement of the property is:
> **The hub cannot open the escrow. The operator can reach a live box. Neither fact substitutes for
> the other, and R is what keeps the first true.**
**[DESIGN] Because R is the only key, the household is ASKED for it from the first login** (controller
v0.245.0, R-543). Off-site backup is enabled by default but does not RUN until the ceremony is done,
so the ask is not a nicety — it is the step that turns the default-on tier into an actual copy. The
volunteer guide asks for it immediately after the dashboard password and before the first app
(`runbooks/VOLUNTEER-first-hour.md` §6), and the product repeats the ask on every page until it is
done (§6.1).
---
## 3. The two lanes (D1)
@@ -299,6 +306,17 @@ never touches off-site snapshots (R-474). A removed app whose unit was kept is l
| **Plane-2** whole-guest, local | `local:` → `/var/lib/vz/dump` on the host | rootfs + `mp0 /var/lib/docker` + `mp1 /mnt/sys_drive` | **24 h** (`backup_cadence_seconds: 0`) | `local_backup_retention: 3` | **no** |
| **Plane-2** whole-guest, offsite | `felhom-pbs:` → ep0 datastore `felhom-offsite`, per-customer namespace, over WireGuard | same contents | **7 days** (`604800`) | server-side prune on ep0, `keep-last 2` at `03:30` — **and the box asks for none** (R-191) | **yes** (per-customer key) |
> **[FACT] Tier-3 has a fifth state the table above does not show: PAUSED (R-543, controller
> v0.245.0).** Off-site is ON by default from hub v0.116.0, and a run does not start until the
> household has performed the escrow ceremony — `tier3State` calls this `escrow_pending` and the page
> says „Kulcsletétre vár". **This is the design, not a defect:** the escrow is zero-knowledge (§2),
> the household's recovery code is the only key, and a run started without one would write a copy
> nobody could ever open. What was wrong until v0.245.0 is that **nothing asked the household for the
> code**, so a fresh box could sit paused indefinitely while its Tier-1 row promised that the off-site
> copy protected the app's files. Since v0.245.0 every dashboard page carries the reminder (the R-241
> bar, second instance) and the Tier-1 sentence renders by state — „védené … szünetel" while paused.
> Measured on a fresh box 2026-09-16: zero snapshots, and the page said the files were protected.
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
@@ -0,0 +1,130 @@
## R-543 live validation — controller v0.245.0, 2026-09-16, demo-hp
## Method: endpoint-level. Every request below is the exact request the browser makes (same handler,
## same session cookie, same CSRF); only rendering is skipped. claude-in-chrome is not available here.
## Reachability note: neither guest answers on its LXC address from DooPlex, and the controller does
## not listen on the guest's 127.0.0.1 — it answers on the container address 172.17.0.2:8080 with the
## mandatory Host header. All requests therefore run INSIDE the guest via `pct exec`.
## Secrets: the dashboard password was read with scripts/read_credential.py (value never printed,
## file->file, 0600) and passed to curl as --data-urlencode password@<file>.
### THE SEAM, STATED
The two halves of this proof are on two boxes, not one box before and after a ceremony:
* PAUSED half — guest 9202 (scratch), off-site configured and escrow NOT complete.
* ESCROWED half — guest 9201 (demo), off-site configured, escrowed, running.
Running a real escrow ceremony on a fresh target would provision off-site storage, and this task's
fences put ep0 out of bounds. So the escrowed side is OBSERVED on a box that is already escrowed
rather than produced here. What is NOT weakened by the seam: both boxes run the same binary
(0.245.0), and the paused box's state was produced through the product's own configuration endpoint.
## ── 1. ESCROWED + ACTIVE — guest 9201 (0.245.0) ────────────────────────────────────────────────
login OK
GET /dashboard -> 200 bytes=60816 escrow-bar-hits=0
GET /launcher -> 200 bytes=44462 escrow-bar-hits=0
GET /backups/apps -> 200 bytes=106304 escrow-bar-hits=0
tier-1 file sentence, as rendered (2 class-A apps):
"Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi — ez a helyi mentés a
beállításokat és az adatbázist tartalmazza."
word counts on /backups/apps: Kulcslet=0 szünetel=0 védi=2 védené=0
=> a box whose recovery code exists is NOT nagged, and the promise it prints is a true one.
## ── 2. PAUSED — guest 9202 (0.245.0) ───────────────────────────────────────────────────────────
The state was produced through the product's own endpoint, POST /backup/offbox/config, with a
throwaway ed25519 key and a pinned known_hosts line. 9202 has no agent local API, so the escrow
STAGE could not run — the handler said so and saved the target anyway:
POST /backup/offbox/config -> 302
flash: "A távoli mentési cél elmentve. — a kulcs letéti előkészítése nem sikerült
(az ügynök nem elérhető); próbáld újra."
Persisted state afterwards (read from the guest's own settings.json):
enabled=True host=192.168.0.162 repo_path=/mnt/nvme-1tb/r543-paused-proof
escrow_state=pending last_run=None last_status=None snapshot_count=None
data/offbox/: known_hosts 95 B (0644) | repo_password 64 B (0600) | ssh_key 411 B (0600)
=> OffboxConfigured() is genuinely true (valid target + both secret files), and the escrow is NOT
complete. This is the state a fresh box lands in on day one.
### 2a. The bar is on EVERY page, not on the three someone remembered
GET /dashboard -> 200 escrow-bar-hits=1
GET /launcher -> 200 escrow-bar-hits=1
GET /backups/apps -> 200 escrow-bar-hits=1
GET /settings -> 200 escrow-bar-hits=1
GET /apps -> 404 escrow-bar-hits=0 (no such route; a 404 carries no bar)
Quoted from /dashboard:
"A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot."
<a href="/backup/escrow"> (the route out, on the bar itself)
### 2b. A manual off-site run, while paused — refused, FOR THE RIGHT REASON
POST /backup/offbox/run -> 302
flash: "A távoli mentés a kulcs letétbe helyezésére vár."
after the attempt: last_run=None last_status=None snapshot_count=None
=> No snapshot, and the refusal came from the fork-4 escrow gate (offbox_handlers.go:224
OffboxRunnable), NOT from an unreachable target — the target was never contacted. This matters:
a run that failed to connect would have proven nothing about the pause.
### 2c. „Most nem" is for the visit only
POST /backup/escrow/banner/dismiss -> 302
Set-Cookie: felhom_escrow_banner=1; Path=/; HttpOnly; SameSite=Lax
^ no Max-Age and no Expires => a browser SESSION cookie, exactly as R-241 does it
same visit, cookie sent -> escrow-bar-hits=0
next visit, no cookie -> escrow-bar-hits=1
=> the off-site tier is still paused tomorrow, so the question is still asked tomorrow.
### 2d. The tier-1 FILE sentence, live, with the tier paused
The first pass of this phase could not show it: 9202 had no class-A app, so the sentence had nothing
to render („védi"=0 AND „védené"=0 — the „szünetel"=1 on that page was the BAR in the layout, not the
sentence). Recorded because it was briefly written down as a limit and it was not one. A throwaway
class-A app was then deployed on the box's own drive and the off-site copy turned on for it:
POST /api/stacks/calibre-web/deploy -> 202 {"ok":true,"message":"Telepítés elindítva ..."}
(HDD_PATH=/mnt/felhom-drives/scratch_hdd — the guest's own registered data drive, 938 G)
POST /backup/offbox/toggle app=calibre-web enabled=true -> 302
Rendered on /backups/apps, with the tier configured and PAUSED:
"Az alkalmazás fájljait a távoli másolat védené — a távoli mentés a helyreállítási kód
létrehozásáig szünetel."
...followed by the route: "Helyreállítási kód létrehozása →"
word counts on /backups/apps: Kulcslet=1 szünetel=3 védi=0 védené=1
Kulcslet=1 is the tier-3 row's own state („Kulcsletétre vár").
szünetel=3 is the bar + the sentence + the tier row.
védi=0 is the point: the page no longer claims a protection that has never run.
=> Both wordings are now observed LIVE on real boxes running 0.245.0: „védi" on the escrowed box
(9201, §1) and „védené … szünetel" on the paused box (9202, here).
## ── 3. TEARDOWN, three layers ──────────────────────────────────────────────────────────────────
Recorded after the evidence above was already written to DooPlex (R-320: evidence leaves the machine
at the END OF THE PHASE, before any revert — not at the end of the session).
MACHINE (guest 9202, scratch):
* the throwaway class-A app removed WITH its data and its backups:
POST /api/stacks/calibre-web/stop -> 200
POST /api/stacks/calibre-web/remove -> 200 {"removed":"calibre-web",
"volumes_removed":["calibre-web_calibre_web_config"],"hdd_paths_removed":[],
"hdd_note":"Az alkalmazás nem tárolt saját adatot…"}
verified after: containers named calibre = 0, drive folders = 0
* the off-site target this proof created is GONE:
before: offbox present=True, data/offbox/ held ssh_key, repo_password, known_hosts
after : offbox present=False, data/offbox/ ABSENT, all three files `shred -u`'d
MY OWN MISTAKE, recorded: the first teardown pass used the CONTAINER's view of the data path
(/opt/docker/felhom-controller/data) from a shell running in the GUEST, where that path does
not exist. It printed „offbox dir now: ABSENT" — which was TRUE of a path that never existed
and FALSE of the thing being claimed. The target was still fully configured. The same wrong
path had already produced three FileNotFoundError tracebacks earlier in this phase; I read
those as noise instead of as the instrument telling me it was pointed at nothing. The re-run
stops the controller first, edits the real file, shreds the secrets, restarts, and RE-READS
the state to confirm — a teardown asserted is not a teardown observed.
* controller restarted and healthy: felhom-controller:0.245.0 Up (healthy)
* temp files: none left (`ls /tmp/.r543*` empty in the guest)
* NOT touched: filebrowser, traefik, the box's registered data drive, its password, its claim state.
HOST (demo-hp): /tmp/.r543pw and /tmp/.r543key shredded; .r543kh, the scripts and one stray
/tmp/.r543out removed. Verified: 0 files matching /tmp/.r543* remain.
HUB: provisioned nothing. Guest 9202 has hub reporting OFF and no tunnel, so no customer, no host
record, no escrow row and no event was created by any of this. Nothing to tear down.
NOT torn down, deliberately: guest 9201 now runs controller 0.245.0. That is the release being
shipped, not a drill artifact, and its own off-site tier was never touched (still escrowed, active).
@@ -0,0 +1,48 @@
## R-543 red-proofs — controller v0.245.0, 2026-09-16, DooPlex
## Each fix was BROKEN first and the test was watched convicting it. A test never seen failing has
## not been shown to test anything.
### RED-PROOF 1 — the reminder bar
Break: delete `s.addEscrowBanner(data, r)` from executeTemplate (internal/web/server.go).
That is the whole wiring: the bar hangs off the single render choke point, so removing one line
returns the product to the measured 2026-09-16 state (a paused off-site tier, and silence).
--- FAIL: TestR543_A_PausedBoxAsksOnEveryPage (0.20s)
r543_escrow_banner_test.go:72: R-543: /dashboard does not tell the household the off-site copy is PAUSED. The tier is on, nothing is running, and the page is silent about it
r543_escrow_banner_test.go:76: R-543: /dashboard states the pause but names no route to end it — a reminder without its door is the shape that left a fresh box waiting indefinitely
r543_escrow_banner_test.go:72: R-543: /launcher does not tell the household the off-site copy is PAUSED. The tier is on, nothing is running, and the page is silent about it
r543_escrow_banner_test.go:76: R-543: /launcher states the pause but names no route to end it — a reminder without its door is the shape that left a fresh box waiting indefinitely
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.208s
Note BOTH pages fail. That is the point of the hook placement: a per-handler helper would have
covered the three pages someone remembered, which is the seam-built-but-never-wired shape.
### RED-PROOF 2 — the tier-1 sentence
Break: return the v0.244.0 wording from driveFilesNoteFor before the state switch, i.e. compute the
sentence from the app's shape alone, exactly as v0.244.0 shipped it.
--- FAIL: TestR543_Tier1Sentence_PausedStateDoesNotPromise (0.00s)
r543_tier1_sentence_test.go:38: R-543: the sentence claims the files ARE protected while the copy is paused for the recovery code. This is the exact promise a fresh box read for its whole first day: "Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi — ez a helyi mentés a beállításokat és az adatbázist tartalmazza."
r543_tier1_sentence_test.go:42: R-543: the paused sentence must say the copy WOULD protect them and that it is waiting; got "Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi — ez a helyi mentés a beállításokat és az adatbázist tartalmazza."
r543_tier1_sentence_test.go:46: R-543: the paused sentence names no route out (link="" text="")
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.007s
### RESTORED
go test ./internal/web/ -run 'R543'
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.895s
### The filter was proven to match (the `-run` trap: a pattern matching nothing prints `ok`, exit 0)
=== RUN TestR543_A_PausedBoxAsksOnEveryPage --- PASS (0.29s)
=== RUN TestR543_B_EscrowedBoxIsNotNagged --- PASS (0.19s)
=== RUN TestR543_C_UnconfiguredBoxIsNotNagged --- PASS (0.20s)
=== RUN TestR543_D_DismissIsForThisVisitOnly --- PASS (0.28s)
=== RUN TestR543_Tier1Sentence_ActiveStateKeepsThePromise --- PASS
=== RUN TestR543_Tier1Sentence_PausedStateDoesNotPromise --- PASS
=== RUN TestR543_Tier1Sentence_NoCopyAtAllSaysSo --- PASS
=== RUN TestR543_Tier1Sentence_SecondDriveCounts --- PASS
=== RUN TestR543_Tier1Sentence_NoFileLegsNoSentence --- PASS
### Fixture validity (an instrument that can lose its precondition measures nothing)
escrowServer asserts backupMgr.OffboxConfigured() itself before any assertion runs: the target is
enabled and valid AND the ssh_key + repo_password files exist on disk. A fixture that silently fell
back to "not configured" would make every one of these tests pass for the wrong reason.
+2 -1
View File
@@ -725,8 +725,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-540** | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-541** | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** |
| **R-542** | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** |
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. | **READY — rank P1-HIGH; owner: CC (controller copy + first-run prompt)** |
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** |
| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
+26 -7
View File
@@ -102,7 +102,26 @@ nem kell újra összekötni.)*
**legalább 12 karakteres** jelszót. Ez lesz a vezérlőpult jelszava.
4. Ha nem jött meg a kód: **„Új kód kérése"** — mindig ugyanarra az e-mail címre érkezik.
## 6. Az első két alkalmazás telepítése (~2 perc)
## 6. A helyreállítási kód (~2 perc) — ezt ne hagyd ki
A doboz a fájljaidról **titkosított** másolatot küld a Felhom távoli tárhelyére. A titkosítás
kulcsát **csak te** kapod meg: ez a **helyreállítási kód**. Amíg nem hozod létre, **a távoli
mentés nem indul el** — a vezérlőpult minden oldalán látszó sáv ezt írja:
„A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot."
1. **Biztonsági mentés → Távoli mentés**, vagy egyszerűen kattints a sávon a
**„Helyreállítási kód létrehozása"** hivatkozásra.
2. Add meg a vezérlőpult jelszavát, és indítsd el. Kb. fél perc.
3. A kód **egyetlen egyszer** jelenik meg a képernyőn. **Írd fel papírra**, és tedd el oda, ahová a
fontos iratokat teszed. Fénykép a telefonról nem elég, ha a telefon is elveszik.
> ⚠ **A Felhom nem tudja visszaszerezni ezt a kódot.** Nem azért, mert nem akarja: a másolataidat
> úgy titkosítják, hogy mi magunk se tudjuk megnyitni őket. Ha a kód elvész, a távoli másolatok
> megmaradnak, de **senki — mi sem — nem tudja többé megnyitni őket**.
Ha ezzel megvagy, a sáv eltűnik, és a távoli mentés magától elindul.
## 7. Az első két alkalmazás telepítése (~2 perc)
1. A vezérlőpulton: **Alkalmazások**. Keresd meg például a **BookStack**-et (családi wiki) és a
**PrivateBin**-t (titkosított jegyzet), és nyomd meg a **Telepítés** gombot.
@@ -110,7 +129,7 @@ nem kell újra összekötni.)*
Az „Automatikusan generált értékek" részt nem kell felírnod.
3. **Telepítés indítása.** A BookStack kb. 1 perc, a PrivateBin kb. 20 másodperc.
## 7. Első belépés az alkalmazásokba
## 8. Első belépés az alkalmazásokba
- **BookStack:** az alkalmazás oldalán az „Első lépések" rész a címet `wiki.DOMAIN` alakban írja —
a DOMAIN helyére a saját domained kerül (ismert hiba). Belépés: `admin@admin.com` / `password`.
@@ -118,7 +137,7 @@ nem kell újra összekötni.)*
- **PrivateBin:** nincs belépés. Írj be szöveget, **Küldés**, és a kapott linket oszd meg — a kulcs a
linkben van, a szerver nem látja a tartalmat.
## 8. Mentések
## 9. Mentések
- **Biztonsági mentés → Áttekintés:** két sárga figyelmeztetést látsz („Csak egy másolat készül",
„ugyanazon a lemezen van") — **ezek igazak**: amíg nincs második meghajtó vagy távoli mentés, egy
@@ -128,23 +147,23 @@ nem kell újra összekötni.)*
teljes rendszermentésben (PBS)" — ez nem minden dobozra igaz (ismert hiba). Az Áttekintés oldal a
pontos.*
## 9. Visszaállítás
## 10. Visszaállítás
**Biztonsági mentés → Visszaállítás:** válaszd az alkalmazást és a mentést, pipáld be a „Megértettem"
négyzetet, **Visszaállítás indítása**. Az alkalmazás kb. fél percre leáll, majd az utolsó mentés
állapotával indul újra.
## 10. Alkalmazás eltávolítása
## 11. Alkalmazás eltávolítása
**Alkalmazások:** előbb **Leállítás**, utána megjelenik az **Eltávolítás**. A párbeszédablak felsorolja,
mi törlődik mindenképp, és bepipálhatod a mentések törlését is.
## 11. Áramszünet
## 12. Áramszünet
Ha elmegy az áram, a doboz magától visszaindul, kb. 2 perc múlva minden alkalmazás ugyanazon a
verzión fut, amelyen előtte. A vezérlőpultba újra be kell jelentkezned.
## 12. Ha elgépelted a kódot
## 13. Ha elgépelted a kódot
A beállító oldal „Hibás vagy lejárt kód" üzenettel visszadobja — írd be újra. **Öt** hibás próbálkozás
után 15 percre zárol („Túl sok próbálkozás — próbáld újra 15 perc múlva."). Új kódot a
+20
View File
@@ -95,6 +95,26 @@ On save the hub generates two credentials:
self-bind mail tells the customer they received it from the Felhom operator.) Treat it like a password.
- **Customer API key** — internal (baked into the generated controller.yaml); never handled manually.
### A.2b REBUILDING a box that already exists — the one press (F-14 ruling, 2026-07-13)
**A rebuilt box does NOT always get its off-site credentials by itself, and this step was missing
from this runbook.** Which of the two happens is decided by how the PREVIOUS box left:
| How the previous box was removed | What the hub does | Operator action |
|---|---|---|
| Deleted through the **acknowledged** delete flow (the escrow acknowledgement was given) | the hub re-issues the off-site credentials **by itself** when the new box enrolls — the box consumes the single-use secret and the tier comes up | **none** |
| The host record was left in place, or the delete was not acknowledged | the hub **reuses** the existing record and mints nothing — the box asks, is refused, and the off-site tier never starts | **one press:** Hub UI → the customer → **„Re-issue PBS credentials"** |
The refusal is correct and deliberate: re-issuing over a live record would orphan the history the
customer's recovery code protects. The mint-once-and-reuse decision lives in `hub/internal/web/pbsdr.go`;
there are **two** Re-issue buttons on that page — the off-site one and the PBS-DR one — and they are
not interchangeable.
> This is the correction to the note that said a rebuilt box needs zero presses. It needs zero
> presses only on the acknowledged path. Measured on a fresh box 2026-09-16: the box reproduced the
> refusal by itself, the re-issue then ADOPTED the record (generation 2) and the box consumed the
> single-use secret.
### A.3 Verify the Day-0 artifact manifest
Hub UI → **Configuration → Day-0 artifacts**. The manifest must vouch an **agent version** and a