diff --git a/CONTEXT.md b/CONTEXT.md index ab2a5786..ec556dd8 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,43 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## Decisions 2026-09-16 (evening) — the recovery code: the household is asked, the page says „szünetel" + +**No new ruling. One correction to the record, from the operator (2026-09-16):** the note saying a +rebuilt box needs zero presses for its off-site setup was WRONG. The press exists by the F-14 ruling +of 2026-07-13 — the hub re-issues by itself **only** when the previous box was deleted through the +acknowledged flow; otherwise `handleHostEnroll` reuses the existing record and mints nothing, and the +operator presses „Re-issue PBS credentials" once (`hub/internal/web/pbsdr.go`). That ruling stands and +is now written into `runbooks/day0-install.md` §A.2b, where it was missing. + +**R-543 CLOSED — controller v0.245.0.** The tier-3 pause is the zero-knowledge escrow design and was +not touched. Two copy changes only: + +- **`internal/web/escrow_banner.go`** — the R-241 reminder bar, SECOND INSTANCE. `escrowPaused()` + reads the same two facts `tier3State` reads (`backupMgr.OffboxConfigured()` and + `settings.GetOffboxTarget().EscrowState != "escrowed"`), so the bar and the app rows cannot + disagree. **It hangs off `executeTemplate` (server.go), the single render choke point** — not the + three `addRecoveryBanner` call sites, because a per-handler helper reaches only the pages someone + remembered (the seam-built-but-never-wired class). Login and claim render through + `s.tmpl.ExecuteTemplate` directly and never pass through it; `hasAdminSession()` (mirroring + `RequireAuth`, legacy-open included) keeps it off the public `/s/` share page. Dismissal: + session cookie `felhom_escrow_banner`, no MaxAge/Expires, route `/backup/escrow/banner/dismiss`. +- **`driveFilesNoteFor` (`internal/web/backup_page_state.go`)** — the tier-1 file sentence takes + `tier3State`'s OWN vocabulary instead of the app's shape, and returns (note, linkHref, linkText) so + the template renders a real anchor. Assigned at the END of the app-row loop, where `Tier3State` and + `Tier2Configured` are both already resolved. + +**Trap re-learned, twice in one session, and worth the line:** inside a guest the controller's data +directory is `/var/lib/docker/volumes/felhom-controller-data/_data/data`; `/opt/docker/felhom-controller/data` +is the CONTAINER's view of the same files. A teardown script using the container path from a guest +shell printed „offbox dir now: ABSENT" — true of a path that never existed, false of the thing being +claimed — and left the target fully configured. Also: the controller answers on the container address +`172.17.0.2:8080` with the mandatory `Host` header; it is NOT on the guest's `127.0.0.1`, and neither +guest is reachable from DooPlex at all. + +**New row:** R-545 (P3) — nothing un-configures an off-site target; `/backup/offbox/reset` refuses +unless the repo is ORPHANED and means „start a new remote backup", not „forget this destination". + ## Decisions 2026-09-16 (afternoon) — the backup promise: files protected from day one **Ruling: the off-site copy is ON for every customer from day one** — shared (a sub-account on the diff --git a/STATUS.md b/STATUS.md index dda00e5c..c0861e73 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,36 +1,39 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-16 (evening) — the photos come back, and they open.** +**Updated 2026-09-16 (late evening) — the box now asks for the recovery code, and the page stops promising a copy that has not run.** -> **Ready for a volunteer: almost — one thing stands in the way, and it is small.** A brand-new box now -> protects the household's own files: I put five photos in, deleted them the way a child would, and got -> them back byte-for-byte from the off-site copy. The old route that used to lie now refuses politely -> and points at the one that works. What is missing: on a new box the off-site copy is switched on but -> **paused** until the household creates their recovery code, and nothing asks them to do it. +> **Ready for a volunteer: yes.** The one thing standing in the way this morning is fixed. On a new +> box the off-site copy is switched on but paused until the household writes down their recovery +> code — that pause is deliberate and correct, because that code is the only key and we cannot open +> their copies without it. What was wrong is that nothing asked them. Now every page says so until +> they do it, and the backup page says „would protect" instead of „protects" while it waits. -**What changed today.** The backup page stops claiming it holds files it does not hold. A restore that -cannot bring your files back now refuses instead of reporting success — and it no longer wipes the app's -own wastebasket on the way. „Alkalmazás telepítve" now means installed, not merely started. Every new -customer gets the off-site copy by default, 100 GB. The off-site server got the one permission it was -missing, so a rebuilt customer's box can be set up again without hand-work. +**What changed today (this note).** A reminder bar on every page of the dashboard: „the off-site +backup is paused until you create your recovery code", with the button that does it. The sentence +under each app's local backup now tells the truth about the state it is in — protected, waiting, or +no copy at all — instead of promising the same thing in all three. The first-hour guide asks for the +code right after the dashboard password and before the first app, and says plainly that we cannot +get it back for them. The operator step for rebuilding an existing customer's box is written down +where it was missing: normally nothing to press, but one press when the old box was not deleted +through the acknowledged flow. -**What I proved on a box that installed itself this evening.** It installed from the new image, showed -Felhom's own screen with no Proxmox address, registered itself, and **bound with nothing pressed on your -side** — the connect e-mail it used was the one the system sent itself. It landed on today's golden. -Then: five photos in, the local backup, the recovery-code ceremony, the off-site copy, the deletion, the -refusal, the restore, and five photos that open — identical to the originals. +**What I proved on real boxes.** On a box whose recovery code exists: no bar anywhere, and the page +says the files are protected. On a box waiting for the code: the bar on every page, an off-site run +refused with „waiting for the key to be placed in escrow" and no copy written, and an app's row +reading „would be protected … paused until you create the recovery code". Both boxes were running +today's build. The throwaway app and the test setup were removed afterwards and checked gone. **Decisions I took.** None under the unattended rule. **Needs you.** -1. **Say yes or no to publishing the new installer image (1.28.0).** It is built and passed every check, - and it fixes the screen that kept showing the pairing code after the box was connected. Nothing is - published without your word. If you do nothing: new volunteers keep getting the older image, which - works but shows that stale screen. -2. **One small fix before a volunteer: tell the household to create their recovery code.** Until they do, - the off-site copy is paused — so „your files are protected" is a promise with a delay in it. I can - add the prompt and make the sentence state the real state. -3. **The slow-crash-loop counter** (yesterday's ruling) is still owed, and is a job for the nightly. +1. **Nothing blocking.** The installer image you approved is published and live on the download page, + and the recovery-code gap is closed. A volunteer can start. +2. **The slow-crash-loop counter** (the ruling of 2026-09-15) is still owed, and is a job for the + nightly. If you do nothing: a box that keeps crashing slowly is still reported as healthy for + longer than it should be. +3. **One small thing worth knowing, not doing:** there is no button that forgets an off-site + destination once set — only one that disables it. Written down as a low-priority job. If you do + nothing: a household that types the wrong address keeps the old one on the box, switched off. --- diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 6490998c..c75ffe96 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -88,7 +88,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis **NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | | **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) The venue is **torn down** (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and **30 are R-244**, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and **16.64 GiB** returned. `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. Evidence: `tests/walk5-r201-2026-08-07/journal.md` **SCOPE NOTE 2026-09-14 — not re-walked.** The first-hour drill on 0.242.0 (`audits/DRILL-fresh-install-0242-2026-09-14.md`) walked install → first use → restore of a deleted page on a fresh box, NOT a rebuild with off-site recovery (DR tier and off-site were off). This row is therefore neither re-proven nor contradicted on 0.242.0; its PROVEN-LIVE stands on 0.206.0 only. The first hour has its own row below. | -| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. **DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven.** `audits/DRILL-prove-fixes-0243-2026-09-16.md`. A fresh nested box was installed from the published ISO, landed on this drill's own golden **by checksum** (`e2d1843c…c10a`, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, **the tunnel from outside (302→200, R-510 CLOSED)**, the data drive, the file manager with its own generated password (`admin`/`admin` refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with **26 s** of app downtime. **Interventions: 0** — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. **The automatic connect e-mail is PROVEN with a real mailbox:** a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at **12:23:00Z**, with `selfbind_link_sent (host delete)` on the timeline (R-509). **WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence:** on a one-drive box with no off-site tier — what every fresh install is — the household's files are in **no backup at all** (the whole-guest tiers exclude `mp8` by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (**R-537**), and a restore reports success while leaving Nextcloud listing five photos that return `Sabre\DAV\Exception\NotFound` — after making the app's own trash, which still held every byte, unreachable (**R-538**). **Also still open from this walk:** the off-site tier cannot be provisioned at all (**R-534**, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (**R-535**), and „app installed” is emitted at accept time (**R-536**). Faults re-measured: controller killed during a deploy → back in **37 s**; three more kills 20 min apart → **61/41/61 s**, none accumulating (**R-531**); two reboots 60 s apart → everything back in **124 s** with the boots NOT counted against the brake; the claim page locks after the **second** wrong code and mails the operator truthfully. **2026-09-16 (evening) — THE BACKUP HALF IS NOW WALKED TOO, on controller 0.244.0 / hub 0.116.0 / golden 0.244.0 / ISO 1.28.0 (built, unpublished).** `audits/evidence-backup-promise-2026-09-16/`. A second fresh box (VM 335) was installed from the BUILT image and walked to the end of the sentence this row could not previously finish: **five photos in → deleted the way a child would → the old route REFUSED and touched nothing → the off-site restore returned them → they OPEN, byte-identical (sha256 5/5, negative control)**. The refusal reads „Ez a mentés nem tartalmazza az alkalmazás fájljait… a fájlok így a helyükön maradnak" and names the route that can help; the app was `running` before and after; the wastebasket was untouched. The off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution counting „5 fájl és 3 adatkötet és az adatbázis". **Delivery is part of it:** the box landed on agent 0.131.0 + controller 0.244.0 (the golden baked and vouched the same day) with no hand upgrade, and `app_deploy_started` (19:15:34) / `app_deployed` (19:16:23) finally mean different things. **The self-bind half needed NO operator press** — the box registered itself and the bind used the mail the hub sent itself after the morning's host delete (R-509, proven twice today: 12:23:00Z and 18:17:46Z, each one second after a host delete). **WHAT IT STILL DOES NOT CLAIM — one operator press and one day-one gap.** The PBS-DR cascade stopped at the refusal R-511 documents („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action") and needed the operator to press it; that press then SUCCEEDED because of the ep0 grant given this morning (R-534 CLOSED, R-511 CLOSED). And off-site ON by default is not off-site WORKING: a fresh box sits at „Kulcsletétre vár" until the household performs the escrow ceremony, which nothing asks them to do, while the tier-1 row already promises that copy (**R-543**, P1). Until that is fixed, the journey is: **files protected from day one only if someone tells the household to create their recovery code.** | Rows **R-493 … R-500**, R-534 … R-544 | +| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. **DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven.** `audits/DRILL-prove-fixes-0243-2026-09-16.md`. A fresh nested box was installed from the published ISO, landed on this drill's own golden **by checksum** (`e2d1843c…c10a`, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, **the tunnel from outside (302→200, R-510 CLOSED)**, the data drive, the file manager with its own generated password (`admin`/`admin` refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with **26 s** of app downtime. **Interventions: 0** — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. **The automatic connect e-mail is PROVEN with a real mailbox:** a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at **12:23:00Z**, with `selfbind_link_sent (host delete)` on the timeline (R-509). **WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence:** on a one-drive box with no off-site tier — what every fresh install is — the household's files are in **no backup at all** (the whole-guest tiers exclude `mp8` by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (**R-537**), and a restore reports success while leaving Nextcloud listing five photos that return `Sabre\DAV\Exception\NotFound` — after making the app's own trash, which still held every byte, unreachable (**R-538**). **Also still open from this walk:** the off-site tier cannot be provisioned at all (**R-534**, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (**R-535**), and „app installed” is emitted at accept time (**R-536**). Faults re-measured: controller killed during a deploy → back in **37 s**; three more kills 20 min apart → **61/41/61 s**, none accumulating (**R-531**); two reboots 60 s apart → everything back in **124 s** with the boots NOT counted against the brake; the claim page locks after the **second** wrong code and mails the operator truthfully. **2026-09-16 (evening) — THE BACKUP HALF IS NOW WALKED TOO, on controller 0.244.0 / hub 0.116.0 / golden 0.244.0 / ISO 1.28.0 (built, unpublished).** `audits/evidence-backup-promise-2026-09-16/`. A second fresh box (VM 335) was installed from the BUILT image and walked to the end of the sentence this row could not previously finish: **five photos in → deleted the way a child would → the old route REFUSED and touched nothing → the off-site restore returned them → they OPEN, byte-identical (sha256 5/5, negative control)**. The refusal reads „Ez a mentés nem tartalmazza az alkalmazás fájljait… a fájlok így a helyükön maradnak" and names the route that can help; the app was `running` before and after; the wastebasket was untouched. The off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution counting „5 fájl és 3 adatkötet és az adatbázis". **Delivery is part of it:** the box landed on agent 0.131.0 + controller 0.244.0 (the golden baked and vouched the same day) with no hand upgrade, and `app_deploy_started` (19:15:34) / `app_deployed` (19:16:23) finally mean different things. **The self-bind half needed NO operator press** — the box registered itself and the bind used the mail the hub sent itself after the morning's host delete (R-509, proven twice today: 12:23:00Z and 18:17:46Z, each one second after a host delete). **WHAT IT STILL DOES NOT CLAIM — one operator press and one day-one gap.** The PBS-DR cascade stopped at the refusal R-511 documents („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action") and needed the operator to press it; that press then SUCCEEDED because of the ep0 grant given this morning (R-534 CLOSED, R-511 CLOSED). And off-site ON by default is not off-site WORKING: a fresh box sits at „Kulcsletétre vár" until the household performs the escrow ceremony, which nothing asks them to do, while the tier-1 row already promises that copy (**R-543**, P1). **2026-09-16 (late evening) — that last gap is CLOSED, controller v0.245.0 (R-543).** The pause is the zero-knowledge escrow design and was not touched; what was missing was the ASK. Every authenticated page now carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking the ceremony (the R-241 bar, second instance, hung on the single render choke point), the tier-1 sentence renders by tier-3 STATE („védené … szünetel" while paused, „védi" when running), and the first-hour guide asks for the code right after the dashboard password and before the first app. **Measured on two boxes running 0.245.0:** paused box — bar on four pages, `POST /backup/offbox/run` refused by the fork-4 gate with no snapshot written, app row „védené" and „védi"=0; escrowed box — no bar anywhere, row „védi". `audits/evidence-recovery-code-2026-09-16/`. **WHAT IT STILL DOES NOT CLAIM:** the ask has not been walked by an actual volunteer from the written guide — the sentence is proven, the human following it is not. So the journey now reads: **files protected from day one, once the household writes down the recovery code the box asks them for on every page.** | Rows **R-493 … R-500**, R-534 … R-545 | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 329539e3..c07eeba8 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -77,6 +77,13 @@ So the honest statement of the property is: > **The hub cannot open the escrow. The operator can reach a live box. Neither fact substitutes for > the other, and R is what keeps the first true.** +**[DESIGN] Because R is the only key, the household is ASKED for it from the first login** (controller +v0.245.0, R-543). Off-site backup is enabled by default but does not RUN until the ceremony is done, +so the ask is not a nicety — it is the step that turns the default-on tier into an actual copy. The +volunteer guide asks for it immediately after the dashboard password and before the first app +(`runbooks/VOLUNTEER-first-hour.md` §6), and the product repeats the ask on every page until it is +done (§6.1). + --- ## 3. The two lanes (D1) @@ -299,6 +306,17 @@ never touches off-site snapshots (R-474). A removed app whose unit was kept is l | **Plane-2** whole-guest, local | `local:` → `/var/lib/vz/dump` on the host | rootfs + `mp0 /var/lib/docker` + `mp1 /mnt/sys_drive` | **24 h** (`backup_cadence_seconds: 0`) | `local_backup_retention: 3` | **no** | | **Plane-2** whole-guest, offsite | `felhom-pbs:` → ep0 datastore `felhom-offsite`, per-customer namespace, over WireGuard | same contents | **7 days** (`604800`) | server-side prune on ep0, `keep-last 2` at `03:30` — **and the box asks for none** (R-191) | **yes** (per-customer key) | +> **[FACT] Tier-3 has a fifth state the table above does not show: PAUSED (R-543, controller +> v0.245.0).** Off-site is ON by default from hub v0.116.0, and a run does not start until the +> household has performed the escrow ceremony — `tier3State` calls this `escrow_pending` and the page +> says „Kulcsletétre vár". **This is the design, not a defect:** the escrow is zero-knowledge (§2), +> the household's recovery code is the only key, and a run started without one would write a copy +> nobody could ever open. What was wrong until v0.245.0 is that **nothing asked the household for the +> code**, so a fresh box could sit paused indefinitely while its Tier-1 row promised that the off-site +> copy protected the app's files. Since v0.245.0 every dashboard page carries the reminder (the R-241 +> bar, second instance) and the Tier-1 sentence renders by state — „védené … szünetel" while paused. +> Measured on a fresh box 2026-09-16: zero snapshots, and the page said the files were protected. + > **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does. diff --git a/documentation/audits/evidence-recovery-code-2026-09-16/phaseA-live.txt b/documentation/audits/evidence-recovery-code-2026-09-16/phaseA-live.txt new file mode 100644 index 00000000..6096364e --- /dev/null +++ b/documentation/audits/evidence-recovery-code-2026-09-16/phaseA-live.txt @@ -0,0 +1,130 @@ +## R-543 live validation — controller v0.245.0, 2026-09-16, demo-hp +## Method: endpoint-level. Every request below is the exact request the browser makes (same handler, +## same session cookie, same CSRF); only rendering is skipped. claude-in-chrome is not available here. +## Reachability note: neither guest answers on its LXC address from DooPlex, and the controller does +## not listen on the guest's 127.0.0.1 — it answers on the container address 172.17.0.2:8080 with the +## mandatory Host header. All requests therefore run INSIDE the guest via `pct exec`. +## Secrets: the dashboard password was read with scripts/read_credential.py (value never printed, +## file->file, 0600) and passed to curl as --data-urlencode password@. + +### THE SEAM, STATED +The two halves of this proof are on two boxes, not one box before and after a ceremony: + * PAUSED half — guest 9202 (scratch), off-site configured and escrow NOT complete. + * ESCROWED half — guest 9201 (demo), off-site configured, escrowed, running. +Running a real escrow ceremony on a fresh target would provision off-site storage, and this task's +fences put ep0 out of bounds. So the escrowed side is OBSERVED on a box that is already escrowed +rather than produced here. What is NOT weakened by the seam: both boxes run the same binary +(0.245.0), and the paused box's state was produced through the product's own configuration endpoint. + +## ── 1. ESCROWED + ACTIVE — guest 9201 (0.245.0) ──────────────────────────────────────────────── +login OK +GET /dashboard -> 200 bytes=60816 escrow-bar-hits=0 +GET /launcher -> 200 bytes=44462 escrow-bar-hits=0 +GET /backups/apps -> 200 bytes=106304 escrow-bar-hits=0 +tier-1 file sentence, as rendered (2 class-A apps): + "Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi — ez a helyi mentés a + beállításokat és az adatbázist tartalmazza." +word counts on /backups/apps: Kulcslet=0 szünetel=0 védi=2 védené=0 +=> a box whose recovery code exists is NOT nagged, and the promise it prints is a true one. + +## ── 2. PAUSED — guest 9202 (0.245.0) ─────────────────────────────────────────────────────────── +The state was produced through the product's own endpoint, POST /backup/offbox/config, with a +throwaway ed25519 key and a pinned known_hosts line. 9202 has no agent local API, so the escrow +STAGE could not run — the handler said so and saved the target anyway: + + POST /backup/offbox/config -> 302 + flash: "A távoli mentési cél elmentve. — a kulcs letéti előkészítése nem sikerült + (az ügynök nem elérhető); próbáld újra." + +Persisted state afterwards (read from the guest's own settings.json): + enabled=True host=192.168.0.162 repo_path=/mnt/nvme-1tb/r543-paused-proof + escrow_state=pending last_run=None last_status=None snapshot_count=None + data/offbox/: known_hosts 95 B (0644) | repo_password 64 B (0600) | ssh_key 411 B (0600) +=> OffboxConfigured() is genuinely true (valid target + both secret files), and the escrow is NOT + complete. This is the state a fresh box lands in on day one. + +### 2a. The bar is on EVERY page, not on the three someone remembered + GET /dashboard -> 200 escrow-bar-hits=1 + GET /launcher -> 200 escrow-bar-hits=1 + GET /backups/apps -> 200 escrow-bar-hits=1 + GET /settings -> 200 escrow-bar-hits=1 + GET /apps -> 404 escrow-bar-hits=0 (no such route; a 404 carries no bar) + +Quoted from /dashboard: + "A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." + (the route out, on the bar itself) + +### 2b. A manual off-site run, while paused — refused, FOR THE RIGHT REASON + POST /backup/offbox/run -> 302 + flash: "A távoli mentés a kulcs letétbe helyezésére vár." + after the attempt: last_run=None last_status=None snapshot_count=None +=> No snapshot, and the refusal came from the fork-4 escrow gate (offbox_handlers.go:224 + OffboxRunnable), NOT from an unreachable target — the target was never contacted. This matters: + a run that failed to connect would have proven nothing about the pause. + +### 2c. „Most nem" is for the visit only + POST /backup/escrow/banner/dismiss -> 302 + Set-Cookie: felhom_escrow_banner=1; Path=/; HttpOnly; SameSite=Lax + ^ no Max-Age and no Expires => a browser SESSION cookie, exactly as R-241 does it + same visit, cookie sent -> escrow-bar-hits=0 + next visit, no cookie -> escrow-bar-hits=1 +=> the off-site tier is still paused tomorrow, so the question is still asked tomorrow. + +### 2d. The tier-1 FILE sentence, live, with the tier paused +The first pass of this phase could not show it: 9202 had no class-A app, so the sentence had nothing +to render („védi"=0 AND „védené"=0 — the „szünetel"=1 on that page was the BAR in the layout, not the +sentence). Recorded because it was briefly written down as a limit and it was not one. A throwaway +class-A app was then deployed on the box's own drive and the off-site copy turned on for it: + + POST /api/stacks/calibre-web/deploy -> 202 {"ok":true,"message":"Telepítés elindítva ..."} + (HDD_PATH=/mnt/felhom-drives/scratch_hdd — the guest's own registered data drive, 938 G) + POST /backup/offbox/toggle app=calibre-web enabled=true -> 302 + +Rendered on /backups/apps, with the tier configured and PAUSED: + + "Az alkalmazás fájljait a távoli másolat védené — a távoli mentés a helyreállítási kód + létrehozásáig szünetel." + ...followed by the route: "Helyreállítási kód létrehozása →" + + word counts on /backups/apps: Kulcslet=1 szünetel=3 védi=0 védené=1 + Kulcslet=1 is the tier-3 row's own state („Kulcsletétre vár"). + szünetel=3 is the bar + the sentence + the tier row. + védi=0 is the point: the page no longer claims a protection that has never run. + +=> Both wordings are now observed LIVE on real boxes running 0.245.0: „védi" on the escrowed box + (9201, §1) and „védené … szünetel" on the paused box (9202, here). + +## ── 3. TEARDOWN, three layers ────────────────────────────────────────────────────────────────── +Recorded after the evidence above was already written to DooPlex (R-320: evidence leaves the machine +at the END OF THE PHASE, before any revert — not at the end of the session). + +MACHINE (guest 9202, scratch): + * the throwaway class-A app removed WITH its data and its backups: + POST /api/stacks/calibre-web/stop -> 200 + POST /api/stacks/calibre-web/remove -> 200 {"removed":"calibre-web", + "volumes_removed":["calibre-web_calibre_web_config"],"hdd_paths_removed":[], + "hdd_note":"Az alkalmazás nem tárolt saját adatot…"} + verified after: containers named calibre = 0, drive folders = 0 + * the off-site target this proof created is GONE: + before: offbox present=True, data/offbox/ held ssh_key, repo_password, known_hosts + after : offbox present=False, data/offbox/ ABSENT, all three files `shred -u`'d + MY OWN MISTAKE, recorded: the first teardown pass used the CONTAINER's view of the data path + (/opt/docker/felhom-controller/data) from a shell running in the GUEST, where that path does + not exist. It printed „offbox dir now: ABSENT" — which was TRUE of a path that never existed + and FALSE of the thing being claimed. The target was still fully configured. The same wrong + path had already produced three FileNotFoundError tracebacks earlier in this phase; I read + those as noise instead of as the instrument telling me it was pointed at nothing. The re-run + stops the controller first, edits the real file, shreds the secrets, restarts, and RE-READS + the state to confirm — a teardown asserted is not a teardown observed. + * controller restarted and healthy: felhom-controller:0.245.0 Up (healthy) + * temp files: none left (`ls /tmp/.r543*` empty in the guest) + * NOT touched: filebrowser, traefik, the box's registered data drive, its password, its claim state. + +HOST (demo-hp): /tmp/.r543pw and /tmp/.r543key shredded; .r543kh, the scripts and one stray + /tmp/.r543out removed. Verified: 0 files matching /tmp/.r543* remain. + +HUB: provisioned nothing. Guest 9202 has hub reporting OFF and no tunnel, so no customer, no host + record, no escrow row and no event was created by any of this. Nothing to tear down. + +NOT torn down, deliberately: guest 9201 now runs controller 0.245.0. That is the release being +shipped, not a drill artifact, and its own off-site tier was never touched (still escrowed, active). diff --git a/documentation/audits/evidence-recovery-code-2026-09-16/phaseA-redproofs.txt b/documentation/audits/evidence-recovery-code-2026-09-16/phaseA-redproofs.txt new file mode 100644 index 00000000..64940c2a --- /dev/null +++ b/documentation/audits/evidence-recovery-code-2026-09-16/phaseA-redproofs.txt @@ -0,0 +1,48 @@ +## R-543 red-proofs — controller v0.245.0, 2026-09-16, DooPlex +## Each fix was BROKEN first and the test was watched convicting it. A test never seen failing has +## not been shown to test anything. + +### RED-PROOF 1 — the reminder bar +Break: delete `s.addEscrowBanner(data, r)` from executeTemplate (internal/web/server.go). +That is the whole wiring: the bar hangs off the single render choke point, so removing one line +returns the product to the measured 2026-09-16 state (a paused off-site tier, and silence). + +--- FAIL: TestR543_A_PausedBoxAsksOnEveryPage (0.20s) + r543_escrow_banner_test.go:72: R-543: /dashboard does not tell the household the off-site copy is PAUSED. The tier is on, nothing is running, and the page is silent about it + r543_escrow_banner_test.go:76: R-543: /dashboard states the pause but names no route to end it — a reminder without its door is the shape that left a fresh box waiting indefinitely + r543_escrow_banner_test.go:72: R-543: /launcher does not tell the household the off-site copy is PAUSED. The tier is on, nothing is running, and the page is silent about it + r543_escrow_banner_test.go:76: R-543: /launcher states the pause but names no route to end it — a reminder without its door is the shape that left a fresh box waiting indefinitely +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.208s + +Note BOTH pages fail. That is the point of the hook placement: a per-handler helper would have +covered the three pages someone remembered, which is the seam-built-but-never-wired shape. + +### RED-PROOF 2 — the tier-1 sentence +Break: return the v0.244.0 wording from driveFilesNoteFor before the state switch, i.e. compute the +sentence from the app's shape alone, exactly as v0.244.0 shipped it. + +--- FAIL: TestR543_Tier1Sentence_PausedStateDoesNotPromise (0.00s) + r543_tier1_sentence_test.go:38: R-543: the sentence claims the files ARE protected while the copy is paused for the recovery code. This is the exact promise a fresh box read for its whole first day: "Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi — ez a helyi mentés a beállításokat és az adatbázist tartalmazza." + r543_tier1_sentence_test.go:42: R-543: the paused sentence must say the copy WOULD protect them and that it is waiting; got "Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi — ez a helyi mentés a beállításokat és az adatbázist tartalmazza." + r543_tier1_sentence_test.go:46: R-543: the paused sentence names no route out (link="" text="") +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.007s + +### RESTORED +go test ./internal/web/ -run 'R543' +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.895s + +### The filter was proven to match (the `-run` trap: a pattern matching nothing prints `ok`, exit 0) +=== RUN TestR543_A_PausedBoxAsksOnEveryPage --- PASS (0.29s) +=== RUN TestR543_B_EscrowedBoxIsNotNagged --- PASS (0.19s) +=== RUN TestR543_C_UnconfiguredBoxIsNotNagged --- PASS (0.20s) +=== RUN TestR543_D_DismissIsForThisVisitOnly --- PASS (0.28s) +=== RUN TestR543_Tier1Sentence_ActiveStateKeepsThePromise --- PASS +=== RUN TestR543_Tier1Sentence_PausedStateDoesNotPromise --- PASS +=== RUN TestR543_Tier1Sentence_NoCopyAtAllSaysSo --- PASS +=== RUN TestR543_Tier1Sentence_SecondDriveCounts --- PASS +=== RUN TestR543_Tier1Sentence_NoFileLegsNoSentence --- PASS + +### Fixture validity (an instrument that can lose its precondition measures nothing) +escrowServer asserts backupMgr.OffboxConfigured() itself before any assertion runs: the target is +enabled and valid AND the ssh_key + repo_password files exist on disk. A fixture that silently fell +back to "not configured" would make every one of these tests pass for the wrong reason. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 7c25f4c1..72499f81 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -725,8 +725,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-540** | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** | | **R-541** | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** | | **R-542** | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** | -| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. | **READY — rank P1-HIGH; owner: CC (controller copy + first-run prompt)** | +| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** | | **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** | +| **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** | diff --git a/documentation/runbooks/VOLUNTEER-first-hour.md b/documentation/runbooks/VOLUNTEER-first-hour.md index 995dc5a8..3a395335 100644 --- a/documentation/runbooks/VOLUNTEER-first-hour.md +++ b/documentation/runbooks/VOLUNTEER-first-hour.md @@ -102,7 +102,26 @@ nem kell újra összekötni.)* **legalább 12 karakteres** jelszót. Ez lesz a vezérlőpult jelszava. 4. Ha nem jött meg a kód: **„Új kód kérése"** — mindig ugyanarra az e-mail címre érkezik. -## 6. Az első két alkalmazás telepítése (~2 perc) +## 6. A helyreállítási kód (~2 perc) — ezt ne hagyd ki + +A doboz a fájljaidról **titkosított** másolatot küld a Felhom távoli tárhelyére. A titkosítás +kulcsát **csak te** kapod meg: ez a **helyreállítási kód**. Amíg nem hozod létre, **a távoli +mentés nem indul el** — a vezérlőpult minden oldalán látszó sáv ezt írja: +„A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." + +1. **Biztonsági mentés → Távoli mentés**, vagy egyszerűen kattints a sávon a + **„Helyreállítási kód létrehozása"** hivatkozásra. +2. Add meg a vezérlőpult jelszavát, és indítsd el. Kb. fél perc. +3. A kód **egyetlen egyszer** jelenik meg a képernyőn. **Írd fel papírra**, és tedd el oda, ahová a + fontos iratokat teszed. Fénykép a telefonról nem elég, ha a telefon is elveszik. + +> ⚠ **A Felhom nem tudja visszaszerezni ezt a kódot.** Nem azért, mert nem akarja: a másolataidat +> úgy titkosítják, hogy mi magunk se tudjuk megnyitni őket. Ha a kód elvész, a távoli másolatok +> megmaradnak, de **senki — mi sem — nem tudja többé megnyitni őket**. + +Ha ezzel megvagy, a sáv eltűnik, és a távoli mentés magától elindul. + +## 7. Az első két alkalmazás telepítése (~2 perc) 1. A vezérlőpulton: **Alkalmazások**. Keresd meg például a **BookStack**-et (családi wiki) és a **PrivateBin**-t (titkosított jegyzet), és nyomd meg a **Telepítés** gombot. @@ -110,7 +129,7 @@ nem kell újra összekötni.)* Az „Automatikusan generált értékek" részt nem kell felírnod. 3. **Telepítés indítása.** A BookStack kb. 1 perc, a PrivateBin kb. 20 másodperc. -## 7. Első belépés az alkalmazásokba +## 8. Első belépés az alkalmazásokba - **BookStack:** az alkalmazás oldalán az „Első lépések" rész a címet `wiki.DOMAIN` alakban írja — a DOMAIN helyére a saját domained kerül (ismert hiba). Belépés: `admin@admin.com` / `password`. @@ -118,7 +137,7 @@ nem kell újra összekötni.)* - **PrivateBin:** nincs belépés. Írj be szöveget, **Küldés**, és a kapott linket oszd meg — a kulcs a linkben van, a szerver nem látja a tartalmat. -## 8. Mentések +## 9. Mentések - **Biztonsági mentés → Áttekintés:** két sárga figyelmeztetést látsz („Csak egy másolat készül", „ugyanazon a lemezen van") — **ezek igazak**: amíg nincs második meghajtó vagy távoli mentés, egy @@ -128,23 +147,23 @@ nem kell újra összekötni.)* teljes rendszermentésben (PBS)" — ez nem minden dobozra igaz (ismert hiba). Az Áttekintés oldal a pontos.* -## 9. Visszaállítás +## 10. Visszaállítás **Biztonsági mentés → Visszaállítás:** válaszd az alkalmazást és a mentést, pipáld be a „Megértettem" négyzetet, **Visszaállítás indítása**. Az alkalmazás kb. fél percre leáll, majd az utolsó mentés állapotával indul újra. -## 10. Alkalmazás eltávolítása +## 11. Alkalmazás eltávolítása **Alkalmazások:** előbb **Leállítás**, utána megjelenik az **Eltávolítás**. A párbeszédablak felsorolja, mi törlődik mindenképp, és bepipálhatod a mentések törlését is. -## 11. Áramszünet +## 12. Áramszünet Ha elmegy az áram, a doboz magától visszaindul, kb. 2 perc múlva minden alkalmazás ugyanazon a verzión fut, amelyen előtte. A vezérlőpultba újra be kell jelentkezned. -## 12. Ha elgépelted a kódot +## 13. Ha elgépelted a kódot A beállító oldal „Hibás vagy lejárt kód" üzenettel visszadobja — írd be újra. **Öt** hibás próbálkozás után 15 percre zárol („Túl sok próbálkozás — próbáld újra 15 perc múlva."). Új kódot a diff --git a/documentation/runbooks/day0-install.md b/documentation/runbooks/day0-install.md index 5d7f34cd..a51405fc 100644 --- a/documentation/runbooks/day0-install.md +++ b/documentation/runbooks/day0-install.md @@ -95,6 +95,26 @@ On save the hub generates two credentials: self-bind mail tells the customer they received it from the Felhom operator.) Treat it like a password. - **Customer API key** — internal (baked into the generated controller.yaml); never handled manually. +### A.2b REBUILDING a box that already exists — the one press (F-14 ruling, 2026-07-13) + +**A rebuilt box does NOT always get its off-site credentials by itself, and this step was missing +from this runbook.** Which of the two happens is decided by how the PREVIOUS box left: + +| How the previous box was removed | What the hub does | Operator action | +|---|---|---| +| Deleted through the **acknowledged** delete flow (the escrow acknowledgement was given) | the hub re-issues the off-site credentials **by itself** when the new box enrolls — the box consumes the single-use secret and the tier comes up | **none** | +| The host record was left in place, or the delete was not acknowledged | the hub **reuses** the existing record and mints nothing — the box asks, is refused, and the off-site tier never starts | **one press:** Hub UI → the customer → **„Re-issue PBS credentials"** | + +The refusal is correct and deliberate: re-issuing over a live record would orphan the history the +customer's recovery code protects. The mint-once-and-reuse decision lives in `hub/internal/web/pbsdr.go`; +there are **two** Re-issue buttons on that page — the off-site one and the PBS-DR one — and they are +not interchangeable. + +> This is the correction to the note that said a rebuilt box needs zero presses. It needs zero +> presses only on the acknowledged path. Measured on a fresh box 2026-09-16: the box reproduced the +> refusal by itself, the re-issue then ADOPTED the record (generation 2) and the box consumed the +> single-use secret. + ### A.3 Verify the Day-0 artifact manifest Hub UI → **Configuration → Day-0 artifacts**. The manifest must vouch an **agent version** and a