diff --git a/CONTEXT.md b/CONTEXT.md index ac2abdc8..bdd0c114 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -20,7 +20,11 @@ > restore test takes only the box's own).** `09` §3. 50: when the customer has off-site, every newly installed app is > in the off-site copy; over quota the page names the largest and the household chooses; history is never deleted to > make room without that choice. 51: the three drill archives in `tester-1`'s ep0 namespace go, nothing else. Also: the -> first real tester is a NEW record `Tester-2`; the volunteer guide is rewritten as measured (R-722). +> first real tester is a NEW record `Tester-2`; the volunteer guide is rewritten as measured (R-722). **Outcome (same day):** controller v0.283.0/v0.283.1 (apps off-site by default + one press + size card; a Stop holds +> at the quiesce resume, the volume dump, the update leg and the crash recovery — v0.283.0 missed the dump's production +> adapter, caught live), agent v0.138.0 (the restore test skips an archive written with another key; signed delivery to +> both demo boxes), hub v0.126.0 („Új linket kérek" on an expired/used bind link; no day-one false mails), ep0's three +> drill archives removed. Floor 0.283.1. Report: `REPORT-fixes-first-tester-2026-09-30.md`. > **2026-09-29 evening — operator rulings 48 (wanderer stays, A) and 49 ("close sign-up now", A).** Controller > **v0.282.0**: `after_setup:` (the app's own sign-up switch, env merged + one `compose up -d`, when the gate opens or on diff --git a/REPORT-fixes-first-tester-2026-09-30.md b/REPORT-fixes-first-tester-2026-09-30.md new file mode 100644 index 00000000..62b86dd0 --- /dev/null +++ b/REPORT-fixes-first-tester-2026-09-30.md @@ -0,0 +1,79 @@ +# REPORT — 2026-09-30: the fixes before the first real tester + +Architecture read first: `07` §6 (the 2026-09-16 ruling), `03` (the restore test's selection), `05` (self-bind), +`04` §3.1 (signed delivery), `09` §3 decisions 45–49; rows R-719 … R-727, R-95, R-240, R-494, R-600, R-688. +Releases: **controller v0.283.0 + v0.283.1**, **agent v0.138.0**, **hub v0.126.0**. Evidence: +`documentation/audits/evidence-fixes-first-tester-2026-09-30/` (part0, partA, partC, partD, partE, release). + +## Tester-2 — read-only checklist (nothing on Tester-2, Cloudflare, ep0 or the Storage Box was changed) + +| # | item | state | where to fix | +|---|---|---|---| +| 1 | Tunnel token pasted | **done** — tunnel `3ce0eccd…` (Peti's original, reused) | — | +| 2 | DNS of `sajatfelhom.hu` | **done** — ONE record `*.sajatfelhom.hu` → that same tunnel; no leftover pointing elsewhere; today 530 (no connector, right with no box) | — | +| 3 | The tunnel's route `*.sajatfelhom.hu` → `https://traefik`, No TLS Verify | **UNKNOWN** — the record's Cloudflare key reads DNS only (`Authentication error` on the tunnel; stopped there) | Cloudflare → Zero Trust → Networks → Tunnels → this tunnel → Published application routes | +| 4 | Off-site | **done** — shared 100 GB, new sub-account 322460 (username `sub2` reused; Peti's was emptied and deleted 2026-09-25) | — | +| 5 | DR tier (ep0) | **done** — ON; nothing of Tester-2 or Peti on ep0 yet (made at the first connection) | — | +| 6 | The connect e-mail | **sent twice at 07:36 UTC** (the customer was created twice, R-728) — **only one of the two links works**; valid until 2026-10-07 07:36 UTC | Tell your friend: if one link says „expired", use the other; or press „Send self-bind link" once just before the install | +| 7 | E-mail language | **English** — mail, bind page and the box start in English | Hub → Tester-2 → Edit, if Hungarian is wanted | +| 8 | Owner passphrase | yours to hand over | in person / by phone | +| 9 | Customer id `Tester-2` has a capital letter | never walked before; no known break | note only | + +## The Parts + +| Part | state | note | +|---|---|---| +| 0 Tester-2 read-only | **done** | checklist above; items 3 and 6 need you | +| A1 measure the over-quota path | **done** | it deletes nothing extra: refuses new pushes, runs only the ruled retention; now pinned | +| A2 apps off-site by default + one press for older apps | **done** | controller v0.283.0; live on 9202 | +| A3 size warning | **done** | page card; unit + parity proven (a household NAS target has no quota to test live) | +| B guide + slips | **done / narrowed** | six stale lines rewritten (the sixth found on the way: auto-reboot); R-724/R-725 partly, the rest narrowed | +| C1 ep0 cleanup | **done** | three archives removed; every other namespace byte-identical | +| C2 agent v0.138.0 | **done** | signed delivery to both demo boxes (340 s); due-check normal on both | +| D fresh link for a returning customer | **changed** | the brief's trigger cannot be built (a registering box is unclaimed); built „Új linket kérek" on the expired/used page | +| E Stop holds | **done, after a live failure** | v0.283.0 was wrong in production (adapter); v0.283.1 fixed and proven live | +| F day-one mails | **done** | unit + red-proof; live proof at Tester-2's first hour | + +## Claims in the brief that turned out wrong, named + +- **"The over-quota path prunes history"** — **wrong.** Over the quota the box refuses new pushes and runs the SAME + retention as every night; with no new snapshots nothing extra ages out. No P1. +- **"Peti's Cloudflare records for `sajatfelhom.hu` may still be there"** — **there is exactly one record, and it + points at the tunnel Tester-2 now carries** (Peti's tunnel, reused). Nothing stale to remove. +- **"A PBS archive carries its box's key fingerprint or host id"** — **half right:** the key fingerprint yes (PVE + content `encrypted`), a host id no (the comment is only „felhom local-api", the owner is the customer's token). +- **"The hub sends no link at registration"** — **right, and it cannot:** the registration carries nothing of a + customer. The fix was changed to a button on the old link's page. +- **"Stop is lost at the backup's resume only"** — **wrong:** also at the nightly volume dump, the update leg and the + startup crash recovery; all four fixed. +- **Mine, from 2026-09-30:** "the ✗ names the wrong tier" — misread; the local-tier heading was the NEXT section. + +## Red-proofs (each seen failing on its assertion, then restored) + +RP31 new app not ON · RP32 earlier OFF overridden · RP33 hook not wired · RP34 over-quota forget differs · RP35 exact +quota "does not fit" · RP36 quiesce restarts a stopped app · RP37 dump restarts it · RP38 update leg presses it · +RP39 (agent) an earlier box's archive picked · RP40 new box "recovered" · RP41 first-hour skip mailed · RP42 no fresh +link · RP43 the production adapter does not answer · RP44 crash recovery restarts it. Outputs: +`evidence-fixes-first-tester-2026-09-30/` and the scratchpad `rp/` copies. + +## Slips of mine, said + +- The first hub/evidence commit went out after my secret-scan script crashed (it looked for last night's shredded + files). Scanned right after: 5 125 files, 7 secrets, **0 hits**, control 1. Nothing leaked. +- v0.283.0 shipped the Stop fix un-wired in production; the live test caught it; v0.283.1 is the second controller + release this session (the one-release rule bent, reason in its CHANGELOG). + +## Rows + +Closed **R-719, R-720, R-721, R-722, R-727**. Fixed pending live proof **R-723**. Narrowed **R-724, R-725**. Opened +**R-728** (customer created twice), **R-729** (no way to remove an off-site target). **Register 359 → 361 rows.** +R-726 (a returning customer's orphaned repository) stays open — a NEW record does not meet it. + +## Teardown + +- **9202:** controller left on 0.283.1 (scratch; the floor does not reach it); the throwaway NAS target switched off + through the form and then removed from `settings.json` with the controller stopped (R-729: no product path); + `glance` removed with its data; paperless-ngx running; the off-site switches back OFF. +- **Demo boxes:** controller 0.283.1 by the floor, agent 0.138.0 by signed jobs — nothing else. +- **ep0:** only the three archives of decision 51. **Hub:** floor 0.283.1 (MinAgent 0.131.0 declared); hub 0.126.0. +- **Tester-2 / Cloudflare:** read only. DooPlex: pushes, builds, the hub deploy, signing. diff --git a/STATUS.md b/STATUS.md index e31563ec..19a8a0e7 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,23 +1,27 @@ # STATUS — what works, what's broken, what's next -**Ready for a first real tester: YES, on a new customer record — a fresh box went from the download to restored apps with no help at all; fix the volunteer guide first, and decide whether their apps go off-site by themselves (today they do not).** +**Ready for the first real tester (Tester-2): yes, once you check two things — the tunnel's route in Cloudflare, and which of the two connect mails your friend uses.** -**Updated 2026-09-30 morning. Both demo boxes run controller 0.282.0 and host agent 0.137.0. Hub 0.125.0. New installs get golden 0.282.0 with agent 0.137.0 (baked and vouched last night).** +**Updated 2026-09-30. Both demo boxes run controller 0.283.1 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.282.0 with agent 0.137.0.** + +**Decisions today** (yours, recorded): every new app goes off-site by itself (50); the three drill backups on ep0 go, and the restore test uses only the box's own backups (51). + +**Tester-2 — what I checked, read only.** +- Done: tunnel key pasted; the domain has exactly one DNS record, and it points at that same tunnel; off-site on (100 GB, a fresh storage account); DR tier on; e-mail language English. +- **Check yourself:** the tunnel's route. My Cloudflare key could read DNS but not tunnels. It must say `*.sajatfelhom.hu` → `https://traefik`, with „No TLS Verify" ticked. +- **Two connect mails went out** (the customer was created twice). Only one link works. Tell your friend: if one says "expired", use the other. The live one is valid until 7 October. **What I did, and it worked.** -- **A fresh box, walked like a volunteer, needed no help.** Download, install on the Hungarian keyboard, the e-mailed link, the dashboard through the internet, the recovery code, three apps, a family member, backup, restore, delete, a power cut, typos, a phone. Last walk needed one intervention; this one needed none. -- **The new safety parts held on a fresh box.** Strangers saw only the gate page. The gate opened by itself after the household's setup. Sign-up stayed closed. The random first password worked, and the old default password was refused. +- **Apps go off-site by themselves now.** Older apps get one button. If the apps do not fit in 100 GB, the page names the biggest. The box never deletes old copies to make room — I checked; it never did. +- **A Stop holds during a backup.** The first version was wrong in real use; my live test caught it, and the second version is proven live. +- **ep0 is clean:** the three old drill backups are gone; the other customers' backups are unchanged. +- **A returning customer gets a new link** from a button on the expired link's page. +- **The volunteer guide is current** (six stale lines fixed). +- **Two false operator mails on a new box's first day are gone.** -**What broke.** -- **No off-site copy on night one.** Every app starts with its off-site copy switched off, and nothing tells the household. On top of that, tester-1's old off-site store belonged to an earlier box, so the box skipped it all night. The household's data stayed safe on the box. -- **The night's restore test tried an old box's backup** and failed on its key. The page shows a bare ✗ and names the wrong place. -- **A Stop pressed during a backup was undone by that backup.** -- **The volunteer guide is out of date in five places.** For example, it still gives BookStack's old default password, which is now refused. - -**Rows.** 9 opened, 1 closed. The list went from 350 to 359 rows. +**Rows.** 5 closed, 2 opened. The list went from 359 to 361 rows. **What needs you.** -1. **Decide: should apps go off-site by themselves?** Option A: yes, every app is switched on when the customer has off-site (my pick — that is what the guide already promises). Option B: no, the guide tells the household to switch each app on. If you do nothing, a new household's apps have no off-site copy. -2. **Approve the guide fix** (five lines, I write them). If you do nothing, a volunteer tries BookStack's old password and fails. -3. **Three whole-machine backups of deleted drill boxes are still in tester-1's space on ep0** (two from 16 September, one from last night). I only looked; I did not touch them. The old ones make the restore test fail. If you do nothing, every new tester-1 box repeats that failure. Say "remove them" and I will do it with before/after checks. -4. **D4, the image copies, the Peti leftovers:** unchanged. If you do nothing, nothing changes. +1. **The Tester-2 tunnel route** (above). If you do nothing and it is missing, your friend's dashboard answers an error page. +2. **A golden before your friend installs?** The rule says bake one before any fresh install; this task did not include it. Option A (my pick): I bake 0.283.1 before the install — the box then starts on today's code. Option B: skip it — the box starts on 0.282.0 and updates itself to 0.283.1 within seconds (proven on both demo boxes today). If you do nothing, B happens. +3. **D4, the image copies:** unchanged. If you do nothing, nothing changes. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 73cd6ced..7602eb13 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -89,7 +89,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis **NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | | **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) The venue is **torn down** (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and **30 are R-244**, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and **16.64 GiB** returned. `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. Evidence: `tests/walk5-r201-2026-08-07/journal.md` **SCOPE NOTE 2026-09-14 — not re-walked.** The first-hour drill on 0.242.0 (`audits/DRILL-fresh-install-0242-2026-09-14.md`) walked install → first use → restore of a deleted page on a fresh box, NOT a rebuild with off-site recovery (DR tier and off-site were off). This row is therefore neither re-proven nor contradicted on 0.242.0; its PROVEN-LIVE stands on 0.206.0 only. The first hour has its own row below. **SCOPE NOTE 2026-09-17 — not re-walked, and this time it could not be.** Chaos night (`audits/DRILL-chaos-night-2026-09-17.md`) ran twelve rounds of household actions under injected accidents on a fresh 0.245.0 box and **did not walk the recovery journey at all**. It could not have: the box was a REBUILD for an existing customer, so its restic repository was orphaned by design — zero readable snapshots, confirmed independently by `restic` (`Fatal: wrong password or no key found`, exit 1) and by the product's own status (`orphaned:true, snapshots:0, status:"error"`). The product surfaced that honestly as a true alarm within seconds of the first off-site run. This row is therefore **neither re-proven nor contradicted on 0.245.0**; its PROVEN-LIVE still stands on 0.206.0 only. | | **A random night of household actions while random things go wrong — twelve rounds, unattended** | controller v0.245.0, agent v0.131.0, hub v0.116.0, golden 0.245.0, ISO 1.28.0 | **PROVEN-LIVE (2026-09-17, chaos night)** | A fresh nested box installed itself from the **published** 1.28.0 ISO and bound with **zero operator presses** (the automatic self-bind mail was already waiting; the acknowledged-delete path re-issued PBS credentials by itself — the F-14 path, measured live for the first time). Twelve rounds were drawn **once** from seed `20260917` by a committed script and written into the findings document **before round 1 began**. Across a power cut mid-restore, a hard reset four seconds into another, a system disk at 96 %, a killed tunnel, a restarted Docker, three severed networks and **the data drive pulled out of a running machine for twenty minutes**: the box healed itself **every time** with no human action — 148 s after the power cut, 97 s tunnel repair by the controller, 150 s after the hard reset, 67 s after the drive returned. **17 alarms fired, all 17 TRUE, none missing**, and the mailbox proves each was **delivered** to the operator, not merely stored. **One intervention** all night (a local backup leg that could never have fit; its off-site leg then succeeded unaided). Evidence: `audits/DRILL-chaos-night-2026-09-17.md` + `audits/evidence-chaos-night-2026-09-17/` | **WHAT IT DOES NOT CLAIM.** (1) **Per-app off-site RESTORE was not tested** — see the scope note on the recovery-journey row above; the repository was orphaned by design and held zero readable snapshots. The WHOLE-GUEST off-site copy did work and is present on ep0 (two intact snapshots, the later one written tonight), but that is a **listing, not a verification** — restorability was not tested. (2) **The dropped-event path was never exercised.** Events pushed while the hub is unreachable are retried 3× then dropped permanently with no queue; three ten-minute hub outages happened and **no event was raised during any of them**, so that behaviour remains unmeasured. What WAS measured is the REPORT path: built, three attempts over 1 m 40.8 s, given up, and the next scheduled report succeeded — a snapshot, so nothing was lost. (3) **Twelve rounds is a sample, not coverage.** (4) The household loop samples each app every two minutes, so ten of the twelve rounds left no mark in it — that silence is the instrument's sampling rate, not proof the household saw nothing. Three findings filed: **R-547**, **R-549**, **R-550**. | -| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. **DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven.** `audits/DRILL-prove-fixes-0243-2026-09-16.md`. A fresh nested box was installed from the published ISO, landed on this drill's own golden **by checksum** (`e2d1843c…c10a`, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, **the tunnel from outside (302→200, R-510 CLOSED)**, the data drive, the file manager with its own generated password (`admin`/`admin` refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with **26 s** of app downtime. **Interventions: 0** — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. **The automatic connect e-mail is PROVEN with a real mailbox:** a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at **12:23:00Z**, with `selfbind_link_sent (host delete)` on the timeline (R-509). **WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence:** on a one-drive box with no off-site tier — what every fresh install is — the household's files are in **no backup at all** (the whole-guest tiers exclude `mp8` by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (**R-537**), and a restore reports success while leaving Nextcloud listing five photos that return `Sabre\DAV\Exception\NotFound` — after making the app's own trash, which still held every byte, unreachable (**R-538**). **Also still open from this walk:** the off-site tier cannot be provisioned at all (**R-534**, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (**R-535**), and „app installed” is emitted at accept time (**R-536**). Faults re-measured: controller killed during a deploy → back in **37 s**; three more kills 20 min apart → **61/41/61 s**, none accumulating (**R-531**); two reboots 60 s apart → everything back in **124 s** with the boots NOT counted against the brake; the claim page locks after the **second** wrong code and mails the operator truthfully. **2026-09-16 (evening) — THE BACKUP HALF IS NOW WALKED TOO, on controller 0.244.0 / hub 0.116.0 / golden 0.244.0 / ISO 1.28.0 (built, unpublished).** `audits/evidence-backup-promise-2026-09-16/`. A second fresh box (VM 335) was installed from the BUILT image and walked to the end of the sentence this row could not previously finish: **five photos in → deleted the way a child would → the old route REFUSED and touched nothing → the off-site restore returned them → they OPEN, byte-identical (sha256 5/5, negative control)**. The refusal reads „Ez a mentés nem tartalmazza az alkalmazás fájljait… a fájlok így a helyükön maradnak" and names the route that can help; the app was `running` before and after; the wastebasket was untouched. The off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution counting „5 fájl és 3 adatkötet és az adatbázis". **Delivery is part of it:** the box landed on agent 0.131.0 + controller 0.244.0 (the golden baked and vouched the same day) with no hand upgrade, and `app_deploy_started` (19:15:34) / `app_deployed` (19:16:23) finally mean different things. **The self-bind half needed NO operator press** — the box registered itself and the bind used the mail the hub sent itself after the morning's host delete (R-509, proven twice today: 12:23:00Z and 18:17:46Z, each one second after a host delete). **WHAT IT STILL DOES NOT CLAIM — one operator press and one day-one gap.** The PBS-DR cascade stopped at the refusal R-511 documents („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action") and needed the operator to press it; that press then SUCCEEDED because of the ep0 grant given this morning (R-534 CLOSED, R-511 CLOSED). And off-site ON by default is not off-site WORKING: a fresh box sits at „Kulcsletétre vár" until the household performs the escrow ceremony, which nothing asks them to do, while the tier-1 row already promises that copy (**R-543**, P1). **2026-09-16 (late evening) — that last gap is CLOSED, controller v0.245.0 (R-543).** The pause is the zero-knowledge escrow design and was not touched; what was missing was the ASK. Every authenticated page now carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking the ceremony (the R-241 bar, second instance, hung on the single render choke point), the tier-1 sentence renders by tier-3 STATE („védené … szünetel" while paused, „védi" when running), and the first-hour guide asks for the code right after the dashboard password and before the first app. **Measured on two boxes running 0.245.0:** paused box — bar on four pages, `POST /backup/offbox/run` refused by the fork-4 gate with no snapshot written, app row „védené" and „védi"=0; escrowed box — no bar anywhere, row „védi". `audits/evidence-recovery-code-2026-09-16/`. **WHAT IT STILL DOES NOT CLAIM:** the ask has not been walked by an actual volunteer from the written guide — the sentence is proven, the human following it is not. So the journey now reads: **files protected from day one, once the household writes down the recovery code the box asks them for on every page.** **2026-09-29/30 — RE-WALKED on golden 0.282.0 (controller v0.282.0, agent v0.137.0, hub v0.125.0, published ISO 1.29.0), customer `tester-1`: the WALK is PROVEN-LIVE with ZERO interventions; the DAY-ONE BACKUP sentence is NARROWED.** `audits/DRILL-new-household-2026-09-30.md`. Download → install on the Hungarian keyboard → the mailed link and self-bind page (one typo refused) → the claim **through the public tunnel** → the recovery code → three apps (a random first password that works while `password` is refused; a gated app whose probe opened it ~12 s after the household's setup; a gated app with both sign-up locks, a family member added through the 15-minute window, a stranger refused before and after) → use → backup-now → remove/reinstall/restore → a byte-identical restore → delete-with-data → power cut (same versions, gates kept, no alarm) → typos in both codes → a phone first. The box landed on the golden's own controller. One operator press the guide says is not needed (**R-719**). **WHAT IT NO LONGER CLAIMS:** „files protected from day one, once the household writes down the recovery code" does NOT hold on 0.282.0 — every app starts with its off-site copy OFF and nothing asks the household to switch it on (**R-720**), and on a customer who had a box before, the off-site repository is orphaned on night one until the household presses a reset (**R-726**); the night's restore test proved nothing (**R-727**). The household's data stayed on the box (database dump, volumes, whole-guest local tier) and in the whole-guest off-site tier from the evening before. | Rows **R-493 … R-500**, R-534 … R-545, R-719 … R-727 | +| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. **DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven.** `audits/DRILL-prove-fixes-0243-2026-09-16.md`. A fresh nested box was installed from the published ISO, landed on this drill's own golden **by checksum** (`e2d1843c…c10a`, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, **the tunnel from outside (302→200, R-510 CLOSED)**, the data drive, the file manager with its own generated password (`admin`/`admin` refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with **26 s** of app downtime. **Interventions: 0** — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. **The automatic connect e-mail is PROVEN with a real mailbox:** a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at **12:23:00Z**, with `selfbind_link_sent (host delete)` on the timeline (R-509). **WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence:** on a one-drive box with no off-site tier — what every fresh install is — the household's files are in **no backup at all** (the whole-guest tiers exclude `mp8` by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (**R-537**), and a restore reports success while leaving Nextcloud listing five photos that return `Sabre\DAV\Exception\NotFound` — after making the app's own trash, which still held every byte, unreachable (**R-538**). **Also still open from this walk:** the off-site tier cannot be provisioned at all (**R-534**, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (**R-535**), and „app installed” is emitted at accept time (**R-536**). Faults re-measured: controller killed during a deploy → back in **37 s**; three more kills 20 min apart → **61/41/61 s**, none accumulating (**R-531**); two reboots 60 s apart → everything back in **124 s** with the boots NOT counted against the brake; the claim page locks after the **second** wrong code and mails the operator truthfully. **2026-09-16 (evening) — THE BACKUP HALF IS NOW WALKED TOO, on controller 0.244.0 / hub 0.116.0 / golden 0.244.0 / ISO 1.28.0 (built, unpublished).** `audits/evidence-backup-promise-2026-09-16/`. A second fresh box (VM 335) was installed from the BUILT image and walked to the end of the sentence this row could not previously finish: **five photos in → deleted the way a child would → the old route REFUSED and touched nothing → the off-site restore returned them → they OPEN, byte-identical (sha256 5/5, negative control)**. The refusal reads „Ez a mentés nem tartalmazza az alkalmazás fájljait… a fájlok így a helyükön maradnak" and names the route that can help; the app was `running` before and after; the wastebasket was untouched. The off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution counting „5 fájl és 3 adatkötet és az adatbázis". **Delivery is part of it:** the box landed on agent 0.131.0 + controller 0.244.0 (the golden baked and vouched the same day) with no hand upgrade, and `app_deploy_started` (19:15:34) / `app_deployed` (19:16:23) finally mean different things. **The self-bind half needed NO operator press** — the box registered itself and the bind used the mail the hub sent itself after the morning's host delete (R-509, proven twice today: 12:23:00Z and 18:17:46Z, each one second after a host delete). **WHAT IT STILL DOES NOT CLAIM — one operator press and one day-one gap.** The PBS-DR cascade stopped at the refusal R-511 documents („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action") and needed the operator to press it; that press then SUCCEEDED because of the ep0 grant given this morning (R-534 CLOSED, R-511 CLOSED). And off-site ON by default is not off-site WORKING: a fresh box sits at „Kulcsletétre vár" until the household performs the escrow ceremony, which nothing asks them to do, while the tier-1 row already promises that copy (**R-543**, P1). **2026-09-16 (late evening) — that last gap is CLOSED, controller v0.245.0 (R-543).** The pause is the zero-knowledge escrow design and was not touched; what was missing was the ASK. Every authenticated page now carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking the ceremony (the R-241 bar, second instance, hung on the single render choke point), the tier-1 sentence renders by tier-3 STATE („védené … szünetel" while paused, „védi" when running), and the first-hour guide asks for the code right after the dashboard password and before the first app. **Measured on two boxes running 0.245.0:** paused box — bar on four pages, `POST /backup/offbox/run` refused by the fork-4 gate with no snapshot written, app row „védené" and „védi"=0; escrowed box — no bar anywhere, row „védi". `audits/evidence-recovery-code-2026-09-16/`. **WHAT IT STILL DOES NOT CLAIM:** the ask has not been walked by an actual volunteer from the written guide — the sentence is proven, the human following it is not. So the journey now reads: **files protected from day one, once the household writes down the recovery code the box asks them for on every page.** **2026-09-29/30 — RE-WALKED on golden 0.282.0 (controller v0.282.0, agent v0.137.0, hub v0.125.0, published ISO 1.29.0), customer `tester-1`: the WALK is PROVEN-LIVE with ZERO interventions; the DAY-ONE BACKUP sentence is NARROWED.** `audits/DRILL-new-household-2026-09-30.md`. Download → install on the Hungarian keyboard → the mailed link and self-bind page (one typo refused) → the claim **through the public tunnel** → the recovery code → three apps (a random first password that works while `password` is refused; a gated app whose probe opened it ~12 s after the household's setup; a gated app with both sign-up locks, a family member added through the 15-minute window, a stranger refused before and after) → use → backup-now → remove/reinstall/restore → a byte-identical restore → delete-with-data → power cut (same versions, gates kept, no alarm) → typos in both codes → a phone first. The box landed on the golden's own controller. One operator press the guide says is not needed (**R-719**). **WHAT IT NO LONGER CLAIMS:** „files protected from day one, once the household writes down the recovery code" does NOT hold on 0.282.0 — every app starts with its off-site copy OFF and nothing asks the household to switch it on (**R-720**), and on a customer who had a box before, the off-site repository is orphaned on night one until the household presses a reset (**R-726**); the night's restore test proved nothing (**R-727**). **2026-09-30 — two of those three are FIXED:** apps go off-site by themselves when the customer has off-site (decision 50, controller v0.283.0, live on 9202 — R-720 closed), and the restore test takes only the box's own archives (agent v0.138.0, the drill archives removed from ep0 — R-727 closed). **Still open: R-726** (a returning customer's orphaned repository). A NEW customer record — what the first real tester gets — does not meet R-726. The household's data stayed on the box (database dump, volumes, whole-guest local tier) and in the whole-guest off-site tier from the evening before. | Rows **R-493 … R-500**, R-534 … R-545, R-719 … R-727 | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | **A restore-test can never fill the box's disk: it sizes the restore first (uncompressed), keeps off the tested guest's pool when it can, refuses what does not fit, and retries a failed clean-up on a timer** | agent **v0.133.0** (RELEASED, NOT DELIVERED), hub **v0.124.0** | **BUILT + red-proofed; the refusal PROVEN on demo-hp (2026-09-24); the full-test path NOT proven live** | `audits/r672-2026-09-24/C7a…C7b`, `audits/r672-2026-09-24/redproofs/C-*` | The 2026-09-24 incident (R-672) filled demo-hp's pool and turned 9201 read-only. Live: both margins refuse 9201's 21.1 GiB restore (22.1 GiB free, needs 30.3). No full restore-test fits demo-hp under 80 % pool use, so a pass after the change is unproven; the scheduled test is OFF on both demo hosts until the agent is delivered (operator ruling). | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 20641492..0a4bcd4f 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -553,12 +553,19 @@ R-636's louder repeated alarm. which a per-app switch starting OFF had undone. If the apps will not fit the customer's quota, the page says so, names the largest, and the household chooses which stay off-site. Apps already installed are not switched by a release (the 2026-09-29 Part 0 rule); the backups page offers one press to switch them all on. - **The box never deletes off-site history to make room without the household's choice.** *(Outcome: filled in by - the 2026-09-30 fixes session.)* + **The box never deletes off-site history to make room without the household's choice.** **Outcome (2026-09-30, + controller v0.283.0):** measured first — over the quota the box already refused NEW pushes and ran only the ruled + retention, and an app whose files would cross the quota went up settings + database only; no history was ever + deleted to make room (now pinned). A fresh install switches the app ON (`DefaultOffboxOnForNewApp`; an earlier + choice is kept); older apps get one press on both backup pages; the size card names the three largest. Live on + 9202: the press, and a fresh app joining by itself. 51. **ep0: the three whole-guest archives of deleted drill boxes in `tester-1`'s namespace are removed, and only those** — *operator ruling 2026-09-30 (R-727).* With before/after controls on every other namespace. The restore test is fixed to take only the current box's own archives, so an archive of an earlier box in the same namespace - can never be tested again. *(Outcome: filled in by the 2026-09-30 fixes session.)* + can never be tested again. **Outcome (2026-09-30):** the three archives (2026-09-16T17:27:32Z, 2026-09-16T21:59:54Z, + 2026-09-29T19:37:07Z), mapped to their drill boxes by key fingerprint and host record, were forgotten on ep0 one by + one; every other namespace byte-identical before and after; chunks go at ep0's own weekly GC. Agent v0.138.0: an + archive carries its key FINGERPRINT (not a host id), and the restore test skips one written with another key. Same day, operator: the first real tester gets a NEW customer record (`Tester-2`, domain `sajatfelhom.hu`), not `tester-1`; and CC rewrites the volunteer guide as measured (R-722). diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/partA/9202-deploy.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partA/9202-deploy.txt new file mode 100644 index 00000000..ea103ba5 --- /dev/null +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partA/9202-deploy.txt @@ -0,0 +1,3 @@ +gitea.dooplex.hu/admin/felhom-controller:0.283.0 +gitea.dooplex.hu/admin/felhom-controller:0.283.0 Up 25 seconds (healthy) +gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 25 seconds (healthy) diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/partA/9202-live.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partA/9202-live.txt new file mode 100644 index 00000000..408c69d0 --- /dev/null +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partA/9202-live.txt @@ -0,0 +1,31 @@ +## 2026-09-30T08:40:33Z 9202: the household's own NAS target, through the page's form (192.0.2.1 = RFC 5737 documentation address, never reachable; throwaway key) + HTTP/2 302 + location: /backups/remote?flash_error=flash.offbox.target_saved_escrow_agent_down +## 2026-09-30T08:41:15Z before: paperless-ngx=OFF privatebin=OFF + /backups/apps shows the offer: 1 + press „Igen, mindegyikre“ → 302 https://192.168.0.114/backups/remote?flash=flash.offbox.enabled_all + after: paperless-ngx=ON privatebin=ON + offer still shown: 0 +## 08:41:27Z fresh install of glance on 9202, fields ['SUBDOMAIN'] + deploy → {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} + state after 18s: running + off-site switches now: glance=ON paperless-ngx=ON privatebin=ON +2026/09/30 08:41:30 [INFO] [backup] glance: off-site copy switched ON at install (decision 50) +## 2026-09-30T08:55:05Z 9202 teardown + paperless-ngx Start → {"ok":true,"data":{"state":"starting"},"message":"Stack paperless-ngx start requ + glance Stop → {"ok":true,"message":"Stack glance stop completed"} + glance remove (data + backups) → {"ok":true,"data":{"removed":"glance","volumes_removed":["glance_glance_config"],"hdd_paths_removed":[],"hdd_paths_preserved":[],"hdd_note":"Az alkalmazás nem + off-site OFF paperless-ngx → 302 + off-site OFF privatebin → 302 + target switched off via the form → 302 https://192.168.0.114/backups/remote?flash=flash.offbox.target_saved + data= + Stopping 'felhom-controller-bootstrap.service', but its triggering units are still active: + felhom-controller-bootstrap.path + Traceback (most recent call last): + File "", line 2, in + FileNotFoundError: [Errno 2] No such file or directory: '/settings.json' + gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 27 seconds (healthy) + /backups/remote now: 0 × „Még nincs beállítva távoli mentési cél“ + file=/var/lib/docker/volumes/felhom-controller-data/_data/data/settings.json + gitea.dooplex.hu/admin/felhom-controller:0.283.1 Up 30 seconds (healthy) + /backups/remote: „Még nincs beállítva távoli mentési cél“ × 1; paperless: "state":"running" diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/partC/agent-0138-live.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partC/agent-0138-live.txt new file mode 100644 index 00000000..8ede911f --- /dev/null +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partC/agent-0138-live.txt @@ -0,0 +1,55 @@ +== demo-hp +felhom-agent 0.138.0 +Sep 30 10:40:28 demo-hp felhom-agent[2355282]: time=2026-09-30T10:40:28.593+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=11d75b6842043927 op=agent_update +Sep 30 10:40:30 demo-hp felhom-agent[2355282]: time=2026-09-30T10:40:30.718+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled" +Sep 30 10:40:31 demo-hp felhom-agent[2135127]: time=2026-09-30T10:40:31.715+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s +Sep 30 10:40:32 demo-hp felhom-agent[2135127]: time=2026-09-30T10:40:32.658+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s +Sep 30 10:41:32 demo-hp felhom-agent[2135127]: time=2026-09-30T10:41:32.681+02:00 level=WARN msg="selfupdate: update committed" version=0.138.0 wrapper="" +== felhom-pve +felhom-agent 0.138.0 +Sep 30 10:37:12 demo-felhom felhom-agent[2880126]: time=2026-09-30T10:37:12.223+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=ccafbfb170d1b714 op=agent_update +Sep 30 10:37:14 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:37:14.960+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s +Sep 30 10:37:16 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:37:16.219+02:00 level=INFO msg="janitor: starting (restore-test scratch retry + stale-lock sweep)" interval=10m0s +Sep 30 10:38:16 demo-felhom felhom-agent[2081460]: time=2026-09-30T10:38:16.243+02:00 level=WARN msg="selfupdate: update committed" version=0.138.0 wrapper="" +== demo-hp 08:58:03 — selftest restore-test-due (READ-ONLY: the scheduler's own verdict) + * demo-hp status=online fp=07:B5:72:5D:5E:C1… + [ ok ] node status up 953h13m36s, load [0.60 0.67 0.58], mem 5.2GiB/29.3GiB, root 23.0GiB/38.6GiB + [ ok ] list lxc 1 guest(s) + - 9201 "demo-hp" status=running + [ ok ] pool read pool "felhom", 1 member(s) + - 9201 type=lxc + [ ok ] storage 4 store(s) + - local-lvm type=lvmthin content=images,rootdir used=32.4GiB/53.9GiB + - local type=dir content=vztmpl,backup,import,iso used=23.0GiB/38.6GiB + - felhom-pbs type=pbs content=backup used=0.0GiB/0.0GiB + - nvme-scratch type=dir content=images,rootdir used=73.5GiB/937.8GiB +=== selftest OK === +== felhom-pve 08:58:04 — selftest restore-test-due (READ-ONLY: the scheduler's own verdict) + * demo-felhom status=online fp=60:8F:4C:50:C3:8E… + [ ok ] node status up 1225h32m30s, load [0.29 0.30 0.26], mem 2.4GiB/15.4GiB, root 27.4GiB/93.9GiB + [ ok ] list lxc 1 guest(s) + - 9201 "demo-felhom" status=running + [ ok ] pool read pool "felhom", 1 member(s) + - 9201 type=lxc + [ ok ] storage 4 store(s) + - felhom-backup type=dir content=backup used=11.9GiB/915.8GiB + - local type=dir content=import,backup,iso,vztmpl used=27.4GiB/93.9GiB + - local-lvm type=lvmthin content=rootdir,images used=11.9GiB/348.8GiB + - felhom-pbs type=pbs content=backup used=0.0GiB/0.0GiB +=== selftest OK === +== demo-hp 08:58:13 — selftest=restore-test-due (READ-ONLY) +time=2026-09-30T10:58:13.922+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/felhom-golden-0.236.0.tar.zst size_bytes=654115664 reason="not a backup of a guest (the storage reports no v +time=2026-09-30T10:58:13.924+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst size_bytes=656970239 reason="guest 9100 does not exist on this n +tier=felhom-pbs due=false archive="felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z" landed=2026-09-24T20:06:25Z proven="felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z" + reason: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven +tier=local due=false archive="local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst" landed=2026-09-29T02:42:11Z proven="local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst" + reason: newest settled archive (landed 2026-09-29T02:42:11Z) is already proven +cost tier=felhom-pbs one_lookup=322ms +cost tier=local one_lookup=45ms +== felhom-pve 08:58:14 — selftest=restore-test-due (READ-ONLY) +tier=felhom-backup due=true archive="felhom-backup:backup/vzdump-lxc-9201-2026_09_29-07_36_43.tar.zst" landed=2026-09-29T05:36:43Z proven="felhom-backup:backup/vzdump-lxc-9201-2026_09_28-07_36_39.tar.zst" + reason: newest settled archive (landed 2026-09-29T05:36:43Z) has not been proven (last proven archive was a different one) +tier=felhom-pbs due=false archive="felhom-pbs:backup/ct/9201/2026-09-29T04:16:43Z" landed=2026-09-29T04:16:43Z proven="felhom-pbs:backup/ct/9201/2026-09-29T04:16:43Z" + reason: newest settled archive (landed 2026-09-29T04:16:43Z) is already proven +cost tier=felhom-backup one_lookup=20ms +cost tier=felhom-pbs one_lookup=306ms diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/partD/live-resend.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partD/live-resend.txt new file mode 100644 index 00000000..04f9b9ec --- /dev/null +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partD/live-resend.txt @@ -0,0 +1,5 @@ +## 2026-09-30T08:38:33Z live, hub 0.126.0 — a made-up link (never real; same page as an expired one by design) + GET /bind/ → 200; page: Felhom — Doboz összekötése | Felhom | doboz | összekötése | Ez a hivatkozás érvénytelen vagy lejárt. | A hivatkozás 7 napig érvényes. Ha lejárt, kérj újat az ügyfélszolgálattól, vagy az összekötést az üzemeltető is elvégezheti. | Ha ez a te linked volt, és a dobozod még nincs összekötve, új linket küldünk arra az e-mail címre, amelyet a Felhomnál megadtál. | Új linket kérek | Felhom.eu + the button's form:
+ POST …/resend → 200; page: Felhom — Doboz összekötése | Felhom | doboz | összekötése | Kész. | Ha ez egy valódi hivatkozás volt, és a dobozod még nincs összekötve, néhány percen belül új e-mailt kapsz a regisztrált címedre. Ha nem jön, szólj az ügyfélszolgálatnak. | Felhom.eu + hub log (self-bind lines, last minute): 0 diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/partE/9202-stop-during-dump.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partE/9202-stop-during-dump.txt new file mode 100644 index 00000000..aec08f69 --- /dev/null +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/partE/9202-stop-during-dump.txt @@ -0,0 +1,29 @@ +## 2026-09-30T08:42:08Z 9202: „Mentés most“, and the household presses Stop on paperless-ngx while the dump holds it down + paperless-ngx before: "state":"running" + „Mentés most“ → {"ok":true,"message":"Mentés elindítva"} + 08:42:13.593 saw: 2026/09/30 08:42:13 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume + 08:42:13.600 Stop pressed → {"ok":true,"message":"Stack paperless-ngx stop completed"} + backup run: "success":true + paperless-ngx after: "state":"starting" + 2026/09/30 08:42:13 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume dump + 2026/09/30 08:42:13 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx + 2026/09/30 08:42:13 router.go:601: [INFO] [api] stop requested for stack: paperless-ngx + 2026/09/30 08:42:13 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "stopped" (was "running") + 2026/09/30 08:42:13 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx + 2026/09/30 08:42:21 backup.go:971: [INFO] [backup] Restarting paperless-ngx after volume dump + 2026/09/30 08:42:21 manager.go:1223: [INFO] [stacks] Starting stack: paperless-ngx +## 2026-09-30T08:54:15Z RETEST on v0.283.1 — „Mentés most“, Stop on paperless-ngx while the dump holds it + paperless-ngx before: "state":"running" + „Mentés most“ → {"ok":true,"message":"Mentés elindítva"} + 08:54:19.823 saw: 2026/09/30 08:54:19 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume + 08:54:19.828 Stop pressed → {"ok":true,"message":"Stack paperless-ngx stop completed"} + backup run: "success":true + paperless-ngx 20 s after the run: "state":"stopped" + 2026/09/30 08:54:14 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "running" (was "stopped") + 2026/09/30 08:54:14 manager.go:1223: [INFO] [stacks] Starting stack: paperless-ngx + 2026/09/30 08:54:19 backup.go:955: [INFO] [backup] Stopping paperless-ngx for safe volume dump + 2026/09/30 08:54:19 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx + 2026/09/30 08:54:19 router.go:601: [INFO] [api] stop requested for stack: paperless-ngx + 2026/09/30 08:54:19 desiredstate.go:81: [INFO] [stacks] desired state for paperless-ngx recorded as "stopped" (was "running") + 2026/09/30 08:54:19 manager.go:1312: [INFO] [stacks] Stopping stack: paperless-ngx + 2026/09/30 08:54:28 backup.go:967: [INFO] [backup] paperless-ngx NOT restarted after the volume dump: the household stopped it meanwhile diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/controller-floor.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/controller-floor.txt index c3ba4edb..30cfee72 100644 --- a/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/controller-floor.txt +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/controller-floor.txt @@ -2,3 +2,6 @@ HTTP/1.1 303 See Other Location: /configuration?flash=floor_set 2026/09/30 10:33:19 [INFO] Global controller-version floor set to "0.283.0" (declared MinAgent "0.131.0") +## 2026-09-30T08:53:29Z floor → 0.283.1 (MinAgent 0.131.0 declared) +Location: /configuration?flash=floor_set +2026/09/30 10:53:29 [INFO] Global controller-version floor set to "0.283.1" (declared MinAgent "0.131.0") diff --git a/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/hub-deploy.txt b/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/hub-deploy.txt new file mode 100644 index 00000000..19b78177 --- /dev/null +++ b/documentation/audits/evidence-fixes-first-tester-2026-09-30/release/hub-deploy.txt @@ -0,0 +1,15 @@ +## 2026-09-30T08:37:05Z hub 0.126.0 via ArgoCD (HEAD 7e5009808856179dea9cafcde472cb462943ae6f) +revision seen: 7e5009808856179dea9cafcde472cb462943ae6f +deployment "hub" successfully rolled out +image: gitea.dooplex.hu/admin/felhom-hub:0.125.0 +argocd: sync=OutOfSync health=Healthy rev=7e5009808856179dea9cafcde472cb462943ae6f +2026/09/30 10:34:43 [INFO] enqueued signed-op job ccafbfb170d1b714 for host demo-felhom-8363b5 (775 bytes) +2026/09/30 10:35:14 [INFO] PBS-DR box refreshed: 18.7% full (18.3 GB of 97.9 GB) +2026/09/30 10:37:11 [INFO] host-report from demo-felhom-8363b5 (1 guests, 4 storage targets, 2 backups, 2 restore-tests, 2 pbs-snapshots, 14458 bytes) +2026/09/30 10:37:11 [INFO] DR-recipe host-half stored for customer demo-felhom (host demo-felhom-8363b5, v1) +2026/09/30 10:37:12 [INFO] host demo-felhom-8363b5 cleared signed-op job ccafbfb170d1b714 (executed or rejected) +08:38:00 image=gitea.dooplex.hu/admin/felhom-hub:0.126.0 argocd=Synced/Healthy +running pod image: gitea.dooplex.hu/admin/felhom-hub:0.126.0 +2026/09/30 10:37:19 [INFO] felhom-hub 0.126.0 starting +2026/09/30 10:37:19 [INFO] Default controller-version floor: 0.120.0 +2026/09/30 10:37:20 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000 diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b88f3a45..93f690cd 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -830,15 +830,17 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-716** | **[P3-LOW] Apps installed before controller 0.281.0 keep their open sign-up — decision 47 closes it only on apps whose gate the box opened.** READ 2026-09-29 on the demo boxes after catalog `6faf432` synced: demo-hp's adventurelog and opengist, demo-felhom's opengist carry `signup_block:` in their synced template and no gate record, so no block (`audits/gate-rollout-2026-09-29/0/P0-3-demo-boxes-after-push.txt`). This is Part 0's rule working as designed (a catalog change never touches an installed app). **Needs an operator word** before anything changes on an installed app: a one-time "close sign-up now" press on the app page for an installed app, or leave them. Only the demo boxes have such installs today. **Operator ruled A (decision 49); built in controller v0.282.0 and pressed** on demo-hp's adventurelog and opengist and demo-felhom's opengist: before, sign-up served; after, refused; adventurelog's own switch on (`audits/signup-lock-2026-09-29/`C). | **CLOSED — 2026-09-29** | | **R-717** | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | | **R-718** | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. **ALSO MEASURED 2026-09-29 (new-household drill):** the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. | **OPEN — P3; owner: CC** | -| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. | **READY — rank P2-MEDIUM; owner: operator (which shape) / CC** | -| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. | **READY — rank P2-MEDIUM; owner: operator (decide) / CC** | -| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). | **READY — rank P2-MEDIUM; owner: CC** | -| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. | **READY — rank P2-MEDIUM; owner: CC (text) / operator (approve)** | -| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. | **READY — rank P3-LOW; owner: CC** | -| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). | **READY — rank P3-LOW; owner: CC** | -| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). | **READY — rank P3-LOW; owner: CC** | +| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | +| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. **BUILT 2026-09-30 (controller v0.283.0, decision 50):** a fresh install on a box with off-site switches the app ON; older apps get one press on both backup pages; over the quota the page names the largest apps. Measured first: over the quota nothing but the ruled retention runs — no history is deleted (now pinned). Live on 9202: the press put both older apps ON; a fresh `glance` joined by itself. The size card is unit + parity proven (a household NAS target has no quota). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partA/`. | **CLOSED 2026-09-30 — controller v0.283.0** | +| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). **FIXED 2026-09-30 (controller v0.283.0 + v0.283.1):** the quiesce resume, the nightly volume dump, the update leg and the startup crash recovery skip an app the household stopped. **v0.283.0 was wrong in production** — its adapter did not answer the question and the dump restarted the app 8 s after the Stop (measured live on 9202); v0.283.1 wires it and pins the PRODUCTION types. Live on 0.283.1: `paperless-ngx NOT restarted after the volume dump: the household stopped it meanwhile`. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partE/`. | **CLOSED 2026-09-30 — controller v0.283.1, proven live** | +| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. **REWRITTEN 2026-09-30** (operator: CC rewrites as measured): both guides — BookStack's generated password on the app page, the ~15 min recovery-code delay, apps off-site by default + the one-press offer, a 4 GB stick, the returning customer's mail and the expired-link button, the pre-selected disk, focus on „Previous", and the auto-reboot (a sixth stale line: with the defaults the „reboot now?" prompt never appears). Operator part: set the e-mail language, check the domain's DNS for an earlier box's records, the tunnel route with No TLS Verify. Every changed line dated. | **CLOSED 2026-09-30** | +| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. **FIXED 2026-09-30 (hub v0.126.0):** no `node_recovered` when the customer's host was enrolled after the outage began; `backup_tier_skipped` in a box's first hour recorded, not mailed. Real recoveries and old boxes' skips still mail (controls). Unit + red-proof (RP40, RP41); the live proof is Tester-2's first hour. | **FIXED — hub v0.126.0; live proof at the first real install; owner: CC** | +| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** | +| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). **FIXED 2026-09-30:** the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). **NARROWED — remaining:** the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). | **NARROWED — three small copy items; owner: CC** | | **R-726** | **[P2-MEDIUM] A new box for a customer who had one before makes NO off-site copy on night one: the old repository is found orphaned, and the fix is a button nobody pointed the household to.** MEASURED 2026-09-30 00:15 UTC on the new-household drill box (`tester-1`, whose previous box was deleted 2026-09-17): `[offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset` → `offbox_repo_orphaned` (warning) to the household's timeline and an operator mail; `offsite-integrity` then checked nothing. The household's page is honest („A távoli tároló másik kulccsal készült mentéseket tartalmaz … Új távoli mentés indítása…", old history set aside, never deleted), but the evening before, the recovery-code ceremony and the off-site page raised nothing, and the guide does not mention it. The hub re-issued the off-site credentials on re-enroll by itself; it could have known the repository would orphan. Customer data was never at risk (the old repository is untouched); the household simply has no off-site copy until someone presses the button. **Fix direction:** offer the reset at the recovery-code ceremony when the repository already holds another key's snapshots, or the hub's re-enroll re-issue sets the old history aside the same way (it is the same move-aside), and the guide says so. | **READY — rank P2-MEDIUM; owner: CC (design first) / operator (which)** | -| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. | **READY — rank P2-MEDIUM; owner: CC** | +| **R-727** | **[P2-MEDIUM] The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the household sees a bare ✗ labelled with the wrong tier.** MEASURED 2026-09-30 01:52 UTC on the drill box: the agent's restore test chose `felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z` („newest settled archive … has not been proven") — written by a drill box of 2026-09-16, still in the customer's ep0 namespace next to tonight's own `2026-09-29T19:37:07Z` — and failed `wrong key - unable to verify signature since manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…`. One operator mail (`restore_test_failed`); the hub then re-logged the stored failure at every 15-minute report for 5 hours. The household's backups page read „✗ Visszaállítás ellenőrizve 2026-09-30 03:52 — Helyi tároló (local)" — the local tier had not failed; the pbs tier had, and the page says neither which nor why. Root: host delete leaves the old box's archives in the namespace (R-526's shape). **Fix direction:** the restore test skips (and reports as foreign) archives whose key fingerprint is not this host's; the page names the tier that failed. **FIXED 2026-09-30 (agent v0.138.0 + decision 51):** measured — a PBS archive carries its key FINGERPRINT (PVE content `encrypted`), not a host id; the storage carries its own (`encryption-key`). The restore test skips an archive written with another key (logged by name). The ✗ card names the tier (controller v0.283.0) — **and the 2026-09-30 claim "the page blames the local tier" was my misreading**: the „Helyi tároló (local)" after the ✗ was the next section's heading. ep0: the three drill archives in `tester-1` removed, other namespaces byte-identical. Delivered by signed jobs to both demo boxes (340 s); their due-check reads normally on 0.138.0. Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partC/`. | **CLOSED 2026-09-30 — agent v0.138.0** | +| **R-728** | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** MEASURED 2026-09-30 on `Tester-2`: the hub logged `Customer config created: Tester-2` twice in the same second and two self-bind mints (hashes `c40df008…`, `6a1cbef4…`); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). **Fix direction:** make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. | **READY — rank P3-LOW; owner: CC (hub)** | +| **R-729** | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** |