From dfd854474e92ae5dfb791b8ab45e76867b514092 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 14:27:04 +0200 Subject: [PATCH] drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The automatic connect e-mail is proven with a real mailbox: the host record was deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second later, with selfbind_link_sent (host delete) on the timeline. The requirement was two minutes. The hub refuses to delete an ONLINE host with no override, so the record had to fall stale first — that wait is part of the proof. Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box. The verdict is still no, for a new reason: a one-drive box with no off-site tier keeps none of the household's own files in any backup, the page says otherwise, and the restore that should save them makes it worse (R-537, R-538). Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing on the off-site server written or removed. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT.md | 119 +++++++++++------- STATUS.md | 47 +++++++ .../architecture/00-capability-map.md | 2 +- .../DRILL-prove-fixes-0243-2026-09-16.md | 6 +- .../phase3-alarm-truth-table.txt | 11 ++ .../phase3-hostdelete.txt | 60 +++++++++ 6 files changed, 196 insertions(+), 49 deletions(-) diff --git a/REPORT.md b/REPORT.md index 9c9a24ab..94ff4dda 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,52 +1,81 @@ -# REPORT — before the volunteer: the big night's P1 fixes, the two rulings, the publish (2026-09-15) +# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16) -Releases: **agent v0.131.0**, **controller v0.243.0** (MinAgent 0.131.0), **hub v0.114.0**, catalog templates (no version), -**installer ISO 1.27.1 published**. Baselines re-verified at start: controller `406755f`, agent `4586f0f`, felhom.eu -`a028a9a`, catalog `882a43e` — all equal to the task table. Architecture read and cited: 03 §4, 05 (new §13/§14), 07 §6, -08 §6.2, 02 (settings after install). +## Claims in the prompt that turned out wrong — first, as asked -## Claims in the prompt that turned out wrong (first) -1. **"The floor carries the agent release to every box"** — false. The hub HOLDS a floor whose declared MinAgent is - above the box's agent; an agent updates only by an operator-signed `agent_update` job (R-530). Delivered to demo-hp - with the operator's keys; demo-felhom and Peti's box stay on 0.130.0. -2. **"`--restart always` does not restart a killed container"** — CORRECT, now measured (Docker 29.8.0, both policies - exited 60 s after `docker kill`). -3. **"Appliance registration knows no customer"** — correct as read (`api/appliance.go`); not a trigger. -4. **"a `node_down` mail would have gone at ~60 min"** — not measured; unchanged as an inference. -5. **"R-110 tag move" for the ISO** — R-110 governs the host installer's `installer-v*` tag, which the ISO does not - change; no tag was moved. **"index page on iso.felhom.eu"** — the bucket serves no index (R-504); the index is the - website page `felhom.eu/letoltes`, published with the ISO. -6. **"`08-alarm-ladder.md` §5 holds the cooldown"** — it is in §6.1/§6.2; the ruling was written there. -7. **"Park note in `RUNBOOK-manual-build.md` §4"** — §4 is the golden image; the park note went to §3.1 (controller). +1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused + exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran, + and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing + written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1). +2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot. + After a successful claim the same URL becomes the password-reset surface with a "request a new code" + button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says + "Túl sok próbálkozás — próbáld újra 15 perc múlva." +3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than + assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's + own bake, and the guest runs controller 0.243.0 with no self-update. +4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box — + there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the + customer's namespace, empty before and after. +5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false + and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was + true. -## Parts -- **A (R-523)** agent supervisor — live on 9201: idle 59 s, parked held + unpark 25 s, swap deferred; crash-loop guard - tripped for real and the hub mailed `controller_crashloop`; after the 30-minute pause the agent restarted it - (09:36:24Z, dashboard 200 at 09:36:30Z) and the hub minted `controller_restarted_by_agent`; mid-deploy timing not measured. -- **B (R-513)** generated FileBrowser password — live on 9202 (generated) and 9201 (operator-set left alone); public 401. -- **C (R-517/R-518)** per-tier page live on 9201; absent-tier skip unit-proven; „0 B" found live and fixed (unreleased). -- **D.1 (R-509)** self-bind auto-send — shipped and red-proofed; **live mail NOT checked** (a throwaway customer's - delete runs the RESET cascade on ep0 — fenced). -- **D.2** `node_*` bypass the quiet hour (operator ruling) — red-proofed; recorded in 08 §6.2 and CONTEXT.md. -- **D.3 (R-511)** re-issue adopts — shipped, red-proofed; not provable on tester-1 (no host). **STOP held**: listing - posted, operator said yes, one snapshot forgotten on ep0 (1 928 820 672 B logical), token kept, nothing else touched. -- **D.4 (R-510)** three GETs → 530 ×3 (no box, no tunnel); stays open; day-0 A.1 names the setting. -- **E (R-512/R-514/R-515)** catalog — live on 9202: stranger 400, 20/20 documents, peak 772 MB; OOM line not proven (R-528). -- **F** ISO 1.27.1 — uploaded; round trip sha256 `25637007…c053`, 1 705 322 496 B; `.sha256` 200; manifest 200; - 1.26.1 kept for rollback; `felhom.eu/letoltes` 200. G11 PASS, G12 not measurable with the token. -- **G** docs: 03, 05 §13–14, 07 §6.4, 08 §6.2, 02, CONTEXT, RUNBOOK §3.1, day0 A.1, VOLUNTEER prerequisites. +## What I exercised -## Extra acts, stated -- The felhom.eu push was blocked by the due-checks gate (R-433 due today): the Gmail-read mailbox holds no Hetzner reply; - re-dated to 2026-09-22 with the reason in the row. No `--no-verify` anywhere. -- The operator's three signing keys arrived mode 664; set to 600 (R-533). +A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first +console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps, +ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the +morning-after checks and a three-layer teardown. + +## What broke — product, and mine + +**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token +lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim), +**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup +page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves +the app listing files it cannot open, after making the app's own wastebasket unreachable). + +**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was +restoring caps after the memory test — the second time this exact mistake has happened, repaired +per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the +controller-injected environment and crash-looped it; repaired through the controller's own API. I measured +my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice +before reading it off the dashboard. ## Rows -Register rows **232 → 232**: closed 9 (R-493, R-495, R-496, R-512, R-513, R-514, R-515, R-517, R-523); opened 9 -(R-525 … R-533); narrowed R-509, R-510, R-511, R-518; R-433 re-dated. -## Teardown, three layers -Machine: 9202 throwaways removed; 9201 controller running again (09:36:24Z); the killed homebox deploy created no container but left its `app.yaml` -(09:04:55Z, keys only read) — moved aside to `/root/homebox-app.yaml.p1fixes-residue`, noted under R-531. -Host: demo-hp park marker removed; agent 0.131.0 stays (the release). ep0: one tester-1 snapshot removed on yes. -Hub: demo-hp floor override 0.243.0 kept (it delivers the release); no customer created or deleted. +Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533, +plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's +first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538. + +## The automatic connect e-mail (R-509) — PASSED + +The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host, +with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the +box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host +delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second +later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The +requirement was two minutes. + +## Verdict + +**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the +supervisor restarts a dead controller, the file manager has its own password, the backup page tells the +truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the +connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are +in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes +things worse. + +## Teardown + +Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state, +token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers +9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host +record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**. +Nothing on the off-site server was written, removed or pruned — its listing is empty before and after. + +## Checks + +`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at +**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by +`head_sha`. diff --git a/STATUS.md b/STATUS.md index b003050d..00842aaf 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,5 +1,51 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-16 (drill on a fresh box) — the fixes hold; the backup promise does not.** + +> **Ready for a volunteer: NO — one reason, and it is new.** On a brand-new box with one drive, the +> household's own files are in **no backup at all**, and the backup page says they are. I deleted five +> photos the way a child would, restored from the box's own backup, and the folder came back listing all +> five photos — none of which opens. The bytes had never been copied. The app's own wastebasket still held +> them, and the restore made that unreachable too. + +**What I proved on a fresh box.** The installer downloads and installs; the box lands on the golden this +drill baked (checked by checksum, not by trust); the connect e-mail and the bind page work; the dashboard +opens through the tunnel from outside; the file manager has its own password and „admin/admin" is refused; +four apps installed and were used; the backup page tells the truth per tier; „Mentés most" stopped the apps +for 26 seconds, inside what the button promises. + +**The five faults.** A controller killed during an install: back in 37 seconds. Two reboots a minute apart: +everything back in 124 seconds, and the box did not count the reboots against its own safety brake. Wrong +passwords five times: the app lets you keep trying, the box's own setup code locks for 15 minutes after two +and e-mails you — correctly. Memory pressure: the box still cannot see it (second box, same result). +The deleted photo folder: see above. + +**The automatic connect e-mail: it works.** I deleted the box's record on the hub and the „connect your +Felhom box" e-mail reached the customer **one second later**, naming the reason. That was the last thing +waiting to be proven with a real mailbox. + +**Decisions I took.** None under the unattended rule. + +**Needs you.** +1. **Say whether the backup page may keep promising what it does not hold.** Today, a new box with one drive + backs up its apps' settings and databases — not the household's own files. The page says otherwise, and a + restore then reports success while the files are gone. If you do nothing: the first volunteer can lose + their photos and be told everything is fine. I can fix the wording and the refusal in the controller; the + real protection needs a second drive or the off-site copy switched on. +2. **Grant the off-site server one permission.** The re-issue fails on a missing grant, so a rebuilt or new + box gets no off-site copy at all. If you do nothing: the third backup level stays unavailable for every + new box, and the fix already written stays dead. +3. **Rule on the restart brake.** The box stops retrying after three restarts in fifteen minutes. I measured + that a controller dying every twenty minutes is restarted forever, and the only trace is a note that + e-mails nobody. Options: leave it (the box heals itself and the timeline records it); add a second, + slower counter that raises a warning; or make the fifth restart in a day a warning. My pick: the second + counter — it keeps the healing and ends the silence. If you do nothing: a slowly failing box stays + invisible until someone reads the timeline. + +--- + +## Previous note + **Updated 2026-09-15 (P1 fixes) — the big night's blockers, fixed and shipped.** > **Ready for a volunteer: almost.** The file manager has a real password, the backup page tells the truth, the @@ -739,3 +785,4 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store). + diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 51b05f6f..3a6801d2 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -88,7 +88,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis **NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | | **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) The venue is **torn down** (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and **30 are R-244**, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and **16.64 GiB** returned. `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. Evidence: `tests/walk5-r201-2026-08-07/journal.md` **SCOPE NOTE 2026-09-14 — not re-walked.** The first-hour drill on 0.242.0 (`audits/DRILL-fresh-install-0242-2026-09-14.md`) walked install → first use → restore of a deleted page on a fresh box, NOT a rebuild with off-site recovery (DR tier and off-site were off). This row is therefore neither re-proven nor contradicted on 0.242.0; its PROVEN-LIVE stands on 0.206.0 only. The first hour has its own row below. | -| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. | Rows **R-493 … R-500** | +| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. **DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven.** `audits/DRILL-prove-fixes-0243-2026-09-16.md`. A fresh nested box was installed from the published ISO, landed on this drill's own golden **by checksum** (`e2d1843c…c10a`, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, **the tunnel from outside (302→200, R-510 CLOSED)**, the data drive, the file manager with its own generated password (`admin`/`admin` refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with **26 s** of app downtime. **Interventions: 0** — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. **The automatic connect e-mail is PROVEN with a real mailbox:** a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at **12:23:00Z**, with `selfbind_link_sent (host delete)` on the timeline (R-509). **WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence:** on a one-drive box with no off-site tier — what every fresh install is — the household's files are in **no backup at all** (the whole-guest tiers exclude `mp8` by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (**R-537**), and a restore reports success while leaving Nextcloud listing five photos that return `Sabre\DAV\Exception\NotFound` — after making the app's own trash, which still held every byte, unreachable (**R-538**). **Also still open from this walk:** the off-site tier cannot be provisioned at all (**R-534**, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (**R-535**), and „app installed” is emitted at accept time (**R-536**). Faults re-measured: controller killed during a deploy → back in **37 s**; three more kills 20 min apart → **61/41/61 s**, none accumulating (**R-531**); two reboots 60 s apart → everything back in **124 s** with the boots NOT counted against the brake; the claim page locks after the **second** wrong code and mails the operator truthfully. | Rows **R-493 … R-500**, R-534 … R-538 | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | diff --git a/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md b/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md index 97677f66..1bc09ee9 100644 --- a/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md +++ b/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md @@ -1,8 +1,8 @@ # DRILL — prove the P1 fixes on a fresh box (2026-09-16) -**Interventions: _pending_** (O1/O2 pre-declared, counted apart). -**Ready for a volunteer: _pending_.** -**The automatic connect e-mail: _pending_.** +**Interventions: 0** (O1 the operator's self-bind press and O2 the PBS re-issue press were pre-declared and are counted apart). +**Ready for a volunteer: NO — and the reason is new, not one of the old ones.** Every P1 fix this drill set out to prove did hold on a fresh box. But a one-drive box with no off-site tier — the state every fresh install starts in — keeps **none of the household's own files in any backup**, while the backup page says it does, and a restore then reports success and leaves the app listing photos it cannot open (R-537, R-538). +**The automatic connect e-mail: PASSED.** The host record was deleted at 12:22:59Z and the mail „Kösd össze a Felhom dobozodat” reached `tester1@felhom.eu` at **12:23:00Z — one second later**, with `selfbind_link_sent … (host delete)` on the customer timeline. The requirement was two minutes. > Baselines at start (re-verified against live Gitea): controller `383a30b3c07b` v0.243.0 (`Unreleased`: the > „0 B" tile fix), agent `e98b857684f4` v0.131.0, felhom.eu `351296114c4d` hub v0.114.0, catalog `94bc5febaca2`. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt index aea4fba3..074981b0 100644 --- a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-alarm-truth-table.txt @@ -61,3 +61,14 @@ ## 12:58 CEST ("Operator email suppressed ... cooldown") and mailed again at 13:18 CEST - one hour after ## the 12:08 mail. So the one-hour operator cooldown both SUPPRESSES and RELEASES correctly, proven from ## the hub log and from the inbox independently. +## +## THE STALENESS ALARM, fired by the teardown itself (and it is TRUE - the box really is gone): +## the box's last report: 13:56:14 CEST (11:56:14Z); the machine destroyed ~13:57:30 CEST. +## 14:22:00 CEST host_stale "tester-1-652049 ok -> stale" -> OPERATOR EMAIL SENT, same second. +## That is 25m46s after the last report, i.e. the 30-minute staleness threshold measured from the +## LAST REPORT, not from the moment of death - exactly as the dead-man's-switch is designed. +## The hub's delete-impact endpoint flipped to {"status":"stale","deletable":true} in the same minute +## (12:21:27Z deletable=False -> 12:22:28Z deletable=True), which is what released the host delete. +## NOTE on R-529 (the widened host_* cooldown bypass): this fired ONCE, so the 5-minute dedupe was still +## not exercised on this box - a second host_* within the hour would be needed, and the box was deleted +## instead. The bypass remains proven only by its red-proofed test and the 2026-09-15 node_* run. diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-hostdelete.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-hostdelete.txt index dede4f9c..fa6dfb85 100644 --- a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-hostdelete.txt +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-hostdelete.txt @@ -4,3 +4,63 @@ form fields: NONE ## waiting for the host record to fall stale (delete is refused while ONLINE, by design) 2026-09-16T11:59:21Z status=ok deletable=False + 2026-09-16T12:00:21Z status=ok deletable=False + 2026-09-16T12:01:21Z status=ok deletable=False + 2026-09-16T12:02:21Z status=ok deletable=False + 2026-09-16T12:03:22Z status=ok deletable=False + 2026-09-16T12:04:22Z status=ok deletable=False + 2026-09-16T12:05:22Z status=ok deletable=False + 2026-09-16T12:06:23Z status=ok deletable=False + 2026-09-16T12:07:23Z status=ok deletable=False + 2026-09-16T12:08:23Z status=ok deletable=False + 2026-09-16T12:09:24Z status=ok deletable=False + 2026-09-16T12:10:24Z status=ok deletable=False + 2026-09-16T12:11:24Z status=ok deletable=False + 2026-09-16T12:12:25Z status=ok deletable=False + 2026-09-16T12:13:25Z status=ok deletable=False + 2026-09-16T12:14:25Z status=ok deletable=False + 2026-09-16T12:15:25Z status=ok deletable=False + 2026-09-16T12:16:26Z status=ok deletable=False + 2026-09-16T12:17:26Z status=ok deletable=False + 2026-09-16T12:18:26Z status=ok deletable=False + 2026-09-16T12:19:27Z status=ok deletable=False + 2026-09-16T12:20:27Z status=ok deletable=False + 2026-09-16T12:21:27Z status=ok deletable=False + 2026-09-16T12:22:28Z status=stale deletable=True + DELETABLE at 2026-09-16T12:22:28Z +## 2026-09-16T12:22:59Z DELETING the host record (customer tester-1 is KEPT; RESET never used) + pre-state: {"deletable":true,"escrow_present":false,"guests":1,"log_bundles":0,"pbs_secret_present":false,"recovery_present":true,"reports":10,"status":"stale","wg_peer_bound":true} + delete POST at 2026-09-16T12:22:59Z + http=303 + + host record after: 404 (404 = gone) + customer record after: 200 (200 = KEPT) + hub log right after the delete: + 2026/09/16 14:22:00 [INFO] Host staleness: tester-1-652049 ok → stale (host_stale) + 2026/09/16 14:22:00 [INFO] Operator email sent for tester-1/host_stale + 2026/09/16 14:22:59 [INFO] host deleted: tester-1-652049 (escrow deleted: false) + 2026/09/16 14:22:59 [INFO] self-bind link emailed to the registered address of tester-1 + 2026/09/16 14:22:59 [INFO] self-bind link (hash 029595d7…, valid 7 days) emailed to the registered address of tester-1 + 2026/09/16 14:22:59 [INFO] self-bind link auto-minted for tester-1 on host delete (the console banner's promised email now exists) + customer timeline, newest events after the host delete: + Sep 16 12:22 | info | selfbind_link_sent | Self-bind link e-mailed (host delete) | hub + Sep 16 12:22 | warning | host_stale | Host tester-1-652049: no report for 30m | hub + Sep 16 11:56 | info | controller_started | Controller elindult (0.243.0) | controller + Sep 16 11:39 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller + Sep 16 11:37 | info | controller_restarted_by_agent | Host tester-1-652049 guest 9201: the agent restarted the controller (controller container exited on 2 consecutive sweeps) — restart #2 since the agent started | hub + Sep 16 11:34 | info | controller_started | Controller elindult (0.243.0) | controller + host list now: 0 occurrences of the deleted host id (0 = gone) +## RESULT - the automatic connect e-mail after a host delete (R-509) - PASSED +## host delete POST 2026-09-16T12:22:59Z -> 303, host record 404, customer record 200 (KEPT) +## hub log, same second: "host deleted: tester-1-652049 (escrow deleted: false)" +## "self-bind link emailed to the registered address of tester-1" +## "self-bind link (hash 029595d7..., valid 7 days) emailed ..." +## "self-bind link auto-minted for tester-1 on host delete (the console banner's +## promised email now exists)" +## MAILBOX, read independently: a NEW message at 2026-09-16T12:23:00Z - ONE SECOND after the delete - +## to tester1@felhom.eu, subject "[Felhom] Kosd ossze a Felhom dobozodat", body "Elkeszult a Felhom +## dobozod, es keszen all az osszekotesre ... https://hub.felhom.eu/bind/..." +## (The 09:59:56Z message in the same thread is the O1 operator-pressed one; this is a second, new one.) +## TIMELINE EVENT: "selfbind_link_sent | Self-bind link e-mailed (host delete) | hub" - the occasion is +## named, as required. +## Well inside the two-minute requirement: 1 second.