From 328c3fc23ca94e2d520c2b00a74a9d4d0be048af Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 15 Sep 2026 00:28:12 +0200 Subject: [PATCH] BIGNIGHT morning note (STATUS), topic report, capability-map annotations; teardown layer 1 done, layer 3 pending --- REPORT-bignight-2026-09-14.md | 66 +++++++++++++++++++ STATUS.md | 50 ++++++++++++++ .../architecture/00-capability-map.md | 4 +- .../evidence-bignight-2026-09-14/journal.md | 15 +++++ 4 files changed, 133 insertions(+), 2 deletions(-) create mode 100644 REPORT-bignight-2026-09-14.md diff --git a/REPORT-bignight-2026-09-14.md b/REPORT-bignight-2026-09-14.md new file mode 100644 index 00000000..428baeec --- /dev/null +++ b/REPORT-bignight-2026-09-14.md @@ -0,0 +1,66 @@ +# REPORT — BIGNIGHT: a household's first month in one night (2026-09-14/15) + +**Unattended drill run from `drills/BIGNIGHT-2026-09-14.md` under `.claude/rules/unprompted-work.md`. No product code +changed.** Findings: `documentation/audits/BIGNIGHT-household-month-2026-09-14.md`. Every observable in order: +`documentation/audits/evidence-bignight-2026-09-14/journal.md`. The alarm truth table: +`…/evidence-bignight-2026-09-14/alarm-truth-table.md`. A parallel session may own root `REPORT.md`; this is a topic sibling. + +## 0. Where the brief and the record disagreed — named first + +1. **„Off-site (Tier 3) is ON for this customer — its own namespace on ep0."** On the record the ep0 namespace is the + **DR tier (PBS)**; restic Tier 3 was **off**, and ticking it provisions a Hetzner Storage Box (money — fenced). Not + ticked. The DR tier then could not provision on the new box (R-511). This box had **no off-site tier of any kind**; + Phase 4's off-site integrity check and Phase 6's off-site restore onto 9202 were therefore not walked. +2. **„Expect the hub to issue a fresh claim, or to require its reset flow."** Neither: the hub treated the box as a + **re-enrolment** and mailed the *reinstall* setup code („újratelepült … A korábbi jelszavad már nem érvényes"). +3. **„The operator fixed and tested the tunnel."** The route now reaches the box, and still returns 502: it lacks „No + TLS Verify" (R-510, with demo-hp's working route as the control). +4. **Faults F10–F12 were not run** — the brief's own stop rule was met at F9 (R-523). + +## 1. Baselines + +controller `406755fa` v0.242.0 · agent `4586f0f7` v0.130.0 · felhom.eu `a4d68441` hub v0.113.0 · catalog `6d6eec30`. +ISO 1.27.1 sha `25637007…` found in the build output, not rebuilt. Venue: VM 333 on demo-hp, 4 cores, 16 GB, 200 G + +100 G qcow2 on `nvme-scratch` (`/mnt/hdd_1` root). + +## 2. What changed in the repos + +| repo | change | +|---|---| +| felhom.eu | 16 register rows **R-509 … R-524**; amendments to R-516, R-517, R-519, R-521, R-523; audit page; evidence directory; capability-map annotations (self-bind row, first-hour row); STATUS morning note; this report. Documents only. | +| app-catalog-felhom.eu | drill bump `d5d91e0` privatebin 2.0.5 → 2.0.6 and its revert `a161ccb` in the same phase; two CHANGELOG entries. Net template change: none (`catalog_since` stays 2026-09-14 by the gate). | +| felhom-controller, felhom-agent | nothing | + +## 3. Results in one table + +| phase | result | +|---|---| +| 2 first hour | install ✓ · Hungarian first screen ✓ · self-bind via real mail ✓ · mailed setup code ✓ · version current ✓ · **tunnel gate FAIL** · data disk: no screen tells a household · **interventions 2** | +| 3 twelve apps | all deployed and seeded through their front doors · memory guard never refused · 12/12 „Naprakész" · **interventions 0** · Paperless lost 20 uploads to OOM · FileBrowser `admin/admin` on every box | +| 4 routines | Tier 1 ✓ · Tier 2 ✓ · whole-system local ✓ but apps down 8 min and a false PBS claim · guarded Update on a real bump ✓ 11 s, data intact · catalog reverted | +| 5 faults | F1 F2 F3 power cuts heal ≈ 4 min · F4 F5 drive pull/return honest, heals 91 s · F6 drive lost in backup: skipped apps reported success, alarm mails silenced · F7 disk 95 % holds, English banner, operator not told · F8 internet gone: LAN works, tunnel self-heals 9 s · **F9 controller killed: dead 33 min, nobody told — STOP** | +| 6 morning after | apps healthy · 1 false label (downgrade offered as update) · local BookStack restore ✓ 24 s | +| 7 teardown | layer 1 done (VM, disks, ISO, harness files); layer 2 host delete + ep0 peer: see the audit's teardown section; customer `tester-1` **kept**; its ep0 data **kept, stated** | + +## 4. Rows (register 221 → 237) + +P1: **R-509** no auto bind mail for an existing customer · **R-510** tunnel route lacks No TLS Verify · **R-513** FileBrowser +admin/admin, demo-hp public · **R-517** backup page claims a failed PBS tier current and present · **R-523** killed +controller never restarts. P2: R-511 DR tier stuck after a rebuild · R-512 Vaultwarden open signup, read-only control · +R-514 Paperless OOM silent · R-518 whole-system backup stops apps 8 min · R-519 torn backup dated by its newest part · +R-524 downgrade offered as update. P3: R-515 Paperless card's wrong login · R-516 English strings · R-520 interrupted +update untestable same-version · R-521 alarm mail noise and cooldown silence · R-522 tunnel tile „Fut" while offline. + +## 5. Harness slips, recorded + +API key printed once into tool output (hub customer page read) · two quoted-string inserts into the register failed +and were redone · first claim POST sent two CSRF tokens · several poll loops read a stale status and stopped early or +ran long (guest backup, Tier 2, restore) · F8's first two attempts cut nothing (nft reserved word; the hub resolves to +the LAN) and the measured cut lasted 17½ min, not 20 · time waiters broke at local midnight (`date -d HH:MMZ`) · +AdventureLog account took five attempts. None changed a finding; each is in the journal where it happened. + +## 6. Security note + +A FileBrowser login (`admin`/`admin`) was tested against demo-hp 9201 and 9202 **over loopback only**; demo-hp's public +login page was checked with a GET and no login. Nothing was changed on either guest. The operator was told by the +morning note; the push notification was not sent because the terminal was active. diff --git a/STATUS.md b/STATUS.md index 79e90856..ae00aff9 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,5 +1,55 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-15 (morning note) — the big night: a household's first month on one fresh box.** + +> **Ready for a volunteer: not yet — three things stop them.** The dashboard link still does not open through +> the tunnel. A box installed for an existing customer gets no bind e-mail. And every box's file manager opens +> with the login „admin" / „admin" — on the HP that login page is reachable from the internet. + +**Decisions I took.** None under the unattended-decision rule. Two things I did not do, each with one reason: I did +not switch on the paid off-site storage for „Tester 1" (it costs money), so this box had no off-site copy and the +off-site restore could not be walked. I stopped injecting faults after the ninth, because the brief's stop rule was met. + +**Interventions (first hour + moving in): 2.** (1) No bind e-mail came for 10 minutes; I pressed the operator's +„send link" button. (2) The tunnel answered 502; I used the home-network address for the rest of the night. + +**Alarm truth table, in five lines.** +1. Every alarm that fired was true. None was false. +2. Missed: the Paperless crash, the broken tunnel, the dead controller, the full disk, and the second drive loss. +3. Three of those were missed because a 30–60 minute mail cooldown silenced a new incident. +4. One drive loss sent five operator mails; the household got no mail for anything all night. +5. The customer's pages were honest about drives, and wrong about backups twice. + +**What broke, and whether the box healed itself.** Power cuts (three, one during a backup, one during an update): healed +in about four minutes, same versions, data intact. Drive pulled and returned: healed in 91 seconds, data intact. Internet +gone: the tunnel came back by itself in 9 seconds. Disk 95 % full: the box kept working. **Controller killed: it did not +heal — no dashboard for 33 minutes, nobody told, only a reboot brought it back. That was the stop.** Also: Paperless +silently lost 20 uploads to memory; the backup page claimed a remote backup that does not exist; the whole-system backup +stopped every app for 8 minutes while promising „a few seconds"; after I reverted a test update in the catalog, the box +offered the downgrade as an update. Nothing lost data that a restore could not bring back. I fixed nothing tonight. + +**Rows.** Opened 16, closed 0. Register: 221 rows before, 237 after. + +**The one sentence for recruiting.** Felhom survived power cuts, a pulled drive and a lost internet by itself tonight, +but do not invite anyone until the tunnel works, the file-manager password is changed on every box, and a dead +controller restarts itself. + +**Needs you.** +1. **Change the file manager's admin password on the HP now** (and on the N100). **If you do nothing:** anyone on + the internet who tries „admin" / „admin" at the HP's files address can read its data drive. +2. **Tick „No TLS Verify" on the „Tester 1" tunnel route in Cloudflare.** **If you do nothing:** no volunteer can open + their dashboard from outside their home. +3. **Decide whether a box installed for an existing customer should get the bind e-mail automatically**, or you press + the button each time. **If you do nothing:** each new install waits for you. +4. **Say whether installer 1.27.1 should be published.** My recommendation: **yes, publish it** — nothing tonight was + the installer's fault; the install, first screen and bind all worked. The three blockers above are on the tunnel, + the controller and the file manager, and they block inviting people, not the installer. **If you do nothing:** the + old installer with the English admin line stays online. +5. **Decide what to do with „Tester 1"'s old off-site data on ep0** (still there, kept tonight) and its stuck DR tier. + **If you do nothing:** the data stays, and every new box for this customer gets a failed whole-system backup alarm. + +--- + **Updated 2026-09-14 (evening) — the doorstep: installer fixed, walked again, NOT published.** > **Ready for a volunteer: not yet — one thing stops them.** On the test customer you chose, the diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 490b9987..51b05f6f 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -50,7 +50,7 @@ | Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row | **Second board, 2026-07-21 (demo-hp, HP t740 / Ryzen V1756B / AMI M42):** the whole chain ran on virgin hardware in one pass — armed install → self-registration as an unclaimed appliance → operator bind → day-0 → running guest 9201 + agent 0.92.1 as `demo-hp-bb76ea`. **The shim loader booted with Secure Boot ENABLED**, which retires the assumption that Felhom installs need SB off — that was an N100-firmware workaround. The exact-serial disk filter took the system SSD and left the box's 1TB NVMe untouched/unenrolled on hardware it had never seen. Two failures filed rather than smoothed over: **R-59** (no DHCP → the installer baked a static fallback instead of aborting) and **R-61** (baked root password unknowable → no console access). | Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (`.rootpw.txt`) | scripts v1.24.0 (`scripts/iso/felhom-bootstrap.sh` `network_gate`/`sweep_nics`, `build-felhom-iso.sh` rootpw emission) | **PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install)** | `audits/SPIKE-firstboot-nic-sweep-2026-07-22.md` — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and **self-registered at the hub unaided in under a minute**; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash | R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (`state.json` gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched | | Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) | controller v0.122, hub v0.50 | **PROVEN-LIVE** (drill VM) | `DRILL-day0-vm-2026-07-12` §10/F-4 (gate ON via real edge; claimed, code consumed) | Never executed by a non-Viktor human → R-3. **Deliverability (R-4), gmail half DONE 2026-07-18:** the rehearsal's claim email was the first sent under the tightened DMARC `p=quarantine` and **landed in the gmail Inbox, not spam** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`). **freemail.hu remains Viktor's open half.** (Dropped mis-cited `CAMPAIGN-4` F-C — that is the escrow-claim 502, not password claim) | -| Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor `/bind/` (console pairing code + retrieval passphrase) → hub stages the bind, no operator | hub v0.66.0 + ISO scripts v1.20.0 | **PROVEN-LIVE** (real customer-zero bind on metal, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md`:** operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); **the customer bound their own box at 16:29:55 with `attempts=0`, `locked=0`** — `appliance_bound` carries source **`customer_selfbind`**, and the credential was delivered **26 s later** with no operator action. Hub-side lifecycle in `hub-state.txt` (`selfbind_tokens` mint→email→consume). Prior unit evidence: hub v0.66.0 (`web/selfbind.go`, `store/selfbind.go`; Scenarios A–F + F1/F2; 4 red-proofs verified red — THE TRAP `/bind/` exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) | R-27 **slice 1**. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. **Live first-run DONE 2026-07-18** (rehearsal; the console banner rendered on the real ISO). **R-27b** (controller second-box dismissable prompt) deferred; **multi-box-per-link** = repeated operator sends | +| Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor `/bind/` (console pairing code + retrieval passphrase) → hub stages the bind, no operator | hub v0.66.0 + ISO scripts v1.20.0 | **PROVEN-LIVE** (real customer-zero bind on metal, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md`:** operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); **the customer bound their own box at 16:29:55 with `attempts=0`, `locked=0`** — `appliance_bound` carries source **`customer_selfbind`**, and the credential was delivered **26 s later** with no operator action. Hub-side lifecycle in `hub-state.txt` (`selfbind_tokens` mint→email→consume). Prior unit evidence: hub v0.66.0 (`web/selfbind.go`, `store/selfbind.go`; Scenarios A–F + F1/F2; 4 red-proofs verified red — THE TRAP `/bind/` exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) **Re-walked 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, hub v0.113.0):** the mailed link, read in the real mailbox, bound a fresh box first try (`self-bind SUCCESS … by customer self-service`, credentials delivered 17 s later) — `audits/BIGNIGHT-household-month-2026-09-14.md`. **Gap measured the same night:** the link is auto-sent only at customer creation and RESET, so a box installed for an existing customer waits for the operator's press (R-509). | R-27 **slice 1**. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. **Live first-run DONE 2026-07-18** (rehearsal; the console banner rendered on the real ISO). **R-27b** (controller second-box dismissable prompt) deferred; **multi-box-per-link** = repeated operator sends | | Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge | controller v0.127, agent v0.88/0.89 | **PROVEN-LIVE** (drill VM, endpoint-exact) | agent v0.88.0 REPORT (ceremony ~4s, one-shot claim 200→410, R absent from every payload); `SPIKE-controller-escrow-2026-07-13` | **Customer-facing browser wizard FIRST LIVE FIRING 2026-07-18** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S6): customer zero drove the wizard on the reborn box — ceremony started 16:56:29, recovery code claimed one-shot 16:56:39 (absent from logs by design), hub-verified and `EscrowState` auto-confirmed 16:56:41, **offsite runs enabled 12 s after the ceremony began**; the v0.138.0 „megerősítésre vár, legfeljebb 15 perc" awaiting card rendered and flipped on the ACK (operator screenshots: Viktor's set). Honest caveat: at a 12-second confirm the awaiting window is so short that catching *both* states on screen is luck, not procedure. Prior: endpoints driven on the drill VM. **agent v0.89.0:** `/escrow/preflight` `pbs_storage_id` row now live-reloads (reads current agent.json) — a pbsdr convergence that seeds the id flips it green with NO service restart. **hub v0.60.0 (data-first retention) — ⚠ the claim as written was FALSE for the offsite tier for two months; **FIXED in hub v0.93.0 (2026-08-04), and the row below states what ships TODAY:** a re-escrow with a DIFFERENT sealed passphrase no longer destroys the old blob — the hub RETAINS it (`host_escrow_superseded`) — and **since v0.93.0 the retained row carries `identity_blob` as well as the K-escrow blob**, so a previous passphrase does now stay recoverable with the recovery code that sealed it — **OPERATOR-ONLY, and the customer-facing half of that sentence is FALSE (R-304, drill 2026-08-12).** The retention was exercised end-to-end for the first time that day and **it works**: the retained row carried the identity blob byte-identically (`sha256 a10032341c8584ed…`), the old recovery code unsealed it, and three planted files — including a Hungarian accented filename verified as raw bytes — restored **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`). **What does not exist is the door.** `ListSupersededEscrow` (`hub/internal/store/store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers**; the product's recovery path (`POST /escrow/recover-offsite-password` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked with the code that demonstrably opens the retained row, the product answers **"the recovery code did not open the sealed bundle"**. So: recoverable by an operator with SQLite, `age` and a shell; **not recoverable by the customer, who is told their correct code is wrong.** — **AMENDED 2026-08-12 evening (R-311, shipped hub v0.103.0 + agent v0.129.0 + controller v0.214.0):** the customer is **no longer told their code is wrong.** The agent now tries the retained packages when the current one refuses (`GET /hosts//escrow/retained` → `ErrCodeOpensRetained` → HTTP 422), and the screen says the code is CORRECT, names the supersession date, says the earlier package is kept and the current backups are unaffected, and routes to support. **What is still true and must not be read away:** there is no in-product ROUTE to the set-aside data (R-312 — every restore entry point resolves its repository from settings and its password from one file; adding an alternative is new surface, not wiring), so the recovery itself remains operator-performed. The status of this capability is therefore **"the customer is told the truth and handed to a human"**, not "the customer can recover their old history". **What was wrong until v0.93.0, recorded because it is the ninth entry in `CLAUDE.md`'s comment-vs-code table and the first that was also customer-facing copy:** `host_escrow_superseded` had no `identity_blob` column and `demoteCurrentEscrowTx` did not copy one, so what survived a supersession was the **PBS datastore key only** — never the restic repository password, which lives in `identity_blob`. The destroying act was the **escrow ceremony** a rebuilt box asks its customer to run. Measured live 2026-08-04, before the fix: both current rows held `blob`=383 B **and** `identity_blob`=572 B; both retained rows held `blob`=383 B only. → **R-198** (SHIPPED), evidence `audits/RECON-offsite-dr-chain-2026-08-04.md` §7. **THREE SCOPE LIMITS THIS ROW MUST NOT BE READ PAST.** (1) **Nothing was backfilled and nothing could be** — rows superseded before v0.93.0 were written without the blob and their source rows are already overwritten; both demo boxes' pre-2026-08-04 repository passwords are gone permanently. (2) **A retained key is not a restore — but as of 2026-08-04 evening it IS a recovered key.** See the row below. (3) The customer-facing orphan card still promises recoverability unconditionally (R-202, gate hit, card untouched). Guided-recovery flow = R-26. Red-proof `TestSaveHostEscrow_RetainsSuperseded`. **hub v0.60.1 — custody survives the host lifecycle:** host deletion (with the escrow ack) DEMOTES the current blob to retained custody (moved into `host_escrow_superseded`, never destroyed; existing superseded rows spared); the customer Danger-zone Delete is the one true purge point (cascades both escrow tables incl. already-deleted hosts). No operator path through host lifecycle can lose a blob. Red-proofs `TestDeleteHost_DemotesEscrowNeverDestroys` + `TestDeleteCustomer_PurgesEscrowCustody`. **agent v0.93.0 (2026-07-21) — recovery codes can no longer contain a hyphenated word.** The EFF large list holds exactly four entries containing the hyphen the words are joined with (`drop-down`, `felt-tip`, `t-shirt`, `yo-yo`); drawing one produced a code that reads as 11 words instead of 10 — ambiguous to transcribe in exactly the situation R exists for. They are now excluded **from GENERATION only**: the draw space goes 7776 → 7772 and a 10-word code 129.248 → 129.241 bits, still well clear of the 128-bit floor. **Every code already issued remains valid** — R is verified as a whole passphrase by the PBS scrypt KDF and is never re-split, so no customer needs to re-run a ceremony. This also retired the long-standing ~1/5 `TestGenerateRecoveryCode_EntropyAndFormat` flake, which was this defect and not a flaky test | | **The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code** | hub v0.94.0, agent v0.125.0, controller v0.195.0 | **PROVEN-LIVE (2026-08-04)** | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: *escrow blob SERVED … 572 opaque bytes, self_scope=true*), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was **byte-identical** to the one on disk — `c60c8bc737a6…`, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 links 6–8 | **WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier.** It covers the KEY, not the DATA. **R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in `audits/RECON-offsite-dr-chain-2026-08-04.md` §3; links 1–8 are walked, 9–11 are not.** **Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half:** the R-201 drill was prepared on demo-hp and **halted before the wipe** — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. **No file has still ever been restored from an off-site backup after a wipe** (`audits/DRILL-r201-offsite-recovery-2026-08-04.md`). The install half of the chain (controller v0.196.0 `--recover-offsite-install`, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been **installed** (the diagnostic compares and refuses to write, by design), no existing repository has ever been **reopened** under one, and **no file has ever been restored** from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | @@ -88,7 +88,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis **NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). | | **The customer's own UNAIDED recovery journey, end to end** | controller v0.206.0, hub v0.98.0, agent v0.127.0 | **PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim** | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with **no command line inside the guest**. **Data: PASS.** All three sentinels byte-identical (`beb9175d…`, `7c8cb0ad…`, `e012e76f…`), including a 12 MB binary and `C11-őrszem-ékezetes-árvíztűrő.txt` — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in **16 s** out of the pre-wipe snapshot `f3d9cd67`, non-destructively. **Journey: FAIL.** Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: `tests/campaign11-evidence-2026-08-05/journal.md` | **WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change.** Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but **fixes are not a journey.** This row goes green only when the walk is repeated end to end and completes with no operator intervention. **Three findings are deliberately still open and each is a live blocker for some flow:** **R-214** (the console never stops showing a stale pairing code), **R-220** (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), **R-221** (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded `escrow.pbs_storage_id` is rewritten out of `agent.json`). **And R-216's fix does not by itself make a NEW box work:** with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is **HELD, not served** — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. **⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS **FAIL**.** **What they add:** eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. **The backup promise strengthened** — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (`snapshot_count` 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. **R-217's and R-215's fixes were proven live under exactly their faults**, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. **What they do NOT add: any progress on the JOURNEY.** These faults are **not a re-walk** — no customer route was walked end to end — and four new findings say the journey got no better where it matters: **R-224** (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own `err` distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers `source=version` so it cannot see reachability — **Phase 1's headline defect relocated from the version channel to the transport**), **R-226** (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), **R-225** (an unread store renders `0 pillanatkép · 0 GB` above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), **R-228** (the set-aside history is recorded in `orphaned_renamed_to` and shown by nobody — `OrphanedRenamedTo` has zero references in any template or handler). **§4.1 is now MEASURED rather than deduced** — the floor IS served, from the box's own rendered `GetFloor()` and a cold-started controller's settle-gate line. **§4.2's positive half is still NOT measured** and needs a rebuild. Evidence: `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md`, `tests/campaign11-evidence-2026-08-05/journal-phase24.md` **2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL.** controller **v0.202.0** + agent **v0.126.0** close R-224, R-226, R-225, R-227, R-228. **What changed:** the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with **unknown defaulting to neutral**. Proven live on the venue with the same wrong code and only the hub's reachability changed: `400 → 502 → 400`. **What it does NOT change: this row.** Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (**R-220** in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). **What still needs proving:** a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because `/recovery` correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: `felhom-controller/REPORT.md` **⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no.** A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. **THE DATA: PASS** — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename **whose NAME BYTES are also identical**, restored in **23 s** out of the pre-destruction snapshot through the customer's own flow. **THE JOURNEY: FAIL — TWO dead ends against Phase 1's four.** (1) **R-218's CONSUME half**: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) **R-220**: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. **The unaided RTO is therefore STILL UNDEFINED**; the attended figure was 30 m 13 s and must not be quoted as the customer number. **What passed and is new:** the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. **⚠ AND IT DOES NOT CLAIM DELIVERY:** the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: `tests/rewalk-r201-2026-08-06/journal.md`   **⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL.** controller **v0.203.0** + agent **v0.127.0** close **R-218's consume half** and **R-220**, and golden **0.203.0** / agent **0.127.0** / `min_agent` **0.127.0** are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. **Both were then proven live on a genuinely rebuilt box** (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with **no downgrade and no hand upgrade** (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), `/disks/candidates` returned **both drives** after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both **re-attached through the customer endpoint**; the off-site credential was re-staged by `offsiteheal` at 13:24:57Z and **collected by the box on a tick, unaided**. **Why the row is still FAIL:** the walk did not finish. It stopped at the restore surface — **R-237**, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in **controller v0.204.0** (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but **no sentinel was restored in that walk**, so the data half is unproven in EITHER direction for this venue. **R-238 was RECLASSIFIED, not fixed as filed:** `mode=full` without `confirm=1` is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry `?full_prep=` forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. **R-236 was WITHDRAWN: not a defect.** **This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel.** Evidence: `tests/part4-rewalk-2026-08-06/journal.md`, `tests/teardown-2026-08-06.md`, `felhom-controller/REPORT.md` **⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before.** A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `f5c53b03`, including a 12 MB binary and an accented Hungarian filename **whose name BYTES are identical too**. **THE JOURNEY: FAIL, and this time the customer has NO route at all** — `/` lands on the launcher with no recovery pointer, `/recovery` **302s away**, and the remote-backup page offers to **CREATE a new recovery code**, which would orphan the very history the customer's code protects. **There is no field anywhere to enter the code they hold**, and the operator's documented remedy refuses too. Cause (**R-241**): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of `OffsiteRecoveryOffer()`'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by `escrow_state: pending`. **R-218's shape one level up.** **What DID pass, and is new:** the whole credential chain ran end to end with zero human action on an unclaimed box — declare, `offsiteheal` re-stages after two reports, the box's 5-minute retry collects it, tier applied — the **first live sighting** of that success line, settling R-218's consume half and R-236's withdrawal. Also new: **R-239** — a fresh install lands on controller **0.203.0** while 0.205.0 is released, so R-234 and R-237 are written but **not delivered**, proven from the customer's side (T2, T3). **This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel.** Evidence: `tests/finalwalk-r201-2026-08-07/journal.md` **⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written.** The question was whether R-241 is a screen-predicate defect or a minting defect. **It is a MINTING defect.** The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. **Three measurements, from the venue rather than from the earlier report:** (1) `WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist — while `OffsiteRecoveryOffer()` and `needsOffsiteCredential()`, both in the same file, consult `GetHubEscrowIdentityPresent()`; the same fact is available on three paths and used on two. (2) That flag was the **precondition of the chain that reached the minting**: the retry job logs only when the declaration is live, and the venue logged `credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** — thirty minutes and six ticks before the mint at 03:18:06Z. (3) **The box computed the answer and discarded it**: at **03:28:03Z**, thirty-five minutes before the customer looked, `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged `the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…)`. It is recomputed every report cycle and never persisted. **And the hub explicitly disclaims doing this** — `offsiteheal`'s package doc: *"credential automatic, key customer-present — the ruling this session implements and must not quietly widen."* **Shape (b) is structurally unreachable on this box** (the escrow gate at `offbox.go:743` sits upstream of `ensureOffboxRepo`, the only producer of `RepoState="orphaned"`; positive control: the scheduler was alive, 241 `agent-channel-health` ticks, and `offbox-backup` is a `sched.Daily` leg whose slot fell before the destruction). **Two by-products:** **R-243** — a box in this state silently stops backing up off-site and **no alarm fires** (`isStale` requires `escrowed`, `offsite_delivery_stuck` skips the `applied` shape, `backup_failed` needs a run that never happens); and the trap in the obvious fix — `ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. **Q7:** the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the `identity_blob` — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md` **⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL.** Three changes, following the spike's ruling rather than the obvious reading: the box **no longer mints a repository key while the hub holds a sealed package** (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as `offsite.state=awaiting_recovery_key` and shown inert to every existing hub reader); the **hub-vs-local key comparison that was computed every cycle and discarded is now persisted** and drives the offer as **shape (c)**; and **abandoning starts a 14-day countdown** whose terminal step removes the set-aside store and the sealed package **together**, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears **once per ENTRY into the offered state**, three dismissal levers with three scopes, and **none removes the entry point**. Q7's trap is closed — „Helyreállítási kód létrehozása" is **unavailable** while a recovery is outstanding. **WHY THE ROW STAYS FAIL: these are fixes, not a walk.** Nothing here walked a customer end to end, and this row goes green only when one completes with **no operator intervention AND a byte-identical sentinel**. Two of Phase 1's blockers are also still open (**R-214**, **R-202**), and **R-240** is untouched. Evidence: `felhom-controller/CHANGELOG.md` v0.206.0, `hub/CHANGELOG.md` v0.98.0. **DELIVERY, separately: R-239 is CLOSED 2026-08-07** — golden **0.205.0** baked, published, round-trip verified (`./etc/felhom-controller-image` read OUT of the downloaded archive) and **VOUCHED**, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: `agent_version` and `min_agent` both stayed 0.127.0. Evidence: `tests/golden-0.205.0-2026-08-07/`. **The row still says FAIL** — delivery is not a journey, and R-241 is diagnosed, not fixed. **✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS.** A brand-new appliance (VM **325**, customer `walk5`) was installed from the published ISO on `demo-hp`, claimed, given three sentinels, escrowed, backed up off-site, then **destroyed on purpose** — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. **THE DATA: PASS** — all three sentinels byte-identical out of snapshot `5b0f20f7` (`11eb7fb2…`, `6b504d1e…`, `0baaf402…`), including a 12 MB binary and `WALK5-őrszem-ékezetes-árvíztűrő.txt` **whose name BYTES are identical too**, read back with `os.listdir` on a bytes path so no decode round trip could launder a `U+FFFD`. **THE JOURNEY: PASS — ZERO guest command lines were needed to progress**, against three on the previous walk. **RTO 71.7 s** from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen **appeared without being sought** (`/` → `/launcher` → `/recovery`), answered all three of its questions, and its sealed-at timestamp matched `host_escrow.created_at` exactly. **WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time:** at **14:58:52Z**, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and **refused to mint a repository password** over the sealed package — `NOT minting a repository password: the hub holds a sealed recovery package for this box`. Sampled every 20 s from T0: **no key at any moment**, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. **AND DELIVERY IS PART OF THE PASS:** the fresh install landed on **controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade**, both times, so the box under test is the box a customer receives. **WHAT THIS ROW STILL DOES NOT CLAIM.** (1) **Putting files back in place is built and worked here** — `reconstitute` placed 6 files — but **only after two obstacles the customer must guess past**: **R-252** (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive *registration*, and nothing on the recovery path says to re-attach) and **R-253** (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared **from the dashboard with no shell** — which is why the journey passes — but neither is signposted, so *unaided* here means *possible without a shell*, not *obvious*. (2) **Shape (c) did NOT fire positively.** With the mint guard holding there is no local key, so the offer comes from **shape (a)**; shape (c) was measured in Phase A in its **negative** half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) **R-214, R-202 and R-240 are untouched.** (4) The venue is **torn down** (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and **30 are R-244**, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and **16.64 GiB** returned. `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. Evidence: `tests/walk5-r201-2026-08-07/journal.md` **SCOPE NOTE 2026-09-14 — not re-walked.** The first-hour drill on 0.242.0 (`audits/DRILL-fresh-install-0242-2026-09-14.md`) walked install → first use → restore of a deleted page on a fresh box, NOT a rebuild with off-site recovery (DR tier and off-site were off). This row is therefore neither re-proven nor contradicted on 0.242.0; its PROVEN-LIVE stands on 0.206.0 only. The first hour has its own row below. | -| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. | Rows **R-493 … R-500** | +| **A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code** | controller **v0.242.0**, agent **v0.130.0**, hub **v0.112.0**, golden **0.242.0**, installer ISO **1.26.1** | **PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app** | `audits/DRILL-fresh-install-0242-2026-09-14.md`, evidence `audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md`. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back **byte-identical with its attachment** 32 s after restore; a hard power cut returned every app on the **same version** with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. **ONE intervention a volunteer could not make:** the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — **R-494**), so the dashboard was reached by LAN address. **And no instruction exists to begin with (R-493).** **WHAT IT DOES NOT CLAIM:** the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. **2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL.** `audits/DOORSTEP-walk-1270-2026-09-14.md`. Held again on a new customer (`tester-1`, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, `pvebanner` masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. **BIGNIGHT 2026-09-14/15 (ISO 1.27.1):** the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: `audits/BIGNIGHT-household-month-2026-09-14.md`. | Rows **R-493 … R-500** | | **Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts** | controller v0.197.0 | **PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it** | On demo-hp, an app on the **system-data fallback** bound `/mnt/sys_drive/userdata/media/books` while the capture set looked in `/mnt/sys_drive/felhom-data/userdata/media/books`: the declared-mandatory directory was in **no** snapshot and the run reported `ok` (R-203). After v0.197.0 the capture log reads `1 mandatory path(s)` and the file is **listed inside the snapshot** — `restic ls -l latest --tag calibre-web` → `-rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt`. Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md` §2 and the v0.197.0 CHANGELOG | **WHAT IT DOES NOT CLAIM.** It covers CAPTURE, not RESTORE: **no file has ever been restored from an off-site backup after a wipe** (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports **`incomplete`** rather than `ok`, so this row's guarantee is one the status can express. **Widened 2026-08-06 (controller v0.205.0, R-234):** the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported `ok`, so the smaller gap moved the verdict and the bigger one did not. **This row still claims CAPTURE, not that a newly-selected app is protected by the next run** — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || **Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE** — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | **PROVEN-LIVE (2026-07-26)** | `audits/SPIKE-r82-phase0-2026-07-26.md`; per-repo CHANGELOGs/REPORTs. **Restore round-trip on demo-hp:** `--selftest=restore-test` against `felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z` → `pass:true`, `verified:"boot+running"`, **`mount_parity:"ok"`** (`mp0=/var/lib/docker 50G`, `mp1=/mnt/sys_drive 20G`, mp8/mp9 throwaway stand-ins for the archived binds), `source_tier:"pbs"`, 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | **This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL.** That tier was PROVEN-LIVE as `applied` since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held **zero, ever** — "applied and empty", the R-39 shape one level quieter. **What earns PROVEN-LIVE here:** (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it **restores into a bootable, mount-complete guest** — `mount_parity` is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) **the multi-tier quiesce ran through the real UI endpoint** (`POST /api/guest-backup/trigger`, authed+CSRF) and produced **exactly ONE stop/start pair with BOTH backups inside it** — `quiescing 1 stack(s)` 17:01:39 → local done 17:02:56 *"next tier may start (app still quiesced)"* → felhom-pbs snapshotted 17:03:06 → `unquiescing` 17:03:06. **App downtime 1m27s for both tiers**, and the app came back healthy. **Known gaps, recorded not hidden:** the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → **R-82** | | **Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard** | agent v0.104.0 → **v0.121.0**, hub v0.77.0 → **v0.91.0** | **PROVEN-LIVE (2026-08-03)** | per-repo CHANGELOGs; `backlog/SPEC-r85-phase4-5-2026-07-26.md`. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | **The row above is earned by a MANUAL `--selftest=restore-test`; this one is about the SCHEDULED path, and the distinction is the whole point.** Before R-85 the scheduler could only ever see `cfg.Backup.BackupTarget()`, so the offsite tier was never a candidate — and a failed restore-test was a `[WARN]` line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — `restore_test_failed` (broken now) and `restore_test_stale` (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. **Why this was NOT PROVEN-LIVE until now:** rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → **R-85**. **R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon**, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. **UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer.** The scheduler's own log: `15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven"` → `proxmox-backup-client restore --crypt-mode=encrypt` under the agent's own token → `15:25:08 gate decision class=guest_destroy guest=990000 allowed=true` → `15:25:14 scratch guest torn down` → **`15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1`**. A **14.5 GB encrypted offsite archive pulled from ep0 over the WAN**, restored, booted, verified and destroyed in 635 s, unattended. **The three things a timer could not show**, all verified after it: the state names THAT archive (`{"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}`); a second evaluation reports `due=false … is already proven` and runs nothing; and **an agent restart runs nothing**, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from `pct list`, **zero** `990000` volumes in `lvs`, and the hub-side `restore_tests[]` entry deliberately RETAINED (it IS the proof the staleness check reads). **R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path:** the proof above reached the hub only because no restart intervened — `restore_tests[]` came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (`0 restore-tests` on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. **THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly.** Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was **unprovable**, and on demo-hp equally: the agent's token had no ACL on `/storage/felhom-backup`, the storage both boxes configure as `local_backup_target`, so the content API answered `{"data":[]}` through the token while root listed three archives. The scheduler skipped it as *"no settled archive yet"* — which is exactly what a brand-new tier reports — so nothing ever said so (**R-185**). **CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0.** The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now **asks whether it may read each tier it depends on** instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on **while the box was still blind**. **THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed.** Four SCHEDULED runs, nothing triggered by hand: **demo-felhom** host tier `…2026_08_02-04_42_14.tar.zst` **passed in 83.8 s** at 00:55, offsite `…2026-07-28T04:49:43Z` **passed in 540.4 s** at 06:55; **demo-hp** host tier `…2026_08_02-04_49_29.tar.zst` **passed in 109.3 s** at 02:05, offsite `…2026-07-28T19:19:45Z` **passed in 300.1 s** at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; `pct list` and `lvs` show zero `990000` afterwards on both, and both boxes' `local-lvm` returned to their pre-run figures (1.95 % and 30.83 %). **Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule:** never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. **The proofs reached the hub**, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds **two** entries, one per tier, and the `local` one can only have come from the persisted state because the in-memory store held only that morning's offsite run. **SCOPE, stated because one box proving something does not make it a fleet property:** this covers `demo-felhom` and `demo-hp`. The tester's box is untested and untouched. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md` | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **R-25b CLOSED (hub v0.69.0, 2026-07-21):** the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | diff --git a/documentation/audits/evidence-bignight-2026-09-14/journal.md b/documentation/audits/evidence-bignight-2026-09-14/journal.md index ea81041b..62128613 100644 --- a/documentation/audits/evidence-bignight-2026-09-14/journal.md +++ b/documentation/audits/evidence-bignight-2026-09-14/journal.md @@ -678,3 +678,18 @@ attachments sha equal.** PASS. (Harness: the poll loop missed the status's `last taken separately at 22:24:57Z.) Phase 6 evidence pulled off the box 22:26Z (`box-logs-phase6/`), before any teardown. + +## Phase 7 — teardown, three layers + +**Layer 1 — the machine** (`teardown-before.txt`, `teardown-layer1-machine.txt`), 22:25:58Z: `qm stop 333`, `qm destroy +333 --purge 1 --destroy-unreferenced-disks 1` (system and data qcow2 on `nvme-scratch` removed; `/mnt/hdd_1/images/` +now holds only `9202`); ISO 1.27.1 removed from `local:iso`; `/root/bn` harness files removed; nft: only `inet +felhom_oob` (the drill's bridge table was already deleted at the end of F8). `pvesm status` before → after: +`nvme-scratch` 59 436 396 → **10 140 556 KiB** (before the drill 10 134 820); `local` 24 728 328 → **23 071 188 KiB** +(before the drill 23 041 736); `local-lvm` 44.17 % unchanged all night. `qm list` empty; **9201 and 9202 running**. +`pct fstrim`: not applied — the drill used only qcow2 files on a dir storage, now deleted; 9201 and 9202 were not +touched and are not trimmed. + +**ep0 read-only before the host delete** (`teardown-ep0-before-host-delete.txt`, 22:26:02Z): WireGuard peer +`10.77.0.5/32` present; namespaces `demo-felhom demo-hp tester-1`; `tester-1/ct/9201` with 1 snapshot directory (the +doorstep walk's data). Nothing was written to ep0 tonight.