Files
felhom.eu/documentation/architecture/00-capability-map.md
T

195 KiB
Raw Blame History

00 — Felhom Capability Map

How to read this document. Two kinds of statement appear, and where this document marks them it marks them like this — the same wording as 07-backup-architecture.md:11-17, carried here on 2026-08-22 (R-376) so a reader meets one convention and not eight:

  • [DESIGN] — a decision taken. Not derived from code; the code may not implement it yet.
  • [FACT] — an observed property, carrying a file:line, a live command output or a citation.

Statements in this document are NOT yet all marked. Marking them wholesale is a large judgement exercise and a wrong mark is worse than none, so only what a session touches is marked (R-376). An unmarked statement therefore means "not yet classified", never "observed". That ambiguity is exactly what cost this project three sessions in August 2026: the hot/bulk placement decision sat unmarked beside a marked [FACT], and was read as an observation and reported as a defect.

What this is: the single cross-component truth table of what the Felhom platform can do today, at what confidence level, with verifiable evidence. Rows are scenarios (user- or operator-visible outcomes), not modules — a scenario spans agent + controller + hub + catalog, and this is the only doc that shows that view.

What this is NOT: a roadmap. Planned work lives in documentation/backlog/ROADMAP.md and is referenced from gap rows by ID (→ R-n). A row here never claims future behavior.

Status enum (strict):

Status Meaning
PROVEN-LIVE Exercised end-to-end on real infrastructure; MUST cite a campaign/drill/validation doc in documentation/audits/ or documentation/tests/. No citation → not PROVEN-LIVE.
IMPLEMENTED Shipped + unit/red-proof tested, but the real flow has not been exercised live (or not on the surface that matters — noted per row).
PARTIAL Some legs live, some missing/unvalidated — the note says which.
MISSING Does not exist. Present-tense fact; if planned, the row points at a roadmap ID.

Update rule (end-of-session checklist item): if a task changed any capability's status, update the row in the same session — with the new evidence citation. A PROVEN-LIVE claim is subject to the cardinal rule like any other claim.

Verified 2026-07-16 against evidence corpus @ felhom.eu tip 4b18cc5 by CC (capability-map audit); see REPORT.md for the per-row verdict table.


A. Provisioning & day-0

Scenario Components Status Evidence Gap / roadmap
Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box installer, agent, hub, golden PROVEN-LIVE (nested VM) DRILL-day0-vm-2026-07-12, DRILL-day0-take2-2026-07-12 First firing on real customer hardware pending → R-1
BYO install: --mode byo, mandatory caps, host-mutation disclosure, coexistence guards installer v1.15+, agent PARTIAL DRILL-GL6-2026-07-08 (demo box); GL-8 coexistence fixes Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1
The installer is PUBLISHED, not pushed — the artifact that runs as root on a virgin box is served from a version-controlled ref, and rolling back is one act scripts v1.23.0 + manifests/webpage.yaml (R-110, operator ruling option (b)) PROVEN-LIVE (2026-08-03) scripts/CHANGELOG.md v1.23.0 + REPORT.md. Proven by HTTP against the real URL, not from a pod's filesystem. Scenario A: a real push to main without moving the tag left the served script byte-identical (sha256 2f859555…), and a marker comment planted in that very commit was absent from the served bytes, while the website tree advanced to the new commit in the same observation — both halves of the split in one measurement. Scenario B: moving the tag published in ~40 s (sha → ea2b4aa9…, marker present) and moving it back restored exactly the pre-publish sha. https://felhom.eu/ returned 200 throughout. P-A, measured BEFORE the manifest was touched because the model rests on it: git-sync v4.4.0 follows a tag and notices a moved one (update required … local:<old> remote:<new> → updated successfully) Two syncs, deliberately: the WEBSITE still tracks main. Pinning both would turn every copy edit into a release, which makes the release meaningless and the site slow to fix. Publish = cut installer-v<SCRIPT_VERSION> + bump the manifest --ref + sync; roll back = move the tag back, which needs no ArgoCD sync and no deploy. Continuity is structural rather than lucky: both trees are seeded by init containers so a fresh pod is not Ready until the tag is checked out, and maxUnavailable rounds to 0 on one replica, so a failed scripts-init leaves the OLD pod serving — the failure direction is no update, never no /scripts/. The URL never carried a ref, so the bootstrap script and the hub's day-0 command follow the tag with no edit and no hub change. The installer's own sixteen run-time fetches are a separate channel pinned to the AGENT's version (R-183), because they are the agent's configs and not this repo's — leaving them on main would have made the whole change cosmetic
Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot host-install); selectable UEFI loader; universal secret-free / operator-bind mode scripts v1.19.0 (scripts/iso/) + hub v0.62.0 + assistant container PROVEN-LIVE on TWO different boards (N100 2026-07-18; HP t740 2026-07-21) tests/VALIDATION-n100-rehearsal-2026-07-18.md — the full chain on real metal in a single pass: the generic reusable pairing ISO (v1.20.0, --loader mkimage, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted (appliance_registrations id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. F1 is closed on physical hardware. Prior nested legs: slice A SPIKE-baremetal-iso-2026-07-16 (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; --loader mkimage boots+installs SB-off; mkimage SB-enforcing FAILS Access Denied; surgery byte-identical); slice C (2026-07-17): the GENERIC secret-free ISO — box self-registers as an unclaimed appliance (POST /api/v1/appliance/register, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) F1 loader caveat: --loader mkimage fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → Secure Boot must be OFF; default shim keeps SB. Slice C bind is operator-password-gated (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17) — see the dedicated self-bind row
Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (<iso>.rootpw.txt) scripts v1.24.0 (scripts/iso/felhom-bootstrap.sh network_gate/sweep_nics, build-felhom-iso.sh rootpw emission) PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install) audits/SPIKE-firstboot-nic-sweep-2026-07-22.md — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and self-registered at the hub unaided in under a minute; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (state.json gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched
Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) controller v0.122, hub v0.50 PROVEN-LIVE (drill VM) DRILL-day0-vm-2026-07-12 §10/F-4 (gate ON via real edge; claimed, code consumed) Never executed by a non-Viktor human → R-3. Deliverability (R-4), gmail half DONE 2026-07-18: the rehearsal's claim email was the first sent under the tightened DMARC p=quarantine and landed in the gmail Inbox, not spam (tests/VALIDATION-n100-rehearsal-2026-07-18.md). freemail.hu remains Viktor's open half. (Dropped mis-cited CAMPAIGN-4 F-C — that is the escrow-claim 502, not password claim)
Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor /bind/<token> (console pairing code + retrieval passphrase) → hub stages the bind, no operator hub v0.66.0 + ISO scripts v1.20.0 PROVEN-LIVE (real customer-zero bind on metal, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md: operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); the customer bound their own box at 16:29:55 with attempts=0, locked=0 — appliance_bound carries source customer_selfbind, and the credential was delivered 26 s later with no operator action. Hub-side lifecycle in hub-state.txt (selfbind_tokens mint→email→consume). Prior unit evidence: hub v0.66.0 (web/selfbind.go, store/selfbind.go; Scenarios A–F + F1/F2; 4 red-proofs verified red — THE TRAP /bind/ exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) Re-walked 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, hub v0.113.0): the mailed link, read in the real mailbox, bound a fresh box first try (self-bind SUCCESS … by customer self-service, credentials delivered 17 s later) — audits/BIGNIGHT-household-month-2026-09-14.md. Gap measured the same night: the link is auto-sent only at customer creation and RESET, so a box installed for an existing customer waits for the operator's press (R-509). R-27 slice 1. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. Live first-run DONE 2026-07-18 (rehearsal; the console banner rendered on the real ISO). R-27b (controller second-box dismissable prompt) deferred; multi-box-per-link = repeated operator sends
Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge controller v0.127, agent v0.88/0.89 PROVEN-LIVE (drill VM, endpoint-exact) agent v0.88.0 REPORT (ceremony ~4s, one-shot claim 200→410, R absent from every payload); SPIKE-controller-escrow-2026-07-13 Customer-facing browser wizard FIRST LIVE FIRING 2026-07-18 (tests/VALIDATION-n100-rehearsal-2026-07-18.md, S6): customer zero drove the wizard on the reborn box — ceremony started 16:56:29, recovery code claimed one-shot 16:56:39 (absent from logs by design), hub-verified and EscrowState auto-confirmed 16:56:41, offsite runs enabled 12 s after the ceremony began; the v0.138.0 „megerősítésre vár, legfeljebb 15 perc" awaiting card rendered and flipped on the ACK (operator screenshots: Viktor's set). Honest caveat: at a 12-second confirm the awaiting window is so short that catching both states on screen is luck, not procedure. Prior: endpoints driven on the drill VM. agent v0.89.0: /escrow/preflight pbs_storage_id row now live-reloads (reads current agent.json) — a pbsdr convergence that seeds the id flips it green with NO service restart. **hub v0.60.0 (data-first retention) — ⚠ the claim as written was FALSE for the offsite tier for two months; FIXED in hub v0.93.0 (2026-08-04), and the row below states what ships TODAY: a re-escrow with a DIFFERENT sealed passphrase no longer destroys the old blob — the hub RETAINS it (host_escrow_superseded) — and since v0.93.0 the retained row carries identity_blob as well as the K-escrow blob, so a previous passphrase does now stay recoverable with the recovery code that sealed it — OPERATOR-ONLY, and the customer-facing half of that sentence is FALSE (R-304, drill 2026-08-12). The retention was exercised end-to-end for the first time that day and it works: the retained row carried the identity blob byte-identically (sha256 a10032341c8584ed…), the old recovery code unsealed it, and three planted files — including a Hungarian accented filename verified as raw bytes — restored byte-identical from a store the box itself could no longer open (negative control first: Fatal: wrong password or no key found). What does not exist is the door. ListSupersededEscrow (hub/internal/store/store.go:2841) is the only reader of a retained identity_blob and has zero production callers; the product's recovery path (POST /escrow/recover-offsite-password → GetHostDRBundle, store.go:3152) selects FROM host_escrow — the CURRENT row only. Asked with the code that demonstrably opens the retained row, the product answers "the recovery code did not open the sealed bundle". So: recoverable by an operator with SQLite, age and a shell; not recoverable by the customer, who is told their correct code is wrong. — AMENDED 2026-08-12 evening (R-311, shipped hub v0.103.0 + agent v0.129.0 + controller v0.214.0): the customer is no longer told their code is wrong. The agent now tries the retained packages when the current one refuses (GET /hosts/<id>/escrow/retained → ErrCodeOpensRetained → HTTP 422), and the screen says the code is CORRECT, names the supersession date, says the earlier package is kept and the current backups are unaffected, and routes to support. What is still true and must not be read away: there is no in-product ROUTE to the set-aside data (R-312 — every restore entry point resolves its repository from settings and its password from one file; adding an alternative is new surface, not wiring), so the recovery itself remains operator-performed. The status of this capability is therefore "the customer is told the truth and handed to a human", not "the customer can recover their old history". What was wrong until v0.93.0, recorded because it is the ninth entry in CLAUDE.md's comment-vs-code table and the first that was also customer-facing copy: host_escrow_superseded had no identity_blob column and demoteCurrentEscrowTx did not copy one, so what survived a supersession was the PBS datastore key only — never the restic repository password, which lives in identity_blob. The destroying act was the escrow ceremony a rebuilt box asks its customer to run. Measured live 2026-08-04, before the fix: both current rows held blob=383 B and identity_blob=572 B; both retained rows held blob=383 B only. → R-198 (SHIPPED), evidence audits/RECON-offsite-dr-chain-2026-08-04.md §7. THREE SCOPE LIMITS THIS ROW MUST NOT BE READ PAST. (1) Nothing was backfilled and nothing could be — rows superseded before v0.93.0 were written without the blob and their source rows are already overwritten; both demo boxes' pre-2026-08-04 repository passwords are gone permanently. (2) A retained key is not a restore — but as of 2026-08-04 evening it IS a recovered key. See the row below. (3) The customer-facing orphan card still promises recoverability unconditionally (R-202, gate hit, card untouched). Guided-recovery flow = R-26. Red-proof TestSaveHostEscrow_RetainsSuperseded. hub v0.60.1 — custody survives the host lifecycle: host deletion (with the escrow ack) DEMOTES the current blob to retained custody (moved into host_escrow_superseded, never destroyed; existing superseded rows spared); the customer Danger-zone Delete is the one true purge point (cascades both escrow tables incl. already-deleted hosts). No operator path through host lifecycle can lose a blob. Red-proofs TestDeleteHost_DemotesEscrowNeverDestroys + TestDeleteCustomer_PurgesEscrowCustody. agent v0.93.0 (2026-07-21) — recovery codes can no longer contain a hyphenated word. The EFF large list holds exactly four entries containing the hyphen the words are joined with (drop-down, felt-tip, t-shirt, yo-yo); drawing one produced a code that reads as 11 words instead of 10 — ambiguous to transcribe in exactly the situation R exists for. They are now excluded from GENERATION only: the draw space goes 7776 → 7772 and a 10-word code 129.248 → 129.241 bits, still well clear of the 128-bit floor. Every code already issued remains valid — R is verified as a whole passphrase by the PBS scrypt KDF and is never re-split, so no customer needs to re-run a ceremony. This also retired the long-standing ~1/5 TestGenerateRecoveryCode_EntropyAndFormat flake, which was this defect and not a flaky test

| The offsite repository password can be RECOVERED from the sealed escrow with the customer's recovery code | hub v0.94.0, agent v0.125.0, controller v0.195.0 | PROVEN-LIVE (2026-08-04) | On demo-felhom, through the real endpoints end to end: the box fetched its own sealed blob from the hub with its own per-host credential (hub log: escrow blob SERVED … 572 opaque bytes, self_scope=true), the agent unsealed it with the customer's recovery code, and the extracted repository password's sha256 was byte-identical to the one on disk — c60c8bc737a6…, which is ALSO the hash the hub had independently stored, so three sources agree. Five minutes earlier the same path with a WRONG code failed closed at age's KDF with nothing written, which proves links 6 and 7 ran independently of the success. R was searched for afterwards and found in 0 log lines and 0 files, with a positive control confirming the search would have found it. Evidence: per-repo CHANGELOGs; audits/RECON-offsite-dr-chain-2026-08-04.md §3 links 6–8 | WHAT THIS ROW DOES NOT CLAIM, stated because the previous over-claim here was struck out four hours earlier. It covers the KEY, not the DATA. R-199 BACK-POINTER (omitted when this row was written): the recovery chain's link inventory and the per-link status live in audits/RECON-offsite-dr-chain-2026-08-04.md §3; links 1–8 are walked, 9–11 are not. Re-confirmed 2026-08-04 evening by an attempt to prove the DATA half: the R-201 drill was prepared on demo-hp and halted before the wipe — the sentinel file was not in the off-site snapshot (R-203), so the wipe would have destroyed it and proven nothing. No file has still ever been restored from an off-site backup after a wipe (audits/DRILL-r201-offsite-recovery-2026-08-04.md). The install half of the chain (controller v0.196.0 --recover-offsite-install, R-200) is likewise unit-proven only — it has never run against a live recovery. A recovered password has never been installed (the diagnostic compares and refuses to write, by design), no existing repository has ever been reopened under one, and no file has ever been restored from an off-site history via a recovered key. Links 9–11 of the chain are open (R-200's remaining half, R-201). The proof also used a box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is exactly what the drill covers and it has not run || DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | PROVEN-LIVE (2026-07-21) | DRILL-day0-take2-2026-07-12 §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) ⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39). On the reborn N100 the descriptor auto-provisioned and the agent reported converged state=applied (16:45:53), yet the storage is dead: pvesm status → felhom-pbs: error fetching datastores - 401 Unauthorized / inactive, and a direct probe with the stored credential returns 401 on every endpoint including /version while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and consumed_at is still NULL; the converged state machine will not re-apply, and the agent's 15-minute verify loop cannot even read the credential to notice (open /etc/pve/priv/storage/felhom-pbs.pw: permission denied — non-root agent reading a file it writes through a root wrapper). A tier that reports applied while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See tests/VALIDATION-n100-rehearsal-2026-07-18.md F2 and pbs-dr-state.txt. agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17: the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick pbsdr: pre-check 403 … self-granting … (R-22) → converged state=adopted in ~3 s, ACLs self-restored, pvesm status felhom-offsite=active, zero operator action. No more one-shot pveum grant 2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end. The three defects that let a box be applied and dead simultaneously are each addressed: the hub stamps a monotonic secret_generation into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow read verb so the non-root agent can read the credential it writes (it never could — /etc/pve/priv is 0700 root:www-data, which made the verify loop blind by construction); and pbs.ProbeAuth turns a 401 into a loud auth_failed that the existing pbsdrheal damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says applied, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (rc=0) and probed successfully (credential probe OK storage=felhom-pbs). STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged. The operator pressed Re-issue PBS credentials; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): 08:39:31Z hub mints a fresh secret, generation 0 → 1, and the descriptor gains "secret_generation": 1 — with token_id and fingerprint byte-identical, i.e. exactly the re-key shape that used to be invisible → 10:39:34 the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → 10:39:38 ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD previous_state=applied (leg c: the exact R-39 failure state, detected out loud for the first time ever) → 10:39:45 one-time token secret consumed secret_len=36 (leg a: NO short-circuit — this is the line that never appeared on 2026-07-18) → 10:39:45 felhom-pbs-apply reconcile (the set-only wrapper, no --server) → 10:39:47 pbsdr: converged state=applied. Corroboration: the agent marker hash moved to afbb3b41… (it was byte-identical to the pre-reissue marker in the failure); consumed_at stamped 08:39:45Z; the on-disk secret's mtime moved 2026-07-18 20:28:52 → 2026-07-21 10:39:45; a live probe with the NEW credential returns 200; three consecutive hub reports trace the whole state machine applied → auth_failed → applied; and zero pbsdr_selfheal escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no consumed-failed.json. Row upgraded to PROVEN-LIVE (2026-07-21). Evidence: felhom-agent/REPORT.md (2026-07-21). |

| A customer's file survives a machine rebuild and comes back — the whole off-site story, end to end | hub v0.94.0, agent v0.125.0, controller v0.197.0 | PROVEN-LIVE (2026-08-04 night drill) | demo-hp's controller data volume was destroyed and the sentinel deleted from disk. The customer's recovery code then produced 8a9e33aa4da6… — byte-identical to the pre-wipe on-disk key and to the hub's independent record — and installed cleanly on the bare box. identity_blob was unchanged across the wipe (572 B, updated_at still 11:11:37). Evidence: audits/DRILL-r201-night-run-2026-08-04.md §2 | WHAT IT DOES AND DOES NOT CLAIM. PROVEN: after a real rebuild the key recovers byte-identical, the EXISTING repository opens (3 snapshots, 42 026 B — the pre-wipe size; not a fresh history), and the customer restore flow returns the file byte-identical. ALL FOUR MANUAL INTERVENTIONS ARE CLOSED, AND THE CUSTOMER IS NOW OFFERED THE RECOVERY (2026-08-05). ONE QUALIFIER REMAINS AND IT IS NOT THE OLD ONE — state it precisely. The 2026-08-04 drill needed four interventions between "the key is recoverable" and "the file is back", none of them in any design document (R-204). Items 1–3 closed in controller v0.198.0 + hub v0.95.0 (the reset code works first time; a Re-issue no longer marks a healthy escrow stale; a unit restore states it returned the app's definition and database and NOT the customer's files). Item 4 closed in controller v0.199.0 + hub v0.96.0 (a rebuilt box DECLARES offsite.state=needs_credential and internal/offsiteheal re-arms the stored credential before minting). AND THE "needs someone who knows to look" HALF IS NOW GONE TOO (controller v0.200.0, R-193): a full-page recovery screen takes over the landing pages while the hub holds a sealed package this box cannot open, explains that nobody can replace a lost recovery code, takes the code, opens the repository and lists what is in it — proven live on demo-felhom, which is genuinely in that state (/launcher → 302 /recovery; a wrong code refused with nothing written, three times, no lockout; the code found in no file, no container log and no debug ring, with a planted-copy positive control that first exposed a mis-aimed sweep). WHAT THE QUALIFIER IS NOW, and it is narrower: (1) the final unlock has never been driven with a CORRECT code through the page — no recovery code was kept for demo-felhom's orphaned history and demo-hp's is operator-held out of band, so the live run exercised the whole chain (handler → agent → hub fetch → age KDF) and stopped at the unseal; the install and listing rest on tests. (2) putting files back in place is deliberately NOT part of this — restore stays per-app, and the step after the listing is R-213. (3) the journey has still not been re-walked end to end since these fixes — the closures are proven individually, not as one uninterrupted run. That re-walk is one more drill and it is what is owed. Scope: demo-hp, a controller-data-volume rebuild — NOT a total host loss, and NOT a guest reprovision. ⚠ THE RE-WALK RAN ON 2026-08-05 (CAMPAIGN-11) AND IT FAILED. THIS ROW STAYS AS IT IS — it records a key recovering and a file returning, which remains true — but the JOURNEY row below is the honest verdict and this row must not be read as covering it. The campaign built a throwaway appliance from the published ISO, destroyed it, and walked the customer's route: the data half PASSED (three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian filename, out of the pre-wipe snapshot in 16 s) and the journey half FAILED — four operator interventions, three needing root on the appliance, and the first of them was the machine telling the customer their perfectly correct recovery code was wrong (R-216). R-198's retention is NO LONGER unit-proven — it was PROVEN IN PRODUCTION on 2026-08-05 (CAMPAIGN-11 Phase 3): the first supersession since the fix retained identity_blob at 572 B, byte-length exact, carrying the OLD key 626e424670248db3 while the current row moved to the newly minted e11a6c542b73477a; offsite_repo_key_changed and escrow_superseded both fired at the instant of supersession, and the next off-site run REFUSED as orphaned rather than starting a fresh history. Evidence: tests/campaign11-evidence-2026-08-05/journal.md RE-WIDENED 2026-08-22, and the narrowing below is now HISTORY — read both, in order. Both defects named in the 2026-08-21 narrowing are closed and proven live: R-354 (controller v0.218.0, volReplay) and R-356 (controller v0.219.0). The 40-class end-to-end story is now WALKED, including the hardest ten of it: audits/DRILL-r356-hot-only-restore-2026-08-22/ proved a driveless app with NO database (privatebin: planted, off-sited, deleted, restored, 15/15 files byte-identical, two Hungarian accented names), and audits/DRILL-r356b-driveless-db-restore-2026-08-22/ proved a driveless app WITH a database on both engines — docmost (Postgres 16) and bookstack (MariaDB 12.3), each planted through the app's own interface, destroyed for real, and returned with accented names byte-identical. Five legs that had never run in any combination all ran and all succeeded: the undo copy, DB-service identification, the volume replay, the DB-only start window, and the dump replay on top. A second claim was walked at the same time: R-164's F17 ordering — the logical dump wins over the volume tar's copy of the same database — was recorded only for the LOCAL path (restore_unit.go:262-266) and is now measured on the off-site path too, by a three-way discriminator (volume tar ORIGINAL-VALUE-A, altered dump ALTERED-VALUE-B, live LIVE-VALUE-C3; result ALTERED-VALUE-B). MEASURED 2026-08-22, AND THE READING THAT PROMPTED IT WAS WRONG — recorded so nobody re-derives it. It was read from source that a HELD app would raise the dead-app banner and a customer e-mail, because it keeps one container (its database) and so is not StateStopped. It does not. A held app aggregates to unhealthy, and aggregateState checks unhealthy > 0 before the mixed-case degraded branch while IsDownState excludes unhealthy entirely. Measured on demo-hp on the shipped v0.220.2: hold created 21:11:19Z, dead-app scans every 30 s ran over it, and the heartbeat reported 0 currently down throughout. No suppression was built, because there was nothing to suppress. The detector itself is sound — classifyRunStates is pure and a degraded/exited stack does raise the banner, pinned by TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm. What the same measurement DID expose is R-384: an app whose database has died is unhealthy too, and is likewise silent. Evidence: audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt.

R-384 CLOSED in controller v0.222.0 (2026-08-23), proven live. The defect was the ORDER of two questions, not the unhealthy exclusion: aggregateState now asks "is a supervised member dead?" before the unhealthy/starting/restarting returns, and "some members are up" counts any member not in the down bucket rather than running alone. IsDownState is byte-identical. Measured on demo-hp 2026-08-23 with the same fixture that read 0 currently down the day before: bookstack-db stopped 05:30:07Z → app_start_failed fired at 05:30:14Z, the banner read „Telepített alkalmazás nem fut: BookStack (degraded)", the stack read state=degraded while its front end was unhealthy, and the heartbeat printed 1 currently down against the previous day's 0. Evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/.

The HELD-app half of the paragraph above is now also covered — a held app keeps its database container, so it is the same shape and reaches the same degraded verdict.

The alarm ladder that decides all of this now has an owning document: see 08-alarm-ladder.md (written 2026-08-23 — before that date no document owned it, and that absence is why the ordering defect was legible only from source).

WHAT IS STILL NOT CLAIMED: the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy no product action can apply (R-379), and on MariaDB it does so behind an app that reports health=healthy (R-380). The success story is proven; the recovery-from-a-bad-restore story is not.

NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above. Proven that night on demo-hp with planted, hash-recorded files: the declared-userdata leg of a drive-declaring app does come back byte-identical (calibre-web, 5/5 including two Hungarian accented filenames). Two legs of the same story do NOT: (a) the off-site restore has no named-volume leg at all, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (R-354); (b) for the 40 of 53 apps that declare no data drive the off-site restore refuses outright, saying a running app „nincs telepítve" (R-356). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is unproven for that class and disproven for the volume leg generally. The escrow/key half of this row is untouched by that and still stands. Evidence: audits/DRILL-backup-truth-2026-08-21/evidence/ and REPORT.md (2026-08-21). | | The customer's own UNAIDED recovery journey, end to end | controller v0.206.0, hub v0.98.0, agent v0.127.0 | PROVEN-LIVE (2026-08-07, the fifth walk) — an unaided customer journey on a real installation, SCOPED: see what it does not claim | A throwaway appliance was built from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed — guest purged and both enrolled drives wiped. The rebuild and the customer's route were then walked with no command line inside the guest. Data: PASS. All three sentinels byte-identical (beb9175d…, 7c8cb0ad…, e012e76f…), including a 12 MB binary and C11-őrszem-ékezetes-árvíztűrő.txt — the Hungarian filename survived disk → restic → SFTP → Storage Box → restore → disk intact. Restored in 16 s out of the pre-wipe snapshot f3d9cd67, non-destructively. Journey: FAIL. Four interventions, three of them root on the appliance: R-216 (a correct code called wrong), R-218 (the store never opened; the box stopped asking), R-219 (the promised listing structurally could not render), R-220 (the app could not be redeployed — its drives were unenrollable). Evidence: tests/campaign11-evidence-2026-08-05/journal.md | WHAT IT DOES NOT CLAIM, and what the v0.201.0/v0.97.1 fixes do and do not change. Six of the nine findings are fixed and each is red-proofed (R-216, R-218, R-219, R-217, R-222, R-215) — but fixes are not a journey. This row goes green only when the walk is repeated end to end and completes with no operator intervention. Three findings are deliberately still open and each is a live blocker for some flow: R-214 (the console never stops showing a stale pairing code), R-220 (after a rebuild the customer's drives cannot be re-enrolled — the deploy refuses and the wizard's list is empty, so no app can be redeployed onto its own data drive), R-221 (a rebuilt box cannot run the escrow ceremony at all, because the pbsdr convergence marker survives on the host while the seeded escrow.pbs_storage_id is rewritten out of agent.json). And R-216's fix does not by itself make a NEW box work: with the hub guard corrected but the Day-0 manifest still vouching agent 0.120.0, a new box is HELD, not served — it stops being lied to, which is better, but the feature does not work for it until agent 0.125.0 is vouched. ⚠ CAMPAIGN-11 PHASES 2 + 4 RAN OVERNIGHT 2026-08-05/06 AND THIS ROW STAYS FAIL. What they add: eleven injected faults against the recovery journey, each with a positive control, and a full unattended scheduled cycle. The backup promise strengthened — the off-site tier ran itself at 04:15 on a box rebuilt twice and set aside hours earlier (snapshot_count 1 → 2), all five daily jobs fired exactly once, and nothing on the must-not list fired. R-217's and R-215's fixes were proven live under exactly their faults, the set-aside proved it does not delete (12 535 KB byte-exact at the far end), and a wrong code was refused three times with nothing written and no lockout. What they do NOT add: any progress on the JOURNEY. These faults are not a re-walk — no customer route was walked end to end — and four new findings say the journey got no better where it matters: R-224 (a hub outage and a stopped agent are BOTH reported to the customer as a bad recovery code, in 0.056 s and 0.030 s — no unseal attempted; the agent's own err distinguishes them and it is discarded at the HTTP boundary, and R-216's gate answers source=version so it cannot see reachability — Phase 1's headline defect relocated from the version channel to the transport), R-226 (M1, the only message telling a customer to check their typing, is unreachable on any box that has re-escrowed), R-225 (an unread store renders 0 pillanatkép · 0 GB above a card saying it holds backups — 1 snapshot and 12 535 KB were really there), R-228 (the set-aside history is recorded in orphaned_renamed_to and shown by nobody — OrphanedRenamedTo has zero references in any template or handler). §4.1 is now MEASURED rather than deduced — the floor IS served, from the box's own rendered GetFloor() and a cold-started controller's settle-gate line. §4.2's positive half is still NOT measured and needs a rebuild. Evidence: audits/CAMPAIGN-11-recovery-journey-2026-08-05.md, tests/campaign11-evidence-2026-08-05/journal-phase24.md 2026-08-06 — THE FIVE PHASE-2 FINDINGS ARE FIXED AND THIS ROW STILL SAYS FAIL. controller v0.202.0 + agent v0.126.0 close R-224, R-226, R-225, R-227, R-228. What changed: the unlock path now classifies WHY it failed, from the value and never the text, and the message that mentions typing is reachable from exactly one class — a real refusal — with unknown defaulting to neutral. Proven live on the venue with the same wrong code and only the hub's reachability changed: 400 → 502 → 400. What it does NOT change: this row. Fixes are not a journey; nothing here walked a customer end to end, and three of the blockers that made Phase 1 fail are still open (R-220 in particular is worked around BY HAND on the venue — without that unmount no app can be deployed on a rebuilt box). What still needs proving: a re-walk with no operator intervention, and specifically the customer-facing messages end-to-end — they were NOT re-driven, because /recovery correctly retires itself once the old data is set aside and restoring that state would be the reconfiguration the fix task forbade. Evidence: felhom-controller/REPORT.md ⚠ THE RE-WALK RAN 2026-08-06 (R-201, attended) AND THIS ROW STAYS FAIL — but it is a nearer no. A brand-new appliance was built from the published ISO, given three sentinels, escrowed, backed up off-site, destroyed (guest purged AND both drives wiped), reinstalled and walked. THE DATA: PASS — all three sentinels byte-identical, including a 12 MB binary and an accented Hungarian filename whose NAME BYTES are also identical, restored in 23 s out of the pre-destruction snapshot through the customer's own flow. THE JOURNEY: FAIL — TWO dead ends against Phase 1's four. (1) R-218's CONSUME half: the hub re-staged the credential and said "the box re-consumes on its next cycle"; a full cycle ran (positive control) and it did not — no customer-reachable action fetches it, and the only lever is a command line INSIDE THE GUEST. (2) R-220: the drives are still unenrollable after a rebuild, needing a Proxmox-host unmount, without which no app can be redeployed and the restore page stays empty. The unaided RTO is therefore STILL UNDEFINED; the attended figure was 30 m 13 s and must not be quoted as the customer number. What passed and is new: the recovery screen appeared WITHOUT being sought, answered all three of its questions with a seal date matching the hub exactly, the emailed reset code worked first try, and the unlock was a real 1.528 s unseal that placed the key. ⚠ AND IT DOES NOT CLAIM DELIVERY: the fresh install landed on controller 0.201.0 / agent 0.125.0 — the vouched versions, neither carrying the fixes — which were installed BY HAND. Fleet delivery needs a golden carrying the fixed controller and a vouched agent, and nothing was vouched. Evidence: tests/rewalk-r201-2026-08-06/journal.md   ⚠ 2026-08-06 (later the same day) — THE TWO DEAD ENDS ARE CLOSED AND VOUCHED, AND THIS ROW STILL SAYS FAIL. controller v0.203.0 + agent v0.127.0 close R-218's consume half and R-220, and golden 0.203.0 / agent 0.127.0 / min_agent 0.127.0 are now VOUCHED — so the delivery gap the re-walk recorded (“the fresh install landed on versions neither carrying the fixes”) is gone. Both were then proven live on a genuinely rebuilt box (Part 4 venue, VM 323): the reinstall landed on agent 0.127.0 with no downgrade and no hand upgrade (the previous re-walk's reinstall downgraded 0.126.0→0.125.0), /disks/candidates returned both drives after the guest was purged with the raw mounts still on the surviving host — before the fix, two empty lists — and both re-attached through the customer endpoint; the off-site credential was re-staged by offsiteheal at 13:24:57Z and collected by the box on a tick, unaided. Why the row is still FAIL: the walk did not finish. It stopped at the restore surface — R-237, the list was keyed on apps that are currently INSTALLED and currently TOGGLED ON for future backups, so a rebuilt box was shown nothing to restore while the repository held its snapshots. Fixed in controller v0.204.0 (the store is now the source of the list) and proven live on that same venue with the toggle switched OFF, but no sentinel was restored in that walk, so the data half is unproven in EITHER direction for this venue. R-238 was RECLASSIFIED, not fixed as filed: mode=full without confirm=1 is step 1 of a deliberate two-step that starts no job by design; the endpoint driver did not carry ?full_prep= forward. The operator drove the same restore to completion in a browser. Its real residue — that step logging nothing at all, including when the headroom gate refuses — is fixed. R-236 was WITHDRAWN: not a defect. This row goes green only when a walk completes end to end with no operator intervention, restoring a byte-identical sentinel. Evidence: tests/part4-rewalk-2026-08-06/journal.md, tests/teardown-2026-08-06.md, felhom-controller/REPORT.md ⚠ THE FINAL WALK RAN OVERNIGHT 2026-08-06/07 AND THIS ROW STAYS FAIL — the data half passed again, the journey half failed further from the line than before. A brand-new appliance was built from the published ISO, fixtured, soaked through a real scheduled cycle, destroyed and rebuilt. THE DATA: PASS — all three sentinels byte-identical out of snapshot f5c53b03, including a 12 MB binary and an accented Hungarian filename whose name BYTES are identical too. THE JOURNEY: FAIL, and this time the customer has NO route at all — / lands on the launcher with no recovery pointer, /recovery 302s away, and the remote-backup page offers to CREATE a new recovery code, which would orphan the very history the customer's code protects. There is no field anywhere to enter the code they hold, and the operator's documented remedy refuses too. Cause (R-241): the credential self-heal — proven working unaided in the same walk, six hours earlier, as its best result — writes a FRESH repository password, which moves the box out of OffsiteRecoveryOffer()'s (a) pristine case while (b) orphan-detected is unreachable because runs are blocked by escrow_state: pending. R-218's shape one level up. What DID pass, and is new: the whole credential chain ran end to end with zero human action on an unclaimed box — declare, offsiteheal re-stages after two reports, the box's 5-minute retry collects it, tier applied — the first live sighting of that success line, settling R-218's consume half and R-236's withdrawal. Also new: R-239 — a fresh install lands on controller 0.203.0 while 0.205.0 is released, so R-234 and R-237 are written but not delivered, proven from the customer's side (T2, T3). This row goes green only when a walk completes with no operator intervention AND a byte-identical sentinel. Evidence: tests/finalwalk-r201-2026-08-07/journal.md ⚠ 2026-08-07 — R-241 WAS DIAGNOSED BY A READ-ONLY SPIKE ON THE STANDING VENUE, AND THE DIAGNOSIS REVERSES THE FIX. This row is unchanged: still FAIL, and no code was written. The question was whether R-241 is a screen-predicate defect or a minting defect. It is a MINTING defect. The screen was telling the truth — there genuinely was nothing recoverable under the key the box held, because the box minted that key itself over the top of a sealed package it already knew the hub was holding. Three measurements, from the venue rather than from the earlier report: (1) WriteOffboxSecrets (offbox.go:411) mints on ONE input — does the file exist — while OffsiteRecoveryOffer() and needsOffsiteCredential(), both in the same file, consult GetHubEscrowIdentityPresent(); the same fact is available on three paths and used on two. (2) That flag was the precondition of the chain that reached the minting: the retry job logs only when the declaration is live, and the venue logged credential retry: … (the box still declares a need; retrying) at 02:48:03Z — thirty minutes and six ticks before the mint at 03:18:06Z. (3) The box computed the answer and discarded it: at 03:28:03Z, thirty-five minutes before the customer looked, EscrowAutoConfirmer.Reconcile (escrow_confirm.go:154) logged the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…). It is recomputed every report cycle and never persisted. And the hub explicitly disclaims doing this — offsiteheal's package doc: "credential automatic, key customer-present — the ruling this session implements and must not quietly widen." Shape (b) is structurally unreachable on this box (the escrow gate at offbox.go:743 sits upstream of ensureOffboxRepo, the only producer of RepoState="orphaned"; positive control: the scheduler was alive, 241 agent-channel-health ticks, and offbox-backup is a sched.Daily leg whose slot fell before the destruction). Two by-products: R-243 — a box in this state silently stops backing up off-site and no alarm fires (isStale requires escrowed, offsite_delivery_stuck skips the applied shape, backup_failed needs a run that never happens); and the trap in the obvious fix — ResetOrphanedRepo clears the orphan flag without a ceremony, so a hash-mismatch discriminator alone would re-offer the screen forever to a customer who declined the old data. Q7: the „Helyreállítási kód létrehozása" button does NOT destroy the data — R-198's retention copies the identity_blob — but it converts a self-service recovery into one needing an unbuilt read path (R-199), and it re-enables the screen while invalidating the code that screen accepts. Evidence: audits/SPIKE-r241-recovery-offer-2026-08-07.md ⚠ 2026-08-07 (later) — R-241 IS FIXED (controller v0.206.0 + hub v0.98.0) AND THIS ROW STILL SAYS FAIL. Three changes, following the spike's ruling rather than the obvious reading: the box no longer mints a repository key while the hub holds a sealed package (a conjunction, so a first-time box is untouched; the refusal is a HOLDING state that still writes the transport, declared as offsite.state=awaiting_recovery_key and shown inert to every existing hub reader); the hub-vs-local key comparison that was computed every cycle and discarded is now persisted and drives the offer as shape (c); and abandoning starts a 14-day countdown whose terminal step removes the set-aside store and the sealed package together, so the offer ends because the state is right rather than because a flag suppresses it. Surface: the full page appears once per ENTRY into the offered state, three dismissal levers with three scopes, and none removes the entry point. Q7's trap is closed — „Helyreállítási kód létrehozása" is unavailable while a recovery is outstanding. WHY THE ROW STAYS FAIL: these are fixes, not a walk. Nothing here walked a customer end to end, and this row goes green only when one completes with no operator intervention AND a byte-identical sentinel. Two of Phase 1's blockers are also still open (R-214, R-202), and R-240 is untouched. Evidence: felhom-controller/CHANGELOG.md v0.206.0, hub/CHANGELOG.md v0.98.0. DELIVERY, separately: R-239 is CLOSED 2026-08-07 — golden 0.205.0 baked, published, round-trip verified (./etc/felhom-controller-image read OUT of the downloaded archive) and VOUCHED, so a fresh install now lands on 0.205.0 carrying R-234 and R-237. In the event it was a ONE-field change: agent_version and min_agent both stayed 0.127.0. Evidence: tests/golden-0.205.0-2026-08-07/. The row still says FAIL — delivery is not a journey, and R-241 is diagnosed, not fixed. ✅ 2026-08-07 — THE FIFTH WALK PASSED, BOTH HALVES, AND THIS ROW TURNS. A brand-new appliance (VM 325, customer walk5) was installed from the published ISO on demo-hp, claimed, given three sentinels, escrowed, backed up off-site, then destroyed on purpose — guest purged, both data drives wiped — and rebuilt through the documented day-0 path. THE DATA: PASS — all three sentinels byte-identical out of snapshot 5b0f20f7 (11eb7fb2…, 6b504d1e…, 0baaf402…), including a 12 MB binary and WALK5-őrszem-ékezetes-árvíztűrő.txt whose name BYTES are identical too, read back with os.listdir on a bytes path so no decode round trip could launder a U+FFFD. THE JOURNEY: PASS — ZERO guest command lines were needed to progress, against three on the previous walk. RTO 71.7 s from login to an open store (12.44 s of it the unseal; ~22 s a harness retry). The recovery screen appeared without being sought (/ → /launcher → /recovery), answered all three of its questions, and its sealed-at timestamp matched host_escrow.created_at exactly. WHAT MADE THE DIFFERENCE — R-241's mint guard, exercised live for the first time: at 14:58:52Z, unaided and before anyone logged in, the rebuilt box collected its re-staged credential, configured the transport and refused to mint a repository password over the sealed package — NOT minting a repository password: the hub holds a sealed recovery package for this box. Sampled every 20 s from T0: no key at any moment, with the scheduler proven alive throughout. At the equivalent moment the previous walk minted one and lost the journey silently. AND DELIVERY IS PART OF THE PASS: the fresh install landed on controller 0.206.0 + agent 0.127.0 — the vouched set, no hand upgrade, both times, so the box under test is the box a customer receives. WHAT THIS ROW STILL DOES NOT CLAIM. (1) Putting files back in place is built and worked here — reconstitute placed 6 files — but only after two obstacles the customer must guess past: R-252 (the restore refuses with „nincs elérhető adatmeghajtó" because a rebuild loses the drive registration, and nothing on the recovery path says to re-attach) and R-253 (the restore refuses because the app is not installed, on a page that says three lines above that the restore reinstalls it). Both were cleared from the dashboard with no shell — which is why the journey passes — but neither is signposted, so unaided here means possible without a shell, not obvious. (2) Shape (c) did NOT fire positively. With the mint guard holding there is no local key, so the offer comes from shape (a); shape (c) was measured in Phase A in its negative half (hub hash == local hash, correctly silent). The mint guard is proven positively; the discriminator only negatively. (3) R-214, R-202 and R-240 are untouched. (4) The venue is torn down (2026-08-08) — full-schema census 168 rows → 67, of which 37 are audit-by-design and 30 are R-244, predicted before the run rather than found after; every layer verified absent against a surviving positive control, and 16.64 GiB returned. tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md. Evidence: tests/walk5-r201-2026-08-07/journal.md SCOPE NOTE 2026-09-14 — not re-walked. The first-hour drill on 0.242.0 (audits/DRILL-fresh-install-0242-2026-09-14.md) walked install → first use → restore of a deleted page on a fresh box, NOT a rebuild with off-site recovery (DR tier and off-site were off). This row is therefore neither re-proven nor contradicted on 0.242.0; its PROVEN-LIVE stands on 0.206.0 only. The first hour has its own row below. SCOPE NOTE 2026-09-17 — not re-walked, and this time it could not be. Chaos night (audits/DRILL-chaos-night-2026-09-17.md) ran twelve rounds of household actions under injected accidents on a fresh 0.245.0 box and did not walk the recovery journey at all. It could not have: the box was a REBUILD for an existing customer, so its restic repository was orphaned by design — zero readable snapshots, confirmed independently by restic (Fatal: wrong password or no key found, exit 1) and by the product's own status (orphaned:true, snapshots:0, status:"error"). The product surfaced that honestly as a true alarm within seconds of the first off-site run. This row is therefore neither re-proven nor contradicted on 0.245.0; its PROVEN-LIVE still stands on 0.206.0 only. | | A random night of household actions while random things go wrong — twelve rounds, unattended | controller v0.245.0, agent v0.131.0, hub v0.116.0, golden 0.245.0, ISO 1.28.0 | PROVEN-LIVE (2026-09-17, chaos night) | A fresh nested box installed itself from the published 1.28.0 ISO and bound with zero operator presses (the automatic self-bind mail was already waiting; the acknowledged-delete path re-issued PBS credentials by itself — the F-14 path, measured live for the first time). Twelve rounds were drawn once from seed 20260917 by a committed script and written into the findings document before round 1 began. Across a power cut mid-restore, a hard reset four seconds into another, a system disk at 96 %, a killed tunnel, a restarted Docker, three severed networks and the data drive pulled out of a running machine for twenty minutes: the box healed itself every time with no human action — 148 s after the power cut, 97 s tunnel repair by the controller, 150 s after the hard reset, 67 s after the drive returned. 17 alarms fired, all 17 TRUE, none missing, and the mailbox proves each was delivered to the operator, not merely stored. One intervention all night (a local backup leg that could never have fit; its off-site leg then succeeded unaided). Evidence: audits/DRILL-chaos-night-2026-09-17.md + audits/evidence-chaos-night-2026-09-17/ | WHAT IT DOES NOT CLAIM. (1) Per-app off-site RESTORE was not tested — see the scope note on the recovery-journey row above; the repository was orphaned by design and held zero readable snapshots. The WHOLE-GUEST off-site copy did work and is present on ep0 (two intact snapshots, the later one written tonight), but that is a listing, not a verification — restorability was not tested. (2) The dropped-event path was never exercised. Events pushed while the hub is unreachable are retried 3× then dropped permanently with no queue; three ten-minute hub outages happened and no event was raised during any of them, so that behaviour remains unmeasured. What WAS measured is the REPORT path: built, three attempts over 1 m 40.8 s, given up, and the next scheduled report succeeded — a snapshot, so nothing was lost. (3) Twelve rounds is a sample, not coverage. (4) The household loop samples each app every two minutes, so ten of the twelve rounds left no mark in it — that silence is the instrument's sampling rate, not proof the household saw nothing. Three findings filed: R-547, R-549, R-550. | | A stranger's FIRST HOUR on a fresh install — download → install → claim → two apps → use → backup → remove → restore, plus a power cut and a mistyped code | controller v0.242.0, agent v0.130.0, hub v0.112.0, golden 0.242.0, installer ISO 1.26.1 | PARTIAL (2026-09-14) — every mechanism PASSED on a fresh box; the UNAIDED journey FAILS before the first app | audits/DRILL-fresh-install-0242-2026-09-14.md, evidence audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md. A nested VM on demo-hp installed from the public ISO landed on the vouched set (controller 0.242.0, agent 0.130.0) with no hand upgrade; BookStack and PrivateBin deployed in 68 s / 21 s and were used through their front doors; backup-now moved both dates to the true time; removal with every delete box ticked left no volume, no backup and a 404; a page deleted in BookStack came back byte-identical with its attachment 32 s after restore; a hard power cut returned every app on the same version with data intact and no alarm; a typo in the code was refused and the right code then accepted, five failures locking for 15 min with an honest Hungarian message. ONE intervention a volunteer could not make: the setup-code mail's dashboard link does not resolve for a new customer (no tunnel/DNS is created — R-494), so the dashboard was reached by LAN address. And no instruction exists to begin with (R-493). WHAT IT DOES NOT CLAIM: the self-bind page and the mailed setup code were not exercised (no mailbox in the harness — the operator bind and a box-printed code stood in); DR tier and off-site were off (ep0 fence), so escrow and off-site screens were not walked; no browser, so script-rendered state was not observed. Goes green when a volunteer, from written instructions alone, reaches a working dashboard and a first app with zero interventions. 2026-09-14 (evening) — WALKED AGAIN on ISO 1.27.0/1.27.1 + hub v0.113.0, STILL PARTIAL. audits/DOORSTEP-walk-1270-2026-09-14.md. Held again on a new customer (tester-1, real domain + tunnel token): install on three disks and on one, the vouched set, deploy, use, backup-now, removal, byte-identical restore, power cut (same versions), typo + lockout. New and proven: the console is Felhom-only from the FIRST boot on 1.27.1 (VM 332, reboot proven, pvebanner masked); the hub tells the operator to hand the passphrase over (live). Still ONE intervention: the tunnel connected but received no routes (R-505), and the record has no e-mail (R-508). The graphical installer was proven only to its password screen (R-507). Goes green on a walk with zero interventions on a published installer. BIGNIGHT 2026-09-14/15 (ISO 1.27.1): the first hour re-walked with the real mailbox — self-bind link and mailed setup code both worked; 2 interventions (no auto bind mail R-509; tunnel 502 R-510); then twelve apps, a month of routines and nine faults: audits/BIGNIGHT-household-month-2026-09-14.md. DRILL 2026-09-16 on controller 0.243.0 / agent 0.131.0 / hub 0.115.0 / golden 0.243.0 / ISO 1.27.1 — the WALK half is now PROVEN-LIVE with ZERO interventions, and the BACKUP half is narrowed, not proven. audits/DRILL-prove-fixes-0243-2026-09-16.md. A fresh nested box was installed from the published ISO, landed on this drill's own golden by checksum (e2d1843c…c10a, no self-update), and was walked as a volunteer: the connect e-mail and self-bind page, the claim, the tunnel from outside (302→200, R-510 CLOSED), the data drive, the file manager with its own generated password (admin/admin refused 401 before any dashboard action), four apps deployed and used, and „Mentés most” with 26 s of app downtime. Interventions: 0 — the two operator presses (self-bind send, PBS re-issue) were pre-declared and counted apart, and the four moments that look like help were my own API-driving errors or my own damage, each named in the drill's interventions table. The automatic connect e-mail is PROVEN with a real mailbox: a host delete at 12:22:59Z produced the „Kösd össze a Felhom dobozodat” mail at 12:23:00Z, with selfbind_link_sent (host delete) on the timeline (R-509). WHAT THIS DOES NOT CLAIM, and it is the backup half of this row's own sentence: on a one-drive box with no off-site tier — what every fresh install is — the household's files are in no backup at all (the whole-guest tiers exclude mp8 by design; the file leg lives at tier 2/3, both unset), the app-backup page nevertheless reads „DB + Konfig + Adatok” (R-537), and a restore reports success while leaving Nextcloud listing five photos that return Sabre\DAV\Exception\NotFound — after making the app's own trash, which still held every byte, unreachable (R-538). Also still open from this walk: the off-site tier cannot be provisioned at all (R-534, P1, an ep0 grant), the console keeps its pairing banner after bind and claim (R-535), and „app installed” is emitted at accept time (R-536). Faults re-measured: controller killed during a deploy → back in 37 s; three more kills 20 min apart → 61/41/61 s, none accumulating (R-531); two reboots 60 s apart → everything back in 124 s with the boots NOT counted against the brake; the claim page locks after the second wrong code and mails the operator truthfully. 2026-09-16 (evening) — THE BACKUP HALF IS NOW WALKED TOO, on controller 0.244.0 / hub 0.116.0 / golden 0.244.0 / ISO 1.28.0 (built, unpublished). audits/evidence-backup-promise-2026-09-16/. A second fresh box (VM 335) was installed from the BUILT image and walked to the end of the sentence this row could not previously finish: five photos in → deleted the way a child would → the old route REFUSED and touched nothing → the off-site restore returned them → they OPEN, byte-identical (sha256 5/5, negative control). The refusal reads „Ez a mentés nem tartalmazza az alkalmazás fájljait… a fájlok így a helyükön maradnak" and names the route that can help; the app was running before and after; the wastebasket was untouched. The off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution counting „5 fájl és 3 adatkötet és az adatbázis". Delivery is part of it: the box landed on agent 0.131.0 + controller 0.244.0 (the golden baked and vouched the same day) with no hand upgrade, and app_deploy_started (19:15:34) / app_deployed (19:16:23) finally mean different things. The self-bind half needed NO operator press — the box registered itself and the bind used the mail the hub sent itself after the morning's host delete (R-509, proven twice today: 12:23:00Z and 18:17:46Z, each one second after a host delete). WHAT IT STILL DOES NOT CLAIM — one operator press and one day-one gap. The PBS-DR cascade stopped at the refusal R-511 documents („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action") and needed the operator to press it; that press then SUCCEEDED because of the ep0 grant given this morning (R-534 CLOSED, R-511 CLOSED). And off-site ON by default is not off-site WORKING: a fresh box sits at „Kulcsletétre vár" until the household performs the escrow ceremony, which nothing asks them to do, while the tier-1 row already promises that copy (R-543, P1). 2026-09-16 (late evening) — that last gap is CLOSED, controller v0.245.0 (R-543). The pause is the zero-knowledge escrow design and was not touched; what was missing was the ASK. Every authenticated page now carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking the ceremony (the R-241 bar, second instance, hung on the single render choke point), the tier-1 sentence renders by tier-3 STATE („védené … szünetel" while paused, „védi" when running), and the first-hour guide asks for the code right after the dashboard password and before the first app. Measured on two boxes running 0.245.0: paused box — bar on four pages, POST /backup/offbox/run refused by the fork-4 gate with no snapshot written, app row „védené" and „védi"=0; escrowed box — no bar anywhere, row „védi". audits/evidence-recovery-code-2026-09-16/. WHAT IT STILL DOES NOT CLAIM: the ask has not been walked by an actual volunteer from the written guide — the sentence is proven, the human following it is not. So the journey now reads: files protected from day one, once the household writes down the recovery code the box asks them for on every page. | Rows R-493 … R-500, R-534 … R-545 | | Off-site app-data capture covers the paths an app declares MANDATORY — on BOTH drive layouts | controller v0.197.0 | PROVEN-LIVE (2026-08-04) — and this row was OPTIMISTIC before it | On demo-hp, an app on the system-data fallback bound /mnt/sys_drive/userdata/media/books while the capture set looked in /mnt/sys_drive/felhom-data/userdata/media/books: the declared-mandatory directory was in no snapshot and the run reported ok (R-203). After v0.197.0 the capture log reads 1 mandatory path(s) and the file is listed inside the snapshot — restic ls -l latest --tag calibre-web → -rw-r--r-- 1000 1000 181 … /DRILL-SENTINEL.txt. Evidence: audits/DRILL-r201-offsite-recovery-2026-08-04.md §2 and the v0.197.0 CHANGELOG | WHAT IT DOES NOT CLAIM. It covers CAPTURE, not RESTORE: no file has ever been restored from an off-site backup after a wipe (R-201, ready to resume). It also does not claim the enrolled-drive layout was ever wrong — it was not, which is precisely why this survived. And a run that still misses a mandatory directory now reports incomplete rather than ok, so this row's guarantee is one the status can express. Widened 2026-08-06 (controller v0.205.0, R-234): the same verdict now also covers an app skipped ENTIRELY — until then a missing declared FOLDER made the run incomplete while an app with no recovery unit at all still reported ok, so the smaller gap moved the verdict and the bigger one did not. This row still claims CAPTURE, not that a newly-selected app is protected by the next run — for a DEPLOYED app it is (the run's own pre-dump phase writes the unit, measured 2026-08-06), and for an undeployed one it is not and the card now says so || Recurring offsite (PBS) whole-guest backups actually LAND, and RESTORE — local daily + offsite weekly as scheduled work | agent v0.97–0.103, controller v0.174/0.175, hub v0.76.0, host-install 1.20.0 | PROVEN-LIVE (2026-07-26) | audits/SPIKE-r82-phase0-2026-07-26.md; per-repo CHANGELOGs/REPORTs. Restore round-trip on demo-hp: --selftest=restore-test against felhom-pbs:backup/ct/9201/2026-07-26T15:42:42Z → pass:true, verified:"boot+running", mount_parity:"ok" (mp0=/var/lib/docker 50G, mp1=/mnt/sys_drive 20G, mp8/mp9 throwaway stand-ins for the archived binds), source_tier:"pbs", 4m5s restore+boot+verify+teardown, scratch band clean afterwards and the live guest untouched. | This row is distinct from the DR-tier row above, which proves ACTIVATION, not ARRIVAL. That tier was PROVEN-LIVE as applied since 2026-07-21 while demo-felhom held ONE snapshot (a healing artifact) and demo-hp held zero, ever — "applied and empty", the R-39 shape one level quieter. What earns PROVEN-LIVE here: (a) demo-hp's FIRST EVER offsite backup landed (4.25 GB into a verifiably empty namespace); (b) it restores into a bootable, mount-complete guest — mount_parity is the non-hollow half, since a boot-only verify cannot see a missing data volume; (c) the multi-tier quiesce ran through the real UI endpoint (POST /api/guest-backup/trigger, authed+CSRF) and produced exactly ONE stop/start pair with BOTH backups inside it — quiescing 1 stack(s) 17:01:39 → local done 17:02:56 "next tier may start (app still quiesced)" → felhom-pbs snapshotted 17:03:06 → unquiescing 17:03:06. App downtime 1m27s for both tiers, and the app came back healthy. Known gaps, recorded not hidden: the SCHEDULED restore-test still only selects the PRIMARY tier, so the offsite tier is never AUTOMATICALLY restore-tested (the manual/selftest path is proven, the unattended one is not); the hub infers "PBS ⇒ weekly" from storage TYPE rather than a reported cadence; the installer-default fleet flip awaits a full weekly cycle. → R-82 | | Restore-proof is UNATTENDED — the scheduler covers EVERY tier, follows the BACKUP rather than the clock, and a failure is heard | agent v0.104.0 → v0.121.0, hub v0.77.0 → v0.91.0 | PROVEN-LIVE (2026-08-03) | per-repo CHANGELOGs; backlog/SPEC-r85-phase4-5-2026-07-26.md. Unit red-proofs for tier rotation, restart-survival, the one-heavy-op gate, failure-emits-an-event, and newborn silence. | The row above is earned by a MANUAL --selftest=restore-test; this one is about the SCHEDULED path, and the distinction is the whole point. Before R-85 the scheduler could only ever see cfg.Backup.BackupTarget(), so the offsite tier was never a candidate — and a failed restore-test was a [WARN] line with no event at all, which was true for the LOCAL tier that WAS being tested. Now: oldest-first rotation across every configured tier (operator ruling 2026-07-26), persisted so it survives a restart; a restore-test joins the host-wide one-heavy-operation gate so it never contends with a backup over the same tunnel; and the hub raises two DISTINCT operator-tier signals — restore_test_failed (broken now) and restore_test_stale (unverified, not known-broken), anchored on R-81 so a newborn box never alarms. Why this was NOT PROVEN-LIVE until now: rotation had not been observed selecting both tiers across consecutive UNATTENDED cadences — at a 24h cadence a multi-day window — and a single passing run proves the code path, not the schedule. → R-85. R-86 (agent v0.121.0 + hub v0.91.0, 2026-08-03) replaced the schedule and the observation became possible in one afternoon, because what has to be observed is no longer a multi-day rotation but a RULE: a tier is due when its newest archive that has settled ~24 h has not been proven. UPGRADE EVIDENCE — a real unattended run on demo-felhom (2026-08-03), triggered by DUE-NESS, not by a timer. The scheduler's own log: 15:14:38 restore-test tier is DUE … target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-07-28T04:49:43Z … reason="newest settled archive … has not been proven" → proxmox-backup-client restore --crypt-mode=encrypt under the agent's own token → 15:25:08 gate decision class=guest_destroy guest=990000 allowed=true → 15:25:14 scratch guest torn down → 15:25:14 backup: scheduled restore-test passed archive=felhom-pbs:… duration_s=635.1. A 14.5 GB encrypted offsite archive pulled from ep0 over the WAN, restored, booted, verified and destroyed in 635 s, unattended. The three things a timer could not show, all verified after it: the state names THAT archive ({"felhom-pbs":{"archive":"…2026-07-28T04:49:43Z","proven_at":"2026-08-03T13:25:14Z"}}); a second evaluation reports due=false … is already proven and runs nothing; and an agent restart runs nothing, which is the defect a person actually noticed — every deploy used to restart the timer. Teardown verified at all three layers: guest absent from pct list, zero 990000 volumes in lvs, and the hub-side restore_tests[] entry deliberately RETAINED (it IS the proof the staleness check reads). R-189 (agent v0.122.0, 2026-08-03) closed the reporting half of this row, and it was a REAL gap in the evidence path: the proof above reached the hub only because no restart intervened — restore_tests[] came solely from an in-memory store, so the 15:25:14 PASS was in fact LOST when the agent restarted 2 m 43 s later for a deploy (0 restore-tests on the next two host-reports). Under per-archive due-ness the box would not have repeated the work for a week. The persisted per-tier proof (with the archive, and now the tier) is merged into the report, so a proof survives a restart — one entry per tier, newest wins, and a record that cannot be described honestly is not emitted. THE HOST TIER'S HALF OF THIS ROW WAS OPTIMISTIC UNTIL 2026-08-03, and it is worth saying plainly. Every live restore-test cited here is on the OFFSITE tier. The HOST tier was not merely unproven on demo-felhom — it was unprovable, and on demo-hp equally: the agent's token had no ACL on /storage/felhom-backup, the storage both boxes configure as local_backup_target, so the content API answered {"data":[]} through the token while root listed three archives. The scheduler skipped it as "no settled archive yet" — which is exactly what a brand-new tier reports — so nothing ever said so (R-185). CLOSED 2026-08-03 — agent v0.123.0 + installer 1.24.0. The grant is applied on both demo boxes (the token now lists 3 and 4 archives), the installer's reuse arm grants on a pre-existing target, and the agent now asks whether it may read each tier it depends on instead of inferring it from an empty list: a critical degraded capability naming the storage and the missing role, which the hub alerted and emailed on while the box was still blind. THE HOST TIER IS NOW PROVEN-LIVE, UNATTENDED, ON BOTH DEMO BOXES (2026-08-04) — this row's optimistic half is cashed. Four SCHEDULED runs, nothing triggered by hand: demo-felhom host tier …2026_08_02-04_42_14.tar.zst passed in 83.8 s at 00:55, offsite …2026-07-28T04:49:43Z passed in 540.4 s at 06:55; demo-hp host tier …2026_08_02-04_49_29.tar.zst passed in 109.3 s at 02:05, offsite …2026-07-28T19:19:45Z passed in 300.1 s at 08:05. Each restored into a scratch guest, booted, verified and destroyed itself; pct list and lvs show zero 990000 afterwards on both, and both boxes' local-lvm returned to their pre-run figures (1.95 % and 30.83 %). Both boxes had BOTH tiers due at once and the ordering was observed live for the first time under R-86's rule: never-proven sorted first, so each box took its HOST tier, deferred the offsite one, and picked it up on the following evaluation six hours later — one heavy operation at a time, per box, without anyone sequencing it. The proofs reached the hub, which is R-189 carrying a host-tier entry for the first time: demo-felhom's report holds two entries, one per tier, and the local one can only have come from the persisted state because the in-memory store held only that morning's offsite run. SCOPE, stated because one box proving something does not make it a fleet property: this covers demo-felhom and demo-hp. The tester's box is untested and untouched. Evidence: felhom-agent/REPORT.md, felhom.eu/REPORT.md | | A restore-test can never fill the box's disk: it sizes the restore first (uncompressed), keeps off the tested guest's pool when it can, refuses what does not fit, and retries a failed clean-up on a timer | agent v0.133.0 (RELEASED, NOT DELIVERED), hub v0.124.0 | BUILT + red-proofed; the refusal PROVEN on demo-hp (2026-09-24); the full-test path NOT proven live | audits/r672-2026-09-24/C7a…C7b, audits/r672-2026-09-24/redproofs/C-* | The 2026-09-24 incident (R-672) filled demo-hp's pool and turned 9201 read-only. Live: both margins refuse 9201's 21.1 GiB restore (22.1 GiB free, needs 30.3). No full restore-test fits demo-hp under 80 % pool use, so a pass after the change is unproven; the scheduled test is OFF on both demo hosts until the agent is delivered (operator ruling). | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | PROVEN-LIVE (external teardown, incl. two real firings) | tests/VALIDATION-n100-rehearsal-2026-07-18.md — two live firings, both host-delete-first, on two different customers (demo-vm-felhom 15:49:57, demo-felhom 16:08:51): every leg ok (claim, db_purge, descriptor, hetzner, pbs), escrow acked separately, each completing in 8–9 s (hub-state.txt customer_resets). The Hetzner sub-account destruction is now verified against the live pool box — and produced the run's sharpest lesson: a sub-account is an access-control object, not a data object. Deleting it left its /home intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (a finding by S7's own criterion) and why RESET now needs a base-dir purge → R-32. Prior: hub v0.61.0 REPORT; ep0 live drill 2026-07-17 (throwaway drill-reset-01 with a real backup: deprovision deleted:true destroyed the namespace + backup group + token, idempotent re-run deleted:false, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. Live-clicked 2026-07-18 (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. R-25b CLOSED (hub v0.69.0, 2026-07-21): the Danger-zone DELETE is now the guided full-teardown cascade that runs this very sequence as its middle leg — see the row below | | Customer DELETE cascade (top lifecycle tier): one guided operator action → hosts → RESET → residue → purge; host rows deleted (custody DEMOTED), full external teardown (Hetzner repo destroyed, PBS namespace/token revoked, tunnel + zone removed), then the customer record and ALL escrow ciphertext purged | hub v0.69.0 | UNIT-PROVEN; live leg PENDING | hub/internal/web/customer_delete_test.go — leg ORDER observed from inside leg 2 (hosts already gone, customer row still present, custody still retained); 9 fail-closed gate cases each asserting zero mutations + zero external calls + no journal row; resume-after-external-failure converges; purgeEscrow custody semantics; preview leaks no secret. 5 red-proofs (ack gate, stale-preview gate, ONLINE-host gate, leg order inverted, purgeEscrow=true) | Three acknowledgements + typed customer-id + stale-preview check + ONLINE-host refusal, ALL before any write. Ruling-3 preserved BY CONSTRUCTION (leg 2 never sees a host row); custody purged EXACTLY ONCE, in leg 3. Coupling: hub-only — no agent/controller/catalog change; the cascade calls the same service paths as manual host-delete and standalone RESET, so their rules move together. v0.70.0 (2026-07-21): added the residue leg — GetCustomers() is REPORT-derived, so before it a fully deleted customer stayed on the Customers list and its report stream kept the staleness/offsite checkers alerting (live: demo-vm-felhom deleted 07-18, still emailing offsite_stale on 07-21). The leg also purges the credential-bearing appliance_registrations + selfbind_tokens. Ghost customers (config row already gone) are now deletable — 404 means "nothing here", not "no config row"; the Hetzner/descriptor legs record skipped_no_config. Gap: the end-to-end live leg on a scratch customer (external Hetzner teardown observed from outside) is not yet run | | Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling | installer | PARTIAL | DRILL-GL6-2026-07-08 Phase 1/5 (KEPT-vs-WIPED printed verbatim; drive data intact ×3); GL-4 code | Secret purge (GL6-F1 .bak residue) fixed v1.12.0; enrolled-drive mnt-*.mount units survive (GL6-F2, open); cluster-aware felhom_guests guard + saferemove cost warning missing → R-9 |

B. Apps & catalog

Scenario Components Status Evidence Gap / roadmap
Deploy an app from the catalog (env config, memory guard, health-aware progress) controller, catalog (~52 apps, images pinned) PROVEN-LIVE CAMPAIGN-2 T-DEPLOY-SET (7 apps, env config, health-aware); RERUN-p1p3 (×4 PASS) Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only
App lifecycle: start/stop/restart/update/logs/remove/redeploy controller PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: CAMPAIGN-3 proved remove removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets [] and a note. CAMPAIGN-2 T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in CAMPAIGN-3; remove WITH data: audits/R442-2026-09-13/; data behaviour: audits/SPIKE-app-update-2026-09-01.md (2026-09-01) Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical
Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor controller v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0, hub v0.112.0 PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery) — afternoon (audits/rulings-r472-r475-2026-09-13/): an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served from declared and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). Morning: scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and the restore walk (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app audits/slice4-2026-09-13/ (live/, redproofs/, gates/); design architecture/09-update-architecture.md §6.1 Tier-2-only precondition — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run. audits/DRILL-update-night-2026-09-21.md. On scratch guest 9202 (controller v0.261.0), against a private drill catalog so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back through its own front door (R-156) with a negative control on every readback: 14 proven, 3 failed, 4 inconclusive. What the PROVEN edges prove, precisely: the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. What the FAILED edges prove, and they are the more valuable half. adventurelog (a real upstream edge that migrates and then never serves), tandoor (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. adventurelog v0.12.1 → v0.13.0 applied nine database migrations successfully and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. That is this row's own promise, exercised on a real upstream edge rather than a staged one. AND THE NARROWING, which this row must carry because it is the same mechanism: the verifying phase trusts the .felhom.yml probe absolutely, and two of the 53 templates name a probe the app does not answer — tandoor (port 8080; it listens on 80) and zipline (/api/health; it answers 404 there, while the compose healthcheck in the same file uses /api/healthcheck and is green). For those apps a successful update is stopped by its own health wait: tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, with docker's own healthcheck green, and was then stopped by failAndHold and the household sent to a restore they did not need. R-618, P1. No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". Still true and unchanged: no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. Not measured on this venue, and named rather than assumed: every event and every customer mail. Guest 9202 runs hub.enabled: false and the notifier returns before it logs (R-620), so the whole "who was told" half of 08 was structurally unobservable tonight. THE NARROWING ABOVE WAS CLOSED THE NEXT DAY, 2026-09-22 — and re-widened the row. audits/PROBE-FIX-2026-09-22.md. All three wrong probes were corrected in the catalog (app-catalog-felhom.eu@793c4fb: tandoor 8080→80, wger 80→8000, zipline /api/health→/api/healthcheck) and red-proofed live on 9202 through the product in both directions: at the live pin all three read Nem egészséges / Not healthy on their own app page while docker reported every container healthy and the front door served a real page; after the real sync all three read Fut / Running with no redeploy. tandoor's edge was then re-walked with nothing else changed and ended done at +41.1 s, seed read back, where the identical edge had ended failed at +361.9 s with the app stopped — so the tally is now 15 proven, 2 failed, 4 inconclusive, and all fifteen are on the live catalog. A --fast catalog gate (check-probe-matches-compose.py) now refuses a probe that does not match the same service's own compose healthcheck, with four red-proofs and ten decoys including the no-PyYAML mode CI actually runs. WHAT THIS ROW STILL CANNOT CLAIM, and the reason is exactly R-96 rule 3: the guard is now shown correct for 47 of 53 templates. paperless-ngx's probe has never run on any box — no container name matches its stack name, so it is silently skipped and its badge can never go red (R-630); and five more cannot be judged statically, one of which (home-assistant) is right only because its check type cannot fail (R-631). An absent alarm is equally consistent with healthy and with never checked. And the sweep's ceiling, counted: 28 of the 53 templates have never been deployed by any drill (R-632). THAT CEILING WAS REMOVED THE SAME NIGHT, 2026-09-22 — all 28 walked (audits/DRILL-the-28-2026-09-22.md), so every template in the catalog has now been attempted at least once. 26 of 28 deployed, 6 proven, 5 inconclusive, 14 with no within-a-major edge upstream, 1 failed honestly and 2 undeployable — one of those (plant-it) by design, refused by the product's lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy, with the seed read back again — 21 restored, and 2 were correctly REFUSED with the sentence 07 §6.2 predicts for a class-A app whose local copy holds no file leg. AND THE NIGHT NARROWED THIS ROW AGAIN, in the place the probe work could not reach. paperless-ngx has no container matching its stack name, so no probe is ever built for it — and verifying does not skip: it waits out the full update.health_timeout and HOLDS, stopping an app whose three containers all read healthy. The controller's own words: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app, at +313.0 s. R-630, raised to P1. So "the update is guarded" is now shown correct for 47 of 53 templates, wrong for none, and actively harmful for the one template that has no probe at all. Two further limits on what this row may claim, both about STATE rather than health: a remove sent while a restore is still running reports success and leaves a container restarting with a live public route (R-633) — while the product already refuses exactly that clash for update and for restore, naming the blocking operation; and an app can be running, healthy and serving while recorded as deployed: false, in which state the product refuses to remove it at all (R-634). In both, a person needed a shell to clear what the product could not. ALL THREE ARE FIXED IN CONTROLLER v0.262.0 (2026-09-22), and the first is PROVEN LIVE. The stopped app: verifying no longer loops on a probe that resolves to nothing — it settles on container state, the way an app declaring no check is judged, and says which it did. Measured on paperless-ngx: the identical Update that ended failed at +313.0 s with the app stopped now ends done at +53.4 s, with no no probe container warning in the log because the explicit healthcheck.container resolved the target. The ghost: RemoveStack consults the backup side's Busy guard — which the product already applied to update and to restore — and then WATCHES the compose project for 25 s after down, removing anything that carries its label and answering verified: true/false, because down returning 0 is a request rather than a result. The unremovable app: the refusal now asks whether anything EXISTS (containers, a compose file, an app.yaml) instead of reading a flag. WHAT THIS ROW STILL MAY NOT CLAIM: R-634's MECHANISM — why deployed goes false while containers run — is not diagnosed; only the consequence is fixed. And a probe can be right about the port and still wrong about what a 200 means: romm answered 200 from nginx for six hours while its workers were OOM-killed behind it (R-635). "The update is guarded" has never meant "the new version runs".
A failed update is UNDONE by the box itself — the previous version back with its data from seconds before the update; the app is held only if that undo fails too controller v0.263.2 (the undo), v0.264.0 + hub v0.120.0 (the mail) PROVEN-LIVE (2026-09-23) audits/undo-live-2026-09-23/README.md (+ audits/undo-bakeoff-2026-09-23/ for the method). On scratch guest 9202, through the endpoints the UI invokes: docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite in a volume), each a real migrating edge failing a deliberately wrong probe, undone in 30–52 s with seeds written before the backup, after it and seconds before the press ALL read back through each app's front door, ledgers equal to before; page line in hu and en. Cut-off copy → HOLD (untouched), prefix in the box's language; power cut (pct stop) during undoing → resumed after boot and undone; a person's press after an undo → done; removal deletes kept copies. NARROWED 2026-09-23 night (audits/DRILL-night-2026-09-23.md Part D): a controller kill -9 in verifying → resumed after restart → undone, seeds read back (round 9); a 1 GB memory hog and a whole-box backup during verifying did not disturb it (rounds 6, 8). Two holes found: after ANY restore the app's volumes lose their compose label and the undo copies NOTHING (R-658, P1); and a held FILE-LEG app is told to restore from a copy the restore then refuses, with no other route on a box without an off-site tier (R-659, P1). Folder copy of NAMED volumes only (09 §3 decision 19); bind folders never touched. The household is told (2026-09-23, audits/undo-fleet-2026-09-23/): on guest 9201 the household and the operator each received ONE mail per app per outcome — undone and held, in Hungarian (vikunja) and, after one language switch, in English (glance) — with the app named in the subject; notification_log rows 937–946 all sent. Not built: the automatic caller (part 7) — this protects the manual button today. Leftovers: R-647.
The catalog holds only TESTED steps — an image move without a proven test record (bench + box, digests the registry still serves) is refused at push time catalog 6db08a5 (scripts/check-test-record*.py, ladder.py, upgrade-test.py --write-ladder) PROVEN (2026-09-23 night) audits/DRILL-night-2026-09-23.md Part B/C 16 decoy cases both ways + 3 red-proofs; the gate judged every one of the night's 12 published steps against the live registry and refused a wrong-digest control. 21 earlier moves backfilled from their records.
A box behind climbs ONE tested step per press, each with its own definition (09 §3 decision 14) controller v0.268.0 (stacks/ladder.go) + catalog 5ed599c (steps/<StepKey>.yml, gate rule 4) PROVEN-LIVE (2026-09-24) audits/ladder-2026-09-24/partD/10-romm-two-steps.json romm 5.3.0/11.4 → 5.3.1/11.4 → 5.3.1/11.8 in two presses on 9202, seeded account read back after each; page count 2 → 1 → none; spike confirmed v0.267.0 jumped (partD/00-spike.json). 6 unit tests + red-proofs.
A file app comes back WHOLE from the second drive — its files by four rules (never delete, never overwrite a newer file, bring back what is missing, keep an older replaced file beside it), then its settings and database (09 §3 decision 26) controller v0.269.0 (backup/tier2_whole.go) PROVEN-LIVE (2026-09-24) audits/night-2026-09-24/A1/21-readback.json, A1/30-held-then-whole.*, E/round-03.json nextcloud on 9202: 1 file brought back, the household's newer edit kept, 137 unchanged, 3/3 volumes + 1/1 database in 38 s; again from a hold that named the second drive, and again after a power cut mid-update (round 3). R-538's refusal of the unit-only restore stays.
A crash-looping or out-of-memory app is STOPPED by the box, the household and operator are told, and Start gives one more try (decision 28) controller v0.269.0, hub v0.123.0 PROVEN-LIVE (2026-09-24) audits/night-2026-09-24/A3/21-gokapi-page-event.txt, A3/30-start-then-trip2.*, E/round-01/07/08.json gokapi stopped at 6 restarts in 10 min; Start lifted it, 9 fast restarts, stopped again with the support-informed sentence (hu + en); chaos hour: OOM storm stopped after a power cut (+185 s), a crash loop under a backup run (+116 s), a storm on a nearly full disk (+102 s).
A held app whose page says support is informed can only be removed KEEPING its data (decision 27) controller v0.269.0 PROVEN-LIVE (2026-09-24) audits/night-2026-09-24/A2/10-held-no-copy-keep-data.* the dialog reads keep_data_only; Remove with data or backups → 409 in both languages; the app stayed installed.
The box runs the exact TESTED image of a floating tag, and an installed app keeps its image until a guarded Update moves it (09 §6.4 part 6, box half) controller v0.269.1 (stacks/digest.go) PROVEN-LIVE (2026-09-24) audits/night-2026-09-24/B/21-floating-tag-0.269.1.*, B/01-compose-accepts-digests.txt redis:7-alpine: a newer tested digest → the „Frissítés elérhető" badge, the sync left the running file alone (v0.269.0's sync did not — fixed), Update pulled exactly that digest. Compose accepts tag@digest for all 25 ladder apps (37 digests).
Automatic updates at night — the update leg (09 §6.4 part 7) controller v0.271.0 PROVEN-LIVE on scratch 9202 (six simulated nights, 2026-09-24/25) — the demo boxes' first real night: see audits/DRILL-night-2026-09-25.md Part D audits/night-2026-09-25/C/, B/redproofs/ after the off-site leg on every path; one step per app per night; needs_person never, files_may_change only with a whole copy; a failed step not re-pressed until the catalog re-tests it (R-680); a power cut mid-step resumed and finished; a controller kill mid-step put back; switch off = nothing pressed; data read back after every night. Not claimed: W+5h with steps left and a FAILING off-site leg live (unit only), the full-system gate waiting live (9202 has no agent), resume of the leg after a restart (R-686).
What restart and update do to a deployed app whose compose file the catalog already moved controller v0.235.0 CHANGED 2026-09-06 — they NO LONGER upgrade it. The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it. Every lifecycle action ends in docker compose up -d, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — Manager.RestartStack says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. audits/SPIKE-app-update-2026-09-01.md §2, §3 No safety copy is taken by any of them — writeSafetyDump is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443.
Whether the box UPGRADES an app by itself, with nobody pressing anything controller PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back. A plain power cut does NOT upgrade: Docker's restart: unless-stopped restores the old containers and the reconciler logs no boot-orphaned apps (nothing to start). When an app does NOT return, Reconciler.Run (bootrecon.go:269) calls StartStack -> compose up -d and the app comes back on the NEW version, unattended (measured). 13 non-API call sites across 9 files reach up -d this way — not the five previously believed. audits/SPIKE-app-update-2026-09-01.md §2, §8 The drive-return gate (intermediary.go:222) and AppStopGuard.Recover (appstop_marker.go:283) call the same function; located by reading, not exercised live — stated as such.
Whether an app UPGRADE can be undone controller + catalog PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it. Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — "the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported". A 3-major jump is refused outright ("only possible to upgrade one major version at a time") and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. audits/SPIKE-app-update-2026-09-01.md §7 The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement.
Protected infra stacks can't be stopped/removed from UI controller PROVEN-LIVE CAMPAIGN-nomercy + RERUN-p1p3 T-SEC-PROTECTED (refuse stop/remove, stay Up) (Cited CAMPAIGN-2 T-SEC-PROTECTED was a stale-dryrun FAIL — corrected to the runs with a real server-side refusal)
What VERSION a box is running, and whether it is behind the catalog controller v0.234.0 + catalog 69761cf PROVEN-LIVE — both the record and the rendered badge. tests/VALIDATION-update-slice12-2026-09-02.md — on demo-hp 0.233.0, through a REAL production caller (bootrecon → StartStack → compose up -d → recordInstalledImages, no hand-set state): bentopdf recorded 1 service and bookstack recorded 2, keyed by compose SERVICE name, and all three digests match the ground truth read independently from the containers before anything was touched. catalog_since reached the box on the normal 15-minute sync. bookstack's two encrypted secrets are byte-identical across the write. Badge evidence: same file §4. v0.234.0 startup backfill, PROVEN LIVE 2026-09-03 on demo-hp: the record was stripped from privatebin (1 service) and romm (3 services) to recreate the pre-0.233.0 shape, the controller restarted, and the backfill re-seeded exactly those two — every digest matching the ground truth read from the containers beforehand — while logging 2 app(s) recorded, 7 already had a record, 0 left unrecorded. All 9 deployed apps then carried „Naprakész" on /stacks (ASCII fragments with a negative control at 0). On demo-felhom the operator's own case, OpenGist, now renders the badge. Why the backfill exists at all: without it the label never reached an app that simply runs, which the operator found the morning after v0.233.0. Unit side: installed_test.go + updatebadge_test.go, incl. a wiring test through a real RestartStack, an AST walk of all four call sites, and three companion red-proofs. THE UNEXERCISED LEG, NAMED: one badge STATE of four. „Frissítés elérhető" WITHOUT an age needs an app whose catalog_since is absent, malformed or future-dated, and all 53 now carry a valid one — unit-tested only (TestGroupF). The other three are live: „Naprakész" ×2 on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost, a deployed app with no record — absent is UNKNOWN and is not rendered as current; and „Frissítés elérhető — 52 napja" on both surfaces, the age being real arithmetic on bentopdf's catalog_since 2026-07-12. The behind state was staged by editing bentopdf's compose tag ONLY — no container restarted, no up -d — then reverted byte-identically (sha256 equal, diff empty); so the RENDER is measured and the syncer's own half stays measured separately in the spike. Searched with grep -oF ASCII fragments plus positive AND negative controls. The Frissítés/Újraindítás/Leállítás buttons are unchanged on the live page in the behind state. Absent means UNKNOWN, never current — a legacy app.yaml renders NOTHING, red-proved. No version number is shown to the customer and no registry is queried, so „Naprakész" CAN BE FALSE for the 23 floating pins (R-446). Nothing about updating changed: R-438, R-440, R-441, R-443 all stand. Reasoning: architecture/09-update-architecture.md; remaining slices R-447..R-452
An app's VERSION is frozen to what the customer has; only a deliberate Update moves it, while template CORRECTIONS still arrive controller v0.235.0 PROVEN-LIVE (2026-09-06) — by two REAL catalog pushes travelling the REAL 15-minute cycle, not a hand-edited file tests/VALIDATION-update-slice3-2026-09-06.md — a non-image catalog change REACHED the pinned app (08:01:51Z) with the container untouched; an image change did NOT (08:20:29Z); and the restart afterwards took 0.1 s, did not recreate the container, and never pulled the new image, against the spike's 18.3 s with a pull for the identical sequence before. The Update button still moved the version (pin advanced 17 s BEFORE the pull completed) and the teardown update returned the container to the baseline digest byte for byte. The freeze holds in BOTH directions. Unit side: internal/sync/render_test.go (the whole render table, incl. self-healing in BOTH branches), internal/stacks/pin_test.go (adoption never guesses; the update advances the pin BEFORE the pull), TestGroupG (the badge reads the catalog, not the frozen file), TestGroupH (an AST walk of cmd/controller/main.go asserting the seam, adoption, and their ORDER against syncer.Start()). Three companion red-proofs, each run, failing, and reverted. Operator ruling 2026-09-06, Option 1 (architecture/09-update-architecture.md §3.4). Nothing was added to the thirteen compose up -d call sites — most are repairs, and a repair that refuses to repair leaves an app down; they were made safe by removing the reason. pinned_images is INTENT, installed_images is an OBSERVATION — never fed from each other (R-166, one field over). Known limitations, all recorded rather than fixed: a frozen app is frozen WHOLE (§8.4); .felhom.yml keeps flowing, so a frozen app can get a probe for a newer version — false alarm, never data loss (R-458); and the Update button is still unguarded (R-448 is slice 4). Closes R-447, R-441, R-438
Catalog sync (git, 15 min) + orphan lifecycle + validation choke point (bad backup: block degrades to legacy, loudly) controller v0.132, catalog PROVEN-LIVE CAMPAIGN-2 T-SYNC-IDEMPOTENT; v0.132 LoadMetadata red-proofs
Lemez-egészség felügyelet: per-disk SMART kártya („Lemezek állapota") + degradáció-riasztás (Rendben/Figyelmeztetés/Hiba/Nincs adat) agent v0.94.0→v0.95.0, controller v0.169.0→v0.171.0→v0.215.0, hub v0.73.1 PROVEN-LIVE (healthy path + delivery + the severity wire). IMPLEMENTED, NOT proven-live: the Hiba-from-counters path (v0.215.0) — it has never fired on real hardware, only against the committed fixture's values in unit tests (R-332) 2026-07-25 (v0.95.0 + v0.171.0 — the SMART-coverage fix): the card on guest 9201 now shows BOTH real disks with real verdicts + human model labels — „AirDisk 512GB SSD" → Rendben (34°C) (the system SSD, via LVM/dm resolution) and „TOSHIBA MQ04ABF100" → Rendben (30°C) (the USB, via union-path SMART). /disks carries smart.health=PASSED + model_name for both. This reverses the 2026-07-24 „Nincs adat on a raw UUID" state (SPIKE-smart-coverage-2026-07-25.md had proven both disks answer smartctl -a -j PASSED but the agent never asked). Prior: verdict table (+≥90 red-proof); check first-run/degradation/recovery/UNKNOWN tests; hub allowlist test. Notification pipeline PROVEN-LIVE 2026-07-24 — a disk_health_degraded POST (the exact notify.PushEvent wire call) was 400-rejected by hub v0.73.0 and 200-accepted + „Operator email sent" by hub v0.73.1 No new smartctl load; feature-detect by payload presence → MinAgent floor unchanged; no sudoers/-d sat change. No global banner (deliberate). Agent v0.95.0 fixes: union-path SMART (Fix B) + LVM/dm whole-disk resolution incl. the builtin local on the LVM root (Fix A, SMART-only — never touches backing/durable_id) + model_name capture. 2026-08-14 — a genuinely failing disk HAS now been seen, and it broke three assumptions (audits/DIAG-smart-passed-trap-2026-08-14.md + two committed fixtures: raw smartctl -a -j and 406 smartd lines from ST3000VX010 S/N Z6A07P2G). (1) smart_status.passed is STRUCTURALLY incapable of failing on unreadable sectors — attrs 187/197/198 all carry thresh: 0 and a normalized value floors at 1 — so the drive read PASSED at 352 pending sectors and 1001 uncorrectable reads. (2) The alert it did produce carried severity "warn", which the hub coerces to info and never emails: the counterfactual is ZERO emails about this drive (R-328, fixed controller v0.215.0, and the warning-vs-warn pair proven side by side in notification_log on 2026-08-14 — sent vs no row at all). (3) The old check spoke once and forgot on restart, so between 8 and 352 sectors it emitted nothing. v0.215.0 adds the sustained/count/heat Hiba rules, persisted state and an hourly cadence. The verdict half of that arm remains unit+red-proof covered only — no live drive has reached Hiba from counters (R-332). SMART history/trending (hub-side) PARKED (ROADMAP R-73)
App crashes → customer notified (one event per transition, no flapping spam) controller v0.120, hub v0.48 IMPLEMENTED controller v0.120.0 (dead-app alerting, app_start_failed, one-event-per-transition red-proofs); CAMPAIGN-3 F11 surfaced the gap End-to-end crash→customer-email delivery never live-confirmed (6B deferred / 6C inconclusive: clean stop ≠ crash); anti-spam unit-proven
Post-deploy optional config (API keys etc.) with restart controller, catalog .felhom.yml IMPLEMENTED feature long-standing; config page renders (CAMPAIGN-2 T-PAGE-ALL is GET-only) The config-save+restart flow is exercised in no campaign (CAMPAIGN-3 explicitly skipped interactive app config). Demoted: T-PAGE-ALL is a page-render smoke test, not this flow
Backup classification: 13 bind-bearing apps carry mandatory/optional/excluded classes catalog, controller v0.132–133 PROVEN-LIVE SPIKE-backup-classification-2026-07-14, CAMPAIGN-6D/6E Remaining ~39 apps are legacy-class by design (unit-only offsite)

C. Protection & recovery (the product promise)

Coupling (2026-07-28, S-1). The failure → recovery matrix in 07-backup-architecture.md §8 is authoritative for which failure has which recovery route, who can invoke it, and what its measured RTO is. This section stays authoritative for per-capability status. Neither restates the other — rows below carry a → 07 §8 row n pointer instead of repeating the route. Where a row's status and the matrix's status differ in wording, the matrix is about the failure and the row is about the mechanism; that is not a contradiction, and both cite the same evidence.

The matrix's blank RTO/RPO cells are deliberate: no number is estimated anywhere. The two counts of Tier-2 app coverage that used to disagree (9/43/1 vs 7/45/1) were settled 2026-08-31 by measurement at catalogue 459766cb1639: A = 7 · B = 45 · C = 1 (07-backup-architecture.md §6.2, which also records which prior count was wrong and why). Take the number from there, not from memory.

Scenario Components Status Evidence Gap / roadmap
Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded DEGRADED rather than silently normal (installer Case A/B) agent v0.113, host-install v1.22.0 PROVEN-LIVE E2D-fresh-vm-2026-07-29 C1 (real 1.22.0 install, rc=0, Day-0 provision SUCCESS) + C2 (both DEGRADED lines verbatim, local_backup_target=local, install did not abort) Case A (a second drive already present at install) has never fired naturally — only Case B has
Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes controller v0.118 PROVEN-LIVE CAMPAIGN-2 T-BAK-FULL (pg+mariadb autodiscovered); atomicity CAMPAIGN-6B P4 + CAMPAIGN-6E B1/B2 (SIGKILL mid-write → only .tar.tmp touched, last-good byte-unchanged); DB restore CAMPAIGN-6D P-FAB (Cited CAMPAIGN-3 F7 is the finding of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 DB replay route → 07-backup-architecture.md §8 row 3
Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary controller v0.135 PROVEN-LIVE CAMPAIGN-6E-2026-07-15 (P-TIER2 deep-4 PASS), CAMPAIGN-6C Route + RTO → 07-backup-architecture.md §8 rows 1, 2, 3b, 4, 5. R-403 (controller v0.230.0, 2026-08-31) changes NO status on this row and that is stated rather than left ambiguous: the nightly copy now refuses to replace a complete unit package with an empty one (07 §8.2), which removes a way the route could be DESTROYED between uses — it does not change what the route can be relied on for, and every leg it covers is the same one it covered yesterday. Proven live on demo-hp: the same state that deleted 120 082 104 B on v0.229.0 preserved all 7 files on v0.230.0. R-102 CLOSED 2026-08-31 (controller v0.229.0): the copy's recovery-unit/ mirror is now restorable — „Teljes visszaállítás a másolatból" / POST /backup/tier2/unit-restore — and was proven live on demo-hp with the primary unit moved aside (audits/DRILL-r102-tier2-unit-2026-08-31/, §8 row 3b, 28.65 s). R-103 closed with it: the refusal that used to name a button on another page now offers the action on the row itself
Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping controller v0.134, agent, hub PROVEN-LIVE CAMPAIGN-6D-2026-07-15 (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); VALIDATION-offbox-storagebox-2026-07-09 (byte-perfect round-trip) Raw-data quota (SP-1) + retention regrouping (SP-2) are SPIKE-restic-snapshot-shape dry-run verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. Reinstall-continuity (controller v0.142.0, 2026-07-17): a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (wrong password or no key found) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. The live leg FIRED on its own during the 2026-07-18 rehearsal (tests/VALIDATION-n100-rehearsal-2026-07-18.md, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard classified it, pushed offbox_repo_orphaned, skipped the run and showed the card (16:58:14) rather than nightly-spamming a raw restic error; the operator-confirmed reset then moved the repo aside (never deleted) to .orphaned-20260718 and re-initialised (16:59:26→16:59:32), and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — the finding is that it had to fire at all (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). DIAGNOSE-offbox-repo-orphaned-2026-07-17 Route + RTO → 07-backup-architecture.md §8 rows 4, 10, 12, 15 (incl. the R-95 delete exposure and the R-104 stale-lock defect). 2026-08-04 (R-193/R-197, audits/SPIKE-offsite-credential-recovery-2026-08-04.md) — the row is NOT overclaiming and the guard is not the gap; the CADENCE is. This row already recorded that a recreated data volume orphans the repo, and the spike confirms the mechanism at source: WriteOffboxSecrets (offbox.go:392) mints a fresh 256-bit repo password whenever <DataDir>/offbox/repo_password is absent, and no automatic path ever consults the escrowed one — InjectOffboxPassword has exactly one caller, a web form a human pastes into. What was NOT recorded is that this now fires on an ordinary, planned, unattended guest rebuild, on every box: measured on BOTH demo boxes 2026-08-03/04 by comparing host_escrow.restic_pw_sha256 against host_escrow_superseded.restic_pw_sha256 (demo-hp 8e03eddf…→8a9e33aa…, demo-felhom 48741892…→c60c8bc7…), orphaning 15 snapshots / 40.9 MB and 36 snapshots / 1.14 GB respectively. demo-felhom is the important half: it kept its DELIVERY (a stale staged secret restored the target in 76 s) and lost its REPOSITORY anyway, with nothing marking the escrow stale for 13 h — escrow_stale is wired to ReissueCredentials, the one path that does NOT change the repo password (R-196), and absent from the rebuild path that does. Status unchanged: the classify-and-move-aside guard remains PROVEN-LIVE and correct, and is predicted (not yet measured) to refuse the 2026-08-05 run on both boxes rather than start a silent fresh history
Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live controller v0.134/134.1/135 PROVEN-LIVE (2026-07-20) CAMPAIGN-6D accept legs (immich end-to-end from offsite alone) 2026-07-19: audits/DIAG-immich-restore-2026-07-19.md finds no offsite path loads a DB dump — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "immich end-to-end from offsite alone" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume (immich_postgres_data is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" overclaimed scope: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → PARTIAL, scope-corrected. Evidence: audits/DIAG-immich-restore-2026-07-19.md (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). 2026-07-19, controller v0.148.0: the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but round 2 found it aborts against a running app (audits/DIAG-immich-restore-round2-2026-07-19.md, H4: the replay races immich's own schema repair; clip_index recreated by the app 2 s before the dump's CREATE INDEX). 2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47) — both restore paths now replay into a DB-ONLY window (StartStackServices brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. (The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard BackupStatus fix. R-47 shipped in v0.153.0.) 2026-07-20: the clean run HAPPENED — endpoint-level supervised reconstitute of immich from snapshot 49e7cb46 (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no already exists, operation reported SUCCESS, immich's own DatabaseService logged No schema drift detected twice, 11 assets active, 4/4 containers healthy, 231 public indexes. Operator confirmed the immich timeline renders correctly after the reconstitute (screenshot held, 2026-07-20). Evidence: felhom-controller/REPORT.md §4b. 2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE. The operator deleted the photos in immich own UI and emptied the trash (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. 40 file(s) placed against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets active, No schema drift detected, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: felhom-controller/REPORT.md 4e Route + RTO → 07-backup-architecture.md §8 rows 3, 4 — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107)
Manual .fab export/import: class-scoped capture, browser up/download, tunnel-proof chunking controller v0.125/128/130/136 PROVEN-LIVE CAMPAIGN-6D P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking CAMPAIGN-6B P2 (100 MiB via real CF edge, 120 MiB→413) Chunking proven at the real CF edge via curl --resolve; the rendered browser file-picker upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B finding; fix verified in 6D
Guest-loss DR: PBS restore with full-fidelity layout from archive, restore-test verification agent v0.75/0.76, PBS PROVEN-LIVE CAMPAIGN-2 T-P9-DESTROY-RESTORE (whole-guest pct restore of 9201 → running+healthy) + T-PBS-VERIFY (verify_state: ok, 13 snapshots); DRILL-GL6-2026-07-08 Phase 0d (restore-test mount_parity: ok) (Cited VALIDATION-newbox-restore is offbox restic file-restore, wrong tier — corrected.) Real offsite guest-loss round-trip still R1-blocked → S5 DR drill Route + RTO → 07-backup-architecture.md §8 rows 6, 8, 9 — measured 84–112 s local / 1101 s PBS into a scratch guest; a restore to a DIFFERENT host is unmeasured
PBS-DR secret self-heal on reused-peer re-provision hub v0.56 IMPLEMENTED hub v0.56.0 (pbsdrheal/reconciler.go, RestageHostPBSSecret, all §10 red-proofs); SPIKE-pbsdr-selfheal-2026-07-15 (root cause) Reconciler is scoped to one host (PBSDRHEAL_ONLY_HOST), not fleet-wide; already fired live hands-free on drill qm300 (07-15) — real-customer firing + fleet-wide widening pending
Box survives an unattended app or guest-network failure (a dead app member, a boot-orphaned app, a dead DHCP client) — it is noticed, and where safe it is repaired controller v0.156.0→v0.190.0, agent v0.92.1 PROVEN-LIVE (2026-07-21; boot-orphan leg rebuilt and re-proven 2026-08-02 — 6 of 6 hard resets, repeat count cited per N.5) All three legs exercised on the live demo box, operator-present, in one session — felhom-controller/REPORT.md + felhom-agent/REPORT.md (2026-07-21). Dead primary: docker stop immich-server 12:50:40 CEST → degraded 13 s later → exactly one app_start_failed + dashboard banner → restart → banner self-cleared (the 2026-07-20 shape that was silent for 18 h). Boot orphan: pct reboot 9201 → [bootrecon] 1 boot-orphaned app(s) found: [bookstack] → started in 1 attempt of 2, zero alerts (success inside the boot grace is silent); StartedAt proves Docker's unless-stopped did NOT resurrect it — only the sweep did, which also answers P1 and confirms the F5 hypothesis. Dead DHCP client: deliberate replay of the incident — kill -9 12:43:18 → detected on process liveness 57 s later while the lease was still live → healed 12:45:18 with the incident's verbatim invocation; the tunnel never dropped (cloudflared Up 29 hours), i.e. the outage was prevented rather than merely observed The gap this validation surfaced → R-55, now FIXED (controller v0.157.0, 2026-07-21). For a drive-backed app a customer's deliberate Stop did NOT survive a reboot — the boot bind gate recreated and started every deployed drive-backed app unconditionally. Pre-existing, not introduced by R-52 (whose own gate was observed correct). The gate now also requires the app to still HAVE containers, which is R-52's own existing-Exited vs absent predicate: a UI Stop is compose down and removes them. So "a stopped app stays stopped" now holds for drive-backed apps too — PROVEN LIVE 2026-07-21 (TASK-F Part 3, operator-present). immich was stopped through the real UI endpoint (compose down → 0 containers), calibre-web and bookstack left running, then pct reboot 9201: the gate recreated calibre-web and logged 1 drive-backed app(s) left stopped — zero containers means the customer stopped them on purpose; immich came back stopped, where the identical fixture had brought it back running hours earlier. Zero alerts, ~15 s to steady state. The static-guest half of the network leg stays deliberately out of scope → R-50. 2026-08-02 — the boot-orphan leg now rests on a RECORDED signal, not an inference (R-166, controller v0.189.0). Both the R-52 sweep and the R-55 gate above decided "the customer stopped this" from zero containers, which is also what a power cut mid-compose and an interrupted deploy leave behind — so two real faults were read as deliberate stops and stranded silently (R-157 mechanism B). The customer's intent is now written to app.yaml (desired_state) by their own action and read directly. PROVEN-LIVE on 9201 in three flows: a UI Stop persisted stopped and survived a controller restart with the app still down and NOT listed as a candidate; an app recorded running whose containers were removed out-of-band was recovered by name ([bootrecon] 1 boot-orphaned app(s) found: [calibre-web] → started in 1 attempt) — the case that was invisible before; and a legacy app.yaml with no field was skipped exactly as before and was never inferred to be stopped. Interrupted app-data operations are covered separately and are NOT proven-live — backup.AppStopGuard restarts apps left stopped by a killed volume dump / offsite reconstitute / .fab export, and that leg is unit-proven + red-proofed only (killing the controller mid-backup on a live box was not exercised): IMPLEMENTED, not PROVEN-LIVE. R-170 and R-171 closed the same day (controller v0.190.0). R-157 mechanism A — the sweep observed ONCE at T+5 s, while docker was still restoring, and never re-checked (3 of 6 hard resets). It is now a settle-then-sweep window: sample every 5 s, settled after 3 identical samples, ONE sweep at the end, terminating on settled or a 50 s budget (sized so settle+budget+one retry stays inside the 90 s dead-app grace; a test rejected 60 s at 95 s). Repeat count, per this map's own rule: 6 of 6 hard resets on the shipped build brought every app back, and an app the customer had stopped stayed down in all 6 (window settle times 10/40/10/10/15/15 s — i.e. it routinely waited 2–8× longer than the old fixed 5 s). A same-app before/after on one box is the sharpest evidence: the pre-fix window logged no boot-orphaned apps for calibre-web at 18:08:35; the fixed one found and recovered it at 18:18:50. R-170 — shouldRecreateOnBoot now reads intent too, so the two boot gates agree; proven live in one reboot (calibre-web running+zero containers recreated, immich stopped left alone). R-171 — a regression v0.189.0 introduced, found by reading the diff and CONFIRMED on hardware before any fix was written: the sweep started an app whose drive was absent, burned both attempts and raised a false dead-app alarm. The write hazard was blocked only by an ACCIDENTAL filesystem permission (host-root-owned mountpoint + unprivileged guest) that no code owns and no test pinned — which is why it was fixed rather than noted. New fail-safe bootrecon.StartGate (cannot determine ⇒ do not start), also covering quiesce and in-flight app-data operations. One defect in the fix itself, found by live validation and not by review: the window sampled the Manager's 10 s-refreshed cache, so "settled" could mean "the cache did not update"; sampleBootFleet now refreshes first. Evidence: audits/DIAG-bootrecon-drive-absent-2026-08-02.md, felhom-controller/REPORT.md
Crash/power-loss mid-backup/mid-migration → self-heal on next run controller, agent PROVEN-LIVE CAMPAIGN-6D P5-REST (SIGKILL mid-offbox → auto-restart ~15s, run marked failed not false-success, no stale lock); CAMPAIGN-6E B1-B3 (Cited CAMPAIGN-2 T-RBT-* legs were empty / auth-hollow — corrected.) Live mid-migration crash→self-heal is the weakest sub-claim (P5-REST is mid-backup)
An app can be withdrawn from the catalog without orphaning the customers running it (available / hidden / abandoned) controller v0.158.1, catalog metadata PROVEN-LIVE (2026-07-21) TASK-F Part 1. Verified on 9201 through the real endpoints: lifecycle: abandoned arrived via the normal catalog sync; plant-it renders 0 times on the Alkalmazások page (control app renders 10); a direct POST /api/stacks/plant-it/deploy → HTTP 409 "Ez az alkalmazás jelenleg nem telepíthető."; the app page carries the permanent notice and offers no Telepítés button. felhom-controller/REPORT.md (2026-07-21) Deployed instances keep FULL function in every state — lifecycle governs what is offered, never what runs. Orphan detection deliberately never sees the field (red-proofed): a withdrawn template stays in the catalog tree, or every deployed instance would read Elavult and be offered deletion. Unknown values fail OPEN; the deploy gate fails CLOSED. R-57
Box survives a site/network change (relocation, different subnet, DHCP re-lease) with the control plane intact agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap PROVEN-LIVE (2026-07-25) R-50 SHIPPED and deployed to the whole fleet. The control plane now rides a host-internal, portless island bridge (vmbr9, 169.254.253.1/30↔.2/30) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original bind: cannot assign requested address daemon-death) + cold-reboot survival (SPIKE-island-bridge-2026-07-25.md), the migration runbook run verbatim (RUNBOOK-island-migration.md), a fresh provision auto-attaching the island net1 (A4), and the live migration of both demo boxes (demo-hp + demo-felhom, 2026-07-25) — island /storage HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), apps served throughout (0 container restarts), hub reporting 0.96.0. Origin: audits/AUDIT-vacation-remote-ops-2026-07-20.md — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box Fleet: DONE. Remaining: R-74 — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped
The customer is warned BEFORE a filesystem fills — per filesystem, in Hungarian, naming the drive and the free space, edge-triggered controller v0.191.0/.1/.2, hub v0.89.0 (R-167, decision D-c) PROVEN-LIVE (2026-08-02) audits/SPIKE-r165-mp1-merge-2026-08-02.md (context) + felhom-controller/REPORT.md. Exercised on guest 9201 against a REAL filesystem (/mnt/sys_drive filled with fallocate): disk_warning at 90% used / 4.7 GB free → hub notification_log `customer disk_warning
A failed per-app Tier-1 backup reaches the OPERATOR — EVERY failing app, in ONE mail per run, and every failure recorded whether or not it is mailed controller v0.194.0, hub v0.90.1 (R-158 → R-167 → R-182) PROVEN-LIVE (2026-08-03) felhom-controller/REPORT.md. Two real capture failures on guest 9201 (mkdir …/backups: permission denied) → both accepted and stored by the hub, `operator recovery_unit_capture_failed
A local backup is bounded by the box's FREE SPACE, not by a partition set at build time — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime golden build-golden.sh v3.0.0, agent v0.120.0, controller v0.193.1 (R-165 / D-a / B2, completed by R-181) PROVEN-LIVE (2026-08-03) — BOTH halves REPORT.md (R-178 reinstalls) + audits/SPIKE-r165-phase0-2026-08-03.md (P1/P2/P3) + the bake transcript. The golden bake is real evidence and is cited as such: build-golden.sh v3.0.0 produced including mount point mp0 ('/var/lib/felhom') with no mp1 line at all, and its own guards printed /var/lib/docker is a real mount, /mnt/sys_drive is a real mount and both paths are ONE filesystem. Archive published (registry HTTP 200, sha 54e2a4c4…). The B2 floor is unit-proven with 3 red-proofs and live on 9201 The row's FIRST clause is now PROVEN-LIVE; its SECOND is not, and they are separated deliberately. Proven (R-178, 2026-08-03): "a local backup is bounded by the box's FREE SPACE, not by a partition set at build time" — both demo boxes reinstalled from this golden, by two different supply paths (demo-hp --golden <local volid>; demo-felhom the normal manifest route with verified sha256 54e2a4c431daf580… matches the hub manifest), each showing mp0 at /var/lib/felhom with no mp1, both consumer paths real mounts on ONE filesystem (stat -c %d = 64519 on all three), 3/3 reboots each, and claim → deploy → backup → restore with a planted marker returning byte-identical. Space available to a recovery unit measured at 65 GiB / 233 GiB, against the 19 GiB / 45 GiB those boxes' mp1 slices offered. NOT proven — and measured FALSE in part: "a capture that would exhaust it is refused per app rather than allowed to stop the container runtime". The floor fired live for the first time (demo-hp 06:40:03) and does refuse per app, delete nothing, and alert — but it is checked only in captureAllRecoveryUnits, while runVolumeDumps writes the bulk with no floor check at all, so the leg that exhausts the volume is the unguarded one; and the refusal's claim that the previous unit is untouched was measured false (a 182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). → R-181, CLOSED THE SAME DAY (controller v0.193.0 + v0.193.1) and the second half is now PROVEN-LIVE TOO. The reserve became a per-app, per-run ADMISSION decision taken before the app's FIRST write and covering all three legs (DB dump, volume dump, capture) — they write under one per-app root, which is what lets one verdict cover them honestly — and it gained a size term, so an app is no longer admitted at 96% and then allowed to write 2 GB. Re-proven by filling demo-hp deliberately, once for EACH term, using the method that found the defect. Headroom @ 08:59:46 (906 MB free / 99%): both apps refused, the whole backups/primary tree byte-identical — TREE_SHA 111d1760c18d3440f700634ab325f8b8 before and after, opengist's tar still at its original 182,272 B; no Stopping <app> for safe volume dump line at all, which is the positive-by-absence observable that matters because that line IS present in the 08:58 baseline run; 0 volume dumps; one alert per app, HTTP 200. Space freed, re-run @ 09:01:33 → both captured normally. Size @ 09:03:00, reproducing the original sequence with a real 2 GiB file in opengist's volume (previous tar 2,147,666,432 B, the exact figure the defect was measured at) and the filesystem at 91% used / 2.9 GB free — both headroom terms deliberately clear: opengist refused (size) while privatebin was ADMITTED and dumped normally, proving the term is per-app rather than a global halt. The refusal's wording was NOT weakened to fit — the behaviour moved so the wording became true, and it is verified by tree fingerprint rather than by reading the log line, which is what lied. The fallocate instrument was re-proven on the rebuilt box before use (5 GiB step moved guest df while thin-pool data_percent held 36.83 → 36.83), and teardown returned the pool to 29.43%, below its own baseline. The golden is now VOUCHED (2026-08-03, hub Artifact manifest set: … golden=0.192.0), so fresh installs pick up the merged layout. Every box in the field that has not been reinstalled is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired -sysdata-grow into the single grow so an older felhom-host-install.sh still provisions the same total capacity
Soft-quota: usage bar, pre-push enlargement block, customer notification controller v0.109/134, hub v0.41/55 PROVEN-LIVE 6D/6E; hub OffsiteChecker
A customer (not the operator) performs a restore via UI alone all MISSING (as evidence) — Alpha will produce this; script it into R-3. 2026-07-19: the C6 evidence attempt ran and found a product gap instead of evidence — audits/DIAG-immich-restore-2026-07-19.md. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING by finding, not by absence of attempt — the rehearsal system working, not failing. 2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0 (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. 2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped (audits/DIAG-immich-restore-round2-2026-07-19.md): a deliberate run from snapshot 49e7cb46 did recover all 11 assets (status=active, files resolve), but the operation reported failure and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. 2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED. (The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.) The C6 drill ran end-to-end through the UI: photos deleted, trash emptied, the full files+database restore pressed on /backups/restore, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: felhom-controller/REPORT.md 4e. Residual: the run was performed by the OPERATOR, not by a customer — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means trash, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash and verify the app's content, not the file count Lane split → 07-backup-architecture.md §3: this row is Lane 1 (customer, unassisted). §8 rows 1–5 are the routes it would exercise

D. Storage & devices

Scenario Components Status Evidence Gap / roadmap
A second drive appearing is OFFERED as the backup target; accepting moves it; registration and the drive gate confer no role by themselves controller v0.186.0, agent v0.113 PROVEN-LIVE SESSION-C-2026-07-29 C4: offer rendered with data-path, decline path proven (target stayed local, no felhom-backup storage, agent.json unchanged), restart_required:true, agent did NOT self-restart, wrapper created the storage at the drive's OWN mountpoint Accept was driven through the endpoint the button POSTs, not a browser click — no browser automation exists on DooPlex
Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts controller, agent v0.87 PROVEN-LIVE DISPOSITION-ia-finding2-systemdisks-2026-07-13 (legacy EFI+LVM host, root not offered, byte-identical); enroll/format live in storage-lifecycle-acceptance-2026-06-15 (E10 re-enroll, data intact); agent fence self-test refuses /dev/sda (Cited CAMPAIGN-2 T-STG-ENROLL/SEC-FORMAT were auth-hollow CSRF-403.) Fresh-USB wizard enroll+format through the customer UI PROVEN-LIVE (controller v0.141.0, 2026-07-17): a 64 GB scratch USB driven through the real /api/storage/init endpoints (login+CSRF) → confirm → detached format (~27 s mkfs) → mount → register → mounted+registered at /mnt/felhom-drives/scratch1. F6 (initialize-to-usable) now covered: the wizard runs the chain as a detached, disconnect-safe, pollable job (3-step progress) with an agent format-status poll for a slow mkfs 2026-07-26 — a SILENT failure class on the channel every agent-backed capability depends on (this row, data migration, USB enrollment, guest RAM, quiesce/PBS) is now DETECTED (controller v0.173.0, R-77). No row status flips. controller.yaml and bootstrap.json could disagree on local_api.endpoint indefinitely with no signal: the R-50 island migration rewrote the latter, the fleet kept dialling the former, and for 17.5 h the only alert was a generic "agent unreachable" that read as an infrastructure blip. Drift now raises its own event type (local_api_endpoint_drift) naming both values. It is DETECTION ONLY — the authority ruling is R-78 — so the class is now loud, not prevented. Evidence: audits/DIAG-agent-channel-2026-07-26.md.
Data migration between drives (all / per-app), crash-safe controller PROVEN-LIVE CAMPAIGN-6C 4P-5 (scope=app round-trip, byte-identical); storage-lifecycle-acceptance-2026-06-15 (two migrate-all runs via dashboard UI, sha256 byte-identical) (Cited CAMPAIGN-2 T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS
NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated controller v0.113–117, agent v0.81/84/85 PROVEN-LIVE SPIKE-nas-verify-2026-07-11, SPIKE-nas-dsm-2026-07-11, CAMPAIGN-3-2026-07-11 (boot/reassert fixes)
Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace controller v0.187.0 PROVEN-LIVE (2026-07-30) audits/R108-network-app-namespace-2026-07-30.md. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (POST /api/storage/migrate-app, authenticated + CSRF): pre-fix v0.186.0 the target was never examined — both a NAS-shaped and an unregistered path passed straight into MigrateApp and failed only on the app name (409); post-fix v0.187.0 both are refused 400 with a Hungarian reason, while a real local drive still reaches MigrateApp (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no migrated_to, nothing decommissioned, app HDD_PATH unchanged, no backups/ on the share Closes R-108 and UNBLOCKS D5 (07-backup-architecture.md §7.3, §10.1). The share-root :rslave FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: /mnt/felhom-drives holds both kinds, so an unregistered path under it is un-classifiable and refused. Not exercised: the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil stackMgr) but were NOT live-fired — only migrate-app was. .fab-export-onto-NAS remains open (→ R-126)
A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5) controller v0.188.0 PROVEN-LIVE (2026-07-30) audits/D5-drive-alone-restore-2026-07-30.md, 07-backup-architecture.md §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (POST /api/stacks/{app}/deploy → POST /api/backup/run → POST /backup/restore): AdventureLog (SECRET_KEY data_key + DB_PASSWORD) restored with the guest's app.yaml moved aside → secrets recovered=2/2, 27.6 s, Restore-from-unit completed. The observable is the DATA, not the exit code: the app itself then read the seeded customer row over TCP with its own credential (connected_as=adventurelog over_TCP=True), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held no .sql dump, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. The withheld half is proven too: Grafana's type: password admin login was live in its container and ENC: in the guest, yet appeared in 0 files anywhere under the backup namespace, and the unit's app.yaml header names it as withheld Ruling (operator, 2026-07-30): type: secret travels, type: password NEVER does, minus the nonPortableSecrets code register (vaultwarden/ADMIN_TOKEN). Plaintext on the drive, like the data — defensible only because the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. The brief's own proposal (data_key-only) was tested and rejected: the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → R-127) and a DB password is not resettable in practice (POSTGRES_PASSWORD is ignored once PGDATA is non-empty). Precedence: the unit wins over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. Not exercised live: the withheld-class O4 regeneration on restore (unit-tested only). Tier-2's own cross-drive copy of a secret-bearing unit — EXERCISED LIVE 2026-08-31 (controller v0.229.0): docmost restored from /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit with the guest's app.yaml moved aside AND the primary unit moved aside, secrets recovered=2/2 (APP_SECRET, DB_PASSWORD) taken from the MIRRORED unit's compose/app.yaml; the guest's app.yaml was rebuilt from it at 0600 and the app then read its own rows over TCP with its own credential. Evidence: audits/DRILL-r102-tier2-unit-2026-08-31/phase8-scenarioD-guest-appyaml-aside.log. Venue was a fixture-class scratch guest, correct per runbooks/target-selection.md since D5's claim is about restore CODE, not the install path
A restore SAYS what it returned, and refuses what it cannot do — the four restore-surface truth defects from the 2026-08-21 drill controller v0.226.0 (R-353, R-357, R-358, R-360, R-396) PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only audits/evidence-r353-r360-live-2026-08-30/live-validation.txt, controller CHANGELOG.md v0.226.0 + REPORT.md. Driven on demo-hp through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). R-353: the sentence read off the customer's own wizard page — A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult. with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). R-360: in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and a planted canary file survived. R-358/R-396: a mode=unit restore wrote {"schema":1,…,"full":false} at mode 0600 with no .tmp left, and the gate logged scratch holds a UNIT-ONLY restore … place-to-live stays closed WHAT IS AND IS NOT CLAIMED, split deliberately. R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (SetOffboxFreeFn, SetOffboxSizer, and the new SetOffboxLatestSnapshotFn) whose central assertion is that StopStack was never called. R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either, and the reason is stated rather than glossed: no app on demo-hp still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. This row is about the MESSAGE and the REFUSALS, not the recovery mechanism — 07-backup-architecture.md §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so
The box PROVES its own off-site copy still HOLDS something — a backup that is intact and EMPTY is caught without a person controller v0.231.0 (R-87) + hub v0.110.0 PROVEN-LIVE (2026-08-31) for the judgement, the alarm, the cleanup, the rotation, the read-only guarantee and the skip-if-busy hazard control; the SCHEDULED FIRING is IMPLEMENTED only documentation/tests/r87-offsite-proof-2026-08-31/. Driven on demo-hp through the endpoint the debug button invokes. The failing case was produced and caught: a hollow unit — compose declaring opengist_data, manifest declaring nothing — was pushed to the live store, and the proof returned verdict:"fail" with volumes_expected_none_captured: opengist_data, emitted exactly one offsite_proof_empty at severity error, accepted by the hub HTTP 200 (which is itself the proof the allowlist entry landed — an unallowlisted type is 400'd and vanishes), and deleted its scratch. The passing case was proven five times over (bookstack, calibre-web, docmost, kimai, opengist), 2.2–4.0 s each, and the customer's own verification copies were untouched throughout. The read-only guarantee was measured with a positively-controlled lock sampler — it saw a lock appear and vanish across a real restic check, and zero across the proof, including a direct 6× test of the snapshot-lookup argv. The skip-if-busy control fired live and unplanned: a proof launched while the off-site backup run held the flag returned skipped:true, duration_ms:0 with no verdict and no alarm ⚠ WHAT A PASS MEANS, AND WHAT IT DOES NOT. It means the newest off-site snapshot of ONE app contains what that app is supposed to have — judged from the unit's own captured compose, not from the live box. It does NOT mean a restore puts data back into a running app: the proof restores to a throwaway folder, looks, and deletes, and 07 §8 matrix row 4 is deliberately NOT moved. It also does not vouch for the BYTES — nothing available can: restic 0.14.0's restore --verify passed a byte-level corruption with size and mtime preserved (measured, 131 ms on a 213 MB tree), and the unit manifest hashes 4 918 B of a 213 231 242 B unit (R-409). 2026-09-01 (R-414, controller v0.232.0): it can now run on a box with NO registered data drive. The first unattended firing, on demo-felhom, REFUSED — "nowhere to restore to" — because that box has storage_paths: [] and the scratch resolver never consulted the system data path. A unit-only restore now falls back there (where a driveless app's unit already lives, 07 §7); a full restore still refuses, because the SSD is a state-only tier. And a proof that cannot start now records cannot_run instead of nothing, so last_proof_result is never ABSENT — absent already means a controller too old to have the feature. PROVEN LIVE on demo-felhom 2026-09-01: verdict:"pass" on 61e9cf30 in 2.117 s, recorded, and the scratch deleted. The nightly firing IS now proven — it ran unattended on demo-hp at 05:30 on 2026-09-01 (bentopdf PASSED on 9d002b38 in 2.315s), which this row previously listed as implemented-only. The job is confirmed REGISTERED on demo-hp (Daily job offsite-proof scheduled for 2026-09-01 05:30 CEST), which is not the same claim, and the fleet is on 0.230.0 until a golden carries 0.231.0
The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable controller v0.228.0 (R-359, R-397, R-399) PROVEN-LIVE (2026-08-30, re-proven at FULL DEPTH 2026-08-31) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only documentation/tests/r359-integrity-2026-08-30/. Driven on demo-hp through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (negative control first), then one pack corrupted; the live store was checked read-only in 35.0 s; and the notifier fired end to end — Event pushed: backup_integrity_ok (info). The hazard control was observed live: a second check fired while the first held the single-writer flag returned skipped:true, duration_ms:0 — it never ran restic at all ⚠ WHAT AN ok MEANS — CHANGED 2026-08-31 (R-399, controller v0.228.0): the check now RE-READS THE DATA. The default is --read-data-subset=100%, so an ok means every stored byte was downloaded and re-hashed, not merely that the catalogue hangs together. The reason is measured, and it is why the default must not be turned back down to save four seconds: a pack corrupted WITHOUT a size change made a structure check return no errors were found, exit 0, while every --read-data* form caught it. Cost curve on 134.3 MB: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s — and those do NOT extrapolate, which is why v0.228.0 ships a slow-check WARN (R-401) rather than a rotation schedule. off returns a box to structure depth. PROVEN-LIVE at the new depth 2026-08-31 on demo-hp, endpoint-level, with the restic argv observed from the guest: default → … check --read-data-subset=100%, 38.7 s; off → … check, 34.7 s. The weekly firing at the new depth is IMPLEMENTED only — the job is confirmed REGISTERED on BOTH demo boxes (Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST), which is not the same claim. demo-felhom reached 0.228.0 by SELF-UPDATE on the 2026-08-31 floor raise and re-registered the job itself, so the depth change is on the fleet and not only on the box that was deployed to by hand. This is a readability check and NOT a restore-test — R-87 remains open and the two are routinely conflated because their register rows are adjacent
USB drive enrollment + unplug detection + recommission controller, agent PROVEN-LIVE storage-lifecycle-acceptance-2026-06-15 E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); CAMPAIGN-4/6A (3 USB re-establish across device-letter reshuffle) (Cited RUNBOOK-usb could NOT complete a wizard enrollment; CAMPAIGN-2 legs were auth-hollow.) Fresh-USB wizard enrollment specifically still unproven
Decommission (migrate-first and anyway-paths), eject agent, controller PROVEN-LIVE storage-lifecycle-acceptance-2026-06-15 E9 (decommission-anyway → bind detached, parent mp untouched, reboot-safe) + E12 (eject drive holding all apps) (Cited CAMPAIGN-2 T-STG-DECOM-* were auth-hollow; SPIKE-decommission was report-only, button still vestigial.)
Boot ordering: automount + networking survive reboot; appliance self-heal watchdog agent v0.85 PROVEN-LIVE CAMPAIGN-4-2026-07-13 (F12 fix HOLDS: demo-host reboot + 5-boot storm, 0 ordering cycles, caps 63/63, WG re-handshake) + CAMPAIGN-6A-2026-07-14 1D (re-arm reboot-survival across 9 guest + 1 host reboots) (CAMPAIGN-3 F10/F11/F12 were the CRITICAL/HIGH failures; fixes shipped in agent v0.85 and were re-validated live in 4/6A — cite the validation, not the finding.) Residual: skip-active on pct reboot carried by the heal path; a NAS outage spanning a guest reboot can strand the share until agent restart (6A)

E. Access, networking & household use

Scenario Components Status Evidence Gap / roadmap
Remote access via Cloudflare Tunnel + Traefik (per-app subdomains) cloudflared, traefik PROVEN-LIVE CAMPAIGN-2 T-FLT-CF Per-customer zone-scoped CF tokens (blast-radius ruling)
LAN access when internet is down (lan_resolver) agent IMPLEMENTED — Never drilled as a customer experience ("net down — can I reach my photos?") → R-19
Phone photo backup immich (classified) PROVEN-LIVE 6D end-to-end restore proof
Documents/OCR paperless-ngx (classified) PROVEN-LIVE CAMPAIGN-6C 4P-1 (deploy paperless-ngx, ingest 3 docs via consume flow, OCR + PDF/A ~90s) + 4P-2/3/5 Consume-folder ingestion awkward without SMB → R-7
Files from Windows Explorer / Mac Finder (SMB server) controller v0.145.0 + felhom-samba:1.0.0 PROVEN-LIVE felhom-controller REPORT.md (v0.144.0) + controller/sharing.md; transport verdict audits/SPIKE-lan-discovery-2026-07-18.md „Megosztás" page: enable + one household password + shares (new folder or picked existing, per-share read-only). Fourth protected infra stack (host-net, smbd+nmbd+wsdd). Live on demo: 445 reachable, NetBIOS FELHOM resolves, write/read byte-compare PASS, write to a read-only share REFUSED, SMB writes land as uid 1000. Explorer leg PASSED 2026-07-18 (Viktor): Network → FELHOM → both shares open; a real Explorer save into dokumentumok landed owned uid 1000, and a write into the read-only filmek was refused by Windows with the folder left untouched. Share data RIDES BOTH BACKUP TIERS (R-7b, controller v0.145.0, Model B′ sibling shares source): tier-2 cross-drive legs + an offsite _shares restic snapshot carrying the share definitions and the credential copy, with a „Megosztások" restore. All four legs PROVEN-LIVE on demo 2026-07-18 — tier-2 tree md5-verified; offsite snapshots e0b9d723 (Viktor 12:18:16Z) and 4e2b15ec both carrying manifest + passdb.tar; restore round-trip returned a deleted probe file byte-identical and a deleted share DEFINITION with its original flags without overwriting live files; samba liveness → hub-accepted health_critical. Remaining human leg: SMB positive auth with the real household password
Media to TV via DLNA — MISSING — Jellyfin app exists; DLNA/SSDP unvalidated → R-6, R-8
File access via browser FileBrowser (infra app, auto-mount sync) IMPLEMENTED FileBrowser runs healthy + userdata-bound (storage-lifecycle-acceptance-2026-06-15, CAMPAIGN-3) Actual browse/download through FileBrowser is exercised in no doc. (Cited CAMPAIGN-2 T-PAGE-ALL renders only the controller dashboard pages, not FileBrowser.) Demoted. 2026-07-26, controller v0.172.0 (R-75) — status DELIBERATELY UNCHANGED. The canonical drop-zone now has its own FileBrowser source („Beolvasás" → /srv/beolvasas, a separate bind of <system namespace>/userdata/import) and the app page carries a per-app deep link into it. Verified live on demo-hp: the source and bind are in the generated config, the app page renders https://files.enkisfelhom.hu/files/Beolvas%C3%A1s/paperless, and a file written through FileBrowser's OWN mount was consumed and deleted by paperless in ~30 s. That is still not a browse. Nothing in this arc drove the FileBrowser HTTP UI — no browser exists on DooPlex — so the row's standing caveat survives intact and the upgrade to PROVEN-LIVE remains unearned. What it would take: a human click-through, or an authenticated /api/resources round-trip against the live instance. See controller/import-and-data-paths.md
Indítópult (app launcher) — one-tap grid of the household's openable apps controller v0.163.0 IMPLEMENTED New FIRST sidebar page /launcher: colored tiles (deterministic slug color or .felhom.yml brand_color) + white glyph/monogram, one per openable app (tile ⟺ „Megnyitás" — subdomain presence is the single criterion; controller excluded). Operational → <a target=_blank> to the public URL; stopped → greyed + state badge, no link. / stays the Vezérlőpult. Endpoint-level + render-test verified; felhom-controller/REPORT.md (2026-07-24) Live operator click-through of a real tile → app pending (browser automation not available on DooPlex). Follow-up: curate brand_color for top catalog apps (R-72). Sharing the launcher outside the household is now the capability-URL guest link — see the row below
Indítópult megosztás (vendég link) — capability URL /s/<token> serves a standalone read-only guest launcher (no account, no admin session); optional per-share password; QR controller v0.165.0 IMPLEMENTED 160-bit crypto/rand token, constant-time match (empty stored = disabled = byte-identical to the mux default 404); guest headers noindex/no-referrer/no-store; optional SEPARATE bcrypt share password + its own per-IP attempt map; signed cookie HMAC(token|passwordHash) keyed with session_secret (rotate-token OR change-password invalidates all cookies); token redacted in logs (/s/<redacted>). Groups A–G (14 tests) + 3 red-proofs; §13 endpoint-level live validation on 9201 all-pass (felhom-controller/REPORT.md 2026-07-24). Design ruling: member accounts SUPERSEDED by this capability-URL model; per-member tile visibility parked under the SSO arc (R-15). Full operator browser click-through + a validation doc pending → then PROVEN-LIVE. Accepted residuals: link-preview crawlers fetch once (noindex prevents indexing); reverse-proxy/CF access logs hold the path (ops-tier); the modal link carries the request Host (LAN-IP admin ⇒ LAN-IP link)
Forgot dashboard password → instant reset code controller v0.123, hub PROVEN-LIVE DRILL-day0-take2-2026-07-12 F-15 (live re-run of the exact failure path: hash applied 1s after request, code accepted first try)
The dashboard in English — the household picks its language; Hungarian unchanged controller v0.247.0 (spike) → v0.250.0 (slice 1, R-556 CLOSED) PARTIAL — PROVEN-LIVE for every dashboard template (all 36 templates carry their copy in the bundles; the switch is on every dashboard page) audits/i18n-slice1-2026-09-17/C/live/ (and A/, B/, audits/i18n-2026-09-17/live/) — on demo-hp 0.250.0 each Hungarian page equals its 0.249.0 fetch once the new switch form is removed (live numbers aside; login, claim, catch-all byte-identical); POST /settings/language made every page English and back; the hub stored hu, en, hu (16:45:10Z, 16:46:06Z, 16:46:43Z). Unit: TestI18nParity (106 states vs fixtures from unconverted templates), TestI18nParityCoversEveryMarker, TestI18nEnglishPages, TestDirectRenderHandlersFollowLanguage, all red-proofed. Design architecture/10-localisation.md Go-side messages and three app-name page titles (R-557 after R-553, R-566), hub e-mails (R-558), console/download page (R-559), catalog copy (R-560), guide + English stranger drill (R-561); ASCII-only Hungarian invisible to the English page test (R-565). Hungarian households see the switch only once the fleet floor reaches 0.250.0
Multiple household users / per-person accounts — MISSING — Single dashboard password; acceptable for alpha → R-15
WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) agent v0.72, hub v0.35 IMPLEMENTED SPIKE-oob-wg-operator-peer-2026-07-05, SPIKE-felhom-sshd-2026-07-05 Mutual-repair desired-state arc not built → R-13. CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: the hub's own OOB health check could not see whether the operator's key was installed. HostOOBRow mirrored five of the agent's eight OOB fields, so operator_key_configured — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by encoding/json on arrival, and oobDegraded returned ok for a box with felhom-sshd active, reachable, a valid config, a configured peer and no operator key at all. operator_peer_configured, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by TestHostOOB_MissingOperatorKey_EndToEnd and TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd
The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page agent v0.119.0, hub v0.85.0 PROVEN-LIVE (2026-07-31) audits/host-addresses-visible-2026-07-31.md Before this the LAN IP was not reportable at all — HostMetrics carried no address of any kind — and the WG IP existed only in /offsite's peer table keyed by pubkey (peer→host, never host→peer). New wire field addresses[], one row per (interface, address); IsGlobalUnicast() is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their ip addr ground truth exactly. Two honesty properties carry the risk and are both red-proofed: WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders UNKNOWN, never "no addresses" — proven live on drill-r50-0a4f9a (0.113.0). Not covered: a two-LAN-bridge box and a real WG drift, neither of which exists to observe
The operator can see whether a managed host's guests still have working networking — and how often the watchdog had to repair them agent v0.92.0 (emitter, 2026-07-21), hub v0.104.0 (reader, 2026-08-13) IMPLEMENTED backlog/OPEN-ITEMS.md R-319; hub/internal/web/hosts_guestnet_test.go (7 tests, fixtures copied verbatim from demo-felhom-8363b5's live host_reports row) The agent emitted guest_net on every heartbeat for twenty-three days while the string occurred nowhere in felhom.eu/hub/ — stored as raw text in report_json, read by nothing (R-260/R-264, the first of that census's readers to be built). The fact that carries the risk is heals_last_hour, not state: a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. heal_succeeded is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. Unknown is never drawn as healthy: three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. Three red-proofs, each mutation asserted applied by grep before its run, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190 as the eight guest_net allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked)
Break-glass management-plane recovery agent v0.71, hub v0.84 IMPLEMENTED runbooks/break-glass.md hub v0.84.0 adds an operator-SESSION retrieval path (host page → Console access → Reveal; POST /hosts/{id}/reveal-recovery-credential, CSRF-gated, writes a customer-visible recovery_credential_revealed event) beside the pre-existing global-key one (GET /api/v1/admin/hosts/{id}/recovery-credential), which is untouched and stays the route for when the hub UI itself is down. The credential half is now PROVEN (2026-07-31): the vaulted demo-hp-bb76ea password was verified against the box's own /etc/shadow hash AND minted a real PVE ticket — POST /api2/json/access/ticket → HTTP 200, root@pam, 367-char ticket, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped disabled until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → R-133

F. Notifications & monitoring

Scenario Components Status Evidence Gap / roadmap
An ABSENT backup-target drive raises its OWN alarm, paired with a matching recovery agent v0.116.0, controller v0.184.1+, hub v0.81.0 PROVEN-LIVE (2026-07-30) audits/R116-v0116-2026-07-30.md. On a fresh box built through the real day-0 on demo-hp, running the agent it installed unaided from the vouched Day-0 manifest (0.116.0), both drives enrolled through the real endpoints and device loss a real hot-detach — the full four-event sequence, two matched pairs, correctly discriminated: 07:20:04 backup_target_absent (error) / 07:22:34 backup_target_restored (info) for the TARGET, and 07:24:04 storage_disconnected (error) / 07:25:34 storage_reconnected (info) for a NON-target drive on the same box minutes apart. Gate fired in 3 s. All four reached the hub — specific alarm, severity, Hungarian copy and hub routing now exercised end-to-end. Discrimination is proven NON-trivially for the first time: both prior runs had the target itself emit the generic event, so the mirror proved nothing. Over-correction guard PASSES with a positive observable — 0 ABSENT lines and 0 drive events over a 2m14s window with both drives present, target degraded:false, while 2 RETURNED lines prove the gate was ticking. The mechanism was isolated from the captured payload first (DIAG-r116-disks-payload-2026-07-30.md), after two fixes aimed at shapes that do not occur. v0.116.0 joins the two records of one drive on the only identity that survives the device — the CONFIGURED path — so one row carries both the flag and the guest path the gate keys on. Both smaller-looking fixes were rejected because they regress R-114 (backup_target_offer.go:79 reads flag+mount_path as healthy). Caveat worth reading: the drill box ran controller 0.185.1 from the golden, which PREDATES R-114 — so its absent-state banner showed the old false "backup is on the system disk" copy. That is the golden being a release behind, not a regression → R-120
Health-degradation email (edge-triggered, cooldowns, Hungarian) via hub → Resend controller, hub IMPLEMENTED delivery pipeline live-proven for the enlarge-block trigger (CAMPAIGN-6D P3-DELIVERY, op+customer "Kedves Ügyfél!"); NotifyHealthChange ok→warn/fail edge-trigger implemented The health-degradation trigger specifically has never fired an email live in any doc. Demoted (pipeline proven for a different event). Deliverability to HU freemail → R-4
Event catalog: app_start_failed, dead-app, offbox_enlarge_blocked, claim/reset codes, critical severity controller, hub v0.31/48/50/55 PROVEN-LIVE live-delivered: CAMPAIGN-6D P3-DELIVERY (enlarge-block, op+customer); DRILL-day0-vm F-4 (claim code); DRILL-day0-take2 F-15 (reset code) app_start_failed/dead-app delivery is unit-only (6C inconclusive) — the pipeline + 3 event families are live, those two are not
Prefs safety: empty-email wipe guard controller v0.137 + hub v0.71.0 IMPLEMENTED controller leg red-proofed 07-15; hub-side no-clobber belt (handleSavePreferences preserves a stored non-empty address on an empty-email push) red-proofed 07-22 Born from a live incident; controller 0.160.0 guards both its push legs, so the hub belt covers older/rogue boxes
Paired recovery notifications + prefs seeding at claim + priority headers (power-outage audit F11/F12/F14-light) hub v0.71.0 IMPLEMENTED (recovery leg PARTIAL until a live staleness cycle fires it) hub/CHANGELOG.md v0.71.0; 17 tests + 4 red-proofs (REPORT.md 2026-07-22); Resend headers mechanism probed live (HTTP 200) pre-implementation; operator+customer test rows live-fired via the controller's own test endpoint Recovery = explicit eventType branch, severity semantics frozen; customer gate = PAIRING (notification_log evidence), not enabled_events. Live legs pending: a natural *_recovered mail (next real staleness cycle or the reboot-drill arc — never fabricated by blocking reports) and seed-at-claim on a real claim (Peti Friday reinstall). F14-full (operator push channel, ntfy/Telegram) stays open → R-69
System + container metrics (SQLite, Chart.js, 30-day downsampling) controller IMPLEMENTED metrics collection + /monitoring render present (page 200) The cited CAMPAIGN-2 T-RES-CPU/T-SOAK-LOOP are H1/H2 harness artifacts (auth-302), not metrics tests; SQLite/Chart.js/30-day downsampling validated in no campaign. Demoted
Always-on debug rings + on-demand log-bundle pulls with TTL/custody controller v0.116, agent v0.83, hub v0.46 PROVEN-LIVE debug rings live-exercised CAMPAIGN-3 fix-6 (1000-cap ring, ~55min horizon under load) The log-bundle-pull TTL/custody half is changelog-only (no dedicated observability audit doc); ring persistence across restart is a known gap
Operator alerting (Healthchecks → monitoring@felhom.eu) k3s, Resend IMPLEMENTED operator infra, stated in production since 02-04; no corpus validation doc
Backup-deadline alerting (expected_backup_missed) is ANCHORED — absence of signal is UNKNOWN, not failure hub v0.75.0 IMPLEMENTED audits/DIAG-backup-missed-2026-07-26.md + red-proofs A/B/C + replay of the real 2026-07-26 03:00 reports (all three silent) No row status flips — this signal had a FALSE-POSITIVE class (three instances: hub v0.12.0, v0.73.0, R-81), now anchored at first contact and read across retained host-report history. Still unit-proven only, not live-fired at a real deadline. The underlying PBS/offsite-DR tier gap it exposed is → R-82.

G. Fleet & operator (hub)

Scenario Components Status Evidence Gap / roadmap
Customer/host management: 8-tab detail, scoped auto-refresh, safe stale-host deletion, capability chips hub v0.47–0.53 PROVEN-LIVE hub v0.53.0 dead-host roll-up live on the Peti cluster (proxmox1 down 23h); CAMPAIGN-4-2026-07-13 (operator UI driven live); DRILL-day0-take2 F-16 (offsite/freeze buttons live) 8-tab render + capability chips are render-test-validated (hub UI is password-gated; CC cannot log in). (Cited "daily operator use" was a no-doc citation; AUDIT-hub-gui-2026-06-30 predates these features at hub v0.25)
Config/state change round-trips in seconds (hub↔box immediacy; 15-min cycle stays the backbone): box→hub out-of-cycle report (Dir 1) + hub→box GET /api/v1/wait long-poll wake (Dir 2) controller v0.139/140, hub v0.58/0.63 PROVEN-LIVE (2026-07-21) Transport proven live through the real DNS-only ingress: SPIKE-immediate-sync-transport-2026-07-16 + hub v0.58.0 / controller v0.140.0 REPORTs — 240 s no-annotation hold (25 s heartbeat defeats nginx's 60 s proxy_read_timeout, no ingress change), 0.047 s wake-on-change, hub rollout restart = 1 WARN + 0-storm reconnect; Dir-1 2 s box→hub round-trip live in controller v0.139.0 The operator-UI save→apply round-trip is not fired end-to-end live (hub UI password-gated; CC can't log in) → R-23; the wake transport and the ACK→config_version→ConfigRefresher delivery chain are each proven, only the UI-triggered bump leg is unexercised. Agent-plane (host-domain desired-state) poke first slice PROVEN-LIVE (Direction-2a, agent v0.89.0 + hub v0.59.0, 2026-07-17): contentless ep0-relayed UDP poke → agent immediate desired-state cycle, per SPIKE-immediate-sync-transport-2026-07-16 P4. Full path live-proven: a real operator manifest save fired poke: sync-poke delivered to 10.77.0.2; the box (0.89.0) received it and logged poke received → triggering an immediate desired-state cycle → out-of-band report triggered — ~31 ms ep0→box, sub-ms to the report cycle (WG-confined, from 10.77.0.1 to the 10.77.0.2-bound socket); save→tick ≈ ~0.45 s (SSH-dominated), well under ≤2–3 s. R-13 first slice (listener+sender only; the rest of the mutual-repair arc stays open). System-initiated immediacy wired (hub v0.63.0, this REPORT): the mutation sites that only OPERATOR actions used to notify now fire the correct plane's notifier when the hub itself mints state — agent-plane pokes at PBSDRAutoProvision (the observed slice-C lag), ReissuePBSDR (also the pbsdrheal escalation), handlePBSDRReissue, and the two admin desired-state api writers; controller-plane bump at reissueOnReenroll. Unit-tested + red-proofed, not yet fired on a real system event (folds into the rehearsal bind sequence). Still PARTIAL: the R-23 operator-UI save→apply leg and the agent fast-tick-until-first-convergence SECONDARY (the WG-registration leg a poke can't reach pre-tunnel) remain unfired live. Fast-tick SHIPPED (agent v0.90.0, R-28): while any desired-state item is unapplied — incl. the pre-tunnel window a poke can't reach — the agent pulses the out-of-band trigger every 30 s and self-disarms on convergence (state-based; four cached sources; LOUD states excluded). LIVE on both demo agents (the fast-tick armed: 30s … startup line verified). REAL-ONBOARDING PROOF DONE — tests/VALIDATION-n100-rehearsal-2026-07-18.md (ledger 8, S5): on a genuine first onboarding on metal, every post-bind leg landed seconds apart with no ~15-minute stall anywhere — bind 16:29:55 → credential delivered 16:30:21 (26 s) → agent 0.90.0 up 16:30:49 → WG registered + tunnel applied 16:30:51 (~2 s) → poke listener 16:30:54 → controller 16:32:28 → floor-lifted and running current 16:32:39. Bind → running-current = 2 min 44 s. The PBS-DR descriptor auto-provisioned on the same cadence (agent converged state=applied 16:45:53) — though see the DR-tier row: the descriptor converged while the credential behind it was already stale (R-39). The pre-tunnel fast-tick window is therefore proven in its real setting; the remaining PARTIAL is the R-23 operator-UI save→apply leg alone 2026-07-21 — THE LAST PARTIAL LEG IS CLOSED (R-23(a) restart leg). The operator saved the global floor to a version the box did NOT run (0.153.0 → v0.154.0) and the managed self-update fired exactly once: 06:57:13Z UpdateState pending (initiated_by=auto-floor) → 06:57:17Z agent controller swap requested → 06:57:21Z container restarted → 06:57:29Z new controller healthy. Save → healthy on the new version = 16 s. Over a 39-minute window: swap requests 1, agent-driven bootstrap restarts 1, rollbacks 0, container RestartCount 0; VerifyStartup confirmed on the next boot and the following periodic check logged Current version 0.154.0 is up to date (the at/above-floor branch doing nothing, as designed). The 2026-07-20 attempt could not prove this because it targeted an already-running version. Evidence: felhom-controller/REPORT.md §6.
Customer right-sizes guest RAM from the controller (agent-enforced bounds, live cgroup apply, no reboot) agent v0.90.0 + controller v0.143.0 (R-24) PROVEN-LIVE (grow and shrink on metal, 2026-07-18) tests/VALIDATION-n100-rehearsal-2026-07-18.md (ledger 9) — the apply is now proven in both directions on a normal-sized box: customer zero shrank 11675 → 8192 MB at 16:50:22 and grew 8192 → 12288 MB at 17:02:17, each a live cgroup apply with no reboot (local-api: guest-memory resized in the agent journal, [web] memory resized in the controller log), and the new total rippled into the deploy page's memory math at 17:05:15 (total=12288MB). F5 auto-sizing had landed the guest at 11675 MB. Controller-direct (R-24's hub-desired-state framing SUPERSEDED, Viktor 2026-07-17). Agent GET/POST /guest/memory enforces every bound FRESH (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)) + verify-after-apply; PVE SetConfig hot-applies (Phase-0 PROVEN on the nested box: maxmem moves with the guest running, /proc/meminfo ripples via lxcfs, no reboot). Controller "Szerver memória (RAM)" card + code→Hungarian map, gated on FeatureGuestMemoryResize (MinAgent 0.90.0). LIVE-validated end-to-end through the real endpoint on the demo (above_max + below_min refusals render the Hungarian, agent English never leaks; SupportYes via the version header) Row complete as of the 2026-07-18 rehearsal — the refusals were proven on the nested demo, the applies on the N100. Cores stay observation
Publish train: MinAgent floors, gated auto-Reissue, version channels, floor-field-LAST rules hub v0.45/0.53, agent PARTIAL runbooks/publish-train-rules.md; demo-fleet updates proven Box-side floor lift PROVEN-LIVE on a fresh install (tests/VALIDATION-n100-rehearsal-2026-07-18.md): the day-0 golden deployed controller 0.143.0 at 16:32:28 and the managed floor lifted it to 0.145.0 by 16:32:34 — a 5-second, fully unattended update inside the first minute of controller life, update-state.json recording initiated_by: auto-floor with controller_updated pushed to the hub. So the mechanism is no longer nested-only. Still never proven on a real REMOTE customer — parked trains RUNBOOK-publish-0.79/0.81/0.85-* await Peti → R-1. Action before first invite: rebuild the golden to 0.145.x now that this evidence is banked, so fresh boxes don't sit two versions stale
Agent self-update: A/B slots, crash-loop auto-rollback, operator-signed agent v0.70+ PROVEN-LIVE (demo) SPIKE-agent-selfupdate-2026-07-05 Remote-customer proof pending → R-1
Controller self-update: anonymous registry, no credentials in guest controller v0.112 PROVEN-LIVE (demo) 07-10 arc
Offsite provisioning: Hetzner API, sub-account per customer, host-key pinning, credential re-issue hub v0.37–0.39 PROVEN-LIVE VALIDATION-offsite-provisioning-e2e-2026-07-09, SPIKE-hetzner-api-provisioning-2026-07-09
Per-customer offsite fill + staleness + freeze lever hub v0.41 IMPLEMENTED OffsiteChecker (hub/internal/monitor/offsite.go): fill 90/95% vs soft quota, staleness >48h No live-fired leg: CAMPAIGN-offsite-overnight-2026-07-10 recorded no quota/fill/staleness emails, and the freeze write-block was inconclusive (only the Hetzner readonly:true API op succeeded). Demoted
Box-level Storage Box aggregate (total fill, Σ quotas, oversubscription alert) hub v0.64.0 (R-5) IMPLEMENTED (data pipeline PROVEN-LIVE) monitor.OffsiteBoxChecker — fetch-throttled Hetzner GET (1/15 min), fill (used/storage_box_type.size, 80/90%) + oversubscription (Σ shared+enabled ConfigJSON quotas / capacity, 2.0×), escalation-only operator alert on the customer-less "pool-box" scope; Offsite-tab panel + dashboard tile. Phase-0-pinned live shape (box 611714) + live-computed in-cluster: 0.2% full (2.6 GB of 1.00 TB), Σ shared quota 150 GB, oversub 0.15x. Tests + 4 red-proofs; hub v0.64.0 REPORT Two open legs: the UI render is unit-verified only (hub UI password-gated → no screenshot); the alert emails are unit + red-proof verified but NOT fired live (real pool nominal — a live-fire emails Viktor). Thresholds pending Viktor's ruling (named config keys). READ-ONLY (GET)
Operator sees PBS DR datastore fill at a glance (Offsite "PBS DR" tab + dashboard gauge) hub v0.65.0 + tenantsync v1.2.0 (R-5) IMPLEMENTED (data pipeline PROVEN-LIVE) The PBS DR datastore (felhom-offsite on ep0) fill — NOT a Hetzner box. Option A: a read-only usage op on the felhom-tenantsync ep0 forced command (twin of fingerprint; df on the datastore path — no customer_id, no admin token, NO mutation), polled by monitor.PBSDRBoxChecker (OffsiteBoxChecker clone; 15-min throttle; states ok/unavailable/degraded; fill 80/90% on the "pbsdr-box" operator scope). /offsite split into Restic + PBS DR tabs; two dashboard gauges. Graceful: hub deploy ⟂ ep0 update (ep0 ≤ v1.1.0 → gauge "n/a" until updated). Phase-0-pinned (df on ep0 PBS 4.2.3) + live-computed in-cluster (ep0 updated to v1.2.0 this session): 19.1% full (7.1 GB of 37.2 GB). 10 Go tests + a bash harness + 3 red-proofs; hub v0.65.0 REPORT Open legs: UI render unit-verified only (hub UI password-gated); the fill alert email is unit + red-proof verified, NOT fired live (datastore nominal at 19%). Separate PBS threshold keys (default 80/90); no oversubscription (namespaces, not quotas). READ-ONLY
The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is) hub v0.106.0 (R-339) IMPLEMENTED — deliberately NOT proven-live Both box checkers count consecutive failed fetch windows and emit pbsdr_box_unreachable / offsite_box_unreachable (severity warning) past a default 3 windows (≈30–45 min), each with a paired *_recovered all-clear routed via recoveredPairedDownTypes — required because the recoveries are severity info, which severityNotifies drops. Scopes stay customer-less (pbsdr-box / pool-box) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: internal/monitor/box_reachability_test.go + the cross-package wiring test in internal/notify/, which asserts an actual operator mail rather than a map entry. Filed BECAUSE of a measured gap, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent The gap that remains is R-340, and it is not small: the ep0 read is the usage op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. No live or constructed outage has exercised the emit path, and one cannot be manufactured against ep0 (Tier 2, protected)
Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store hub v0.53, conventions IMPLEMENTED 07-13 closing bundle
Operator login password changeable from UI hub v0.54 IMPLEMENTED 07-13
An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps controller v0.259.0 + hub v0.119.0 + ISO 1.29.0 + the whole catalog PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED audits/DRILL-first-hour-en-0258-2026-09-20.md — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then audits/i18n-closing-2026-09-21/live/ — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries four plain-ASCII English words where the drill's carried képző-szkítia-ásatás, one day apart in the same inbox. R-596, R-597 and R-598 are CLOSED. What this row still does NOT claim: the fixed journey has not been walked end to end by a stranger on a fresh install. Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). Also not walked: the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page warnings themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".
A deletion of a customer's off-site history is NOTICED within a day hub v0.111.0 (R-431) IMPLEMENTED — not yet PROVEN-LIVE 09-01 hub/internal/monitor/offsite.go — third signal beside FILL and STALENESS. On the hub deliberately: a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by StatsKnown (R-331), the declared State (R-204) and run success (R-100). Threshold reasoned, not invented: over 12 898 reports every decrease lands on ZERO and predates stats_known; in the 380-report stats_known window there are none. ACCEPTANCE: 9 009 real points replayed → ZERO alarms (offsite_r431_test.go, fixture committed). What PROVEN-LIVE would need and this does NOT have: a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught.