Record work, banked first: shrink the E-2d row, create the missing capability-map rows
Unconditional and three sessions overdue, so it commits before any code is
touched — E-2d itself stopped at Phase 0 and banked nothing.
E-2d row: 822 words -> 121, and the contradiction resolved. Its State read
CLOSED — PARTIALLY PROVEN while the cell's final sentence read "This row stays
OPEN only for the residue"; a reader could not tell which. It is CLOSED, with
R-116 the single named open leg.
Nothing unique was binned. Three facts existed ONLY in that cell and are moved
into audits/E2D-fresh-vm-2026-07-29.md as a new §1a: the local-lvm fence figures
with the 888 GB nvme alternative, the exactMount subdirectory caveat and why the
subdirectory is nonetheless the safe placement (no durable_id collision), and
the ISO/PAIRING -> DIRECT fall-through derived at source with its line
citations. drill-r50's blocked status was already in both audits.
Capability map: it had ZERO rows for the backup-target work — grep gives 0 hits
for backup_target and one for "E-2" that is a campaign date string. Three
scenario rows added, at today's honest status, not the value hoped for later:
C. Protection & recovery — installer Case A/B, DEGRADED recorded not hidden
PROVEN-LIVE, cites E2D-fresh-vm C1+C2
D. Storage & devices — the offer, and that registration confers no role
PROVEN-LIVE, cites SESSION-C C4 + the decline path
F. Notifications & monitoring — the absent-target alarm and its pairing
PARTIAL, cites SESSION-C C5, leg named, -> R-116
Row F is PARTIAL today per the doc's own strict enum (a leg not exercised live
is PARTIAL with the leg named, never PROVEN-LIVE). A later session may flip it;
this commit must not.
This commit is contained in:
@@ -72,6 +72,7 @@
|
||||
|
||||
| Scenario | Components | Status | Evidence | Gap / roadmap |
|
||||
|---|---|---|---|---|
|
||||
| Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has |
|
||||
| Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** |
|
||||
| Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 4, 5.** The matrix records that the copy's `recovery-unit/` mirror is read by no path (→ R-102) |
|
||||
| Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect) |
|
||||
@@ -90,6 +91,7 @@
|
||||
|
||||
| Scenario | Components | Status | Evidence | Gap / roadmap |
|
||||
|---|---|---|---|---|
|
||||
| A second drive appearing is OFFERED as the backup target; accepting moves it; registration and the drive gate confer no role by themselves | controller v0.186.0, agent v0.113 | **PROVEN-LIVE** | `SESSION-C-2026-07-29` C4: offer rendered with `data-path`, decline path proven (target stayed `local`, no `felhom-backup` storage, agent.json unchanged), `restart_required:true`, agent did NOT self-restart, wrapper created the storage at the drive's OWN mountpoint | Accept was driven through the endpoint the button POSTs, not a browser click — no browser automation exists on DooPlex |
|
||||
| Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts | controller, agent v0.87 | **PROVEN-LIVE** | `DISPOSITION-ia-finding2-systemdisks-2026-07-13` (legacy EFI+LVM host, root not offered, byte-identical); enroll/format live in `storage-lifecycle-acceptance-2026-06-15` (E10 re-enroll, data intact); agent fence self-test refuses `/dev/sda` | (Cited `CAMPAIGN-2` T-STG-ENROLL/SEC-FORMAT were auth-hollow CSRF-403.) **Fresh-USB wizard enroll+format through the customer UI PROVEN-LIVE (controller v0.141.0, 2026-07-17):** a 64 GB scratch USB driven through the real `/api/storage/init` endpoints (login+CSRF) → confirm → detached format (~27 s mkfs) → mount → register → mounted+registered at `/mnt/felhom-drives/scratch1`. **F6 (initialize-to-usable) now covered:** the wizard runs the chain as a detached, disconnect-safe, pollable job (3-step progress) with an agent format-status poll for a slow mkfs **2026-07-26 — a SILENT failure class on the channel every agent-backed capability depends on (this row, data migration, USB enrollment, guest RAM, quiesce/PBS) is now DETECTED (controller v0.173.0, R-77). No row status flips.** `controller.yaml` and `bootstrap.json` could disagree on `local_api.endpoint` indefinitely with no signal: the R-50 island migration rewrote the latter, the fleet kept dialling the former, and for 17.5 h the only alert was a generic "agent unreachable" that read as an infrastructure blip. Drift now raises its own event type (`local_api_endpoint_drift`) naming both values. It is DETECTION ONLY — the authority ruling is R-78 — so the class is now loud, not prevented. Evidence: `audits/DIAG-agent-channel-2026-07-26.md`. |
|
||||
| Data migration between drives (all / per-app), crash-safe | controller | **PROVEN-LIVE** | `CAMPAIGN-6C` 4P-5 (scope=app round-trip, byte-identical); `storage-lifecycle-acceptance-2026-06-15` (two migrate-all runs via dashboard UI, sha256 byte-identical) | (Cited `CAMPAIGN-2` T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS |
|
||||
| NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated | controller v0.113–117, agent v0.81/84/85 | **PROVEN-LIVE** | `SPIKE-nas-verify-2026-07-11`, `SPIKE-nas-dsm-2026-07-11`, `CAMPAIGN-3-2026-07-11` (boot/reassert fixes) | |
|
||||
@@ -119,6 +121,7 @@
|
||||
|
||||
| Scenario | Components | Status | Evidence | Gap / roadmap |
|
||||
|---|---|---|---|---|
|
||||
| An ABSENT backup-target drive raises its OWN alarm, paired with a matching recovery | agent v0.114.0, controller v0.184.1+, hub v0.81.0 | **PARTIAL** | `SESSION-C-2026-07-29` C5: the drive-absent gate fires (4 s) and an alarm reaches the hub — but it is the **generic** `storage_disconnected`, while the return fires the **specific** `backup_target_restored`, so the pair cannot be matched. `backup_target_absent` never fired (count 0) | The specific alarm and its severity, Hungarian copy and hub routing are all still unexercised end-to-end → **R-116** |
|
||||
| Health-degradation email (edge-triggered, cooldowns, Hungarian) via hub → Resend | controller, hub | **IMPLEMENTED** | delivery pipeline live-proven for the **enlarge-block** trigger (`CAMPAIGN-6D` P3-DELIVERY, op+customer "Kedves Ügyfél!"); `NotifyHealthChange` ok→warn/fail edge-trigger implemented | The **health-degradation** trigger specifically has never fired an email live in any doc. Demoted (pipeline proven for a different event). Deliverability to HU freemail → R-4 |
|
||||
| Event catalog: app_start_failed, dead-app, offbox_enlarge_blocked, claim/reset codes, critical severity | controller, hub v0.31/48/50/55 | **PROVEN-LIVE** | live-delivered: `CAMPAIGN-6D` P3-DELIVERY (enlarge-block, op+customer); `DRILL-day0-vm` F-4 (claim code); `DRILL-day0-take2` F-15 (reset code) | `app_start_failed`/`dead-app` delivery is unit-only (6C inconclusive) — the pipeline + 3 event families are live, those two are not |
|
||||
| Prefs safety: empty-email wipe guard | controller v0.137 + hub v0.71.0 | **IMPLEMENTED** | controller leg red-proofed 07-15; hub-side no-clobber belt (`handleSavePreferences` preserves a stored non-empty address on an empty-email push) red-proofed 07-22 | Born from a live incident; controller 0.160.0 guards both its push legs, so the hub belt covers older/rogue boxes |
|
||||
|
||||
@@ -37,6 +37,36 @@ manual-installer fallback was used.
|
||||
**Operator STOP: not required and now retired.** `HUB_PW` is in `~/.config/credentials`; CC created
|
||||
the customer and performed the bind itself. The one human step that *was* needed is new — see §6.
|
||||
|
||||
## 1a. Phase 0 answers, preserved from the OPEN-ITEMS row
|
||||
|
||||
Moved here when the E-2d register row was rewritten (2026-07-29) — the row had grown to ~820 words and
|
||||
these were the facts that existed nowhere else. They are inputs to any future drill on this host, not
|
||||
narrative.
|
||||
|
||||
- **Storage fence.** `local-lvm` on demo-hp is a thin pool, ~144 GB allocated against ~54 GB real,
|
||||
38.8 % used, on a box running a live customer guest — a full thin pool corrupts every guest on it.
|
||||
`local` has only 23.7 GB and sits on `pve-root`. **Use `/mnt/nvme-1tb` (888 GB free).**
|
||||
- **The `exactMount` caveat, and the placement decision it forces.** A dir storage created at a
|
||||
SUBDIRECTORY of `/mnt/nvme-1tb` fails the agent's `exactMount` check and reports `disconnected` in the
|
||||
host report. Both E-2d and Session C accepted that: hub-side it is a WARN log line only — no event,
|
||||
no email — and the alternative (a second storage at the live backup target's own mountpoint) risks
|
||||
perturbing the drive-role resolution on a production box. The agent deliberately falls back to a
|
||||
stable store id rather than borrowing the nvme's fs-UUID in this case, so there is **no durable_id
|
||||
collision** with `felhom-backup`; that is what makes the subdirectory the safe choice.
|
||||
- **The ISO/PAIRING → DIRECT fall-through, derived at source.** A fresh VM with no baked customer-id
|
||||
lands in PAIRING mode (`scripts/iso/felhom-bootstrap.sh:537-541`), not DIRECT (`:312`), and only
|
||||
DIRECT passes `--customer-id / --mode / --passphrase-file`. On a 200 from `/api/v1/appliance/poll`
|
||||
the pairing loop writes the hub-delivered credentials into the 0600 env, re-sources it and calls
|
||||
`run_direct` **in the same invocation** (`:495-499`), which is the single site that fetches
|
||||
`$INSTALL_URL` (`:322-330`), builds the args (`:334`) and invokes
|
||||
`bash "$SCRIPT_TMP" "${args[@]}"` (`:343`). So the ISO route reaches the identical installer
|
||||
invocation and yields a claimable customer — which is why it is the spine and no manual 1.22.0 run
|
||||
is needed as a separate scenario.
|
||||
- **`drill-r50` stays blocked.** Unblocking it means the fixture stops representing anything real
|
||||
(R-93).
|
||||
|
||||
---
|
||||
|
||||
## 2. Timeline (VM 9300 `e2d-fresh` on demo-hp, nested PVE)
|
||||
|
||||
| UTC | Event |
|
||||
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user