diff --git a/CONTEXT.md b/CONTEXT.md index 23e373e..ec41cfc 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -3,6 +3,17 @@ > Created with the REUSE.md rollout (2026-07-03). Authoritative history: `hub/CHANGELOG.md` (hub), > `website/CHANGELOG.md`, `scripts/CHANGELOG.md`; end-of-task detail in `REPORT.md`. +- **2026-07-12 — Day-0 VM DRILL COMPLETE (auto-confirm FIRST LIVE FIRING): full arc proven on a + fresh nested-PVE box** — appliance Day-0 → floor-at-first-report → ceremony → **auto-confirm + pending→escrowed in ~7.5 min, zero clicks** → offsite backup + restore round-trip. Two HIGH gaps: + **F-4 no operator password-set path exists (G10 unclosable, dashboards born OPEN — blocks tester + gate)** and **F-6 identity-only ceremony never implemented (no-PBS appliance can't escrow — drill + forked to PBS DR tier = full Peti-sequence rehearsal, all green)**. Installer fresh-box gaps: + felhom-pbs-apply not shipped (F-7), `age` missing (F-10), root-owned guests/ parents (F-3 — + check demo for the latent copy), silent root@pam rotation UX (F-8). Runbook fixes committed + (day0 A.2 anonymous-fetch; escrow-ceremony identity-only claim CORRECTED + age prereq). Report: + `documentation/audits/DRILL-day0-vm-2026-07-12.md`. Drill VM qm 300 kept (3 snapshots) for + re-drills; teardown list in report §9. - **2026-07-12 — CAMPAIGN-3 Task A SHIPPED: agent v0.85.0 boot/recovery plane + appliance self-heal.** Fixes F12 (CRITICAL boot ordering cycle — templates drop network-online, `MigrateNetworkUnits` repairs installed units), F11/F10/F9 (reassert: fstype-driven classify, reset-failed+rearm, per-unit diff --git a/REPORT.md b/REPORT.md index b297510..92a0ac4 100644 --- a/REPORT.md +++ b/REPORT.md @@ -2,19 +2,14 @@ > **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md). -## CAMPAIGN-3 — unattended "no mercy" night run (NAS · deploy · backup · chaos · observability) — 2026-07-11/12 +## DRILL — Day-0 on the Demo-VM (nested PVE) + escrow/auto-confirm first firing — 2026-07-12 -**Full report: [`documentation/audits/CAMPAIGN-3-2026-07-11.md`](documentation/audits/CAMPAIGN-3-2026-07-11.md).** Run 22:09 → 04:27 CEST against the demo box (controller 0.117.0 / agent 0.84.0), operator-unattended, real surfaces only, no code changes. Ledger (final incl. morning RCA/recovery): **31 PASS · 17 FAIL · 12 FINDING · 1 DISCREPANCY** (70+ scenario entries, evidence at `180:~/campaign3/`). +**Full report: [`documentation/audits/DRILL-day0-vm-2026-07-12.md`](documentation/audits/DRILL-day0-vm-2026-07-12.md).** Supervised drill (Viktor + CC), ~15:15–17:09 CEST, on a throwaway nested PVE 9.2.2 VM (qm 300 on felhom-pve) against hub customer `demo-vm-felhom` / `enkisfelhom.hu`. No code changes; two runbook docs corrected (day0-install.md A.2 anonymous-fetch wording; RUNBOOK-escrow-ceremony.md's false "identity-only ≥0.80.0" claim + `age` prereq). ### Headlines -- **CRITICAL F12 — the overnight host loss, RCA closed (hardware exonerated):** the agent's network-storage **automount** template (`After=`+`Wants=network-online.target`, implicitly `Before=local-fs.target`) creates a boot **ordering cycle**; systemd breaks it by deleting an arbitrary job. Boot at 23:31 sacrificed `networking.service` → host up **7 hours with no network**; the 06:45 power-cycle boot hit the same cycle and sacrificed the **automount** instead (NAS dead, healed manually). **Every boot of a host with an enrolled share is a coin flip until the template drops the network-online ordering** (`_netdev` on the `.mount` suffices). 4e caught exactly what it was designed to catch. -- **CRITICAL F10 + HIGH F11/F9 — the reboot/recovery plane around NAS automounts is broken:** `mount-start-limit-hit` is never re-armed by any heal path (once even blocked guest start → guest DOWN); guest reboot with an idle share leaves the autofs trigger unpropagated into the container — the post-start reassert **logs its own WARNING and then skips** the automount restart that provably heals ("skip-active" branch). Every such reboot = 4 NAS apps dead-at-boot. Reproduced on 3 of 3 guest reboots. -- **HIGH F7 — backup dumps are written in place (no tmp+rename):** a mid-backup NAS cut left a 0-byte tar *replacing* the last good 247M dump; in that window restore = empty volume. Next run self-heals; run-level `success:false` is the only signal. -- **The data plane held:** all 5 refusal categories ×2 correct + fast (2–5 s, retry=0), verify/rollback/single-flight/orphan flows clean, deploy-view truth holds, restore round-trips **byte-identical**, EIO same-second under outage, organic stub → badge + deploy-409 live-validated, hardlinks work on NFSv4.1. -- **Fix-6 answered with numbers:** ring cap horizon = ~55 min idle but **~6.5 min under load**; every restart/reboot wipes both rings — persistence, not just size, is the gap. -- **Policy discovery (docs):** tier-1 = volumes+config only (NAS media userdata excluded by design); NAS apps' tier-1 lands *on the NAS*, tier-2 is what gets it off; volume-only apps back up to sys_drive with blank drive label and **no tier-2 copy**. - -### Box state / cleanup (final, 06:53) - -DooPlex NAS restored **md5-identical to baseline** (exports + smb.conf; campaign user/share/dirs/creds removed); services never touched (exportfs-only rail held). Host recovered post-power-cycle; guest 9201 in **defined state**: 6 wave apps deployed + healthy with data, `privatebin` stop+removed via the real flow, verification backup `success:true`, nas-media `ok`, stub 0. Hub untouched throughout. ⚠ Next host reboot re-rolls the F12 dice until the agent template is fixed (interim: systemd drop-in on the automount units). +- **The whole arc ran to completion on a fresh box:** appliance-mode Day-0 (agent 0.85.0 + golden 0.120.0, both sha-verified, "Day-0 provision SUCCESS"), guest at floor 0.120.0 at FIRST report with zero manual steps, ActualBudget deployed through the tunnel, escrow ceremony sealed (R with Viktor only), **AUTO-CONFIRM FIRST LIVE FIRING: pending → escrowed in ~7.5 min with zero clicks**, offsite backup (1 snapshot, 14 s) **and the verification-restore round-trip proven** (48K payload decrypted back from the Storage Box). +- **F-4 (HIGH): G10 is unclosable** — no operator-set dashboard password path exists in shipped code (hub has no per-customer `password_hash` UI/API; controller's open-state page defers to the operator; Day-0 preseeded setup skips the wizard's password form). Every fresh box's dashboard stays OPEN on the internet. Blocks the tester onboarding gate. +- **F-6 (HIGH): identity-only escrow ceremony was never implemented** — on a no-PBS box (the documented appliance standard) the offsite escrow chain can never complete. Mid-drill fork (Viktor): attached the **PBS DR tier (ep0)** instead, which live-rehearsed Peti's exact pending sequence — wrapper+`age` prep → `wg_tunnel.enabled` → hands-free WG peer registration → hub PBS-DR enable (WG-first dependency guard works) → apply-bridge fresh path (consume-once → K born → escrow seed) → ceremony → auto-confirm. +- Fresh-box installer gaps found live: `felhom-pbs-apply` binary not shipped (F-7), `age` not installed (F-10), root-owned `guests/` parents break the non-root agent's lanresolver (F-3, live-fixed; **demo host likely has the same latent state**), silent root@pam rotation surprises the operator (F-8), version strings disagree ×3 (F-1). +- Snapshots on qm 300 for re-drills: `pre-day0-clean` / `post-install` / `post-drill`. Demo 9201, Peti's hub entry, demo offbox untouched. diff --git a/documentation/audits/DRILL-day0-vm-2026-07-12.md b/documentation/audits/DRILL-day0-vm-2026-07-12.md new file mode 100644 index 0000000..82b9417 --- /dev/null +++ b/documentation/audits/DRILL-day0-vm-2026-07-12.md @@ -0,0 +1,174 @@ +# DRILL — Day-0 on the Demo-VM (nested PVE on felhom-pve), 2026-07-12 + +Executed per `RUNBOOK` (Day-0 drill on the Demo-VM). Actors: Viktor (console install, dry-run +go/no-go, fork decisions, the ceremony/R-moment, dialog clicks) + CC (everything scriptable over +SSH + the hub/dashboard browser tracks). **Outcome: the full arc ran to completion — Day-0 +appliance install, first live escrow ceremony + AUTO-CONFIRM FIRST FIRING, offsite backup + +restore round-trip — at the cost of one mid-drill fork (PBS DR tier attached) forced by the +drill's headline findings (F-4, F-6).** + +## 1. Baselines recorded (§1 of the runbook, verified LIVE before Phase 1) + +| Item | Recorded | +|---|---| +| Day-0 artifact manifest | Agent **0.85.0** sha256 `31babb2c4f8fa5a1f961428519a0dfbb3a19b6e536fba01c2664afca339da93d`; Golden **0.120.0** sha256 `f7d7d02c76b49891719df9cf68624c9ae193e438e2a3a63b948a2a9d12887596`; MinAgent **0.81.0** — matches the 0.85/0.120 publish train, no bump needed | +| Global controller floor | Effective **v0.120.0**, source DB (hub_settings); env fallback also v0.120.0 | +| Customer `demo-vm-felhom` | exists, domain `enkisfelhom.hu`, CF tunnel + API tokens saved (API token perms incl. Zone WAF:Edit); git credentials EMPTY (correct per G3; hub edit-form already labels Git Sync "Opcionális" — only day0-install.md A.2 was stale, fixed in this commit); DEBUG MÓD was ON | +| Wildcard DNS | `*.enkisfelhom.hu` → CF proxied edge (104.21.3.175 / 172.67.130.252 + AAAA), verified via 1.1.1.1 on a real and a random subdomain | +| felhom-pve capacity | local-lvm thin pool 348.82g @ 7.47%; RAM available 12.9 GiB; nested virt Y | +| Installer | felhom-host-install.sh **v1.14.0** per `-h` (but see F-1) | +| Golden-baked controller | 0.120.0 (= floor; see 4.3 note) | + +## 2. Per-phase gates + +| Phase | Gate | Result | Evidence (abridged) | +|---|---|---|---| +| 0 — build VM | P0 | **PASS** | qm 300 `drill-day0` (8G/4c/host/250G thin); PVE **9.2.2** ISO 9.2-1; key SSH proven; snapshot `pre-day0-clean` 15:12:41. ISO detached + boot→scsi0 BEFORE the snapshot (clean config) | +| 1 — hub verify | — | **PASS** | §1 table above; nothing created/changed in the hub | +| 2 — box prereqs | P2 | **PASS** | single node; nested local-lvm **149.88 GiB** free (≥120 ✓ — note: a 250G VM disk yields ~150G nested pool after the installer's root/swap split); vmid 9201 absent; hub 302 / gitea 200 / felhom.eu 200 | +| 3 — install | P3 | **PASS** | dry-run reviewed by Viktor → GO → real run 15:25–15:28: **"Day-0 provision SUCCESS — vmid=9201 host_id=demo-vm-felhom-2482b0"**; both artifacts sha-verified vs the hub manifest; both expected ANONYMOUS-fetch warns; 4b break-glass vaulted; passphrase file 0600, shredded after (never in CC's transcript — Viktor wrote it himself) | +| 4 — post-install | D.1–D.4 | **PASS** (1 finding) | see §3 | +| 5 — hub tracks | G9/G10 | **PREMISE COLLAPSED** — F-4/F-5; geo exercised via the shipped path | see §4 | +| 6 — escrow/auto-confirm/offsite | fork-4 | **PASS after fork** (PBS DR attached) — F-6/F-7/F-10/F-11 | see §5 | + +## 3. Phase 4 detail (post-install verification) + +- **D.1:** agent active as non-root `felhom-agent`; full read-only selftest ALL-OK incl. pool read + (pool `felhom`, member 9201); guest running, onboot=1, mounts mp0 docker 50G / mp1 sys 20G / + mp8 felhom-drives / rootfs 32G; in-guest containers controller 0.120.0 (healthy) + filebrowser + + cloudflared + traefik; dashboard 200 in-guest (traefik https + Host header; :80 → 301). +- **4.2 appliance gating (first live appliance box):** agent.json `deployment_mode="appliance"`; + journal `selfheal: node watchdog starting mode=appliance interval_s=60` + storage watchdog + armed; felhom-mgmt-watchdog.timer enabled+active. NOT byo-defaulted. +- **4.3 floor self-update (tester-recruitment gate):** guest landed **0.120.0 = floor at its FIRST + report, zero manual steps, elapsed ≈ 0**. LIMITATION: baked == floor on this train, so the + floor-driven *update* path was not stressed — re-prove on the next train where golden < floor. +- **4.4 hub:** customer PENDING→ok ~2 min after provision; host ONLINE agent 0.85.0; guest 9201 + appears at the next agent report (900 s cadence — the empty Guests panel in between is timing, + not a bug); only-degraded capabilities = the 3 `pbsdr-*` ("binary not found") — expected no-PBS + shape (and see F-7). +- **4.5 customer-visible via the REAL edge** (method: `curl --resolve` on both CF anycast IPs — + the split-horizon-proof variant): felhom.enkisfelhom.hu → 200 "Vezérlőpult", **OPEN, no auth** + (G10 "before", ~15:30). +- **4.6 first-app smoke:** ActualBudget via the open dashboard → "Telepítés sikeres"; real-edge + 200 + `Actual` on budget.enkisfelhom.hu. (The old POST-via-public-URL no-op + gotcha did NOT reproduce on 0.120.0 for normal forms.) +- Snapshot `post-install` 15:35:37 (live, no fs-freeze — no qemu-guest-agent in the drill VM). + +## 4. Phase 5 detail — the G9/G10 premise vs shipped code + +- **5.1 operator password-set: IMPOSSIBLE (F-4).** No hub UI/API sets per-customer + `web.password_hash`; the controller's security page in the open state says "Kérd az + üzemeltetőt"; `settingsPasswordHandler` requires a current-password match (no initial-set); + the hub-preseeded Day-0 setup path skips the only wizard form having a password field. +- **5.2 geo-restriction:** apply lives in the CUSTOMER dashboard (controller + `settings_security.html` + `api/geo.go`), not the hub (hub has only `handleGeoDisable`) — F-5. + Exercised via the shipped path: HU-only enabled → controller created WAF rule + **"[felhom-geo] Global"** (`(not ip.src.country in {"HU"})`, action block) on the zone, sync + reported 1 active rule; HU access still 200 through the real edge. Non-HU block not testable + from an HU vantage. +- **5.3 G10 statement:** BEFORE — dashboard OPEN (proven in-guest + real-edge, ~15:30). AFTER — + **STILL OPEN; closure impossible until F-4 ships.** The tester-agreement onboarding gate that + points at this line CANNOT currently be satisfied. + +## 5. Phase 6 detail — offsite → ceremony → auto-confirm → tier proof + +Pre-ceremony state (all verified): offbox pre-provisioned at customer-create +(u629488-sub3@…your-storagebox.de:23, /home/felhom-repo, pinned host fingerprint, 0/50 GB); +controller `EscrowState=pending` — the /backups banner "a mentés addig nem fut" (the F6 +no-single-copy guard, LIVE); staged `escrow-stage/restic_repo_password` on the agent (0600, +staged at provision — fork-4 stage-FIRST ordering held); ActualBudget toggled for NAS. + +**The ceremony blocked → the drill's second headline (F-6):** `--selftest=escrow-create --upload` +refuses without a PBS storage/key; **identity-only mode does not exist in any shipped agent** +(the ≥0.80.0 claim in RUNBOOK-escrow-ceremony.md was false — fixed in this commit). On a no-PBS +appliance box the offsite arc can NEVER complete. **Fork decision (Viktor): attach the PBS DR +tier (ep0)** — which converted the drill into a full rehearsal of Peti's pending sequence: + +1. Ship `configs/felhom-pbs-apply` → /usr/local/sbin (the binary is missing from host-install — + F-7; the FELHOM_PBSDR sudoers alias DOES ship). Agent restart → zero capability-DEGRADED. +2. Hub PBS-DR enable failed correctly: "host has not reported a WG key yet — the tunnel peer must + exist before the PBS DR tier" — the dependency is hub-enforced (good), and the drill runbook's + "no WG" scope was inconsistent with any PBS path. +3. WG: `wg_tunnel.enabled=true` in agent.json (installer never sets it — decide the appliance + default together with the F-6 spec) → keygen → **hands-free hub registration** (10.77.0.3/32, + gen 1, no operator vouch) → conf applied → handshake + ping 10.77.0.1 (33 ms). +4. Hub PBS-DR enable → "Configuration updated" → descriptor gen 2; apply-bridge first 403'd + (`Datastore.Allocate` on /storage/felhom-pbs) — consequence of the drill's narrowed + `--acl-storages "local local-lvm"`; the installer DEFAULT includes felhom-pbs exactly for + this. Retrofit dual-grant → next tick: **token consumed (single-use) → entry + K created → + `escrow.pbs_storage_id` seeded → state=applied** (16:17). pvesm ACTIVE; felhom-pbs.{enc,pw} + present. +5. **Ceremony** (Viktor, R on paper, nothing in CC's transcript): attempt 1 failed — `age` + missing (F-10, `apt-get install -y age` → 1.2.1); attempt 2 SUCCESS ~16:35; the bundle + auto-captured `+wg_private_key +restic_repo_password`; ceremony wiped the staged secret; hub + host page flipped to **DR RECIPE: present / KEY ESCROW: present**. +6. **AUTO-CONFIRM FIRST LIVE FIRING:** hands off, "Letét megerősítése" untouched → 16:42:36 + controller log: *"hub-verified: the escrow covers the current repo password (hash + 99c16e8d84e7…) — EscrowState auto-confirmed escrowed; offsite runs enabled."* + **Elapsed ≈ 7.5 min, zero clicks.** /backups gate banner gone. +7. **Tier proof:** "NAS-mentés most" → repo initialized on the Storage Box → 1 snapshot, 14 s, + "✓ Rendben". Restore-to-verify (native confirm() blocked automation twice and Viktor missed + the popup — F-11; also live operator confusion: he first launched the FULL local tier-1 + restore, which itself completed healthy in 8.7 s) → offbox verification restore → + *"restored actualbudget → …/offbox-restore/actualbudget"*, 48K payload (app.yaml, .felhom.yml, + docker-compose.yml, manifest.json) decrypted from the box — **the offsite round-trip is + proven**. Bonus: the full local tier-1 restore was ALSO proven the same afternoon. + +## 6. Headline metrics + +| Metric | Value | +|---|---| +| Wall-clock Phase 2 → 6 complete | ~15:15 → ~17:09 (**~1 h 55 m**, including the mid-drill fork, 3 fix-and-continue stops, and two supervised STOPs) | +| Install run itself (Phase 3) | ~3 min | +| Floor self-update (4.3) | at-floor at FIRST report, 0 manual steps (baked == floor; update path not stressed) | +| **Auto-confirm (6.4)** | **~7.5 min ceremony→escrowed, zero manual clicks (FIRST LIVE FIRING)** | +| Offsite backup / restore | 14 s backup (1 snapshot) / verification restore round-trip proven | + +## 7. Findings + +| # | Sev | Phase | Finding | Disposition | +|---|---|---|---|---| +| F-1 | LOW | 3 | Version-string mismatches: `-h` v1.14.0 vs run banner v1.13.0 vs hub Setup-tab copy "1.12.0" | fix strings (installer + hub template) | +| F-2 | COSMETIC | 3 | dry-run prints `curl -u ` on the anonymous-fetch branch | fix placeholder | +| F-3 | MEDIUM | 4 | Root-run provision leaves `/var/lib/felhom-agent/guests{,/9201}` root:root 0700 inside the agent-owned state dir → non-root agent lanresolver "permission denied". LIVE-FIXED (chown the two parent dirs; the guest-root-owned bootstrap subtree untouched) | agent/installer: create parents agent-owned at provision; **check demo/felhom-pve for the same latent state** | +| **F-4** | **HIGH** | 5 | **No operator-set dashboard password path exists anywhere** (hub has no UI/API for per-customer `password_hash`; controller open-state page defers to the operator; Day-0 preseeded path skips the wizard's password form) → **G10 unclosable; every fresh box's dashboard stays OPEN on the internet** | hub feature task (operator set → config-delivered hash → controller re-pull); blocks tester onboarding gate | +| F-5 | MEDIUM | 5 | Geo-restriction APPLY is customer-dashboard-side; hub only disables. Runbook premise stale; also: the open dashboard (F-4) exposes the geo toggle unauthenticated | doc/design decision; folded into F-4's arc | +| **F-6** | **HIGH** | 6 | **Identity-only escrow ceremony was never implemented** (`escrow-create` hard-requires PBS storage + key; the ≥0.80.0 runbook claim was false) → on a no-PBS box (the documented appliance standard!) the offsite escrow chain can never complete | agent feature task (identity-only mode); ceremony runbook corrected in this commit | +| F-7 | MEDIUM | 6 | host-install ships the FELHOM_PBSDR sudoers alias but NOT the `felhom-pbs-apply` binary → pbsdr capabilities born DEGRADED on every fresh box | installer: ship the wrapper (like mkfs/selfupdate wrappers) | +| F-8 | LOW/UX | 3 | Step 4b rotates root@pam + vaults silently — operator surprised by 401 at the PVE GUI (live: Viktor) | installer: print "root@pam rotated + vaulted — retrieve at hub → host page" | +| F-9 | NOTE | 6 | Installer never sets `wg_tunnel.enabled`; WG registration itself is hands-free once enabled | decide appliance default alongside the F-6 spec | +| F-10 | MEDIUM | 6 | `age` (ceremony identity-wrap dependency) not installed by host-install — fresh-box ceremony dies; demo host masked it (spike-era install) | installer: add `age` to package set; runbook prereq added in this commit | +| F-11 | LOW | 6 | Native `confirm()` on the offbox restore-verify form: blocks browser automation, easy to miss (operator missed it twice live; meanwhile launched the full tier-1 restore from the adjacent form — the two restore controls invite confusion) | convert to the design-system inline confirm pattern; consider renaming | + +Observations (not F-numbered): hub "REGISTRY LATEST v0.120.0 — up to date" vs the dashboard's own +"Új controller verzió elérhető: 0.121.0" banner disagree on "latest"; the hub dashboard row for +Peti shows "OK · minutes ago" (controller-derived, via his proxmox2 migration) while his agent +HOST is DOWN 23h — the roll-up masks a dead host; retrofit-ACL note: adding the PBS tier to a box +installed with narrowed `--acl-storages` needs the /storage/ dual-grant (documented default +avoids it). + +## 8. What this proved for Peti — and what it did not + +**Proven live on a fresh box (his exact pending sequence):** wrapper+age prep → WG enable → +hands-free peer registration → hub PBS-DR enable (dependency guard works) → apply-bridge fresh +path (consume → K → grant → escrow seed) → ceremony (with K) → **auto-confirm** → gated offsite +run → restore round-trip. Prep list for his box: `felhom-pbs-apply` + `age` + `wg_tunnel.enabled` ++ (if his ACL set was narrowed) the /storage dual-grant. + +**Deliberately NOT covered:** WG OOB operator peer (separate arc); S5 DR restore from total loss +(identity-consume — queued as its own drill); byo-mode re-run from `pre-day0-clean` (queued); +non-HU geo-block verification (needs a non-HU vantage). + +## 9. Snapshot inventory (qm 300 on felhom-pve, at drill end) + +| Snapshot | When | State | +|---|---|---| +| `pre-day0-clean` | 15:12 | PVE 9.2.2 + keys, nothing Felhom — the universal re-drill zero | +| `post-install` | 15:35 | Day-0 SUCCESS + ActualBudget + guests-dir chown | +| `post-drill` | 17:11 | full drill end state (WG + PBS-DR + escrowed + geo HU + dashboard OPEN per F-4) | + +Blast radius honored: guest demo 9201 on felhom-pve, Peti's hub entry, and the demo offbox were +never touched. Drill leftovers to tidy at teardown (deliberately kept for re-drills now): hub +customer `demo-vm-felhom` (its api_key + CF tokens surfaced in the operator UI during the drill — +delete/rotate at teardown), DEBUG MÓD on, geo HU rule on the drill zone, VM 300 running. diff --git a/documentation/runbooks/RUNBOOK-escrow-ceremony.md b/documentation/runbooks/RUNBOOK-escrow-ceremony.md index 8f6806b..dfad157 100644 --- a/documentation/runbooks/RUNBOOK-escrow-ceremony.md +++ b/documentation/runbooks/RUNBOOK-escrow-ceremony.md @@ -18,8 +18,15 @@ ## Prerequisites (check BEFORE scheduling with the customer) - **Agent version:** ≥ v0.79.0 (the ceremony records `restic_pw_sha256` — older agents produce a blob - auto-confirm can never match). **No-PBS hosts** (BYO without the PBS tier): ≥ **v0.80.0** - (identity-only mode; below that the ceremony hard-requires the PBS key and refuses). + auto-confirm can never match). +- **⚠ No-PBS hosts: the ceremony CANNOT run.** The previously documented "identity-only mode + (≥ v0.80.0)" was NEVER implemented — v0.80.0's actual feature was seeding `escrow.pbs_storage_id` + on PBS hosts. `escrow-create` hard-requires a PBS storage id + its key file (drill-proven + 2026-07-12, finding F-6 of DRILL-day0-vm-2026-07-12.md). Until identity-only ships, a box MUST + have the PBS DR tier (which itself requires the WG tunnel peer first) before any escrow/offsite + arc can complete. +- **Host packages:** `age` must be installed (identity wrap dependency; NOT installed by + host-install as of v1.14.0 — drill finding F-10). `apt-get install -y age`. - **K gate (PBS hosts):** `escrow.pbs_storage_id` set and the key file present (`cfg.Backup.PBSEncKeyPath()`). - **Staged secret (offsite):** offsite enabled → `EscrowState="pending"` on the controller and the staged @@ -31,8 +38,9 @@ ```bash felhom-agent --selftest=escrow-create --upload ``` -- `--storage ` only if `escrow.pbs_storage_id` isn't configured. No-PBS hosts (agent - ≥0.80.0): omit — identity-only engages automatically. +- `--storage ` only if `escrow.pbs_storage_id` isn't configured (the PBS DR + apply-bridge seeds it automatically). No-PBS hosts: see the prerequisite warning above — the + ceremony refuses without a PBS key; identity-only mode does not exist yet. - What it does, in order: generates a fresh **R** (EFF-wordlist passphrase; entropy printed) → seals K (if present, via proxmox-backup-client re-key) and the IdentityBundle (age-under-R; the staged restic password auto-injected; the live WG key auto-captured if present) → **self-verifies by recovering its diff --git a/documentation/runbooks/day0-install.md b/documentation/runbooks/day0-install.md index 551cea0..b120be2 100644 --- a/documentation/runbooks/day0-install.md +++ b/documentation/runbooks/day0-install.md @@ -63,8 +63,8 @@ Hub UI (`https://hub.felhom.eu`, operator password) → **Customers → New**: | Email | customer's email | used for customer-tier notifications (Hungarian) | | CF tunnel token | from A.1 | → `infrastructure.cf_tunnel_token` | | CF API token | from A.1 (optional) | → `infrastructure.cf_api_token` (geo rules) | -| Git username | Gitea read account | → `git.username` — **required for Day-0** | -| Git token | Gitea read token | → `git.token` — **required for Day-0**: the install script and the controller fetch artifacts from Gitea with this credential; without it the install dies at step 5/8 | +| Git username | Gitea read account | → `git.username` — **optional** (only for a private app catalog) | +| Git token | Gitea read token | → `git.token` — **optional** since installer v1.11.2 (the G3 ruling): artifacts are world-readable and fetched ANONYMOUSLY, sha256-verified against the hub manifest. Empty is the normal customer shape; expect the installer's "fetching artifacts ANONYMOUSLY" warn line in steps 5/8 (drill-verified 2026-07-12 on the VM Day-0 drill). | On save the hub generates two credentials: @@ -405,7 +405,7 @@ recovery credential). | byo dies: "acl storage(s) not found on this box" | `--acl-storages` names a storage the box lacks | pass the box's real storages (check `pvesm status`) | | step 1 dies: "this is a N-node cluster" | multi-node cluster | re-run with `--node ` | | step 5 dies: "hub artifact manifest has no agent version" | Day-0 manifest unset/incomplete | Part A.3 — set it in the operator UI | -| step 5 dies: "no git token in controller.yaml" | customer created without git credentials | Part A.2 — add `git.username`/`git.token`, regenerate config | +| step 5 warns: "no git credential — fetching artifacts ANONYMOUSLY" | customer has no git credentials — the NORMAL shape since installer v1.11.2 | nothing to do; sha256 verification is unchanged. (Pre-v1.11.2 installers die here instead — upgrade the script.) | | step 1: passphrase REJECTED (401) | typo / wrong customer | re-check with the hub UI's printed curl command | | step 8 fails: "CT already exists" | vmid collision with a hub-invisible guest | pick from `pct list` + `qm list` (Part B); the agent destroys nothing on collision — re-run with a free vmid and `--resume` | | controller container missing in-guest after provision (docker ps empty) | golden older than 0.98.3 (no baked path unit) + pre-v1.9.1 script: the controller-bootstrap unit's boot-time condition lost the race with the bootstrap-mount attach | `pct reboot ` — the unit runs on the next boot (v1.9.1 reboots itself; goldens ≥ 0.98.3 bake a path unit that starts the controller on the mount hot-plug, no reboot needed) |