Files
felhom.eu/documentation/runbooks/RUNBOOK-onboarding-draft-v4.md
T
admin a12c6f9730 docs(v1.24.0): R-59/R-60/R-61 SHIPPED — CHANGELOG, README, ROADMAP (+R-62), runbook, capability map, drill evidence, REPORT
Virgin-ISO nested drill closed the train: dead-NIC install baked the
fallback (incl. the dead default gateway), the R-59 screen painted
(capture committed beside the spike doc), the cable move healed +
registered at the hub in 23s unaided, and the build's rootpw file
matched the installed box's shadow hash. R-59 SHIPPED with the recorded
deviation (first-boot gate; installer-initrd abort out of scope by
operator ack). R-60 SHIPPED (spike + drill cited; F-P9 route-flush fix
included). R-61 slice 1 SHIPPED. New R-62 row (hub delete-dialog
cosmetics, XS). Capability map: new PROVEN-LIVE row (nested != metal,
said so). Cleanup verified: felhom-pve interfaces byte-identical,
bridge/VMs/ISO removed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
2026-07-22 11:45:55 +02:00

246 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DRAFT R-3 — Friend-alpha onboarding runbook (operator)
> **Status: DRAFT, v4 — refreshed 2026-07-22.** Generalized from
> `runbooks/RUNBOOK-rehearsal-v3.md` (S1S8) and the shipped claim/escrow/self-bind flows, refined
> against the N100 rehearsal that executed it (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`),
> and now carrying a **second full onboarding datapoint on virgin hardware** — the HP t740
> (`demo-hp`, 2026-07-21; `operations/nodes.md`, `CONTEXT.md`). This refresh deleted the R-36 and
> R-39 interim workarounds (both shipped and proven) and updated the freemail/loader/golden notes.
> Every step here has been walked on real hardware **except C7**, the customer-performed restore.
>
> **It graduates to `RUNBOOK-onboarding.md` when C7 has evidence** — that step is simultaneously the
> last unwalked step of this script and the capability map's last MISSING customer-facing row, so
> one execution closes both. Until then this stays a DRAFT.
>
> Scope: **appliance hardware only** (BYO track is R-1/Peti, explicitly out of alpha scope).
> One step below still carries an **interim workaround for a known defect** (R-31) — it is marked
> inline and should be deleted from this runbook the moment that item ships.
## A. Before the visit (operator, ~30 min, all remote-doable)
1. **Hub:** create the customer (name, e-mail, domain `<név>.felhom.eu`) — verify the CF tunnel +
zone token provisioning completes. Confirm the Day-0 manifest is current (agent, golden,
MinAgent) and the managed-update floor is at the intended release — a stale golden is fine
(the floor lifts the box post-install), a stale FLOOR is not. *The current golden (0.153.0)
carries all four infra images, and the floor-lift is now proven on TWO shapes:* the rehearsal
(0.143.0 → 0.145.0 in **5 s**, unattended — `tests/VALIDATION-n100-rehearsal-2026-07-18.md`)
AND a virgin board on first contact (0.153.0 → 0.156.0 during day-0, HP t740, 2026-07-21 —
`CONTEXT.md`). Golden freshness before the first real tester is the standing rule in **F**
apply it there, don't re-derive it here.
**Verify the self-bind link went out.** Since hub v0.67.0 (R-36) the link is **auto-minted at
customer creation and at RESET completion**, so the console banner's „e-mailben kapott link" is
already true by construction; "Send self-bind link" on the Setup tab is the manual re-mint if
the customer lost it or the 7-day token expired. See B3.
2. **Hardware:** the standard appliance + its external HDD. One-time bench prep list:
memtest pass, SMART clean, BIOS queue for the single visit in B2.
3. **Install media:** the **reusable generic pairing ISO** (secret-free, slice C). Only rebuild it
when `scripts/` ships a new version — never per-customer. Flash via dd/Rufus-DD.
**Console credential for bench work (R-61, since v1.24.0):** every build emits the baked root
password into the 0600 sibling **`<iso>.rootpw.txt`** next to the ISO in the build output — that
file is how you log into the console of any box installed from that build (bench checks, dead-
network diagnosis). Operator-only; never commit or paste it anywhere; the G1 break-glass vault
remains the mechanism for a *lost* password on an enrolled box.
Loader per firmware — record the board model + loader choice in the customer record; full
fleet inventory in `operations/nodes.md`:
| Board / firmware | Loader | Secure Boot |
|---|---|---|
| N100 (AMI AN3PLUS) | `mkimage` | **OFF** — that firmware can't relocate the signed GRUB (F1, `tests/VALIDATION-n100-baremetal-2026-07-16.md`) |
| **HP t740 (AMI M42)** | **`shim`** | **ENABLED** (proven 2026-07-21, `operations/nodes.md`) |
**SB-off is a per-firmware workaround, not a Felhom requirement** — the shim loader boots with
Secure Boot enforcing on firmware that handles it (the HP proved it).
4. **Paper:** the tester one-pager (R-11, printed or PDF), the tester agreement
(PETI-agreement sibling), and the recovery-code envelope step rehearsed in your head — the
escrow ceremony is the one moment that must not be rushed.
5. **Agent train ≥ v0.93.0 on the customer's box before the FIRST real tester's escrow ceremony.**
v0.93.0 removes the four hyphenated EFF words from the recovery-code generator
(`felhom-agent/CHANGELOG.md` v0.93.0, 2026-07-21) — ceremonies run **box-side**, so the fix
protects a ceremony only once the train is **BUILT + DEPLOYED to that box**; source-only does
not count. Already-issued codes stay valid.
**DISCHARGED 2026-07-22** — agent 0.93.0 published (sha `a68b2ff7…`), Day-0-manifest-vouched,
and deployed to both fleet boxes (`demo-felhom-8363b5` + `demo-hp-bb76ea`), clean-restart
verified; any future manifest-driven install also lands ≥ 0.93.0. Record:
`pilot/RUNBOOK-publish-agent-0.93-2026-07-22.md`.
## B. Install + bind (~15 min wall-clock)
1. Boot the stick. ONE BIOS visit: boot order → USB, Secure Boot per loader, **restore-after-AC-loss
→ Power On** (a power blip must never strand a household).
2. Unattended install (~3 min) → reboot → pairing mode: the console shows the Hungarian
**pairing-code banner**; the box appears under **Unclaimed appliances** at the hub
(verify uuid/MAC/SSH fingerprints match the physical box before binding — never bind a
fingerprint you didn't read off the console).
> **Multi-NIC gotcha: verify the CABLED port got the DHCP lease during the install.** On the
> HP t740 the cable was on the 4-port expansion card, which gets **no lease** — and instead of
> aborting, the installer baked its `192.168.100.2` fallback as a **static** config and
> completed: a box that looked installed and could never call home
> (`operations/nodes.md` — the NIC map and the trap). **Since ISO v1.24.0 the box handles this
> itself (R-59+R-60):** if the hub is unreachable on first boot it sweeps every NIC (bounded
> DHCP + hub probe) and keeps the first that works — a cable in the wrong port just costs the
> sweep a minute — and when nothing works it paints a legible Hungarian screen with the NIC
> table (név/MAC/kábel/sebesség) and retries every minute. Nested-drill proven
> (`audits/SPIKE-firstboot-nic-sweep-2026-07-22.md`); on a pre-v1.24.0 stick the old manual
> check still applies: confirm at the installer's network step that the leasing interface is
> the cabled one.
3. **Bind** to the customer — **the rehearsal used the SELF-BIND path, and it is now the default for
this runbook.** The link is already in the customer's inbox (auto-minted at creation, A1); the
customer opens it and completes the public two-factor `/bind/<token>` page (console pairing code
+ retrieval passphrase). The 2026-07-18 run bound in **60 s from mint to consume, with zero failed
attempts**, and the credential was delivered **26 s later** with no operator action. **Second
datapoint, virgin hardware:** the 2026-07-21 HP t740 onboarding walked the full chain — register
→ pairing banner → unclaimed list → self-bind (code + passphrase) → day-0 → dashboard — on
hardware the system had never seen (`CONTEXT.md` 2026-07-21 entry, `operations/nodes.md`).
> **The self-bind link exists BEFORE the customer sees the console.** The console banner tells
> them to use „az e-mailben kapott link" — since hub v0.67.0 (R-36) the link is auto-minted at
> customer creation and RESET completion, so the banner is true by construction (in the
> pre-auto-mint rehearsal the box sat in pairing mode for **~11.7 minutes** waiting on the
> operator, printing failures the whole time — R-33). Verify the mint in A1; the token is valid
> for **7 days**.
Operator-bind on the Hosts page remains the fallback when the customer is not present or the
token has expired.
4. Watch the event sequence compose (expected: **seconds** between post-enroll legs — enroll →
claim-code issue → WG → fast-tick window → tunnel → PBS-DR descriptor poke → controller up →
claim e-mail lands → floor lift to current). **Any ~15-minute stall between two legs is a
finding** — capture which leg + wall-clocks.
**Reference wall-clocks — measured on the N100, 2026-07-18** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`).
Compare against these; a leg that takes minutes where this table says seconds is the finding:
| Leg | Expect |
|---|---|
| Blank-metal boot → registered as an unclaimed appliance | **< 1 s** (same second) |
| Bind → credential delivered | **26 s** |
| Bind → agent enrolled + WG tunnel up | **56 s** |
| Bind → controller running the CURRENT version | **2 min 44 s** |
| Credential delivered → day-0 provision SUCCESS | **2 min 11 s** |
| Managed-floor lift (golden version → current) | **5 s**, unattended |
| Drive enrol: destructive-confirm → usable (incl. mkfs) | **~30 s** |
| Escrow ceremony started → offsite runs enabled | **12 s** |
The **PBS-DR descriptor** leg no longer needs special distrust: the DR tier now self-detects a
dead credential (a loud `auth_failed`) and self-heals via damped re-issue, and the hub gauge is
trustworthy (**R-39 CLOSED, proven live 2026-07-21** — hub 0.68.1 + agent 0.91.2, the 13 s
re-issue chain `applied → auth_failed → applied`; `backlog/ROADMAP.md` R-39).
## C. Customer zero script (the customer drives, operator narrates — ~30 min)
This is the friction lens: the customer's hands on the keyboard, the operator's hands OFF.
1. **Claim:** open the claim e-mail → claim the dashboard → set the dashboard password.
(R-4 gate: the e-mail must land in inbox, not spam. **gmail PROVEN 2026-07-18** — the rehearsal's
claim mail was the first sent under the tightened DMARC `p=quarantine` and landed in the **Inbox**.
**freemail.hu PROVEN 2026-07-21** — an operator test-send was received at a freemail.hu address;
R-4 is COMPLETE (`backlog/ROADMAP.md` R-4). Keep the habit anyway: if a mail seems missing, have
them check spam before you conclude it was lost.)
2. **Escrow ceremony:** run the recovery-code wizard → the awaiting card
(„megerősítésre vár, legfeljebb 15 perc") → flips on ACK. The code goes on paper, into the
envelope, into the customer's drawer — **state plainly: without this code, a dead box means
the remote backup is unrecoverable; we cannot read it for them.** That sentence is the product.
3. **Storage:** enroll/format the external HDD via the wizard (customer clicks, operator explains
the destructive-confirm). Expect **~30 s** from the confirm to a usable drive — the wizard runs
mkfs detached and polls, so a slow format is not a hang.
4. **Offsite:** enable „Távoli mentés". One operator behaviour here is a **workaround for a known
defect** — follow it until the item ships, and do not improvise:
> **Click Save exactly ONCE, then wait, then verify.** Provisioning runs synchronously inside the
> request, so on a slow Hetzner leg the request can return an nginx **504 while succeeding
> server-side**. A second click races the first attempt. Wait, then confirm on the hub's
> Offsite → Restic tab rather than believing the browser. (**R-31**, still open)
The old re-onboarding manual-Save check is gone: since hub v0.67.0 (R-36 SHIPPED,
`backlog/ROADMAP.md`) an enabled-but-unprovisioned offsite target raises an **amber banner on
the customer page** naming the state and the fix — trust the banner.
> **Expect the first run to be clean on a NEW customer.** If an **orphan card** appears
> („a távoli mentési tároló elárvult"), the repo is not new — that is the R-32 case, which fires
> when a customer's pool directory outlived a RESET. It is safe (the guard moves the old repo
> aside, never deletes it) but it means the box is writing into a reused directory. Record it.
5. **First app:** deploy one app the household actually wants (paperless or immich recommended) —
through the catalog UI, no shortcuts.
6. **Sharing:** enable „Megosztás", set the household SMB password, create one share; open
`\\NÉV` from their own laptop. (Set expectations per the one-pager: device list may lag on
some routers; the address always works.)
7. **Protection proof — the customer performs a restore, unassisted** (this step IS the
capability-map evidence, per customer): drop a file on the share or into the app → „Távoli
mentés most" → wait for the snapshot → delete the file → the customer restores it via
**Visszaállítás** with the operator silent. If they can't without help, that's an R-11/UX
finding, not a customer failure — write it down verbatim.
> **⚠ THIS IS THE ONE STEP THE REHEARSAL DID NOT EXECUTE.** Every other step in this document was
> walked on real hardware on 2026-07-18; C7 was not, so it carries **no friction notes and no
> evidence** — the capability-map row *"a customer (not the operator) performs a restore via UI
> alone"* is still **MISSING**. This runbook stays a **DRAFT** until C7 has been performed once.
> It is remote-doable on the reborn demo box (the dashboard is remote) and does not need the
> hardware in hand.
8. **Alerts sanity:** show them what an alert e-mail looks like (the health-degradation mail);
confirm their address receives it.
## D. Handover (10 min)
- Walk the one-pager together: what it does, what it doesn't do yet, how to report.
- Reporting channel: **support@felhom.eu** (catch-all already delivers it) + the personal channel
for the friend circle. „Ha bármi furcsa: e-mail vagy üzenet, screenshot ha van."
- Sign the tester agreement. Leave the one-pager with them.
- Operator record: board model + loader, drive serials, claim date, escrow ACK timestamp,
friction notes.
- If this relationship ever ends, the levers are **RULED and shipped** (R-25b, hub v0.69.0/v0.70.0,
TASK-I): full teardown = the guided **DELETE cascade**; identity-preserving re-onboarding =
**RESET** (the Friday flow). Which lever and why: **§G below**.
## E. Operator post-checks (same day, remote)
1. Hub: customer row green; Offsite → Restic tab shows the first bytes within a report cycle;
PBS-DR gauge nominal (first PBS run on its own schedule — note, don't wait). The gauge is
trustworthy: since R-39 closed (hub 0.68.1 + agent 0.91.2, proven live 2026-07-21 —
`backlog/ROADMAP.md` R-39), a dead DR credential surfaces as a loud `auth_failed` and self-heals
via damped re-issue; no manual `pvesm status` check is needed.
2. Floor convergence: controller at current, agent at MinAgent+.
3. Route friction notes: UX items → ROADMAP; one-pager corrections → R-11; runbook corrections →
this doc.
## F. Standing rules
- Never perform steps C1C7 *for* the customer. The script's value is watching them do it.
- Never handle the recovery code — see it into the envelope, never into your hands or a photo.
- One onboarding per day maximum during alpha — the friction notes are the point, and they
degrade after the first.
- Post-rehearsal checklist item: rebuild the golden to the current controller before the FIRST
real tester (fresh boxes shouldn't sit in a stale window longer than needed), and re-verify the
Day-0 manifest after every controller release during alpha.
## G. Offboarding — which lever to pull (hub v0.69.0, R-25b)
Two levers, and they are not interchangeable:
- **Ügyfél-visszaállítás (RESET)** — *identity-preserving re-onboarding*. Every operational trace
dies (offsite repository, PBS namespace, DR recipe, one-time password, claim state); the customer
record, the basic config and the audit history survive. This is Peti's Friday flow: the same
customer coming back on new hardware. RESET still **refuses while any host row exists** — delete
the host(s) first, on purpose.
- **Delete customer… (the DELETE cascade)** — *full teardown, the customer is leaving*. One guided
dialog shows the live inventory, takes **three acknowledgements** (1: the host(s) will be deleted,
custody demoted not destroyed; 2: the customer will be RESET incl. external teardown and offsite
data destruction; 3: the record and ALL escrow ciphertext are purged) plus the **typed
customer-id**, then runs `hosts → RESET → purge` in that order.
Notes for the operator:
- An **ONLINE host refuses the whole cascade** — decommission the box first (a live agent would get
permanent 401s). The dialog says so and disables the button.
- If a leg fails (a slow Hetzner call), **nothing is left half-forgotten**: the journal is retained,
the error names the leg, and re-opening the dialog offers **Resume**. Completed legs are no-ops.
- The **audit event stream survives** the delete — the customer_deleted event and the host-deletion
provenance outlive the record, by design.
- **A deleted customer disappears from the Customers list immediately** (hub v0.70.0). If one lingers
with a `MANUAL` config badge, it is a **ghost** left by a pre-v0.70.0 delete: the record is gone but
its report stream survives, which both keeps it on the list and keeps the staleness/offsite checkers
emailing you about it. Open **Delete customer…** on it and confirm — the dialog labels it as a ghost
and the cascade clears the residue. (With no config row the Hetzner leg cannot be determined and is
recorded as `skipped_no_config` — verify the Hetzner side by hand for that customer.)