Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GzammAMzsJTgpQHqxwM2bC
16 KiB
RUNBOOK — Peti's return: convergence, the parked train, the first real-customer onboarding, and the credential rotations
0. Scope & standing rules
- In scope: convergence verification, the publish train's parked D/E/G at CURRENT versions, claim, WG consent, DR-tier migration, ceremony + auto-confirm, offsite opt-in, motioneye remediation, CF token rotation (BOTH boxes — see the trap in Phase 4), hub bearer rotation.
- Out of scope: cluster/agent-follows-guest work (his proxmox1/proxmox2 HA shape is a known roadmap item — if the guest migrates mid-runbook, STOP and reassess); the old-box archive (u629193-sub1) retirement decision — separate.
- Custody rule for the whole runbook: Peti's password and R are PETI'S. They never appear on Viktor's or CC's screen, in transcripts, or in the report. Where a step displays them, Peti drives.
- No gate-lifts anywhere (standing rule). Diagnose-before-fix on every anomaly.
Phase 0 — TODAY, before anything technical (Viktor, ~15 min)
0a. Hub → Peti's customer page, record: current controller version (self-update may be mid-flight or done), claim state (code issued? issued-at? claimed?), the notification/event log — specifically whether the motioneye 100% warning ever emailed him. That answer is the notification pipeline's first real-customer test; record it either way, and if it never fired, that is a FINDING with its own follow-up task.
0b. Message Peti (before the machine talks to him): fan congratulations; "your box is updating itself over the next hours — you'll get (or already have) an email with a setup code, it's real, don't delete it"; camera drive is full, we'll fix it together; propose the onboarding call (~90 min, needs him at a keyboard with root on his own box).
0c. Sign nothing yet. If the claim email was never issued or landed in spam, note it — the resend button is Phase 3 ammunition, not a Phase 0 action.
Gate P0: Peti has acknowledged and a call slot exists.
Phase 1 — passive convergence check (CC, strictly read-only)
Confirm from hub data only: controller reached 0.122.0 (floor); the claim-code hash was ACK-delivered; the box transitioned legacy-banner → gate armed; reports healthy throughout; the roll-up row honest (it will read WARN "storage" territory only if the storage health warns — motioneye's full drive may legitimately color it). If the controller is NOT converging: diagnose first (report errors? MinAgent mismatch? update loop?), report, and STOP for a ruling — do not push anything.
Gate P1: controller ≥ 0.122.0, gate armed, box healthy. Record the self-update elapsed time (fleet's second floor-proof, first on customer hardware).
Phase 2 — the parked train phases, at current versions (CC executes, Viktor signs)
Execute RUNBOOK-publish-0.85-0.120 Phases D/E/G + gates 0b/0g against peti-felhom-86d37d,
adapted: agent target is 0.87.0 (already published + manifest-vouched — re-run the
immutability and live-bytes gates for 0.87.0 rather than 0.85.0; cite the shas in the report).
- STOP — Viktor signs the agent op (0.81.0 → 0.87.0) only after the gates pass. Never sign for a host that has gone offline again (gate 0a re-checked at signing time).
- Post-update: clean restart, self-check; expected capability state:
pbsdr-*will read DEGRADED (binary missing) until Phase 3c's migration — record it, don't chase it. Everything else green. - Re-run the 0b/0g creds+key gates per the train runbook.
Gate P2: agent 0.87.0 live on his host, self-check clean modulo the expected pbsdr gap.
Phase 3 — the onboarding call (Viktor + Peti; CC assists; Peti-paced throughout)
3a. Claim. Peti opens felhom.sajatfelhom.hu, enters his code (resend from the hub if
needed — registered address only), sets HIS password. Verify: login works, code reuse refused.
Record the G10 closure line for the first real customer (timestamped).
3b. WG consent. Walk the tester agreement's WireGuard disclosure (base-infrastructure tunnel, outbound, hub-toggleable). Peti acknowledges explicitly; record the ack in the pilot doc. If he declines: DR-tier and offsite are off the table today — the runbook continues at 3f; no pressure mechanics.
3c. DR-tier migration (supervised SSH on his host — the documented pre-v1.15 one-liner):
felhom-pbs-apply into /usr/local/sbin, age installed, wg_tunnel.enabled: true, ACL restored
to the default set (incl. felhom-pbs + /storage grant). Then hub: DR flag ON → watch the cascade
(WG peer → descriptor → applied) exactly as drilled. Capabilities go clean (pbsdr inactive→active
states per the 0.86 semantics).
3d. The ceremony — operator-blind R. PETI runs
felhom-agent --selftest=escrow-create --upload in HIS terminal on HIS box; Viktor coaches from
the ceremony runbook without screen-sharing the output. Peti writes R on paper, confirms the
command's self-verify line appeared. This is the escrow model's promise performed for real: the
operator never sees R. CC verifies only the OUTCOMES on the hub (blob present, hash staged).
3e. Auto-confirm, hands off. Drilled twice at 6-8 minutes; watch pending → escrowed on the hub; nobody clicks anything. A hash-mismatch → do NOT manual-confirm; finding + re-ceremony.
3f. Offsite opt-in — Peti's decision, with the honest pitch (encrypted recovery units to the Storage Box; media stays local; quota; what R recovers). If yes: hub provisioning → he toggles his chosen apps → first run; the restore-to-verify can run same-call or that evening.
3g. motioneye triage. Options in order of preference: retention settings inside motioneye (cap days/size), move its HDD_PATH to a larger drive via the migration feature, or manual cleanup. His call; verify recordings resume and the storage warning clears (and that the CLEARED state also behaves correctly in notifications).
Gate P3: claimed + (if consented) DR applied + escrowed + offsite per his choice + camera recording again. Capture every moment of confusion VERBATIM — these notes are the alpha one-pager's raw material.
Phase 4 — credential rotations (Viktor; CC assists)
4a. CF token — mind the shared-token trap: the old "Edit zone DNS" token covers BOTH remaining zones and lives on BOTH boxes — revoking it before re-tokening the demo kills the demo's ACME too. Order is mandatory:
- Mint
sajatfelhom.hu-tokenANDdemo-felhom.eu-token(each: Zone DNS:Edit + Zone WAF:Edit- Zone:Read, single zone).
- Deliver each to its box (controller.yaml
infrastructure.cf_api_token+ traefik env refresh — the manual path; the absence of a hub-pushed rotation flow is a KNOWN gap, note it again in the report). - Verify cert issuance on both (force one renewal or confirm a clean DNS-01 in traefik logs).
- Only then revoke the old multi-zone token; confirm "Last used" stops moving.
4b. Hub bearer VALUE rotation — the pending STOP from the closing bundle: secrets.md §"Operator/global bearer key": mint → Secret update → rollout → new-key 200 / old-key 401. The git-history copy is dead only after this.
Gate P4: three single-zone tokens live (enkisfelhom already done), zero multi-zone customer- resident credentials, bearer rotated.
Phase 5 — wrap
Hub screenshots (Peti's row: ONLINE, claimed, DR state per his consent, escrowed, storage recovering); close the parked rows in the train runbook; pilot doc updated with the onboarding record (consent acks, G10 line, ceremony outcome); REPORT overwrite + CONTEXT; findings table (expected candidates: the notification-pipeline answer from 0a, any claim-flow friction, the capability display during the DEGRADED window, anything Peti says that a stranger would also say).
§15 report must include
Per-phase gate evidence; the P1 floor-proof timing; the signed-op record (sha, signature, timestamps); the G10 + ceremony + auto-confirm lines for the first real customer; the 4a before/after token inventory ("Last used" evidence on the revoked token); the bearer rotation transcript (values redacted); the verbatim friction notes; explicit list of what was deliberately NOT done (cluster work, archive retirement) with their queue positions.
EXECUTION RECORD — 2026-07-13 (CC) — STOPPED AT GATE P1 for a ruling
Everything below is hub-side evidence only (API reads with the bearer + a read-only hub.db
snapshot queried on 180 and deleted afterwards). Nothing was pushed, signed, requested, or
changed on any box or in the hub. Per Phase 1's own rule ("if the controller is NOT
converging: diagnose first, report, and STOP for a ruling"), the runbook is halted at P1.
Phase 0a — the hub-side record (CC-gathered; Viktor's 0b/0c messaging still open)
| Item | Value (evidence: hub API + hub.db, 2026-07-13 ~14:10Z) |
|---|---|
| Controller version | 0.115.0 (NOT converged; floor is 0.122.0) |
| Claim state | code issued gen 1 2026-07-12 16:49:08Z, emailed 16:49:09Z (claim_claim, channel=customer, status=sent), NOT claimed, 0 resets |
| Reports | controller reporting healthy every 15 min, unbroken through the whole window; health_status=ok, last seen 14:04:08Z |
| Agent | 0.81.0, host-reports resumed 2026-07-13 10:29:12Z after a ~40h gap (07-11 ~14:00 → 07-13 10:29) |
| motioneye emailed him? | NO — FINDING P1-F1. Every storage_fill_critical notification (16 firings since 07-12) went to channel operator only. The ONLY customer-channel email peti-felhom has ever received is the claim email. There is no customer_notifications row for peti-felhom (no prefs exist pre-claim) — the customer tier of the pipeline never engaged. Follow-up task needed: decide the intended pre-claim customer-notification behavior. |
Phase 1 — convergence check: GATE P1 FAILS — controller is NOT converging
Facts (each with source):
- The floor IS being served. DB
hub_settings.min_controller_version = 0.122.0(set 2026-07-12 17:06:08Z); manifest MinAgent 0.81.0; hosts row agent_version 0.81.0 →ResolveManagedFloorserves the floor (confirmed by absence of any "managed floor HELD for peti-felhom" line in live hub logs while reports flow every 15 min). Manifest also confirmed vouching agent 0.87.0 sha2447d4a3…dc7a+ golden 0.120.0 (runbook ground state matches). - The self-update mechanism itself worked while the channel was up:
controller_updatedevents 0.110→0.112 (07-10 16:01), 0.112→0.113 (07-11 10:41), 0.113→0.115 (07-11 13:38). - The controller→agent channel died 23 minutes after the last successful update and is STILL
dead:
agent_channel_unreachable07-11 14:01:53 (connection refused), thenagent_channel_unknown"dial tcp 192.168.1.170:8443: no route to host" (07-11 14:02–14:06), and a FRESHagent_channel_unreachable(connection refused) at 2026-07-13 16:15:08 local — after the agent's host-reports had already resumed. The controller cannot reach the agent's local :8443 API even now. - Why that blocks convergence (code-verified,
controller/internal/selfupdate/updater.go): the floor auto-update pulls the image in-guest, then delegates the container swap to the host agent over that same channel. A failed swap persistsUpdateState{TargetVersion: <floor>, Status: failed}andMaybeAutoUpdatepermanently skips a floor that has a persisted failed attempt (anti-flapping, by design). Floors 0.120.0 (07-12 11:19) and 0.122.0 (07-12 17:06) were both set while the channel was down. So either (a) an attempt at 0.122.0 failed at the swap (or pull) step and is now persistently skipped, or (b) the registry pre-check fails each cycle and it defers forever — both stall exactly as observed, and both are unresolvable from the hub side. The controller's own log (hub log-pull button, or the call) will say which. controller_startedcount unchanged since 07-11 14:04 — no restart/swap ever happened after the channel died. Consistent with all of the above.
Root cause underneath it all — the out-of-scope trip-wire is ALREADY TRIPPED:
- The agent's node (
host.node = "proxmox") reports uptime 12,663 s (~3.5 h — rebooted ~12:35Z today), 128 GB RAM, 0 guests. - The guest's controller reports uptime 365,289 s (~4.2 days), i5-2500, 8 GB — the guest has been up longer than the node the agent runs on, on visibly different hardware.
- ⇒ The felhom guest is NOT running on the node the agent is enrolled on. Peti's
proxmox1/proxmox2 cluster shape has the guest on one node and
peti-felhom-86d37d's agent on the other. Only ONE host row is enrolled for peti-felhom (verified). The runbook's own scope rule — "if the guest migrates mid-runbook, STOP and reassess" — describes the CURRENT state, before the runbook even starts.
Collateral findings (recorded, not chased):
- P1-F2: agent local backup fails every cycle:
backup FAILED: target=local vmid=9201 err= "proxmox: not a UPID: \"OK\""→expected_backup_missed(newest backup >40 h). Almost certainly the cross-node vzdump of a vmid that is not on this node (and/or agent 0.81 vs his newer PVE — kernel 6.17.2-1-pve). Re-evaluate after the cluster ruling; don't debug in place. - P1-F3:
wg-handshake-readcapability DEGRADED ("binary not found", critical=true) on his node — WireGuard tools missing; matters for Phase 3b/3c.pbsdr-*degraded is the EXPECTED pre-3c state (per Phase 2 note). Hostcloudflared: inactivealso noted. - P1-F4 (design observation): while the agent was down 40 h, the MinAgent conditional floor
kept being served because
ResolveManagedFlooruses the LAST-KNOWNhosts.agent_versionwith no freshness check — the floor was served to a box whose agent was dead, guaranteeing the failed-swap attempt in (4). Worth a ruling on whether the floor should be held when the host is stale/down. - motioneye storage (his own LVM VG, 1.2 TB): 99.66 % full, 4.09 GB free — Phase 3g stands.
- Storage roll-up honesty:
storage_fill_criticalfires from the HOST report; the customer-row coloring behaved as designed.
The ruling needed before anything continues
- Cluster split (the blocker): migrate guest 9201 back to the agent's node, or move/reinstall the agent where the guest lives, or accelerate agent-follows-guest. Viktor + Peti decision — this reshapes Phases 2 and 3c–3e (all assume agent and guest share a host).
- After the channel heals, convergence still needs one of: floor bump (0.125.0 is the earmarked candidate and carries the .fab strand fix; NEVER halt above 0.124.0 without it), a manual update trigger, or clearing the persisted update state during the call — because of the once-per-floor guard in (4).
- Optional pre-call diagnostic: hub controller-log pull (operator UI button) to confirm which stall variant (persisted-failed vs registry-defer) — CC did NOT trigger it (write action; Phase 1 is read-only).
- Phase 0b/0c (message Peti, call slot) are unaffected and remain open — Viktor.
- Phase 4 (CF token + bearer rotation) is independent of the cluster question and could proceed on Viktor's GO at any time.
State: STOPPED at Gate P1. Phases 2–5 not started. Nothing signed, nothing pushed.