RUNBOOK-peti-return: Phase 0a/1 execution record — STOPPED at Gate P1 (controller not converging; cluster split: guest not on the agent's node)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GzammAMzsJTgpQHqxwM2bC
This commit is contained in:
2026-07-13 16:24:33 +02:00
parent 39144a5844
commit fa0ecb2cd7
2 changed files with 293 additions and 157 deletions
@@ -0,0 +1,261 @@
# RUNBOOK — Peti's return: convergence, the parked train, the first real-customer onboarding, and the credential rotations
<!--
Class: operational runbook (GL pattern). Actors: Viktor (operator, signs ops, leads the call),
CC (read-only verification + supervised host steps over SSH), and — for the first time — PETI
as the customer performing his own custody steps. Every finding is captured verbatim: this run
IS the alpha-onboarding rehearsal, and its friction notes seed the tester one-pager.
Ground state (hub, 2026-07-13): peti-felhom-86d37d ONLINE, agent 0.81.0, controller 0.115.0,
motioneye storage 100% full. Global floor 0.122.0; MinAgent 0.81.0 (satisfied → his controller
self-updates WITHOUT any operator action — Phase 0 exists because the machine is already moving).
Current vouched artifacts: agent 0.87.0 (sha 2447d4a3…dc7a), hub 0.53.0.
-->
---
## 0. Scope & standing rules
- **In scope:** convergence verification, the publish train's parked D/E/G at CURRENT versions,
claim, WG consent, DR-tier migration, ceremony + auto-confirm, offsite opt-in, motioneye
remediation, CF token rotation (BOTH boxes — see the trap in Phase 4), hub bearer rotation.
- **Out of scope:** cluster/agent-follows-guest work (his proxmox1/proxmox2 HA shape is a known
roadmap item — if the guest migrates mid-runbook, STOP and reassess); the old-box archive
(u629193-sub1) retirement decision — separate.
- **Custody rule for the whole runbook:** Peti's password and R are PETI'S. They never appear on
Viktor's or CC's screen, in transcripts, or in the report. Where a step displays them, Peti
drives.
- No gate-lifts anywhere (standing rule). Diagnose-before-fix on every anomaly.
---
## Phase 0 — TODAY, before anything technical (Viktor, ~15 min)
0a. **Hub → Peti's customer page, record:** current controller version (self-update may be
mid-flight or done), claim state (code issued? issued-at? claimed?), the notification/event log —
**specifically whether the motioneye 100% warning ever emailed him.** That answer is the
notification pipeline's first real-customer test; record it either way, and if it never fired,
that is a FINDING with its own follow-up task.
0b. **Message Peti** (before the machine talks to him): fan congratulations; "your box is
updating itself over the next hours — you'll get (or already have) an email with a setup code,
it's real, don't delete it"; camera drive is full, we'll fix it together; propose the onboarding
call (~90 min, needs him at a keyboard with root on his own box).
0c. Sign nothing yet. If the claim email was never issued or landed in spam, note it — the
resend button is Phase 3 ammunition, not a Phase 0 action.
**Gate P0:** Peti has acknowledged and a call slot exists.
---
## Phase 1 — passive convergence check (CC, strictly read-only)
Confirm from hub data only: controller reached **0.122.0** (floor); the claim-code hash was
ACK-delivered; the box transitioned legacy-banner → gate armed; reports healthy throughout; the
roll-up row honest (it will read WARN "storage" territory only if the storage health warns —
motioneye's full drive may legitimately color it). If the controller is NOT converging: diagnose
first (report errors? MinAgent mismatch? update loop?), report, and STOP for a ruling — do not
push anything.
**Gate P1:** controller ≥ 0.122.0, gate armed, box healthy. Record the self-update elapsed time
(fleet's second floor-proof, first on customer hardware).
---
## Phase 2 — the parked train phases, at current versions (CC executes, Viktor signs)
Execute RUNBOOK-publish-0.85-0.120 **Phases D/E/G + gates 0b/0g** against `peti-felhom-86d37d`,
adapted: agent target is **0.87.0** (already published + manifest-vouched — re-run the
immutability and live-bytes gates for 0.87.0 rather than 0.85.0; cite the shas in the report).
- **STOP — Viktor signs the agent op** (0.81.0 → 0.87.0) only after the gates pass. Never sign
for a host that has gone offline again (gate 0a re-checked at signing time).
- Post-update: clean restart, self-check; **expected capability state:** `pbsdr-*` will read
DEGRADED (binary missing) until Phase 3c's migration — record it, don't chase it. Everything
else green.
- Re-run the 0b/0g creds+key gates per the train runbook.
**Gate P2:** agent 0.87.0 live on his host, self-check clean modulo the expected pbsdr gap.
---
## Phase 3 — the onboarding call (Viktor + Peti; CC assists; Peti-paced throughout)
3a. **Claim.** Peti opens `felhom.sajatfelhom.hu`, enters his code (resend from the hub if
needed — registered address only), sets HIS password. Verify: login works, code reuse refused.
Record the **G10 closure line for the first real customer** (timestamped).
3b. **WG consent.** Walk the tester agreement's WireGuard disclosure (base-infrastructure
tunnel, outbound, hub-toggleable). Peti acknowledges explicitly; record the ack in the pilot doc.
If he declines: DR-tier and offsite are off the table today — the runbook continues at 3f; no
pressure mechanics.
3c. **DR-tier migration** (supervised SSH on his host — the documented pre-v1.15 one-liner):
felhom-pbs-apply into /usr/local/sbin, `age` installed, `wg_tunnel.enabled: true`, ACL restored
to the default set (incl. felhom-pbs + /storage grant). Then hub: DR flag ON → watch the cascade
(WG peer → descriptor → applied) exactly as drilled. Capabilities go clean (pbsdr inactive→active
states per the 0.86 semantics).
3d. **The ceremony — operator-blind R.** PETI runs
`felhom-agent --selftest=escrow-create --upload` in HIS terminal on HIS box; Viktor coaches from
the ceremony runbook without screen-sharing the output. Peti writes R on paper, confirms the
command's self-verify line appeared. This is the escrow model's promise performed for real: the
operator never sees R. CC verifies only the OUTCOMES on the hub (blob present, hash staged).
3e. **Auto-confirm, hands off.** Drilled twice at 6-8 minutes; watch pending → escrowed on the
hub; nobody clicks anything. A hash-mismatch → do NOT manual-confirm; finding + re-ceremony.
3f. **Offsite opt-in** — Peti's decision, with the honest pitch (encrypted recovery units to the
Storage Box; media stays local; quota; what R recovers). If yes: hub provisioning → he toggles
his chosen apps → first run; the restore-to-verify can run same-call or that evening.
3g. **motioneye triage.** Options in order of preference: retention settings inside motioneye
(cap days/size), move its HDD_PATH to a larger drive via the migration feature, or manual
cleanup. His call; verify recordings resume and the storage warning clears (and that the
CLEARED state also behaves correctly in notifications).
**Gate P3:** claimed + (if consented) DR applied + escrowed + offsite per his choice + camera
recording again. Capture every moment of confusion VERBATIM — these notes are the alpha
one-pager's raw material.
---
## Phase 4 — credential rotations (Viktor; CC assists)
4a. **CF token — mind the shared-token trap:** the old "Edit zone DNS" token covers BOTH
remaining zones and lives on BOTH boxes — revoking it before re-tokening the demo kills the
demo's ACME too. Order is mandatory:
1. Mint `sajatfelhom.hu-token` AND `demo-felhom.eu-token` (each: Zone DNS:Edit + Zone WAF:Edit
+ Zone:Read, single zone).
2. Deliver each to its box (controller.yaml `infrastructure.cf_api_token` + traefik env
refresh — the manual path; the absence of a hub-pushed rotation flow is a KNOWN gap, note it
again in the report).
3. Verify cert issuance on both (force one renewal or confirm a clean DNS-01 in traefik logs).
4. Only then revoke the old multi-zone token; confirm "Last used" stops moving.
4b. **Hub bearer VALUE rotation** — the pending STOP from the closing bundle: secrets.md
§"Operator/global bearer key": mint → Secret update → rollout → new-key 200 / old-key 401. The
git-history copy is dead only after this.
**Gate P4:** three single-zone tokens live (enkisfelhom already done), zero multi-zone customer-
resident credentials, bearer rotated.
---
## Phase 5 — wrap
Hub screenshots (Peti's row: ONLINE, claimed, DR state per his consent, escrowed, storage
recovering); close the parked rows in the train runbook; pilot doc updated with the onboarding
record (consent acks, G10 line, ceremony outcome); REPORT overwrite + CONTEXT; findings table
(expected candidates: the notification-pipeline answer from 0a, any claim-flow friction, the
capability display during the DEGRADED window, anything Peti says that a stranger would also
say).
---
## §15 report must include
Per-phase gate evidence; the P1 floor-proof timing; the signed-op record (sha, signature,
timestamps); the G10 + ceremony + auto-confirm lines for the first real customer; the 4a
before/after token inventory ("Last used" evidence on the revoked token); the bearer rotation
transcript (values redacted); the verbatim friction notes; explicit list of what was deliberately
NOT done (cluster work, archive retirement) with their queue positions.
---
---
# EXECUTION RECORD — 2026-07-13 (CC) — **STOPPED AT GATE P1 for a ruling**
Everything below is hub-side evidence only (API reads with the bearer + a read-only `hub.db`
snapshot queried on 180 and deleted afterwards). **Nothing was pushed, signed, requested, or
changed on any box or in the hub.** Per Phase 1's own rule ("if the controller is NOT
converging: diagnose first, report, and STOP for a ruling"), the runbook is halted at P1.
## Phase 0a — the hub-side record (CC-gathered; Viktor's 0b/0c messaging still open)
| Item | Value (evidence: hub API + hub.db, 2026-07-13 ~14:10Z) |
|---|---|
| Controller version | **0.115.0** (NOT converged; floor is 0.122.0) |
| Claim state | code **issued gen 1 2026-07-12 16:49:08Z, emailed 16:49:09Z** (`claim_claim`, channel=customer, status=sent), **NOT claimed**, 0 resets |
| Reports | controller reporting healthy every 15 min, unbroken through the whole window; `health_status=ok`, last seen 14:04:08Z |
| Agent | 0.81.0, host-reports resumed **2026-07-13 10:29:12Z** after a ~40h gap (07-11 ~14:00 → 07-13 10:29) |
| motioneye emailed him? | **NO — FINDING P1-F1.** Every `storage_fill_critical` notification (16 firings since 07-12) went to channel **operator only**. The ONLY customer-channel email peti-felhom has ever received is the claim email. There is **no `customer_notifications` row** for peti-felhom (no prefs exist pre-claim) — the customer tier of the pipeline never engaged. Follow-up task needed: decide the intended pre-claim customer-notification behavior. |
## Phase 1 — convergence check: **GATE P1 FAILS — controller is NOT converging**
**Facts (each with source):**
1. **The floor IS being served.** DB `hub_settings.min_controller_version = 0.122.0` (set
2026-07-12 17:06:08Z); manifest MinAgent 0.81.0; hosts row agent_version 0.81.0 →
`ResolveManagedFloor` serves the floor (confirmed by absence of any "managed floor HELD for
peti-felhom" line in live hub logs while reports flow every 15 min). Manifest also confirmed
vouching agent **0.87.0** sha `2447d4a3…dc7a` + golden 0.120.0 (runbook ground state matches).
2. **The self-update mechanism itself worked while the channel was up:** `controller_updated`
events 0.110→0.112 (07-10 16:01), 0.112→0.113 (07-11 10:41), 0.113→0.115 (07-11 13:38).
3. **The controller→agent channel died 23 minutes after the last successful update and is STILL
dead:** `agent_channel_unreachable` 07-11 14:01:53 (connection refused), then
`agent_channel_unknown` "dial tcp **192.168.1.170:8443**: no route to host" (07-11 14:0214:06),
and a FRESH `agent_channel_unreachable` (connection refused) at **2026-07-13 16:15:08 local**
— after the agent's host-reports had already resumed. The controller cannot reach the agent's
local :8443 API even now.
4. **Why that blocks convergence (code-verified,** `controller/internal/selfupdate/updater.go`**):**
the floor auto-update pulls the image in-guest, then **delegates the container swap to the
host agent** over that same channel. A failed swap persists `UpdateState{TargetVersion:
<floor>, Status: failed}` and `MaybeAutoUpdate` **permanently skips a floor that has a
persisted failed attempt** (anti-flapping, by design). Floors 0.120.0 (07-12 11:19) and
0.122.0 (07-12 17:06) were both set while the channel was down. So either (a) an attempt at
0.122.0 failed at the swap (or pull) step and is now persistently skipped, or (b) the registry
pre-check fails each cycle and it defers forever — both stall exactly as observed, and both
are unresolvable from the hub side. The controller's own log (hub log-pull button, or the call)
will say which.
5. **`controller_started` count unchanged since 07-11 14:04** — no restart/swap ever happened
after the channel died. Consistent with all of the above.
**Root cause underneath it all — the out-of-scope trip-wire is ALREADY TRIPPED:**
- The agent's node (`host.node = "proxmox"`) reports **uptime 12,663 s (~3.5 h — rebooted ~12:35Z
today), 128 GB RAM, 0 guests**.
- The guest's controller reports **uptime 365,289 s (~4.2 days), i5-2500, 8 GB** — the guest has
been up longer than the node the agent runs on, on visibly different hardware.
-**The felhom guest is NOT running on the node the agent is enrolled on.** Peti's
proxmox1/proxmox2 cluster shape has the guest on one node and `peti-felhom-86d37d`'s agent on
the other. Only ONE host row is enrolled for peti-felhom (verified). The runbook's own scope
rule — "if the guest migrates mid-runbook, STOP and reassess" — describes the CURRENT state,
before the runbook even starts.
**Collateral findings (recorded, not chased):**
- **P1-F2:** agent local backup fails every cycle: `backup FAILED: target=local vmid=9201 err=
"proxmox: not a UPID: \"OK\""` → `expected_backup_missed` (newest backup >40 h). Almost
certainly the cross-node vzdump of a vmid that is not on this node (and/or agent 0.81 vs his
newer PVE — kernel 6.17.2-1-pve). Re-evaluate after the cluster ruling; don't debug in place.
- **P1-F3:** `wg-handshake-read` capability DEGRADED ("binary not found", critical=true) on his
node — WireGuard tools missing; matters for Phase 3b/3c. `pbsdr-*` degraded is the EXPECTED
pre-3c state (per Phase 2 note). Host `cloudflared: inactive` also noted.
- **P1-F4 (design observation):** while the agent was down 40 h, the MinAgent conditional floor
kept being served because `ResolveManagedFloor` uses the LAST-KNOWN `hosts.agent_version` with
no freshness check — the floor was served to a box whose agent was dead, guaranteeing the
failed-swap attempt in (4). Worth a ruling on whether the floor should be held when the host is
stale/down.
- motioneye storage (his own LVM VG, 1.2 TB): 99.66 % full, 4.09 GB free — Phase 3g stands.
- Storage roll-up honesty: `storage_fill_critical` fires from the HOST report; the customer-row
coloring behaved as designed.
## The ruling needed before anything continues
1. **Cluster split** (the blocker): migrate guest 9201 back to the agent's node, or move/reinstall
the agent where the guest lives, or accelerate agent-follows-guest. Viktor + Peti decision —
this reshapes Phases 2 and 3c3e (all assume agent and guest share a host).
2. **After the channel heals**, convergence still needs one of: floor bump (0.125.0 is the
earmarked candidate and carries the .fab strand fix; NEVER halt above 0.124.0 without it), a
manual update trigger, or clearing the persisted update state during the call — because of the
once-per-floor guard in (4).
3. Optional pre-call diagnostic: hub controller-log pull (operator UI button) to confirm which
stall variant (persisted-failed vs registry-defer) — CC did NOT trigger it (write action;
Phase 1 is read-only).
4. Phase 0b/0c (message Peti, call slot) are unaffected and remain open — Viktor.
5. Phase 4 (CF token + bearer rotation) is independent of the cluster question and could proceed
on Viktor's GO at any time.
**State: STOPPED at Gate P1. Phases 25 not started. Nothing signed, nothing pushed.**