RUNBOOK-peti-return: Phase 0a/1 execution record — STOPPED at Gate P1 (controller not converging; cluster split: guest not on the agent's node)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GzammAMzsJTgpQHqxwM2bC
This commit is contained in:
@@ -0,0 +1,261 @@
|
||||
# RUNBOOK — Peti's return: convergence, the parked train, the first real-customer onboarding, and the credential rotations
|
||||
|
||||
<!--
|
||||
Class: operational runbook (GL pattern). Actors: Viktor (operator, signs ops, leads the call),
|
||||
CC (read-only verification + supervised host steps over SSH), and — for the first time — PETI
|
||||
as the customer performing his own custody steps. Every finding is captured verbatim: this run
|
||||
IS the alpha-onboarding rehearsal, and its friction notes seed the tester one-pager.
|
||||
|
||||
Ground state (hub, 2026-07-13): peti-felhom-86d37d ONLINE, agent 0.81.0, controller 0.115.0,
|
||||
motioneye storage 100% full. Global floor 0.122.0; MinAgent 0.81.0 (satisfied → his controller
|
||||
self-updates WITHOUT any operator action — Phase 0 exists because the machine is already moving).
|
||||
Current vouched artifacts: agent 0.87.0 (sha 2447d4a3…dc7a), hub 0.53.0.
|
||||
-->
|
||||
|
||||
---
|
||||
|
||||
## 0. Scope & standing rules
|
||||
|
||||
- **In scope:** convergence verification, the publish train's parked D/E/G at CURRENT versions,
|
||||
claim, WG consent, DR-tier migration, ceremony + auto-confirm, offsite opt-in, motioneye
|
||||
remediation, CF token rotation (BOTH boxes — see the trap in Phase 4), hub bearer rotation.
|
||||
- **Out of scope:** cluster/agent-follows-guest work (his proxmox1/proxmox2 HA shape is a known
|
||||
roadmap item — if the guest migrates mid-runbook, STOP and reassess); the old-box archive
|
||||
(u629193-sub1) retirement decision — separate.
|
||||
- **Custody rule for the whole runbook:** Peti's password and R are PETI'S. They never appear on
|
||||
Viktor's or CC's screen, in transcripts, or in the report. Where a step displays them, Peti
|
||||
drives.
|
||||
- No gate-lifts anywhere (standing rule). Diagnose-before-fix on every anomaly.
|
||||
|
||||
---
|
||||
|
||||
## Phase 0 — TODAY, before anything technical (Viktor, ~15 min)
|
||||
|
||||
0a. **Hub → Peti's customer page, record:** current controller version (self-update may be
|
||||
mid-flight or done), claim state (code issued? issued-at? claimed?), the notification/event log —
|
||||
**specifically whether the motioneye 100% warning ever emailed him.** That answer is the
|
||||
notification pipeline's first real-customer test; record it either way, and if it never fired,
|
||||
that is a FINDING with its own follow-up task.
|
||||
|
||||
0b. **Message Peti** (before the machine talks to him): fan congratulations; "your box is
|
||||
updating itself over the next hours — you'll get (or already have) an email with a setup code,
|
||||
it's real, don't delete it"; camera drive is full, we'll fix it together; propose the onboarding
|
||||
call (~90 min, needs him at a keyboard with root on his own box).
|
||||
|
||||
0c. Sign nothing yet. If the claim email was never issued or landed in spam, note it — the
|
||||
resend button is Phase 3 ammunition, not a Phase 0 action.
|
||||
|
||||
**Gate P0:** Peti has acknowledged and a call slot exists.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — passive convergence check (CC, strictly read-only)
|
||||
|
||||
Confirm from hub data only: controller reached **0.122.0** (floor); the claim-code hash was
|
||||
ACK-delivered; the box transitioned legacy-banner → gate armed; reports healthy throughout; the
|
||||
roll-up row honest (it will read WARN "storage" territory only if the storage health warns —
|
||||
motioneye's full drive may legitimately color it). If the controller is NOT converging: diagnose
|
||||
first (report errors? MinAgent mismatch? update loop?), report, and STOP for a ruling — do not
|
||||
push anything.
|
||||
|
||||
**Gate P1:** controller ≥ 0.122.0, gate armed, box healthy. Record the self-update elapsed time
|
||||
(fleet's second floor-proof, first on customer hardware).
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — the parked train phases, at current versions (CC executes, Viktor signs)
|
||||
|
||||
Execute RUNBOOK-publish-0.85-0.120 **Phases D/E/G + gates 0b/0g** against `peti-felhom-86d37d`,
|
||||
adapted: agent target is **0.87.0** (already published + manifest-vouched — re-run the
|
||||
immutability and live-bytes gates for 0.87.0 rather than 0.85.0; cite the shas in the report).
|
||||
|
||||
- **STOP — Viktor signs the agent op** (0.81.0 → 0.87.0) only after the gates pass. Never sign
|
||||
for a host that has gone offline again (gate 0a re-checked at signing time).
|
||||
- Post-update: clean restart, self-check; **expected capability state:** `pbsdr-*` will read
|
||||
DEGRADED (binary missing) until Phase 3c's migration — record it, don't chase it. Everything
|
||||
else green.
|
||||
- Re-run the 0b/0g creds+key gates per the train runbook.
|
||||
|
||||
**Gate P2:** agent 0.87.0 live on his host, self-check clean modulo the expected pbsdr gap.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — the onboarding call (Viktor + Peti; CC assists; Peti-paced throughout)
|
||||
|
||||
3a. **Claim.** Peti opens `felhom.sajatfelhom.hu`, enters his code (resend from the hub if
|
||||
needed — registered address only), sets HIS password. Verify: login works, code reuse refused.
|
||||
Record the **G10 closure line for the first real customer** (timestamped).
|
||||
|
||||
3b. **WG consent.** Walk the tester agreement's WireGuard disclosure (base-infrastructure
|
||||
tunnel, outbound, hub-toggleable). Peti acknowledges explicitly; record the ack in the pilot doc.
|
||||
If he declines: DR-tier and offsite are off the table today — the runbook continues at 3f; no
|
||||
pressure mechanics.
|
||||
|
||||
3c. **DR-tier migration** (supervised SSH on his host — the documented pre-v1.15 one-liner):
|
||||
felhom-pbs-apply into /usr/local/sbin, `age` installed, `wg_tunnel.enabled: true`, ACL restored
|
||||
to the default set (incl. felhom-pbs + /storage grant). Then hub: DR flag ON → watch the cascade
|
||||
(WG peer → descriptor → applied) exactly as drilled. Capabilities go clean (pbsdr inactive→active
|
||||
states per the 0.86 semantics).
|
||||
|
||||
3d. **The ceremony — operator-blind R.** PETI runs
|
||||
`felhom-agent --selftest=escrow-create --upload` in HIS terminal on HIS box; Viktor coaches from
|
||||
the ceremony runbook without screen-sharing the output. Peti writes R on paper, confirms the
|
||||
command's self-verify line appeared. This is the escrow model's promise performed for real: the
|
||||
operator never sees R. CC verifies only the OUTCOMES on the hub (blob present, hash staged).
|
||||
|
||||
3e. **Auto-confirm, hands off.** Drilled twice at 6-8 minutes; watch pending → escrowed on the
|
||||
hub; nobody clicks anything. A hash-mismatch → do NOT manual-confirm; finding + re-ceremony.
|
||||
|
||||
3f. **Offsite opt-in** — Peti's decision, with the honest pitch (encrypted recovery units to the
|
||||
Storage Box; media stays local; quota; what R recovers). If yes: hub provisioning → he toggles
|
||||
his chosen apps → first run; the restore-to-verify can run same-call or that evening.
|
||||
|
||||
3g. **motioneye triage.** Options in order of preference: retention settings inside motioneye
|
||||
(cap days/size), move its HDD_PATH to a larger drive via the migration feature, or manual
|
||||
cleanup. His call; verify recordings resume and the storage warning clears (and that the
|
||||
CLEARED state also behaves correctly in notifications).
|
||||
|
||||
**Gate P3:** claimed + (if consented) DR applied + escrowed + offsite per his choice + camera
|
||||
recording again. Capture every moment of confusion VERBATIM — these notes are the alpha
|
||||
one-pager's raw material.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — credential rotations (Viktor; CC assists)
|
||||
|
||||
4a. **CF token — mind the shared-token trap:** the old "Edit zone DNS" token covers BOTH
|
||||
remaining zones and lives on BOTH boxes — revoking it before re-tokening the demo kills the
|
||||
demo's ACME too. Order is mandatory:
|
||||
1. Mint `sajatfelhom.hu-token` AND `demo-felhom.eu-token` (each: Zone DNS:Edit + Zone WAF:Edit
|
||||
+ Zone:Read, single zone).
|
||||
2. Deliver each to its box (controller.yaml `infrastructure.cf_api_token` + traefik env
|
||||
refresh — the manual path; the absence of a hub-pushed rotation flow is a KNOWN gap, note it
|
||||
again in the report).
|
||||
3. Verify cert issuance on both (force one renewal or confirm a clean DNS-01 in traefik logs).
|
||||
4. Only then revoke the old multi-zone token; confirm "Last used" stops moving.
|
||||
|
||||
4b. **Hub bearer VALUE rotation** — the pending STOP from the closing bundle: secrets.md
|
||||
§"Operator/global bearer key": mint → Secret update → rollout → new-key 200 / old-key 401. The
|
||||
git-history copy is dead only after this.
|
||||
|
||||
**Gate P4:** three single-zone tokens live (enkisfelhom already done), zero multi-zone customer-
|
||||
resident credentials, bearer rotated.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — wrap
|
||||
|
||||
Hub screenshots (Peti's row: ONLINE, claimed, DR state per his consent, escrowed, storage
|
||||
recovering); close the parked rows in the train runbook; pilot doc updated with the onboarding
|
||||
record (consent acks, G10 line, ceremony outcome); REPORT overwrite + CONTEXT; findings table
|
||||
(expected candidates: the notification-pipeline answer from 0a, any claim-flow friction, the
|
||||
capability display during the DEGRADED window, anything Peti says that a stranger would also
|
||||
say).
|
||||
|
||||
---
|
||||
|
||||
## §15 report must include
|
||||
|
||||
Per-phase gate evidence; the P1 floor-proof timing; the signed-op record (sha, signature,
|
||||
timestamps); the G10 + ceremony + auto-confirm lines for the first real customer; the 4a
|
||||
before/after token inventory ("Last used" evidence on the revoked token); the bearer rotation
|
||||
transcript (values redacted); the verbatim friction notes; explicit list of what was deliberately
|
||||
NOT done (cluster work, archive retirement) with their queue positions.
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
# EXECUTION RECORD — 2026-07-13 (CC) — **STOPPED AT GATE P1 for a ruling**
|
||||
|
||||
Everything below is hub-side evidence only (API reads with the bearer + a read-only `hub.db`
|
||||
snapshot queried on 180 and deleted afterwards). **Nothing was pushed, signed, requested, or
|
||||
changed on any box or in the hub.** Per Phase 1's own rule ("if the controller is NOT
|
||||
converging: diagnose first, report, and STOP for a ruling"), the runbook is halted at P1.
|
||||
|
||||
## Phase 0a — the hub-side record (CC-gathered; Viktor's 0b/0c messaging still open)
|
||||
|
||||
| Item | Value (evidence: hub API + hub.db, 2026-07-13 ~14:10Z) |
|
||||
|---|---|
|
||||
| Controller version | **0.115.0** (NOT converged; floor is 0.122.0) |
|
||||
| Claim state | code **issued gen 1 2026-07-12 16:49:08Z, emailed 16:49:09Z** (`claim_claim`, channel=customer, status=sent), **NOT claimed**, 0 resets |
|
||||
| Reports | controller reporting healthy every 15 min, unbroken through the whole window; `health_status=ok`, last seen 14:04:08Z |
|
||||
| Agent | 0.81.0, host-reports resumed **2026-07-13 10:29:12Z** after a ~40h gap (07-11 ~14:00 → 07-13 10:29) |
|
||||
| motioneye emailed him? | **NO — FINDING P1-F1.** Every `storage_fill_critical` notification (16 firings since 07-12) went to channel **operator only**. The ONLY customer-channel email peti-felhom has ever received is the claim email. There is **no `customer_notifications` row** for peti-felhom (no prefs exist pre-claim) — the customer tier of the pipeline never engaged. Follow-up task needed: decide the intended pre-claim customer-notification behavior. |
|
||||
|
||||
## Phase 1 — convergence check: **GATE P1 FAILS — controller is NOT converging**
|
||||
|
||||
**Facts (each with source):**
|
||||
|
||||
1. **The floor IS being served.** DB `hub_settings.min_controller_version = 0.122.0` (set
|
||||
2026-07-12 17:06:08Z); manifest MinAgent 0.81.0; hosts row agent_version 0.81.0 →
|
||||
`ResolveManagedFloor` serves the floor (confirmed by absence of any "managed floor HELD for
|
||||
peti-felhom" line in live hub logs while reports flow every 15 min). Manifest also confirmed
|
||||
vouching agent **0.87.0** sha `2447d4a3…dc7a` + golden 0.120.0 (runbook ground state matches).
|
||||
2. **The self-update mechanism itself worked while the channel was up:** `controller_updated`
|
||||
events 0.110→0.112 (07-10 16:01), 0.112→0.113 (07-11 10:41), 0.113→0.115 (07-11 13:38).
|
||||
3. **The controller→agent channel died 23 minutes after the last successful update and is STILL
|
||||
dead:** `agent_channel_unreachable` 07-11 14:01:53 (connection refused), then
|
||||
`agent_channel_unknown` "dial tcp **192.168.1.170:8443**: no route to host" (07-11 14:02–14:06),
|
||||
and a FRESH `agent_channel_unreachable` (connection refused) at **2026-07-13 16:15:08 local**
|
||||
— after the agent's host-reports had already resumed. The controller cannot reach the agent's
|
||||
local :8443 API even now.
|
||||
4. **Why that blocks convergence (code-verified,** `controller/internal/selfupdate/updater.go`**):**
|
||||
the floor auto-update pulls the image in-guest, then **delegates the container swap to the
|
||||
host agent** over that same channel. A failed swap persists `UpdateState{TargetVersion:
|
||||
<floor>, Status: failed}` and `MaybeAutoUpdate` **permanently skips a floor that has a
|
||||
persisted failed attempt** (anti-flapping, by design). Floors 0.120.0 (07-12 11:19) and
|
||||
0.122.0 (07-12 17:06) were both set while the channel was down. So either (a) an attempt at
|
||||
0.122.0 failed at the swap (or pull) step and is now persistently skipped, or (b) the registry
|
||||
pre-check fails each cycle and it defers forever — both stall exactly as observed, and both
|
||||
are unresolvable from the hub side. The controller's own log (hub log-pull button, or the call)
|
||||
will say which.
|
||||
5. **`controller_started` count unchanged since 07-11 14:04** — no restart/swap ever happened
|
||||
after the channel died. Consistent with all of the above.
|
||||
|
||||
**Root cause underneath it all — the out-of-scope trip-wire is ALREADY TRIPPED:**
|
||||
|
||||
- The agent's node (`host.node = "proxmox"`) reports **uptime 12,663 s (~3.5 h — rebooted ~12:35Z
|
||||
today), 128 GB RAM, 0 guests**.
|
||||
- The guest's controller reports **uptime 365,289 s (~4.2 days), i5-2500, 8 GB** — the guest has
|
||||
been up longer than the node the agent runs on, on visibly different hardware.
|
||||
- ⇒ **The felhom guest is NOT running on the node the agent is enrolled on.** Peti's
|
||||
proxmox1/proxmox2 cluster shape has the guest on one node and `peti-felhom-86d37d`'s agent on
|
||||
the other. Only ONE host row is enrolled for peti-felhom (verified). The runbook's own scope
|
||||
rule — "if the guest migrates mid-runbook, STOP and reassess" — describes the CURRENT state,
|
||||
before the runbook even starts.
|
||||
|
||||
**Collateral findings (recorded, not chased):**
|
||||
|
||||
- **P1-F2:** agent local backup fails every cycle: `backup FAILED: target=local vmid=9201 err=
|
||||
"proxmox: not a UPID: \"OK\""` → `expected_backup_missed` (newest backup >40 h). Almost
|
||||
certainly the cross-node vzdump of a vmid that is not on this node (and/or agent 0.81 vs his
|
||||
newer PVE — kernel 6.17.2-1-pve). Re-evaluate after the cluster ruling; don't debug in place.
|
||||
- **P1-F3:** `wg-handshake-read` capability DEGRADED ("binary not found", critical=true) on his
|
||||
node — WireGuard tools missing; matters for Phase 3b/3c. `pbsdr-*` degraded is the EXPECTED
|
||||
pre-3c state (per Phase 2 note). Host `cloudflared: inactive` also noted.
|
||||
- **P1-F4 (design observation):** while the agent was down 40 h, the MinAgent conditional floor
|
||||
kept being served because `ResolveManagedFloor` uses the LAST-KNOWN `hosts.agent_version` with
|
||||
no freshness check — the floor was served to a box whose agent was dead, guaranteeing the
|
||||
failed-swap attempt in (4). Worth a ruling on whether the floor should be held when the host is
|
||||
stale/down.
|
||||
- motioneye storage (his own LVM VG, 1.2 TB): 99.66 % full, 4.09 GB free — Phase 3g stands.
|
||||
- Storage roll-up honesty: `storage_fill_critical` fires from the HOST report; the customer-row
|
||||
coloring behaved as designed.
|
||||
|
||||
## The ruling needed before anything continues
|
||||
|
||||
1. **Cluster split** (the blocker): migrate guest 9201 back to the agent's node, or move/reinstall
|
||||
the agent where the guest lives, or accelerate agent-follows-guest. Viktor + Peti decision —
|
||||
this reshapes Phases 2 and 3c–3e (all assume agent and guest share a host).
|
||||
2. **After the channel heals**, convergence still needs one of: floor bump (0.125.0 is the
|
||||
earmarked candidate and carries the .fab strand fix; NEVER halt above 0.124.0 without it), a
|
||||
manual update trigger, or clearing the persisted update state during the call — because of the
|
||||
once-per-floor guard in (4).
|
||||
3. Optional pre-call diagnostic: hub controller-log pull (operator UI button) to confirm which
|
||||
stall variant (persisted-failed vs registry-defer) — CC did NOT trigger it (write action;
|
||||
Phase 1 is read-only).
|
||||
4. Phase 0b/0c (message Peti, call slot) are unaffected and remain open — Viktor.
|
||||
5. Phase 4 (CF token + bearer rotation) is independent of the cluster question and could proceed
|
||||
on Viktor's GO at any time.
|
||||
|
||||
**State: STOPPED at Gate P1. Phases 2–5 not started. Nothing signed, nothing pushed.**
|
||||
Reference in New Issue
Block a user