docs: R-23(a) restart leg BANKED; "round-trips in seconds" PARTIAL -> PROVEN-LIVE
The operator moved the global floor to a version the box did NOT run (0.153.0 -> v0.154.0) and the managed self-update fired exactly once: 06:57:13Z UpdateState pending, initiated_by=auto-floor 06:57:17Z agent: controller swap requested 0.153.0 -> 0.154.0 06:57:21Z container restarted 06:57:29Z agent: new controller healthy (16 s save -> healthy) Over a 39-minute window: swap requests 1, agent-driven bootstrap restarts 1, rollbacks 0, container RestartCount 0. VerifyStartup confirmed on the next boot; the following periodic check logged "Current version 0.154.0 is up to date" — the at/above-floor branch correctly doing nothing. The 2026-07-20 attempt proved nothing because it targeted an already-running version; that was the whole reason this leg stayed open. Disclosed in both rows: a hand-deploy of v0.155.0 at 07:17:10 falls inside the observation window and is what StartedAt shows afterwards. It never goes through SwapController, so the swap-count assertions hold across the full window — and it incidentally re-confirmed the at/above-floor branch (0.155.0 running against a 0.154.0 floor -> updater did nothing). R-23(b) (cosmetic Waiter "recovered" log timing) remains open.
This commit is contained in:
@@ -110,7 +110,7 @@
|
||||
| Scenario | Components | Status | Evidence | Gap / roadmap |
|
||||
|---|---|---|---|---|
|
||||
| Customer/host management: 8-tab detail, scoped auto-refresh, safe stale-host deletion, capability chips | hub v0.47–0.53 | **PROVEN-LIVE** | hub v0.53.0 dead-host roll-up live on the Peti cluster (proxmox1 down 23h); `CAMPAIGN-4-2026-07-13` (operator UI driven live); `DRILL-day0-take2` F-16 (offsite/freeze buttons live) | 8-tab render + capability chips are **render-test-validated** (hub UI is password-gated; CC cannot log in). (Cited "daily operator use" was a no-doc citation; `AUDIT-hub-gui-2026-06-30` predates these features at hub v0.25) |
|
||||
| Config/state change round-trips in **seconds** (hub↔box immediacy; 15-min cycle stays the backbone): box→hub out-of-cycle report (Dir 1) + hub→box `GET /api/v1/wait` long-poll wake (Dir 2) | controller v0.139/140, hub v0.58/0.63 | **PARTIAL** | Transport proven live through the real DNS-only ingress: `SPIKE-immediate-sync-transport-2026-07-16` + hub v0.58.0 / controller v0.140.0 REPORTs — 240 s no-annotation hold (25 s heartbeat defeats nginx's 60 s `proxy_read_timeout`, no ingress change), 0.047 s wake-on-change, hub `rollout restart` = 1 WARN + 0-storm reconnect; Dir-1 2 s box→hub round-trip live in controller v0.139.0 | The operator-UI **save→apply** round-trip is not fired end-to-end live (hub UI password-gated; CC can't log in) → R-23; the wake transport and the ACK→config_version→`ConfigRefresher` delivery chain are each proven, only the UI-triggered bump leg is unexercised. Agent-plane (host-domain desired-state) poke **first slice PROVEN-LIVE** (Direction-2a, agent v0.89.0 + hub v0.59.0, 2026-07-17): contentless ep0-relayed UDP poke → agent immediate desired-state cycle, per `SPIKE-immediate-sync-transport-2026-07-16` P4. Full path live-proven: a real operator manifest save fired `poke: sync-poke delivered to 10.77.0.2`; the box (0.89.0) received it and logged `poke received → triggering an immediate desired-state cycle` → `out-of-band report triggered` — **~31 ms ep0→box, sub-ms to the report cycle** (WG-confined, from 10.77.0.1 to the 10.77.0.2-bound socket); save→tick ≈ ~0.45 s (SSH-dominated), well under ≤2–3 s. R-13 first slice (listener+sender only; the rest of the mutual-repair arc stays open). **System-initiated immediacy wired (hub v0.63.0, this REPORT):** the mutation sites that only OPERATOR actions used to notify now fire the correct plane's notifier when the hub itself mints state — agent-plane pokes at `PBSDRAutoProvision` (the observed slice-C lag), `ReissuePBSDR` (also the pbsdrheal escalation), `handlePBSDRReissue`, and the two admin desired-state api writers; controller-plane bump at `reissueOnReenroll`. Unit-tested + red-proofed, not yet fired on a real system event (folds into the rehearsal bind sequence). Still PARTIAL: the R-23 operator-UI save→apply leg and the agent **fast-tick-until-first-convergence** SECONDARY (the WG-registration leg a poke can't reach pre-tunnel) remain unfired live. **Fast-tick SHIPPED (agent v0.90.0, R-28):** while any desired-state item is unapplied — incl. the pre-tunnel window a poke can't reach — the agent pulses the out-of-band trigger every 30 s and self-disarms on convergence (state-based; four cached sources; LOUD states excluded). LIVE on both demo agents (the `fast-tick armed: 30s …` startup line verified). **REAL-ONBOARDING PROOF DONE — `tests/VALIDATION-n100-rehearsal-2026-07-18.md` (ledger 8, S5):** on a genuine first onboarding on metal, **every post-bind leg landed seconds apart with no ~15-minute stall anywhere** — bind 16:29:55 → credential delivered 16:30:21 (26 s) → agent 0.90.0 up 16:30:49 → **WG registered + tunnel applied 16:30:51 (~2 s)** → poke listener 16:30:54 → controller 16:32:28 → floor-lifted and running current 16:32:39. **Bind → running-current = 2 min 44 s.** The PBS-DR descriptor auto-provisioned on the same cadence (agent `converged state=applied` 16:45:53) — though see the DR-tier row: the descriptor converged while the credential behind it was already stale (R-39). The pre-tunnel fast-tick window is therefore proven in its real setting; the remaining PARTIAL is the R-23 operator-UI save→apply leg alone |
|
||||
| Config/state change round-trips in **seconds** (hub↔box immediacy; 15-min cycle stays the backbone): box→hub out-of-cycle report (Dir 1) + hub→box `GET /api/v1/wait` long-poll wake (Dir 2) | controller v0.139/140, hub v0.58/0.63 | **PROVEN-LIVE** (2026-07-21) | Transport proven live through the real DNS-only ingress: `SPIKE-immediate-sync-transport-2026-07-16` + hub v0.58.0 / controller v0.140.0 REPORTs — 240 s no-annotation hold (25 s heartbeat defeats nginx's 60 s `proxy_read_timeout`, no ingress change), 0.047 s wake-on-change, hub `rollout restart` = 1 WARN + 0-storm reconnect; Dir-1 2 s box→hub round-trip live in controller v0.139.0 | The operator-UI **save→apply** round-trip is not fired end-to-end live (hub UI password-gated; CC can't log in) → R-23; the wake transport and the ACK→config_version→`ConfigRefresher` delivery chain are each proven, only the UI-triggered bump leg is unexercised. Agent-plane (host-domain desired-state) poke **first slice PROVEN-LIVE** (Direction-2a, agent v0.89.0 + hub v0.59.0, 2026-07-17): contentless ep0-relayed UDP poke → agent immediate desired-state cycle, per `SPIKE-immediate-sync-transport-2026-07-16` P4. Full path live-proven: a real operator manifest save fired `poke: sync-poke delivered to 10.77.0.2`; the box (0.89.0) received it and logged `poke received → triggering an immediate desired-state cycle` → `out-of-band report triggered` — **~31 ms ep0→box, sub-ms to the report cycle** (WG-confined, from 10.77.0.1 to the 10.77.0.2-bound socket); save→tick ≈ ~0.45 s (SSH-dominated), well under ≤2–3 s. R-13 first slice (listener+sender only; the rest of the mutual-repair arc stays open). **System-initiated immediacy wired (hub v0.63.0, this REPORT):** the mutation sites that only OPERATOR actions used to notify now fire the correct plane's notifier when the hub itself mints state — agent-plane pokes at `PBSDRAutoProvision` (the observed slice-C lag), `ReissuePBSDR` (also the pbsdrheal escalation), `handlePBSDRReissue`, and the two admin desired-state api writers; controller-plane bump at `reissueOnReenroll`. Unit-tested + red-proofed, not yet fired on a real system event (folds into the rehearsal bind sequence). Still PARTIAL: the R-23 operator-UI save→apply leg and the agent **fast-tick-until-first-convergence** SECONDARY (the WG-registration leg a poke can't reach pre-tunnel) remain unfired live. **Fast-tick SHIPPED (agent v0.90.0, R-28):** while any desired-state item is unapplied — incl. the pre-tunnel window a poke can't reach — the agent pulses the out-of-band trigger every 30 s and self-disarms on convergence (state-based; four cached sources; LOUD states excluded). LIVE on both demo agents (the `fast-tick armed: 30s …` startup line verified). **REAL-ONBOARDING PROOF DONE — `tests/VALIDATION-n100-rehearsal-2026-07-18.md` (ledger 8, S5):** on a genuine first onboarding on metal, **every post-bind leg landed seconds apart with no ~15-minute stall anywhere** — bind 16:29:55 → credential delivered 16:30:21 (26 s) → agent 0.90.0 up 16:30:49 → **WG registered + tunnel applied 16:30:51 (~2 s)** → poke listener 16:30:54 → controller 16:32:28 → floor-lifted and running current 16:32:39. **Bind → running-current = 2 min 44 s.** The PBS-DR descriptor auto-provisioned on the same cadence (agent `converged state=applied` 16:45:53) — though see the DR-tier row: the descriptor converged while the credential behind it was already stale (R-39). The pre-tunnel fast-tick window is therefore proven in its real setting; the remaining PARTIAL is the R-23 operator-UI save→apply leg alone **2026-07-21 — THE LAST PARTIAL LEG IS CLOSED (R-23(a) restart leg).** The operator saved the global floor to a version the box did NOT run (0.153.0 → **v0.154.0**) and the managed self-update fired **exactly once**: `06:57:13Z` UpdateState pending (`initiated_by=auto-floor`) → `06:57:17Z` agent `controller swap requested` → `06:57:21Z` container restarted → `06:57:29Z` `new controller healthy`. **Save → healthy on the new version = 16 s.** Over a 39-minute window: swap requests **1**, agent-driven bootstrap restarts **1**, rollbacks **0**, container `RestartCount` **0**; `VerifyStartup` confirmed on the next boot and the following periodic check logged `Current version 0.154.0 is up to date` (the at/above-floor branch doing nothing, as designed). The 2026-07-20 attempt could not prove this because it targeted an already-running version. Evidence: `felhom-controller/REPORT.md` §6. |
|
||||
| Customer right-sizes guest RAM from the controller (agent-enforced bounds, live cgroup apply, no reboot) | agent v0.90.0 + controller v0.143.0 (R-24) | **PROVEN-LIVE** (grow **and** shrink on metal, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` (ledger 9) — the apply is now proven in both directions on a normal-sized box:** customer zero **shrank 11675 → 8192 MB at 16:50:22** and **grew 8192 → 12288 MB at 17:02:17**, each a live cgroup apply with **no reboot** (`local-api: guest-memory resized` in the agent journal, `[web] memory resized` in the controller log), and the new total rippled into the deploy page's memory math at 17:05:15 (`total=12288MB`). F5 auto-sizing had landed the guest at 11675 MB. Controller-direct (R-24's hub-desired-state framing SUPERSEDED, Viktor 2026-07-17). Agent `GET`/`POST /guest/memory` enforces every bound FRESH (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)) + verify-after-apply; PVE `SetConfig` hot-applies (**Phase-0 PROVEN** on the nested box: maxmem moves with the guest running, /proc/meminfo ripples via lxcfs, no reboot). Controller "Szerver memória (RAM)" card + code→Hungarian map, gated on `FeatureGuestMemoryResize` (MinAgent 0.90.0). LIVE-validated end-to-end through the real endpoint on the demo (above_max + below_min refusals render the Hungarian, agent English never leaks; SupportYes via the version header) | Row complete as of the 2026-07-18 rehearsal — the refusals were proven on the nested demo, the **applies** on the N100. Cores stay observation |
|
||||
| Publish train: MinAgent floors, gated auto-Reissue, version channels, floor-field-LAST rules | hub v0.45/0.53, agent | **PARTIAL** | `runbooks/publish-train-rules.md`; demo-fleet updates proven | **Box-side floor lift PROVEN-LIVE on a fresh install** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`): the day-0 golden deployed controller **0.143.0** at 16:32:28 and the managed floor lifted it to **0.145.0 by 16:32:34 — a 5-second, fully unattended update inside the first minute of controller life**, `update-state.json` recording `initiated_by: auto-floor` with `controller_updated` pushed to the hub. So the *mechanism* is no longer nested-only. **Still never proven on a real REMOTE customer** — parked trains `RUNBOOK-publish-0.79/0.81/0.85-*` await Peti → R-1. **Action before first invite: rebuild the golden to 0.145.x** now that this evidence is banked, so fresh boxes don't sit two versions stale |
|
||||
| Agent self-update: A/B slots, crash-loop auto-rollback, operator-signed | agent v0.70+ | **PROVEN-LIVE** (demo) | `SPIKE-agent-selfupdate-2026-07-05` | Remote-customer proof pending → R-1 |
|
||||
|
||||
@@ -53,7 +53,7 @@
|
||||
| R-17 | Old-box archive (u629193-sub1) retirement decision — 9/9 byte-identical restores verified | XS | awaiting-decision | Viktor ruling |
|
||||
| R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE |
|
||||
| R-20 | ~~Verify operator-key pinning is fully in the day-0 install flow~~ | XS | **closed** (2026-07-16) | Confirmed against `scripts/felhom-host-install.sh` source (not changelog): keys resolve at L1181–1219 (script constants `OPERATOR_KEY_*`, populated, `--operator-pubkey-file` override), pinned automatically by `step_agent_config()` "STEP 6/8" (L2044; python builds `authz.signers` L2146–2156, reinstall preserves existing), verified at L2332–2337 ("authz signers: N … operator-signed self-update armed"). No interactive prompt or post-install hand-edit — fully automatic. Doc-drift note: the L193–197 "EMPTY by default" comment is stale vs the now-populated constants (→ R-16 hygiene) |
|
||||
| R-23 | **Immediate-sync Direction-2 follow-ups** (hub v0.58 / controller v0.140, 2026-07-16): **(a) — MOSTLY BANKED 2026-07-20.** The operator-UI save->apply round trip is PROVEN: the STOP-2 global-floor save (hub `18:56:27 CEST`) released the controller's held wait in the **same second** (`16:56:27Z wait woke: generation=1 - firing out-of-cycle report`), with the report built 2 s later; the ring also shows `wait baseline generation=0` at startup (baseline recorded WITHOUT firing, as designed) then `generation=1`, so the generation advanced past 0. **STILL OPEN: the self-restart single-fire leg** - the floor was set to a version the box ALREADY ran, so there was no work and no restart. Finish by bumping the floor to a version the box does NOT run, debug ring open, asserting EXACTLY ONE restart. **Trap found while banking this: the wake is `logx.Debugf`, so it is INVISIBLE in `docker logs` at INFO** and lives only in the debug ring (`GET /api/debug/logs?level=DEBUG`) - a hunter looking at stdout wrongly concludes the box never woke; arguably (b) generalised. (b) cosmetic: the Waiter's "recovered" INFO logs on the next hold completion (`pollOnce` blocks ~240 s), not at reconnect | S | **(a) mostly banked; restart leg + (b) open** | Flips the map row "config/state change round-trips in seconds" PARTIAL->PROVEN-LIVE once the restart leg lands. Evidence: `felhom-controller/REPORT.md` 4f |
|
||||
| R-23 | **Immediate-sync Direction-2 follow-ups** (hub v0.58 / controller v0.140, 2026-07-16): **(a) — BANKED 2026-07-21 (both legs).** The operator-UI save->apply round trip is PROVEN: the STOP-2 global-floor save (hub `18:56:27 CEST`) released the controller's held wait in the **same second** (`16:56:27Z wait woke: generation=1 - firing out-of-cycle report`), with the report built 2 s later; the ring also shows `wait baseline generation=0` at startup (baseline recorded WITHOUT firing, as designed) then `generation=1`, so the generation advanced past 0. **RESTART LEG BANKED 2026-07-21 — the floor was moved to a version the box did NOT run, and the swap fired EXACTLY ONCE.** Operator saved global floor 0.153.0 → **v0.154.0** (a real version boundary, unlike the 2026-07-20 attempt which targeted an already-running version and therefore proved nothing). Timeline (guest UTC): `06:57:13` UpdateState `pending` written `initiated_by=auto-floor` → `06:57:17` agent `controller swap requested 0.153.0 -> 0.154.0` → `06:57:19` `image file written, restarting bootstrap` → `06:57:21` container StartedAt + UpdateState `completed_at` → `06:57:29` agent `new controller healthy`. **16 s end to end.** Assertions over the whole window (06:50 → 07:29, 39 min): `controller swap requested` = **1**, agent-driven bootstrap restarts = **1**, `new controller healthy` = **1**, rollback/swap-failed/unhealthy = **0**, container `RestartCount` = **0**. `VerifyStartup` banked it on the next boot (`Post-update startup: update successful (0.153.0 → 0.154.0)`) and the `06:57:52` periodic check logged `Current version 0.154.0 is up to date` — the at/above-floor branch correctly doing nothing. **No storm, no rollback, no second attempt.** *Caveat, disclosed: a hand-deploy of v0.155.0 at `07:17:10` sits inside the observation window and is what StartedAt shows after that point; it never touches `SwapController`, so the swap-count assertions above are uncontaminated across the full window. A second, unplanned confirmation of the at/above-floor branch came with it — after the hand-deploy the box ran 0.155.0 against a 0.154.0 floor and the updater logged `Current version 0.155.0 is up to date` and did nothing.* Evidence: `felhom-controller/REPORT.md` §6 (2026-07-21). *(Superseded note:)* **the self-restart single-fire leg was** - the floor was set to a version the box ALREADY ran, so there was no work and no restart. Finish by bumping the floor to a version the box does NOT run, debug ring open, asserting EXACTLY ONE restart. **Trap found while banking this: the wake is `logx.Debugf`, so it is INVISIBLE in `docker logs` at INFO** and lives only in the debug ring (`GET /api/debug/logs?level=DEBUG`) - a hunter looking at stdout wrongly concludes the box never woke; arguably (b) generalised. (b) cosmetic: the Waiter's "recovered" INFO logs on the next hold completion (`pollOnce` blocks ~240 s), not at reconnect | S | **(a) BANKED in full; only (b) cosmetic remains** | Map row "config/state change round-trips in seconds" flipped PARTIAL->PROVEN-LIVE 2026-07-21 on this evidence. Evidence: `felhom-controller/REPORT.md` 4f |
|
||||
| R-34 | **Backup data lifecycle management.** An "inactive backups" section on „Távoli mentés": apps that have snapshots but no active backup — **disabled OR uninstalled** — listed with name / size / last snapshot / restorable, plus an explicit **double-confirmed per-app delete** via `restic forget --tag` + nightly prune. | M | idea | **RULING: the offsite toggle NEVER offers deletion — policy and destruction stay decoupled.** Turning backups off must never be a data-destroying act, and deletion must never hide behind a toggle. Origin: 2026-07-18 rehearsal. Pairs with R-32 (that one is the operator's view of dead bytes; this one is the customer's) |
|
||||
| R-35 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. | S | idea | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` |
|
||||
| R-36 | **Post-RESET re-enroll leaves offsite "enabled but unprovisioned" — silently.** The hub knows the state and says nothing on the customer page. | S | **SHIPPED (hub v0.67.0, 2026-07-18)** | Both halves delivered. **(1) The warning:** the customer page now names the state and the fix — enabled-but-unprovisioned raises an amber banner saying provisioning is *Save*-triggered (press Save once, then verify), reusing the exact `enabled && type == ""` predicate the offsite re-issue handler already refuses on. **(2) The related sub-item, also done:** the self-bind link is now **auto-minted at customer creation AND at RESET completion**, so the console banner's „e-mailben kapott link" is already true instead of true-once-the-operator-remembers. Extracting the shared `mintAndSendSelfBindLink` core keeps the button and the auto-mint callers on the same F1/F2 honesty rules, and the auto-mint never fails the operation it rides on. **Gap found and closed while wiring it:** `PurgeCustomerResetDBState` does NOT clear `selfbind_tokens`, so a link minted BEFORE a reset would have stayed live across it — the skip paths now clear stale tokens, giving the invariant "after auto-mint, the only live link is one we just issued, or none". Tests assert the banner is ABSENT in all three nominal cases too; red-proofed. — Original analysis: Source-cited behaviour, confirmed live in the rehearsal: **provisioning is Save-triggered** (`configs.go` `applyOffsite`) — which also answers S6's open question — and the re-enroll auto-re-issue **correctly** skips unprovisioned targets (`handler.go`). So nothing is broken; the gap is that nobody is told. Direction: flash it on the customer page. **Interim: an R-3 step.** **Related sub-item:** auto-mint the **self-bind link on customer create/RESET**, so the console banner's „e-mailben kapott link" is always already true instead of true-once-the-operator-remembers |
|
||||
|
||||
Reference in New Issue
Block a user