The undo, built and proven live: controller v0.263.2 (09 decision 15)
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two live-only defects, §6.4 part 1 SHIPPED. - Capability map: a failed update is undone by the box - PROVEN-LIVE. - Live evidence on 9202: three apps undone by the product with seeds before the backup, after it and seconds before the press read back; cut-off copy held honestly; power cut during the undo resumed; manual press after undo. - Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643 ruled; R-646 opened. STATUS asks the floor question. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -335,3 +335,7 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-539** | **The restart brake caught a FAST crash loop and was blind to a SLOW one (P3).** Closed by operator ruling 3 of 2026-09-16 in agent **v0.132.0** (`18d03bd`, tag `v0.132.0`, sha256 `4afe8157…`) + hub **v0.117.0** (`37ae31f`): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth `slow_crashloop_since` moves (at most once per 24 h) and the hub mints `controller_slow_crashloop` (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed `agent_update` to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`. **Proven live with the PRODUCTION window (no test-only interval):** five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 `SLOW CRASH-LOOP … restarts_24h=5`; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). **Reasoning kept:** the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. `audits/evidence-chaos-fixes-2026-09-17/partC-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 3c1882a:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-637** | **Build the undo (`09` §3 decision 15) — the box puts a failed update back by itself (P2).** Closed in controller **v0.263.2** (`8fc2b4a` v0.263.0, `5d38573` v0.263.1, `2cd6666` v0.263.2). Copy method chosen by the bake-off (decision 19): a folder copy of every NAMED volume, `cp -a` into `<vol>.pre-update-<stamp>` after the pull, finished-marker last; bind folders never touched. Undo in `failAndHold`: copies validated first, volumes refilled, definition + pin + the pinned version's `.felhom.yml` (`applied-meta/`) put back from the job's own copies, old probe, `undone` or a HOLD whose sentence says the undo failed and the data state. **Proven live on 9202 (0.263.2):** docmost, romm, vikunja undone by the product with seeds before the backup, after it, and seconds before the press all read back, ledgers equal, page line hu/en; a cut-off copy → HOLD `untouched`; a power cut (`pct stop`) during `undoing` → resumed and undone; a manual press after an undo → `done`, note cleared; removal deleted kept copies. **Two defects the live proof found and the unit tests could not:** the undo's probe was gated on `running` while the current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` was already the new one because it flows in on every sync (v0.263.2). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. **Reasoning kept:** *a copy counts only with its finished-marker, judged by the helper container's own exit — killing `docker run` does not stop the copy; the undo reads its own copies, never the recovery unit.* | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-639** | **After a held update the previous definition survived only in the recovery unit (P3).** Closed in controller **v0.263.0/v0.263.2**: the pre-update compose, applied definition, pin and the pinned version's `.felhom.yml` are kept until the undo is over; the undo never reads the unit (R-645). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — SHIPPED** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-641** | **An app with no database server had no last-second copy (P2).** Closed in controller **v0.263.0**: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-642** | **Start answered 200 "start completed" over a crash loop (P3).** Closed in controller **v0.263.0**: start/restart answer `requested — state now: <state>`, never "completed" (`startAnswer`, pinned by `TestR642_*`, red-proofed); live on 9202: `Stack romm start requested — state now: running`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user