From c1c8fe2a7d2b1d6734be04ee8b7f9a5c6b053c42 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 14 Jun 2026 15:43:07 +0200 Subject: [PATCH] docs: F9/F20-BUG2/F20-BUG3 FIXED (agent v0.31.0, live-validated); queue golden controller-tag follow-up --- .../audits/live-drive-fixspec-2026-06-14.md | 11 ++++++- .../FOLLOWUP-golden-default-controller-tag.md | 32 +++++++++++++++++++ documentation/backlog/README.md | 4 +++ 3 files changed, 46 insertions(+), 1 deletion(-) create mode 100644 documentation/backlog/FOLLOWUP-golden-default-controller-tag.md diff --git a/documentation/audits/live-drive-fixspec-2026-06-14.md b/documentation/audits/live-drive-fixspec-2026-06-14.md index 06e48b3..4e8e3c3 100644 --- a/documentation/audits/live-drive-fixspec-2026-06-14.md +++ b/documentation/audits/live-drive-fixspec-2026-06-14.md @@ -38,7 +38,16 @@ key fixes live-verified. (F9, F20-BUG2, F20-BUG3 remain for the SUPERVISED agent | F5 (catalog) | **FIXED** | app-catalog `main` | uptime-kuma healthy → route 302 (was 404) | | F5 (dashboard) | **FIXED** | `803ce50` | funcmap + template render tests | | F17 | **FIXED** | `0b9450e` | **marker DB round-trip PASSED live** (escape hatch cleared) | -| F9, F20-BUG2, F20-BUG3 | DEFERRED | — | SUPERVISED next session (agent/golden) | +| F9 | **FIXED** | agent `4cd1d02` (+`a2a76e7`), ctrl `0.63.0` | live: agent-restart auto-re-asserted the bind (no manual call); `guest_attached=True`; HDD app deployed onto `/dev/sdb1` | +| F20-BUG2 | **FIXED** | agent `a2a76e7`, ctrl `0.63.0` | live: confirmed wipe with `/api/disks` `wipe_durable_id` accepted (no binding_mismatch) | +| F20-BUG3 | **FIXED** | agent `4777f8a` | live on 916GB felhom-usb: 2s client timeout → ~30s mkfs ran to a clean ext4; agent restart mid-format recovered + completed clean | + +> **2026-06-14 supervised session (agent v0.31.0 + controller v0.63.0):** F9 / F20-BUG2 / F20-BUG3 +> implemented + deployed + live-validated. Approach was **attach-to-existing** (no re-provision — 9201 +> ran v0.62.0; the golden bakes a stale controller `:0.43.0`, so re-provision would regress it). The +> golden's stale default controller tag is queued as a separate follow-up (`documentation/backlog/`). +> Note: the controller has no automated SSD↔HDD data-migration feature (de-privileging removed it) — F9 +> unblocks HDD-app deployment with data on the real HDD, which is what was proven. Note on F1: the FIXSPEC's "read the cgroup limit" approach proved a **no-op on the demo** — the controller container's own cgroup is unlimited (the 2 GB cap is on the LXC ancestor, hidden) and there is diff --git a/documentation/backlog/FOLLOWUP-golden-default-controller-tag.md b/documentation/backlog/FOLLOWUP-golden-default-controller-tag.md new file mode 100644 index 0000000..2e68eac --- /dev/null +++ b/documentation/backlog/FOLLOWUP-golden-default-controller-tag.md @@ -0,0 +1,32 @@ +# FOLLOW-UP — bump the golden's default controller tag + validate the full provision path + +**Status:** OPEN (queued 2026-06-14). Surfaced during the F9/F20 supervised session (agent v0.31.0). +**Class:** provisioning correctness / customer-onboarding. **Risk:** SUPERVISED (golden + a real destroy→provision). + +## The problem +`felhom-agent/configs/build-golden.sh:43` bakes the controller image into the golden as: + +``` +CONTROLLER_IMAGE="${6:-gitea.dooplex.hu/admin/felhom-controller:0.43.0}" +``` + +`:0.43.0` is ~20 versions stale (current is `0.63.0`). The bootstrap writes it to +`/etc/felhom-controller-image` and `docker run`s that tag on first boot. So the **next time a golden is +actually baked for a real provision, it would stand a customer guest up on an ancient controller** — +missing every fix since 0.43.0 (incl. F1 memory guard, F17 DB restore, the M18/M19 fixes, etc.). It only +hasn't bitten because the running demo guest 9201 had its `/etc/felhom-controller-image` updated in place +post-provision; a fresh provision would not. + +This is why the F9 session chose **attach-to-existing** over destroy+re-provision (a re-provision would +have regressed 9201's controller 0.62.0 → 0.43.0). + +## The fix (separate task) +1. Bump the `build-golden.sh` default `CONTROLLER_IMAGE` to the current released controller tag (and + establish a convention so it tracks releases — e.g. read a `LATEST_CONTROLLER` pin, or pass it + explicitly from the build pipeline). +2. **Validate the full path end-to-end**, which *does* warrant a supervised destroy + re-provision (it is + the real customer-onboarding flow, not coverable by attach-to-existing): bake golden → provision a + scratch guest → first boot stands up the **current** controller → controller ONLINE on the hub → + enrolled user-data drive is bound (F9 `ReassertGuestBinds` / provision bind) → an app deploys. +3. While there: confirm the provision path itself asserts known user-data binds (F9 part A for the + re-provision case, complementing the agent-startup re-assert already shipped in v0.31.0). diff --git a/documentation/backlog/README.md b/documentation/backlog/README.md index b88fd24..40b7621 100644 --- a/documentation/backlog/README.md +++ b/documentation/backlog/README.md @@ -9,6 +9,10 @@ Verified-LIVE findings with implementable fix plans that are **not yet implement - **FIX-M19-NOTES.md** — `deriveStackName` misattribution edge (low-incidence correctness). **FIXED** in controller v0.62.0 @ `6bab68b` (2026-06-14). (was on the deleted branch `fix/m19-stackname-crossref`.) +- **FOLLOWUP-golden-default-controller-tag.md** — the golden bakes a stale controller `:0.43.0` + (`build-golden.sh:43`); a fresh provision would stand up an ancient controller. Bump it + validate the + full destroy→provision→first-boot path (warrants a supervised re-provision). Queued 2026-06-14. + Related: the live-drive fixspec (`../audits/live-drive-fixspec-2026-06-14.md`) carries the **deferred supervised items** F9 (HDD provisioning/guest-attach), F20-BUG2 (durable_id scheme), F20-BUG3 (async mkfs) — to be implemented in the agent/golden supervised session.