Files
felhom-controller/SESSION-2026-06-14-fixbatch.md
T

14 lines
3.9 KiB
Markdown

# SESSION 2026-06-14 — Fix Batch 1 (+F17) → controller v0.61.0
Branch: `audit/2026-06-14-live-drive`. Implementing controller-side live-drive fixes per `LIVE-DRIVE-FIXSPEC-2026-06-14.md`.
Exclusions: F9, F20-BUG2, F20-BUG3 (agent/golden — supervised next session).
Operator decisions applied: F17 `.sql` wins; F5 gate route but surface "route unpublished"; F8 0600; F11 warn-and-allow (note: F11 not in this batch's fix list — only F1, F20-BUG1, F4, F6, F7, F8, F5, F17).
## Progress log
- **t0** — Read CLAUDE.md (build workflow), FIXSPEC, target source files. Setting up.
- **Test approach (documented assumption):** controller code is largely `//go:build linux`; dev host is Windows. Per-commit gate = `CGO_ENABLED=0 GOOS=linux go build ./... && go vet ./...` locally (compiles + type-checks code AND `_test.go`). Authoritative `go test ./...` runs on the Linux build server (192.168.0.180) against the pushed branch BEFORE the image build/deploy. Windows cannot execute linux test binaries, so this is the faithful interpretation of the per-commit green gate.
- **t2 — quick wins DONE** (separate commits): F20-BUG1 (agentapi.FormatDisk surfaces non-2xx/ok:false instead of zero-value success; test passes, fails on old code), F4 (405 for non-POST /stacks/rescan), F6 (deploy POST → 202 "Telepítés elindítva"; UI checks data.ok so 202 safe), F7 (status-refresh 30s→10s), F8 (writeConfig0600 helper enforces 0600 even on a pre-existing file; test skips on Windows, asserts on Linux). F5 catalog: uptime-kuma healthcheck fixed (→ `extra/healthcheck` binary + 180s start_period) committed+pushed to app-catalog main. F5 dashboard: `routeUnpublished` funcmap helper + indicator on stacks/dashboard cards + CSS; tests (func + real-template parse + fragment render) pass.
- **t3 — F17 DONE.** Reuse decision: put `ImportDump`/`waitDBReady` in `appbackup` (the DB-domain home with DiscoverDatabases/DumpOne/getMariaDBPassword) rather than reusing appexport's unexported copies — appexport→appbackup already exists, so the reverse would CYCLE; appbackup is the clean shared home and `DiscoveredDB` already carries the live container's creds (no env threading). `backup.reimportDBDumps` (injectable discover/import seams) runs after volume restore in both RestoreFromRecoveryUnit and RestoreApp → `.sql` WINS. Volume-restore + DB-import failures now SURFACE (restore returns error). Unit tests pass; full backup/appbackup/agentapi/web suites green locally. **F17 escape hatch:** shipping it in v0.61.0 contingent on the LIVE DB round-trip passing post-deploy; if it fails, revert F17 to branch fix/f17-restore-db-reimport and ship the rest.
- **t4 — DEPLOY + LIVE VERIFY.** Full `go test ./...` GREEN on Linux build server (incl. linux `system`/`api` tests). Built+pushed `:0.61.0`, deployed to 9201 (golden/bootstrap). Healthy, /api/health + dashboard 200. **F17 escape hatch CLEARED** — live romm marker round-trip PASSED (drop after backup → `/backup/restore` → "replayed 1 DB dump(s)" → row returned). **F1 reroot:** discovered the cgroup approach was a no-op (container cgroup=max; cap is on the LXC ancestor; no lxcfs) → reworked to `docker info` MemTotal + committed-memory guard → rebuilt+redeployed → **`total_mem_mb=2048`** live (was 15771). F4 → 405 live. F5 → uptime-kuma healthy → route 302 live (was 404). Docs updated (CHANGELOG, README, FIXSPEC statuses, REPORT). DONE — all Batch-1 + F17 shipped in v0.61.0; F9/F20-BUG2/F20-BUG3 deferred to supervised session.
- **t1 — F1 (cgroup memory)** DONE. `internal/system/info_linux.go`: `readMemInfo` now prefers the cgroup memory LIMIT (v2 `memory.max`, v1 `memory.limit_in_bytes`, "max"/near-uint64-max = unlimited→fallback) when finite and below the host `/proc/meminfo` total; used = `memory.current`/`usage_in_bytes`. Restores the deploy OOM guard (deploy.go:162-185 reads GetMemoryMB). Test `info_cgroup_test.go` (4 cases) — fails on pre-fix code (ignored cgroup). GOOS=linux build+vet OK.