4.7 KiB
4.7 KiB
LIVE-DRIVE-FINDINGS — felhom-controller end-to-end live drive
- Date: 2026-06-14
- Branch:
audit/2026-06-14-live-drive - Versions under test: controller v0.60.0, agent v0.30.0, hub v0.11.0 (per runbook)
- Target: demo guest 9201 (
demo-felhom) on Proxmox hostfelhom-pve(192.168.0.162) - Mode: unattended discovery drive — find bugs, do NOT fix. Destructive ops authorized (demo disposable).
Control-plane access method used
The demo dashboard has no password set, so the controller's RequireAuth and CsrfProtect middleware both skip (auth.go / csrf.go). The full REST API + web forms are therefore open over the public URL https://felhom.demo-felhom.eu. I drive the controller as a user would via that URL (NOT the container IP, which 404s on a network/mux nuance). Host SSH (root@felhom-pve) + pct exec/docker exec are used only for baseline setup and ground-truth verification — never to perform the operation under test.
- Verified:
GET /api/health→{"ok":true};GET /settingsandGET /→ 200 without auth.
Progress log (timestamped, Europe/Budapest)
- ~start — Baseline captured. Controller
:0.60.0 Up (healthy), agent0.30.0 active. 55 stacks in catalog; deployed/running: actualbudget (only customer app) + protected infra traefik, cloudflared, filebrowser. RomM is currently not_deployed (memory/CONTEXT said it was the live HDD app — state has since changed). Headroom: rootfs 32G (29G free), docker-data 252G (238G free). Disks: felhom-usb (user-data, data_bearing, /dev/sdb1, uuid:277a2179…), local + local-lvm (system), felhom-pbs (backup, 192.168.0.180:felhom-spike). Hub-side online status TBD.
BASELINE
| Item | Value |
|---|---|
| Controller image/status | gitea.dooplex.hu/admin/felhom-controller:0.60.0 Up (healthy) |
| Agent | felhom-agent 0.30.0, systemd active |
| Guest rootfs | 32G total, 29G avail (4% used) |
| Guest docker-data (/var/lib/docker) | 252G total, 238G avail (1% used) |
| Guest RAM (LXC cgroup) | 2048 MB (config memory: 2048, swap 512) |
| Deployed customer apps | actualbudget |
| Protected infra running | traefik, cloudflared, filebrowser, felhom-controller |
| Disks | felhom-usb (HDD, user-data), local/local-lvm (system), felhom-pbs (offsite backup) |
FINDINGS
F1 — Memory metric reports HOST RAM (16GB), not the guest's 2GB cap — HIGH
- Area: §11 monitoring accuracy / deploy headroom guard
- Action:
GET /api/system/info; cross-checked withfreeinside the LXC and the LXC config. - Expected: memory total ≈ the guest's 2048 MB cgroup limit.
- Actual:
/api/system/info→total_mem_mb: 15771,avail_mem_mb: 13577. The LXC is capped at 2048 MB (/etc/pve/lxc/9201.conf: memory: 2048);free -minside the LXC correctly shows 2048. The controller Docker container reads the host/proc/meminfo(MemTotal: 16150380 kB≈ 15772 MB) because lxcfs is not bind-mounted into the container and its cgroupmemory.max = max(no Docker memory limit). - Evidence:
- LXC
free -m:Mem: 2048 ... available 1829 - controller container
head -1 /proc/meminfo:MemTotal: 16150380 kB;cat /sys/fs/cgroup/memory.max→max /etc/pve/lxc/9201.conf:memory: 2048,swap: 512
- LXC
- Impact: The monitoring memory bar and (critically) the deploy memory-headroom guard believe the guest has ~13.5 GB free when only ~1.8 GB is usable. An operator deploying apps by that figure will OOM-kill the 2 GB guest. The 3-app bound in the runbook exists precisely because the real ceiling is 2 GB, which the UI hides.
- Verdict: broken (misleading metric + unsafe headroom basis). Severity: HIGH.
F2 — hdd_configured: false despite an attached user-data HDD — LOW (verify)
/api/system/inforeportshdd_configured: falseeven though felhom-usb (HDD, role user-data, data_bearing) is attached and mounted at /mnt/felhom-usb. To confirm whether this gates any HDD-app deploy UI. Severity TBD.
F3 — Hungarian catalog text double-UTF-8-encoded in API JSON — LOW (verify)
GET /api/stacksreturns descriptions likeSzemĂ©lyes pĂ©nzĂĽgyek(mojibake for "Személyes pénzügyek"). Looks like double-encoding. Need to confirm whether the rendered HTML UI is affected or only the JSON API. Severity TBD.
F4 — /api/stacks/rescan returns "stack not found: rescan" — LOW
- The runbook's documented rescan endpoint
GET /api/stacks/rescanis routed as a stack name lookup →{"ok":false,"error":"stack not found: rescan"}. Either the route was removed/renamed or the runbook is stale. (Sync/rescan is reachable viaPOST /api/sync.) Cosmetic but documents a stale/missing endpoint.