8.3 KiB
8.3 KiB
LIVE-DRIVE-FINDINGS — felhom-controller end-to-end live drive
- Date: 2026-06-14
- Branch:
audit/2026-06-14-live-drive - Versions under test: controller v0.60.0, agent v0.30.0, hub v0.11.0 (per runbook)
- Target: demo guest 9201 (
demo-felhom) on Proxmox hostfelhom-pve(192.168.0.162) - Mode: unattended discovery drive — find bugs, do NOT fix. Destructive ops authorized (demo disposable).
Control-plane access method used
The demo dashboard has no password set, so the controller's RequireAuth and CsrfProtect middleware both skip (auth.go / csrf.go). The full REST API + web forms are therefore open over the public URL https://felhom.demo-felhom.eu. I drive the controller as a user would via that URL (NOT the container IP, which 404s on a network/mux nuance). Host SSH (root@felhom-pve) + pct exec/docker exec are used only for baseline setup and ground-truth verification — never to perform the operation under test.
- Verified:
GET /api/health→{"ok":true};GET /settingsandGET /→ 200 without auth.
Progress log (timestamped, Europe/Budapest)
- ~start — Baseline captured. Controller
:0.60.0 Up (healthy), agent0.30.0 active. 55 stacks in catalog; deployed/running: actualbudget (only customer app) + protected infra traefik, cloudflared, filebrowser. RomM is currently not_deployed (memory/CONTEXT said it was the live HDD app — state has since changed). Headroom: rootfs 32G (29G free), docker-data 252G (238G free). Disks: felhom-usb (user-data, data_bearing, /dev/sdb1, uuid:277a2179…), local + local-lvm (system), felhom-pbs (backup, 192.168.0.180:felhom-spike). Hub-side online status TBD.
BASELINE
| Item | Value |
|---|---|
| Controller image/status | gitea.dooplex.hu/admin/felhom-controller:0.60.0 Up (healthy) |
| Agent | felhom-agent 0.30.0, systemd active |
| Guest rootfs | 32G total, 29G avail (4% used) |
| Guest docker-data (/var/lib/docker) | 252G total, 238G avail (1% used) |
| Guest RAM (LXC cgroup) | 2048 MB (config memory: 2048, swap 512) |
| Deployed customer apps | actualbudget |
| Protected infra running | traefik, cloudflared, filebrowser, felhom-controller |
| Disks | felhom-usb (HDD, user-data), local/local-lvm (system), felhom-pbs (offsite backup) |
FINDINGS
F1 — Memory metric reports HOST RAM (16GB), not the guest's 2GB cap — HIGH
- Area: §11 monitoring accuracy / deploy headroom guard
- Action:
GET /api/system/info; cross-checked withfreeinside the LXC and the LXC config. - Expected: memory total ≈ the guest's 2048 MB cgroup limit.
- Actual:
/api/system/info→total_mem_mb: 15771,avail_mem_mb: 13577. The LXC is capped at 2048 MB (/etc/pve/lxc/9201.conf: memory: 2048);free -minside the LXC correctly shows 2048. The controller Docker container reads the host/proc/meminfo(MemTotal: 16150380 kB≈ 15772 MB) because lxcfs is not bind-mounted into the container and its cgroupmemory.max = max(no Docker memory limit). - Evidence:
- LXC
free -m:Mem: 2048 ... available 1829 - controller container
head -1 /proc/meminfo:MemTotal: 16150380 kB;cat /sys/fs/cgroup/memory.max→max /etc/pve/lxc/9201.conf:memory: 2048,swap: 512
- LXC
- Impact: The monitoring memory bar and (critically) the deploy memory-headroom guard believe the guest has ~13.5 GB free when only ~1.8 GB is usable. An operator deploying apps by that figure will OOM-kill the 2 GB guest. The 3-app bound in the runbook exists precisely because the real ceiling is 2 GB, which the UI hides.
- Verdict: broken (misleading metric + unsafe headroom basis). Severity: HIGH.
F2 — hdd_configured: false despite an attached user-data HDD — LOW (verify)
/api/system/inforeportshdd_configured: falseeven though felhom-usb (HDD, role user-data, data_bearing) is attached and mounted at /mnt/felhom-usb. To confirm whether this gates any HDD-app deploy UI. Severity TBD.
F3 — Hungarian catalog text double-UTF-8-encoded in API JSON — LOW (verify)
GET /api/stacksreturns descriptions likeSzemĂ©lyes pĂ©nzĂĽgyek(mojibake for "Személyes pénzügyek"). Looks like double-encoding. Need to confirm whether the rendered HTML UI is affected or only the JSON API. Severity TBD.
F5 — uptime-kuma: broken catalog healthcheck → permanently unhealthy → total 404 outage via Traefik — HIGH
- Area: §2 deploy / §2b health detection / §3 routing — a cascade.
- Action: deployed uptime-kuma (
POST /api/stacks/uptime-kuma/deploy), observed state, then traced the 404. - Expected: deploy → healthy → status.demo-felhom.eu serves the app (200).
- Actual / evidence — the cascade:
- Catalog bug:
/opt/docker/stacks/uptime-kuma/docker-compose.ymldefineshealthcheck.test: ["CMD","node","/app/extra/healthcheck.mjs"], but in thelouislam/uptime-kuma:2image that file does not exist →docker inspecthealth log:Error: Cannot find module '/app/extra/healthcheck.mjs',Health=unhealthy FailingStreak=5. - The app process is actually fine —
curl http://uptime-kuma:3001/from a peer container → 302 (serving). - Traefik gates route registration on Docker health. On the websecure (443) entrypoint,
status.demo-felhom.eu → 404, while every healthy app (tasks/recipes/share/paste) → 200. The HTTP→HTTPS redirect on :80 returns 301 for all hosts (global catch-all), which masks the missing 443 router. So the unhealthy container's TLS router is never published. - Net result: uptime-kuma is completely unreachable at its URL (404) for the customer, despite the app running — purely because of a wrong healthcheck path in the catalog.
- Catalog bug:
- Verdict: broken. Severity: HIGH. Two issues to file: (a) catalog healthcheck wrong for uptime-kuma:2; (b) design risk — any app with a broken/too-slow healthcheck doesn't just show "unhealthy", it becomes a hard 404 outage. The controller's deploy returns success and the dashboard shows "deployed (unhealthy)", giving no hint that the URL is dead. Consider surfacing "route not published because unhealthy" to the operator.
- Good part: the controller did correctly detect and surface
unhealthy(GET /api/stacks/uptime-kuma→state=unhealthy) — health detection itself works.
F6 — Deploy POST returns "deployed" optimistically, before compose completes / before health is known — MEDIUM (API contract)
- Action: timed
POST /api/stacks/<app>/deployvs controller logs. - Evidence: vikunja POST returned ~0s but log shows compose took 3.4s; uptime-kuma POST returned ~0s but compose pull took 31.7s and the app ended unhealthy. The response
{"ok":true,"message":"Stack <app> deployed"}is sent before the container is up and regardless of eventual health. - Impact: This is the documented in-memory
Deployed=true-before-compose uppattern (avoids the card flipping back mid-pull), and the UI compensates by pollingGET /api/stacks/<app>. But the API message "deployed" is misleading — an API consumer (or a script) that trusts the POST result will think a broken/unhealthy app succeeded (see F5). Verdict: clunky/misleading message; not a data-integrity bug. Severity: MEDIUM.
F7 — API state lags Docker health by ~10s after deploy — LOW
- mealie's container reported
(healthy)indocker ps~12s beforeGET /api/stacks/mealieflipped fromstartingtorunning. Cosmetic polling lag; transient. Severity: LOW.
F8 — cloudflared TUNNEL_TOKEN stored in plaintext in docker-compose.yml — LOW/INFO
- The cloudflared infra stack's compose holds
TUNNEL_TOKEN=<redacted>in plaintext on disk (notenc:-wrapped like app.yaml secrets). May be acceptable for an infra/base-bringup stack, but worth confirming against the "secrets must be encrypted at rest" posture. Token redacted here. Severity: LOW/INFO.
F4 — /api/stacks/rescan returns "stack not found: rescan" — LOW
- The runbook's documented rescan endpoint
GET /api/stacks/rescanis routed as a stack name lookup →{"ok":false,"error":"stack not found: rescan"}. Either the route was removed/renamed or the runbook is stale. (Sync/rescan is reachable viaPOST /api/sync.) Cosmetic but documents a stale/missing endpoint.