7f6b00375b
Reproduced twice, two independent runs, same cycle (the runner's RNG is seeded so both drew the same permutation - deterministic, not coincidence). A hard reset mid-backup brought everything back except the app half of the DB-backed stack: rallly left Exited 255, oom=false, restarts=0, its own log ending "Ready" - it died healthy - while rallly-postgres returned healthy. 20:28:13 Status refresh: 8 containers across 55 stacks <-- docker ps -a shows NINE 20:28:18 [bootrecon] Boot reconciliation: no boot-orphaned apps 20:28:25 Status refresh: 7 ... 8 containers <-- still churning AFTER the sweep 20:39:14 [deadapp] 20 scans, 5 deployed evaluated, 1 currently down The predicate is sound: once settled the controller reports rallly state=degraded containers=2, and IsDownState includes StateDegraded, so len>0 && IsDownState holds. The SNAPSHOT was wrong. bootrecon fires as a goroutine ~5s after start while docker is still restoring containers, and is start-once by design, so it never re-checks. Consequence: the app stays down indefinitely. Detection is perfect and recovery never happens - R-52's original shape, an alarm with no recovery. Not fixed. Distinguished from this campaign's two earlier HARNESS defects: both drives bound, every other app returned incl. the drive-backed one, only the app half of a two-container stack missing while its DB is healthy, and it surfaced through the fixed check written for exactly this.
Felhom — Documentation
Felhom is a managed home-server service for Hungarian households, built on a three-component model over Proxmox:
- Hub — operator backend on k3s (
hub.felhom.eu). Repo:felhom.eu/hub/. - Host agent — one per Proxmox host; operator-tier; owns all Proxmox interaction. Repo:
felhom-agent/. - In-guest controller — one per customer LXC; Docker-only; manages the customer's apps. Repo:
felhom-controller/.
This directory is the central, code-verified documentation home for all three components plus the platform and the security-audit record.
Sections
Controller (in-guest) — controller/
The Docker-only app-domain controller. Full per-area docs grounded in current source (v0.59.0).
→ controller/README.md: module map, deploy & stack lifecycle, backup
architecture, storage/monitoring/metrics, auth/hub/sync/integrations.
Host agent & platform — architecture/, proxmox-platform.md
The operator-tier agent and the Proxmox platform.
architecture/01-topology-and-trust.md— topology & trust modelarchitecture/03-host-agent.md— the host agent (Go; v0.29.1)architecture/04-control-plane-authorization.md— signing, escrow, authzarchitecture/02-controller-module-map.md— historical v0.33 planning map; the live map iscontroller/module-map.mdproxmox-platform.md— Proxmox platform reference
Hub (operator backend) — architecture/05
architecture/05-hub-architecture.md— hub architecture (v0.11.0)
Security audits & remediation — audits/
audits/deep-sweep-2026-06-13.md— cross-repo deep audit (controller + agent) with remediation statusaudits/bughunt-reconcile-2026-06-13.md— reconciliation of the v0.30.3 BUGHUNT against current code + merged fix list
Spike & test findings — tests/
Per-slice spike/validation findings (phases 0–5, slices 7–10). See tests/.
Conventions
- Code-verified, not memory-derived. Architectural claims here are checked against the actual current source; if a claim can't be verified it is omitted and flagged, not guessed.
- Per-repo operational working files (
CLAUDE.md,CONTEXT.md,CHANGELOG.md,BUGHUNT.md,REPORT.md,TASK.md) live in their own repos — they are operational, not published docs. - Authoritative versions at last refresh: controller v0.59.0, agent v0.29.1, hub v0.11.0.