docs(v0.218.0): README names the volume leg and the compose-project attribution; CONTEXT records the blocker
gates / gates (push) Successful in 12s
gates / gates (push) Successful in 12s
The README's reconstitution sequence gains the volume replay it never had, and the AppBackup row states that DiscoverDatabases now prefers the compose project label. CONTEXT records the thing an operator most needs next: R-354's fix cannot reach the 40 apps that need it most until R-356 is closed, because the off-site restore still refuses outright for every app that declares no data drive. The live confirmation was therefore done on calibre-web and paperless-ngx.
This commit is contained in:
+366
-339
@@ -7,7 +7,372 @@
|
||||
>
|
||||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||||
|
||||
Last updated: 2026-08-14 (v0.215.0 — R-328..R-333: the disk alert that never sent)
|
||||
Last updated: 2026-08-22 (v0.218.0 — R-354/R-355: the abandoned database, and the volumes the off-site restore never returned)
|
||||
|
||||
> **2026-08-22 — v0.218.0 (R-354/R-355). The database nobody backed up, and the restore that
|
||||
> returned most apps nothing.**
|
||||
>
|
||||
>
|
||||
> Both fixes came out of the 2026-08-21 backup-truth drill. **R-355 went first because it is the only
|
||||
> place in the product where one customer action causes permanent total loss:** paperless-ngx's database
|
||||
> was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the
|
||||
> off-site copy or the restore — and the same misattribution meant a destructive restore of that app took
|
||||
> NO undo copy, then told the customer the app has no database. Fixed by reading the compose project
|
||||
> label, which is the stack name by construction. **One app of 53 affected**, established with a sweep
|
||||
> proven able to convict by planting a second mismatch.
|
||||
>
|
||||
> **R-354:** the off-site restore had no named-volume leg at all. The archives live inside the unit, whose
|
||||
> placement is correctly skipped, and the comment beside that skip said the dump is replayed from the
|
||||
> scratch "so nothing is lost" — true of the database, false of the volumes. `restoreDockerVolumesFrom`
|
||||
> now replays them from the scratch unit; `VolumesReplayed` reaches the message.
|
||||
>
|
||||
> **WHAT THIS DOES NOT FIX, and it is the blocking item for the apps that need it most:** the off-site
|
||||
> restore still REFUSES outright for the 40 of 53 apps that declare no data drive (**R-356**), saying a
|
||||
> running app „nincs telepítve". Those are exactly the apps whose entire dataset is a named volume, so
|
||||
> R-354's fix cannot reach them until R-356 is closed. Proven again on hardware 2026-08-21. The live
|
||||
> confirmation of R-354 was therefore done on `calibre-web` and `paperless-ngx`, which declare a drive
|
||||
> and can reach the restore.
|
||||
>
|
||||
> Also still open from the drill: the empty-restore success message (R-353), where the 40 apps' data
|
||||
> lives (R-352), and the remaining rows R-357..R-366.
|
||||
>
|
||||
> ## THE RESTORE'S OWN MEMORY (v0.217.0, 2026-08-21) — R-351 / R-352 / R-353
|
||||
>
|
||||
> > **Both sides of a written fact must be checked, not just the writing side.** Every recovery unit
|
||||
> > has recorded `drive` and `namespace_root` since schema 1. **No non-test code in the repository ever
|
||||
> > read either back.** A restore into a different destination than the backup recorded therefore
|
||||
> > succeeded silently under a green message. This is the same shape as several defects closed this
|
||||
> > month, and the cheap test for it is one grep: *who reads this field?*
|
||||
> >
|
||||
> > **`IsRunning()` is still the wrong flag, in one more place than we knew.** v0.154.0 fixed the wizard
|
||||
> > and left a comment explaining why. The **seven handlers** were never moved over, so a second press
|
||||
> > genuinely started a second run and reported „…elindult". A comment explaining a trap does not fix
|
||||
> > the other call sites — grep for them.
|
||||
> >
|
||||
> > **A result nobody can see is the same defect as no result.** The banner gated its terminal state on
|
||||
> > a page-local `sawRunning`. The 8.666 s OpenGist restore finished before any poll saw it, so no
|
||||
> > screen said it had completed. Fixed with `RestoreOpStatus.LastRecent` — and the window now lives in
|
||||
> > `internal/backup` as ONE expression that both surfaces read.
|
||||
>
|
||||
> **State, and what is next.**
|
||||
>
|
||||
> - **Shipped:** placement comparison + named mismatch + `ack_placement`; not-installed refusal names
|
||||
> the recorded drive; deploy prefill from the app's own backup; `restoreOpBlocked()`; `LastRecent`;
|
||||
> off-site listing bounded-concurrent (measured 16.1 s → two waves).
|
||||
> - **R-352 partly closed.** 40 of 53 catalogue templates declare no data path, and
|
||||
> `GetDefaultStoragePath()` is read by nothing that places data — its comment `// new apps use this by
|
||||
> default` has never been true. **Only visibility shipped**; the deploy page now states where the data
|
||||
> will live. **No placement changed, nothing migrated.** Specification:
|
||||
> `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
|
||||
> - **R-353 is the next session's first item.** A restore whose unit carries no `db_dumps` and no
|
||||
> `volume_dumps` reports a bare completion. OpenGist's unit held configuration and nothing else, and
|
||||
> the restore said only that it had finished. Fix the outcome first; *then* prove the off-site
|
||||
> coverage of a named-volume app by running a dump cycle — do not close the first on the second.
|
||||
> - **Open and unproven:** whether the 40-class reaches the off-site tier at all has **not been
|
||||
> observed**. `runVolumeDumps` covers them on paper; every unit on the box read `volume_dumps: None`
|
||||
> because no nightly run had happened yet.
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> ## THE TWO RULES THE RECOVERY JOURNEY LEANS ON (v0.203.0, 2026-08-06)
|
||||
>
|
||||
> > **1. A credential the hub stages is collected by the box, not waited for.** The reconcile that
|
||||
> > collects runs on a tick for exactly as long as the box's own declaration says it needs one — and
|
||||
> > stops the instant a target exists. It is driven from `OffboxReportStatus().State`, the same statement
|
||||
> > the hub acts on, so the two can never disagree about whether a retry is wanted.
|
||||
> >
|
||||
> > **2. A mount Felhom itself made is not "something else".** Enrolment mounts a drive twice — the
|
||||
> > managed path and a raw `/mnt/<name>` on the host — and the host survives a guest rebuild while the
|
||||
> > guest's registry does not. The claimed check forgives a non-managed mount **only when corroborated**
|
||||
> > by the same device also being mounted under the managed path. **A genuinely foreign mount is still
|
||||
> > refused, and that fence has its own test.**
|
||||
>
|
||||
> **Why both are stated here rather than left in the code:** each was a dead end that kept the unaided
|
||||
> recovery journey failing, and each looked correct in isolation. R-218's declaration half shipped and
|
||||
> worked while nothing consumed what it asked for; R-220's check was right about foreign disks and wrong
|
||||
> about our own. **Neither is a bug in the thing it guards — both are about what runs, and when.**
|
||||
>
|
||||
> Two things that must not be "simplified" back:
|
||||
> - **The settle gate stays.** The retry goes through `ReconcileWhenSettled`, so the day-0 floor race is
|
||||
> unchanged. A retry that skipped it would trade one defect for another.
|
||||
> - **The R-220 exemption is corroborated, never a prefix.** Widening it to any `/mnt/*` path offers a
|
||||
> disk another system is using for formatting — the red-proof shows exactly that.
|
||||
>
|
||||
> ## THE UNLOCK PATH'S RULE (v0.202.0, 2026-08-06) — state it before changing anything there
|
||||
>
|
||||
> > **On the recovery unlock path the customer is blamed only after a real attempt REFUSED their code.
|
||||
> > Every other outcome — including one that cannot be classified — says something else.**
|
||||
>
|
||||
> This is the rule, and it outlives the bug that produced it. It was learned twice, because fixing it
|
||||
> once was not enough:
|
||||
>
|
||||
> - **v0.201.0** stopped an agent that is too OLD from being reported as a wrong code (R-216).
|
||||
> - **v0.202.0** found the same defect through a different door: an agent that is **stopped**, and a hub
|
||||
> that cannot be **reached**, still fell through to a message about the code. Measured with a
|
||||
> **correct** code at 0.0299 s and 0.0556 s, against ~1.0 s for a real unseal — the machine accused the
|
||||
> customer of something it had not tried (R-224).
|
||||
> - And the inverse: the one message that says *"check your ten words"* was unreachable on any box that
|
||||
> had re-escrowed, which is exactly the box a customer has just recovered (R-226).
|
||||
>
|
||||
> **How it is enforced.** `agentapi.ClassifyRecoveryFailure` maps the failure to one of five classes
|
||||
> **from the value, never the text**; the typing message is reachable from **one** of them
|
||||
> (`RecoveryAskedAndRefused`, i.e. HTTP 400, i.e. the bundle was fetched and `age` refused it); and the
|
||||
> zero value is `RecoveryUnknown`, which renders **neutral**. **The safe default is the load-bearing
|
||||
> part** — an unrecognised status must not fall into an accusation.
|
||||
>
|
||||
> **Two things that are deliberately NOT how it works, and must not be "fixed" into it:**
|
||||
>
|
||||
> 1. **Elapsed time is never a classifier.** It is what diagnosed this, it is logged for the operator,
|
||||
> and that is all. A duration guard would be a second thing that can be wrong.
|
||||
> 2. **The error's TEXT is never read.** A string match is a defect waiting for a rewording. When the
|
||||
> distinction was not available as a value, the **agent was changed to provide one**
|
||||
> (`escrow.ErrBundleFetch` → HTTP 502, agent v0.126.0, `MinAgent 0.126.0`) rather than parsed for.
|
||||
>
|
||||
> **The coupling degrades safely and silently:** an agent below 0.126.0 answers 400 for both causes, so
|
||||
> `FeatureRecoveryFailureClass` withholds the refusal reading and the 400 becomes neutral. The gate
|
||||
> blocks nothing; it only decides whether the customer may be told to check their typing.
|
||||
>
|
||||
> ## CAMPAIGN 11 — what changed in v0.201.0 (2026-08-05)
|
||||
>
|
||||
> **The off-site key recovery is a COUPLED feature and now declares it.** It needs agent **0.125.0**
|
||||
> (`POST /escrow/recover-offsite-password`). `FeatureOffsiteKeyRecovery` has a `featureProbes` row, a
|
||||
> `featureMinAgent` row and a `Supports` gate at the unlock entry point.
|
||||
>
|
||||
> ⚠ **That gate FAILS CLOSED — alone in that table.** The package default is fail-open, and that default
|
||||
> is what produced R-216: an agent that could not answer 404'd, the unlock was attempted anyway, and the
|
||||
> customer was told their correct recovery code was wrong. Anything but `SupportYes` now says *the
|
||||
> machine* cannot ask yet, and no attempt is made. Do not "fix" it back to the package default.
|
||||
>
|
||||
> **The box declares `needs_credential` until the TIER WORKS, not until a key exists.** The old
|
||||
> short-circuit on "a repository password is present" is deleted: installing one is the recovery
|
||||
> screen's whole job, so it made succeeding at recovery switch off the mechanism that delivers the
|
||||
> coordinates to use it. A disabled target still short-circuits at the first line (Scenario E).
|
||||
>
|
||||
> **The unlock finishes the job**: place the key → bring the tier up (`offsiteapply.Bridge.Reconcile`,
|
||||
> wired via `SetRecoveryTierUp`) → list. Without the middle step the promised listing can never render on
|
||||
> shape (a), because no key ⇒ no target ⇒ no inventory.
|
||||
>
|
||||
> **Four messages, not one.** Wrong code (the only one mentioning typing) · the machine cannot ask ·
|
||||
> the store could not be read / the connection details have not arrived · the code belongs to a RETAINED
|
||||
> earlier package. The last one is driven by the ACK's `superseded_present`/`superseded_at` (hub
|
||||
> v0.97.0) and **promises nothing** — no read path for a superseded package exists.
|
||||
>
|
||||
> **Still open from the campaign:** R-214 (console pairing banner), R-220 (drives unenrollable after a
|
||||
> rebuild), R-221 (a rebuilt box cannot run the escrow ceremony), R-223 (the Day-0 manifest still vouches
|
||||
> agent 0.120.0 — operator decision).
|
||||
>
|
||||
> ## R-203 (v0.197.0) — the namespace-root contract, and what `ok` now means
|
||||
>
|
||||
> **The contract, in one line:** `appbackup`'s path helpers (`UserdataDir`, `PrimaryBackupPath`,
|
||||
> `RecoveryUnitPath`, `AppDataDir`) take a **NAMESPACE ROOT**. Anything that came out of `HDD_PATH` or a
|
||||
> `StoragePath` is a **DRIVE path** — put it through `appbackup.NamespaceRootFor(drive, systemDataPath)`
|
||||
> first. `UserdataDir(bareDrivePath)` still compiles and is still wrong; five callers proved it.
|
||||
>
|
||||
> **The rule now has ONE expression.** `NamespaceRootFor` / `IsEnrolledDrive` in `appbackup`;
|
||||
> `backup.Manager.namespaceRoot` and `stacks.Manager.inGuest` delegate. There were two copies before and
|
||||
> **they differed** — one compared without `filepath.Clean`, the other with it.
|
||||
>
|
||||
> **Why it was invisible:** on an enrolled drive the namespace root IS the drive path. The two diverge
|
||||
> only on the system-data fallback, which `paths.go:26` names as a supported arrangement.
|
||||
>
|
||||
> **`last_status` gains `incomplete`.** A run that could not capture a directory an app declares
|
||||
> MANDATORY is not a successful run. **Not `error`** — the rest of the run worked, so `SnapshotCount`
|
||||
> and `LastSuccess` still record what WAS captured. It reaches the operator via the existing
|
||||
> `backup_run_failures` digest (a new event type is a two-repo change; the hub drops unlisted types).
|
||||
> The Hungarian customer warning is unchanged; the page renders `! Hiányos`.
|
||||
>
|
||||
> **Still open, and NOT fixed here:** `resolveAbs` resolves `RootHDD` and `RootUserdata` against the same
|
||||
> root. Both callers now pass the namespace root so the export and the backup agree with each other, but
|
||||
> whether `${HDD_PATH}` should mean the namespace root on the system drive touches every deployed app's
|
||||
> binds and needs a decision, not a patch.
|
||||
>
|
||||
> **Blast radius, measured before changing anything:** exactly one app in the fleet had
|
||||
> `HDD_PATH == system_data_path` (`calibre-web` on demo-hp, the R-201 drill fixture). Its data was
|
||||
> migrated and its sentinel re-verified byte-identical.
|
||||
>
|
||||
> ## R-203 (2026-08-04) — a MANDATORY userdata directory can be absent from the off-site snapshot while the run says `ok`
|
||||
>
|
||||
> Found live on demo-hp while staging the R-201 drill, and it **halted that drill**.
|
||||
>
|
||||
> `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment
|
||||
> **when the drive IS the system data path** — `m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath !=
|
||||
> m.systemDataPath)` (`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is
|
||||
> `<HDD_PATH>/userdata` (`stacks/classify_binds.go:14`).
|
||||
>
|
||||
> With `system_data_path: /mnt/sys_drive` and `calibre-web` deployed at `HDD_PATH=/mnt/sys_drive`:
|
||||
>
|
||||
> live bind (files land here): /mnt/sys_drive/userdata/media/books ← exists
|
||||
> capture set looked for: /mnt/sys_drive/felhom-data/userdata/media/books ← does not
|
||||
>
|
||||
> **The same compose used BOTH roots** — `${IMPORT_PATH}` resolved *with* the segment,
|
||||
> `${USERDATA_PATH}` *without*. The run logged one `[WARN] mandatory data path missing on disk, skipped
|
||||
> from offsite`, then `0 mandatory path(s)` and **`backup OK: 3 app(s), 3 snapshot(s)`**, with
|
||||
> `last_status: ok`. Nothing customer-visible or hub-visible said the directory was dropped.
|
||||
>
|
||||
> **Not established:** whether `HDD_PATH == system_data_path` is a supported deploy. It was accepted
|
||||
> (HTTP 202) one call after the NAS path was correctly refused (R-108). **Either branch is a defect** —
|
||||
> broken resolution, or a missing refusal.
|
||||
>
|
||||
> **Two things a fix must do:** make the two roots one function, and make a skipped **MANDATORY** path a
|
||||
> customer/hub-visible signal rather than a container-log WARN. `opengist`/`privatebin` declare no
|
||||
> mandatory userdata paths and are unaffected.
|
||||
>
|
||||
> ## R-200 (v0.195.0) — the offsite key recovery diagnostic
|
||||
>
|
||||
> `--recover-offsite-check` is a `docker exec` escape hatch (the `--print-reset-code` shape), NOT a page
|
||||
> or an API a browser can reach. R comes from **STDIN** — never argv, never `ps`, never shell history,
|
||||
> never a transcript. It asks the agent (>= v0.125.0) to fetch this host's sealed bundle and open it,
|
||||
> then reports whether the recovered repository password matches the on-disk one **by sha256**.
|
||||
>
|
||||
> docker exec -i felhom-controller /usr/local/bin/felhom-controller --recover-offsite-check < /path/to/code
|
||||
>
|
||||
> **IT COMPARES AND NEVER INSTALLS.** `CheckOffsiteKeyRecoverable` must stay free of any write — if a
|
||||
> future change makes it place the recovered password, it stops being a diagnostic and needs the drill's
|
||||
> supervision (that is link 9, R-200's remaining half). Pinned by
|
||||
> `TestCheckOffsiteKeyRecoverable_WritesNothing`, whose red-proof is adding the install call.
|
||||
>
|
||||
> Exit codes are load-bearing: **0** match, **2** a clean MISMATCH, **1** a step failed. A mismatch is a
|
||||
> finding about the system; a failure is a finding about the run, and they must never share a status.
|
||||
>
|
||||
> **Proven live on demo-felhom 2026-08-04** — recovered sha256 == on-disk sha256 == the hub's stored
|
||||
> hash. Nothing customer-facing ships with it: no card, no form, no preview.
|
||||
>
|
||||
> ## About Viktor (project owner)
|
||||
>
|
||||
> - Works at Deutsche Telekom (Budapest), building Felhom.eu as a side business
|
||||
> - Felhom.eu: managed home-server service for Hungarian households
|
||||
> - Technical but prefers pragmatic solutions over over-engineering
|
||||
> - Runs all infrastructure on Gitea (gitea.dooplex.hu), k3s cluster for management
|
||||
> - Customer deployments use Docker Compose (not Kubernetes) for simplicity
|
||||
>
|
||||
> ### felhom-controller (this repo)
|
||||
> - **Version:** v0.16.1
|
||||
> - **Phase 1:** ✅ COMPLETE — Stack Manager + Deploy Flow
|
||||
> - **Phase 2:** ✅ COMPLETE — Monitoring & Health (scheduler, CPU/temp, healthchecks.io pings)
|
||||
> - **Phase 3:** ✅ COMPLETE — Backups (DB dumps, restic integration, manual trigger, **dedicated backup page**)
|
||||
> - **Phase 4:** ✅ COMPLETE — Monitoring Page with Metrics Store (SQLite, Chart.js, system + container metrics)
|
||||
> - **Phase 5:** ✅ COMPLETE — Authentication, Persistence & Settings Page (settings.json, password change, session management)
|
||||
> - **Phase 6:** ✅ COMPLETE — Monitoring Warnings, Dashboard Alerts & Notification System
|
||||
> - **Phase 7:** ✅ COMPLETE — Storage Overview, Per-App Backup Toggles & Limited Restore
|
||||
> - **Phase A:** ✅ COMPLETE — Storage Paths Foundation (registry, auto-discovery, per-app HDD_PATH, deploy dropdown, health monitoring)
|
||||
> - **Phase B:** ✅ COMPLETE — Storage Management UI Polish & Health Severity Fix (flash messages, label editing, app details, FS info, deploy free space, backup context)
|
||||
> - **Phase C:** ✅ COMPLETE — Storage Init Wizard, Data Migration & Startup Fix (disk scan/format/mount wizard, rsync-based migration, startup pings)
|
||||
> - **v0.11.1 bugfix:** ✅ COMPLETE — Storage Scan: system disk detection via host fstab + blkid UUID resolution; FSType enrichment via `blkid -o export`
|
||||
> - **v0.11.2 bugfix:** ✅ COMPLETE — /host-dev mount for block device access; `HostDevicePath()` helper; all format/scan/safety ops use /host-dev
|
||||
> - **v0.11.3 bugfix:** ✅ COMPLETE — Added `fdisk` package to Dockerfile (provides `sfdisk`; not in `util-linux` on Debian bookworm)
|
||||
> - **v0.11.4 bugfix:** ✅ COMPLETE — FormatAndMount: fixed sfdisk (wipefs+force+`,,`), mount (explicit device path), mount propagation (rshared), ASCII label, smart partition skip, findmnt verification
|
||||
> - **v0.11.6:** ✅ COMPLETE — FileBrowser auto-mount sync (`syncFileBrowserMounts()`) + 3 UI fixes (badge color, progress bar, button text)
|
||||
> - **v0.11.7:** ✅ COMPLETE — Stale data cleanup + FileBrowser sync after migration + deploy page title fix
|
||||
> - **v0.11.8:** ✅ COMPLETE — Per-App Cross-Drive Backup (3-2-1 rule): rsync/restic to secondary drive, deploy page UI, backup page summary, scheduler jobs, API endpoints
|
||||
> - **v0.11.9:** ✅ COMPLETE — UI Polish Fixes: spacing, tooltip on "Módszer", status dot instead of disabled checkbox, progressive disclosure, emoji cleanup
|
||||
> - **First app deployed:** Paperless-ngx on demo-felhom.eu (2026-02-13)
|
||||
> - **Running on:** demo-felhom (N100 mini PC) at 192.168.0.162:8080, felhotest (Proxmox VM) at router.abonet.hu:33022
|
||||
> - **All Phase 1-5 features working:** deploy, start/stop/restart/update, logs, health-aware states, auth, monitoring, backups, backup detail page, system monitoring page, settings page
|
||||
>
|
||||
> ## Architecture decisions
|
||||
>
|
||||
> | Decision | Rationale |
|
||||
> |----------|-----------|
|
||||
> | Go stdlib for web (no Gin/Echo) | Minimal dependencies, single binary, easy to embed templates |
|
||||
> | Templates as go:embed HTML/CSS files | Zero runtime file dependencies (compiled into binary), but each template is a separate editable file |
|
||||
> | Docker Compose for customers (not k8s) | Simpler troubleshooting, customers don't need k8s knowledge |
|
||||
> | k3s for management infra only | Viktor's own services (gitea, monitoring, website) run on k3s |
|
||||
> | Cloudflare Tunnel for remote access | No port forwarding needed, works behind any NAT |
|
||||
> | app.yaml per stack | Separates deploy config from compose files, survives git pulls |
|
||||
> | Password fields require explicit input | Prevents accidental empty-password deployments |
|
||||
> | Health-aware state from Docker Status field | Docker's State says "running" even for unhealthy containers |
|
||||
> | Memory limits via deploy.resources.limits | Prevents runaway containers; ~50% headroom over expected usage |
|
||||
> | System info from /proc/meminfo + statfs | No external dependencies, cheap to read on each page load |
|
||||
> | mem_request vs mem_limit (K8s-inspired) | Requests = expected usage (hard block), limits = peak (overcommit OK) |
|
||||
> | 384MB reserved for system | Prevents deploying apps that would starve the OS/controller |
|
||||
> | Logo SVG embedded as Go constant | Same approach as CSS/HTML — zero external file deps |
|
||||
> | Git sync via os/exec git CLI | No Go git library needed, git is in the container image |
|
||||
> | SHA-256 for content comparison | Only copy changed files, avoid unnecessary disk writes |
|
||||
> | 30s debounce on manual sync | Prevents spamming the git server |
|
||||
> | Orphan = deployed but not in catalog | Safe lifecycle: remove from catalog → mark orphaned → user deletes via UI |
|
||||
> | FileBrowser as infra (not catalog) | Needed even after apps deleted (user browses HDD data); deployed by setup script |
|
||||
> | Protected HDD paths | Safety net: never delete top-level HDD dirs (media, storage, Dokumentumok, appdata) |
|
||||
> | Central scheduler (not ad-hoc goroutines) | Single place to register/monitor all periodic tasks, graceful shutdown, skip-if-running |
|
||||
> | CPU sampling via background goroutine | /proc/stat delta needs two readings — collector runs every 5s, GetInfo() reads cached value |
|
||||
> | Temperature from /host/sys (Docker mount) | Container can't read host /sys directly — mount /sys:/host/sys:ro, try /host/sys first |
|
||||
> | Restic password auto-generated | No manual setup needed — generated on first backup run, stored in named volume |
|
||||
> | DB discovery via docker inspect | No config needed — discovers postgres/mariadb containers by image name + env vars |
|
||||
> | Backup orchestrator with running flag | Prevents concurrent backups, supports both scheduled and manual trigger |
|
||||
> | modernc.org/sqlite (pure Go) | No CGO/gcc needed in Docker build stage — keeps `CGO_ENABLED=0` static binary |
|
||||
> | AlertManager state-based refresh | Alerts regenerated every 5min from health report — no persistent storage needed, always reflects current state |
|
||||
> | Notification relay via hub | Controller → hub → Resend → email. Hub acts as central relay: knows customer email, handles Resend API. Controller only needs hub URL + API key |
|
||||
> | In-memory notification cooldowns | Per-event-type cooldown map (default 6h). Lost on restart = acceptable (better to re-notify than miss). No persistence needed |
|
||||
> | Health status change detection | Only notify on degradation (ok→warn, ok→fail, warn→fail). Avoids spam on flapping. First run records baseline, doesn't notify |
|
||||
> | Resend HTTP API (no SMTP) | Direct POST to api.resend.com — same pattern as website contact-mailer. Simpler than SMTP setup, good deliverability |
|
||||
> | Preferences sync on save + startup | Controller pushes prefs to hub (not pull). Startup sync handles hub DB rebuild. Local save always succeeds even if sync fails |
|
||||
> | Chart.js embedded locally | Customer hardware may not have internet — CDN not reliable for offline environments |
|
||||
> | StackDataProvider interface | backup package needs stack data but can't import stacks (circular). Interface in backup, thin adapter in main.go |
|
||||
> | Password sync to hub via report | Restic password in Docker named volume on SSD. Hub sync provides redundancy for disaster recovery |
|
||||
> | App backup via HDD mounts only | Docker volumes at /var/lib/docker/volumes/ not mounted in controller. HDD data is the important user data; DB in volumes covered by nightly dump |
|
||||
> | Restore uses running mutex | Prevents concurrent backup+restore on same restic repo. Reuses existing `m.running` flag |
|
||||
> | Storage paths registry in settings.json | Multi-storage support: each app's HDD_PATH from app.yaml is authoritative. Auto-discovery on startup avoids manual config. Registry enables UI management + health monitoring per path |
|
||||
> | /mnt:/mnt:rw mount in controller | Replaces per-path HDD_PATH mount. Enables multi-storage + restore writes. All customer HDD mounts are under /mnt/ by convention |
|
||||
> | Per-app HDD_PATH resolution (app.yaml > global) | App's own env HDD_PATH is Priority 1, registered storage paths as fallback. Eliminates dependency on global controller.yaml hdd_path |
|
||||
> | Mount-point detection via syscall.Stat_t.Dev | Compares device ID of path vs parent dir — reliable check that path is on separate filesystem. Prevents data writes to SSD |
|
||||
> | Health severity: mount-point = warning | Non-mount-point is informational, not a service failure. FAIL reserved for genuinely broken things. Avoids false alarms on demo/test environments |
|
||||
> | FS info via findmnt + sysfs | `findmnt -n -o SOURCE,FSTYPE --target <path>` for filesystem type/device. `/sys/block/<dev>/device/model` for disk model. Best-effort, returns nil on failure |
|
||||
> | Query param flash messages | Stateless, no session store needed. Consistent with backup page pattern. `?storage_msg=success&storage_detail=...` |
|
||||
> | StorageLabels map on stacks page | Separate map passed to template (not modifying Stack struct). Built from deployed apps' HDD_PATH → registered path label lookup |
|
||||
> | Metrics downsampling via SQL | Bucket-based AVG in GROUP BY keeps Chart.js responsive with up to 30 days of data |
|
||||
> | 60s metrics collection interval | Good balance of resolution vs. storage — ~44K rows/month for system metrics |
|
||||
> | /etc/os-release mounted read-only | Container can't read host OS info directly — mount to /host/etc/os-release:ro |
|
||||
>
|
||||
> ## Key file locations on demo-felhom
|
||||
>
|
||||
> ```
|
||||
> /opt/docker/felhom-controller/ # Controller compose + config
|
||||
> ├── controller.yaml # Customer config (domain, auth, paths)
|
||||
> ├── docker-compose.yml # Controller's own compose
|
||||
> └── data/ # Controller persistent data (named volume)
|
||||
>
|
||||
> /opt/docker/stacks/ # All app stacks
|
||||
> ├── traefik/ # Reverse proxy (protected)
|
||||
> ├── cloudflared/ # Tunnel (protected)
|
||||
> ├── paperless-ngx/ # First deployed app ✅
|
||||
> │ ├── docker-compose.yml
|
||||
> │ ├── .felhom.yml # App metadata
|
||||
> │ └── app.yaml # Deploy config (env vars, locked fields)
|
||||
> └── whoami/ # Test stack (not deployed)
|
||||
>
|
||||
> /mnt/hdd_placeholder/storage/ # HDD storage for apps
|
||||
> └── paperless/
|
||||
> ├── consume/ # Drop files here for OCR
|
||||
> ├── media/ # Processed documents
|
||||
> └── export/ # Backup exports
|
||||
> ```
|
||||
>
|
||||
> ## Related repositories and their state
|
||||
>
|
||||
> | Repository | Status | Notes |
|
||||
> |------------|--------|-------|
|
||||
> | felhom-controller | Active | This repo. Controller code + deploy scripts |
|
||||
> | app-catalog-felhom.eu | Active | 10 app templates, all with .felhom.yml metadata + memory limits |
|
||||
> | felhom.eu | Active | Website + hub/ subfolder (felhom-hub service) + k8s manifests |
|
||||
> | homelab-manifests | Stable | k3s cluster running (dooplex.hu services) |
|
||||
> | misc-scripts | Utility | collect-repo.sh, backup helpers |
|
||||
>
|
||||
> ## Gotchas & lessons learned
|
||||
>
|
||||
> - `docker compose restart` ≠ `docker compose up -d` — restart doesn't pick up new images
|
||||
> - Go maps have random iteration order — always sort slices before displaying
|
||||
> - Docker `.State`="running" doesn't mean healthy — check `.Status` for "(health: starting)" / "(unhealthy)"
|
||||
> - Paperless-ngx needs `PAPERLESS_OCR_LANGUAGES` (plural) to install language packs, `PAPERLESS_OCR_LANGUAGE` (singular) to select
|
||||
> - In-memory Deployed flag must be set BEFORE `docker compose up -d` (not after) — compose can take 30-60s for image pulls, during which the UI would show a stale "Telepítés" button
|
||||
> - Cloudflare Tunnel handles *.demo-felhom.eu → Traefik handles Host()-based routing to containers
|
||||
> - BIOS "AC Power Recovery" must be enabled on N100 for auto-restart after power outage
|
||||
> - `docker compose up -d` returns exit 0 even when containers immediately crash-loop — need post-start status check to detect this
|
||||
> - When logging env vars for debugging, only log keys (not values) to avoid leaking secrets in log files
|
||||
> - Mealie image (`ghcr.io/mealie-recipes/mealie`) doesn't include wget/curl — use Python TCP socket check for healthcheck
|
||||
> - Mealie DB migrations on first start take ~40s (alembic) — use `start_period: 60s` to avoid premature unhealthy status
|
||||
> - Alpine-based images (filebrowser, vaultwarden) have wget via BusyBox — healthchecks with `wget --spider` work fine
|
||||
> - Deploy `sed` command to update image version must target only the `image:` line — naive `sed 's|name:OLD|name:NEW|'` also matches the service name line (e.g., `felhom-controller:` → `felhom-controller:0.2.12`), breaking YAML. Use `sudo sed -i 's|image:.*felhom-controller:[^ ]*|image: ...felhom-controller:NEW|'` or similar scoped pattern
|
||||
> - Hungarian quotation marks `„"` in YAML: `„` (U+201E) is safe inside YAML double-quoted strings, but the closing `"` must NOT be ASCII `"` (0x22) — it terminates the YAML string. Use `\"` escape or Unicode `"` (U+201D). This caused a silent parse failure for the entire `.felhom.yml` file
|
||||
> - Never silently swallow parse errors — always log them. Silent failures make debugging impossible (took a dedicated debug session to find a simple quoting issue)
|
||||
|
||||
> **2026-08-14 — v0.215.0 (R-328..R-333). Disk health, phase 1: the alert that reached nobody.**
|
||||
>
|
||||
@@ -1949,341 +2314,3 @@ Last updated: 2026-06-13 (v0.60.0 backlog-Medium cleanup)
|
||||
> archive `local:backup/vzdump-lxc-9100-2026_06_12-09_53_03.tar.zst`, build guest 9100 purged.
|
||||
|
||||
---
|
||||
|
||||
## THE RESTORE'S OWN MEMORY (v0.217.0, 2026-08-21) — R-351 / R-352 / R-353
|
||||
|
||||
> **Both sides of a written fact must be checked, not just the writing side.** Every recovery unit
|
||||
> has recorded `drive` and `namespace_root` since schema 1. **No non-test code in the repository ever
|
||||
> read either back.** A restore into a different destination than the backup recorded therefore
|
||||
> succeeded silently under a green message. This is the same shape as several defects closed this
|
||||
> month, and the cheap test for it is one grep: *who reads this field?*
|
||||
>
|
||||
> **`IsRunning()` is still the wrong flag, in one more place than we knew.** v0.154.0 fixed the wizard
|
||||
> and left a comment explaining why. The **seven handlers** were never moved over, so a second press
|
||||
> genuinely started a second run and reported „…elindult". A comment explaining a trap does not fix
|
||||
> the other call sites — grep for them.
|
||||
>
|
||||
> **A result nobody can see is the same defect as no result.** The banner gated its terminal state on
|
||||
> a page-local `sawRunning`. The 8.666 s OpenGist restore finished before any poll saw it, so no
|
||||
> screen said it had completed. Fixed with `RestoreOpStatus.LastRecent` — and the window now lives in
|
||||
> `internal/backup` as ONE expression that both surfaces read.
|
||||
|
||||
**State, and what is next.**
|
||||
|
||||
- **Shipped:** placement comparison + named mismatch + `ack_placement`; not-installed refusal names
|
||||
the recorded drive; deploy prefill from the app's own backup; `restoreOpBlocked()`; `LastRecent`;
|
||||
off-site listing bounded-concurrent (measured 16.1 s → two waves).
|
||||
- **R-352 partly closed.** 40 of 53 catalogue templates declare no data path, and
|
||||
`GetDefaultStoragePath()` is read by nothing that places data — its comment `// new apps use this by
|
||||
default` has never been true. **Only visibility shipped**; the deploy page now states where the data
|
||||
will live. **No placement changed, nothing migrated.** Specification:
|
||||
`felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
|
||||
- **R-353 is the next session's first item.** A restore whose unit carries no `db_dumps` and no
|
||||
`volume_dumps` reports a bare completion. OpenGist's unit held configuration and nothing else, and
|
||||
the restore said only that it had finished. Fix the outcome first; *then* prove the off-site
|
||||
coverage of a named-volume app by running a dump cycle — do not close the first on the second.
|
||||
- **Open and unproven:** whether the 40-class reaches the off-site tier at all has **not been
|
||||
observed**. `runVolumeDumps` covers them on paper; every unit on the box read `volume_dumps: None`
|
||||
because no nightly run had happened yet.
|
||||
|
||||
---
|
||||
|
||||
## THE TWO RULES THE RECOVERY JOURNEY LEANS ON (v0.203.0, 2026-08-06)
|
||||
|
||||
> **1. A credential the hub stages is collected by the box, not waited for.** The reconcile that
|
||||
> collects runs on a tick for exactly as long as the box's own declaration says it needs one — and
|
||||
> stops the instant a target exists. It is driven from `OffboxReportStatus().State`, the same statement
|
||||
> the hub acts on, so the two can never disagree about whether a retry is wanted.
|
||||
>
|
||||
> **2. A mount Felhom itself made is not "something else".** Enrolment mounts a drive twice — the
|
||||
> managed path and a raw `/mnt/<name>` on the host — and the host survives a guest rebuild while the
|
||||
> guest's registry does not. The claimed check forgives a non-managed mount **only when corroborated**
|
||||
> by the same device also being mounted under the managed path. **A genuinely foreign mount is still
|
||||
> refused, and that fence has its own test.**
|
||||
|
||||
**Why both are stated here rather than left in the code:** each was a dead end that kept the unaided
|
||||
recovery journey failing, and each looked correct in isolation. R-218's declaration half shipped and
|
||||
worked while nothing consumed what it asked for; R-220's check was right about foreign disks and wrong
|
||||
about our own. **Neither is a bug in the thing it guards — both are about what runs, and when.**
|
||||
|
||||
Two things that must not be "simplified" back:
|
||||
- **The settle gate stays.** The retry goes through `ReconcileWhenSettled`, so the day-0 floor race is
|
||||
unchanged. A retry that skipped it would trade one defect for another.
|
||||
- **The R-220 exemption is corroborated, never a prefix.** Widening it to any `/mnt/*` path offers a
|
||||
disk another system is using for formatting — the red-proof shows exactly that.
|
||||
|
||||
## THE UNLOCK PATH'S RULE (v0.202.0, 2026-08-06) — state it before changing anything there
|
||||
|
||||
> **On the recovery unlock path the customer is blamed only after a real attempt REFUSED their code.
|
||||
> Every other outcome — including one that cannot be classified — says something else.**
|
||||
|
||||
This is the rule, and it outlives the bug that produced it. It was learned twice, because fixing it
|
||||
once was not enough:
|
||||
|
||||
- **v0.201.0** stopped an agent that is too OLD from being reported as a wrong code (R-216).
|
||||
- **v0.202.0** found the same defect through a different door: an agent that is **stopped**, and a hub
|
||||
that cannot be **reached**, still fell through to a message about the code. Measured with a
|
||||
**correct** code at 0.0299 s and 0.0556 s, against ~1.0 s for a real unseal — the machine accused the
|
||||
customer of something it had not tried (R-224).
|
||||
- And the inverse: the one message that says *"check your ten words"* was unreachable on any box that
|
||||
had re-escrowed, which is exactly the box a customer has just recovered (R-226).
|
||||
|
||||
**How it is enforced.** `agentapi.ClassifyRecoveryFailure` maps the failure to one of five classes
|
||||
**from the value, never the text**; the typing message is reachable from **one** of them
|
||||
(`RecoveryAskedAndRefused`, i.e. HTTP 400, i.e. the bundle was fetched and `age` refused it); and the
|
||||
zero value is `RecoveryUnknown`, which renders **neutral**. **The safe default is the load-bearing
|
||||
part** — an unrecognised status must not fall into an accusation.
|
||||
|
||||
**Two things that are deliberately NOT how it works, and must not be "fixed" into it:**
|
||||
|
||||
1. **Elapsed time is never a classifier.** It is what diagnosed this, it is logged for the operator,
|
||||
and that is all. A duration guard would be a second thing that can be wrong.
|
||||
2. **The error's TEXT is never read.** A string match is a defect waiting for a rewording. When the
|
||||
distinction was not available as a value, the **agent was changed to provide one**
|
||||
(`escrow.ErrBundleFetch` → HTTP 502, agent v0.126.0, `MinAgent 0.126.0`) rather than parsed for.
|
||||
|
||||
**The coupling degrades safely and silently:** an agent below 0.126.0 answers 400 for both causes, so
|
||||
`FeatureRecoveryFailureClass` withholds the refusal reading and the 400 becomes neutral. The gate
|
||||
blocks nothing; it only decides whether the customer may be told to check their typing.
|
||||
|
||||
## CAMPAIGN 11 — what changed in v0.201.0 (2026-08-05)
|
||||
|
||||
**The off-site key recovery is a COUPLED feature and now declares it.** It needs agent **0.125.0**
|
||||
(`POST /escrow/recover-offsite-password`). `FeatureOffsiteKeyRecovery` has a `featureProbes` row, a
|
||||
`featureMinAgent` row and a `Supports` gate at the unlock entry point.
|
||||
|
||||
⚠ **That gate FAILS CLOSED — alone in that table.** The package default is fail-open, and that default
|
||||
is what produced R-216: an agent that could not answer 404'd, the unlock was attempted anyway, and the
|
||||
customer was told their correct recovery code was wrong. Anything but `SupportYes` now says *the
|
||||
machine* cannot ask yet, and no attempt is made. Do not "fix" it back to the package default.
|
||||
|
||||
**The box declares `needs_credential` until the TIER WORKS, not until a key exists.** The old
|
||||
short-circuit on "a repository password is present" is deleted: installing one is the recovery
|
||||
screen's whole job, so it made succeeding at recovery switch off the mechanism that delivers the
|
||||
coordinates to use it. A disabled target still short-circuits at the first line (Scenario E).
|
||||
|
||||
**The unlock finishes the job**: place the key → bring the tier up (`offsiteapply.Bridge.Reconcile`,
|
||||
wired via `SetRecoveryTierUp`) → list. Without the middle step the promised listing can never render on
|
||||
shape (a), because no key ⇒ no target ⇒ no inventory.
|
||||
|
||||
**Four messages, not one.** Wrong code (the only one mentioning typing) · the machine cannot ask ·
|
||||
the store could not be read / the connection details have not arrived · the code belongs to a RETAINED
|
||||
earlier package. The last one is driven by the ACK's `superseded_present`/`superseded_at` (hub
|
||||
v0.97.0) and **promises nothing** — no read path for a superseded package exists.
|
||||
|
||||
**Still open from the campaign:** R-214 (console pairing banner), R-220 (drives unenrollable after a
|
||||
rebuild), R-221 (a rebuilt box cannot run the escrow ceremony), R-223 (the Day-0 manifest still vouches
|
||||
agent 0.120.0 — operator decision).
|
||||
|
||||
## R-203 (v0.197.0) — the namespace-root contract, and what `ok` now means
|
||||
|
||||
**The contract, in one line:** `appbackup`'s path helpers (`UserdataDir`, `PrimaryBackupPath`,
|
||||
`RecoveryUnitPath`, `AppDataDir`) take a **NAMESPACE ROOT**. Anything that came out of `HDD_PATH` or a
|
||||
`StoragePath` is a **DRIVE path** — put it through `appbackup.NamespaceRootFor(drive, systemDataPath)`
|
||||
first. `UserdataDir(bareDrivePath)` still compiles and is still wrong; five callers proved it.
|
||||
|
||||
**The rule now has ONE expression.** `NamespaceRootFor` / `IsEnrolledDrive` in `appbackup`;
|
||||
`backup.Manager.namespaceRoot` and `stacks.Manager.inGuest` delegate. There were two copies before and
|
||||
**they differed** — one compared without `filepath.Clean`, the other with it.
|
||||
|
||||
**Why it was invisible:** on an enrolled drive the namespace root IS the drive path. The two diverge
|
||||
only on the system-data fallback, which `paths.go:26` names as a supported arrangement.
|
||||
|
||||
**`last_status` gains `incomplete`.** A run that could not capture a directory an app declares
|
||||
MANDATORY is not a successful run. **Not `error`** — the rest of the run worked, so `SnapshotCount`
|
||||
and `LastSuccess` still record what WAS captured. It reaches the operator via the existing
|
||||
`backup_run_failures` digest (a new event type is a two-repo change; the hub drops unlisted types).
|
||||
The Hungarian customer warning is unchanged; the page renders `! Hiányos`.
|
||||
|
||||
**Still open, and NOT fixed here:** `resolveAbs` resolves `RootHDD` and `RootUserdata` against the same
|
||||
root. Both callers now pass the namespace root so the export and the backup agree with each other, but
|
||||
whether `${HDD_PATH}` should mean the namespace root on the system drive touches every deployed app's
|
||||
binds and needs a decision, not a patch.
|
||||
|
||||
**Blast radius, measured before changing anything:** exactly one app in the fleet had
|
||||
`HDD_PATH == system_data_path` (`calibre-web` on demo-hp, the R-201 drill fixture). Its data was
|
||||
migrated and its sentinel re-verified byte-identical.
|
||||
|
||||
## R-203 (2026-08-04) — a MANDATORY userdata directory can be absent from the off-site snapshot while the run says `ok`
|
||||
|
||||
Found live on demo-hp while staging the R-201 drill, and it **halted that drill**.
|
||||
|
||||
`NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment
|
||||
**when the drive IS the system data path** — `m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath !=
|
||||
m.systemDataPath)` (`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is
|
||||
`<HDD_PATH>/userdata` (`stacks/classify_binds.go:14`).
|
||||
|
||||
With `system_data_path: /mnt/sys_drive` and `calibre-web` deployed at `HDD_PATH=/mnt/sys_drive`:
|
||||
|
||||
live bind (files land here): /mnt/sys_drive/userdata/media/books ← exists
|
||||
capture set looked for: /mnt/sys_drive/felhom-data/userdata/media/books ← does not
|
||||
|
||||
**The same compose used BOTH roots** — `${IMPORT_PATH}` resolved *with* the segment,
|
||||
`${USERDATA_PATH}` *without*. The run logged one `[WARN] mandatory data path missing on disk, skipped
|
||||
from offsite`, then `0 mandatory path(s)` and **`backup OK: 3 app(s), 3 snapshot(s)`**, with
|
||||
`last_status: ok`. Nothing customer-visible or hub-visible said the directory was dropped.
|
||||
|
||||
**Not established:** whether `HDD_PATH == system_data_path` is a supported deploy. It was accepted
|
||||
(HTTP 202) one call after the NAS path was correctly refused (R-108). **Either branch is a defect** —
|
||||
broken resolution, or a missing refusal.
|
||||
|
||||
**Two things a fix must do:** make the two roots one function, and make a skipped **MANDATORY** path a
|
||||
customer/hub-visible signal rather than a container-log WARN. `opengist`/`privatebin` declare no
|
||||
mandatory userdata paths and are unaffected.
|
||||
|
||||
## R-200 (v0.195.0) — the offsite key recovery diagnostic
|
||||
|
||||
`--recover-offsite-check` is a `docker exec` escape hatch (the `--print-reset-code` shape), NOT a page
|
||||
or an API a browser can reach. R comes from **STDIN** — never argv, never `ps`, never shell history,
|
||||
never a transcript. It asks the agent (>= v0.125.0) to fetch this host's sealed bundle and open it,
|
||||
then reports whether the recovered repository password matches the on-disk one **by sha256**.
|
||||
|
||||
docker exec -i felhom-controller /usr/local/bin/felhom-controller --recover-offsite-check < /path/to/code
|
||||
|
||||
**IT COMPARES AND NEVER INSTALLS.** `CheckOffsiteKeyRecoverable` must stay free of any write — if a
|
||||
future change makes it place the recovered password, it stops being a diagnostic and needs the drill's
|
||||
supervision (that is link 9, R-200's remaining half). Pinned by
|
||||
`TestCheckOffsiteKeyRecoverable_WritesNothing`, whose red-proof is adding the install call.
|
||||
|
||||
Exit codes are load-bearing: **0** match, **2** a clean MISMATCH, **1** a step failed. A mismatch is a
|
||||
finding about the system; a failure is a finding about the run, and they must never share a status.
|
||||
|
||||
**Proven live on demo-felhom 2026-08-04** — recovered sha256 == on-disk sha256 == the hub's stored
|
||||
hash. Nothing customer-facing ships with it: no card, no form, no preview.
|
||||
|
||||
## About Viktor (project owner)
|
||||
|
||||
- Works at Deutsche Telekom (Budapest), building Felhom.eu as a side business
|
||||
- Felhom.eu: managed home-server service for Hungarian households
|
||||
- Technical but prefers pragmatic solutions over over-engineering
|
||||
- Runs all infrastructure on Gitea (gitea.dooplex.hu), k3s cluster for management
|
||||
- Customer deployments use Docker Compose (not Kubernetes) for simplicity
|
||||
|
||||
### felhom-controller (this repo)
|
||||
- **Version:** v0.16.1
|
||||
- **Phase 1:** ✅ COMPLETE — Stack Manager + Deploy Flow
|
||||
- **Phase 2:** ✅ COMPLETE — Monitoring & Health (scheduler, CPU/temp, healthchecks.io pings)
|
||||
- **Phase 3:** ✅ COMPLETE — Backups (DB dumps, restic integration, manual trigger, **dedicated backup page**)
|
||||
- **Phase 4:** ✅ COMPLETE — Monitoring Page with Metrics Store (SQLite, Chart.js, system + container metrics)
|
||||
- **Phase 5:** ✅ COMPLETE — Authentication, Persistence & Settings Page (settings.json, password change, session management)
|
||||
- **Phase 6:** ✅ COMPLETE — Monitoring Warnings, Dashboard Alerts & Notification System
|
||||
- **Phase 7:** ✅ COMPLETE — Storage Overview, Per-App Backup Toggles & Limited Restore
|
||||
- **Phase A:** ✅ COMPLETE — Storage Paths Foundation (registry, auto-discovery, per-app HDD_PATH, deploy dropdown, health monitoring)
|
||||
- **Phase B:** ✅ COMPLETE — Storage Management UI Polish & Health Severity Fix (flash messages, label editing, app details, FS info, deploy free space, backup context)
|
||||
- **Phase C:** ✅ COMPLETE — Storage Init Wizard, Data Migration & Startup Fix (disk scan/format/mount wizard, rsync-based migration, startup pings)
|
||||
- **v0.11.1 bugfix:** ✅ COMPLETE — Storage Scan: system disk detection via host fstab + blkid UUID resolution; FSType enrichment via `blkid -o export`
|
||||
- **v0.11.2 bugfix:** ✅ COMPLETE — /host-dev mount for block device access; `HostDevicePath()` helper; all format/scan/safety ops use /host-dev
|
||||
- **v0.11.3 bugfix:** ✅ COMPLETE — Added `fdisk` package to Dockerfile (provides `sfdisk`; not in `util-linux` on Debian bookworm)
|
||||
- **v0.11.4 bugfix:** ✅ COMPLETE — FormatAndMount: fixed sfdisk (wipefs+force+`,,`), mount (explicit device path), mount propagation (rshared), ASCII label, smart partition skip, findmnt verification
|
||||
- **v0.11.6:** ✅ COMPLETE — FileBrowser auto-mount sync (`syncFileBrowserMounts()`) + 3 UI fixes (badge color, progress bar, button text)
|
||||
- **v0.11.7:** ✅ COMPLETE — Stale data cleanup + FileBrowser sync after migration + deploy page title fix
|
||||
- **v0.11.8:** ✅ COMPLETE — Per-App Cross-Drive Backup (3-2-1 rule): rsync/restic to secondary drive, deploy page UI, backup page summary, scheduler jobs, API endpoints
|
||||
- **v0.11.9:** ✅ COMPLETE — UI Polish Fixes: spacing, tooltip on "Módszer", status dot instead of disabled checkbox, progressive disclosure, emoji cleanup
|
||||
- **First app deployed:** Paperless-ngx on demo-felhom.eu (2026-02-13)
|
||||
- **Running on:** demo-felhom (N100 mini PC) at 192.168.0.162:8080, felhotest (Proxmox VM) at router.abonet.hu:33022
|
||||
- **All Phase 1-5 features working:** deploy, start/stop/restart/update, logs, health-aware states, auth, monitoring, backups, backup detail page, system monitoring page, settings page
|
||||
|
||||
## Architecture decisions
|
||||
|
||||
| Decision | Rationale |
|
||||
|----------|-----------|
|
||||
| Go stdlib for web (no Gin/Echo) | Minimal dependencies, single binary, easy to embed templates |
|
||||
| Templates as go:embed HTML/CSS files | Zero runtime file dependencies (compiled into binary), but each template is a separate editable file |
|
||||
| Docker Compose for customers (not k8s) | Simpler troubleshooting, customers don't need k8s knowledge |
|
||||
| k3s for management infra only | Viktor's own services (gitea, monitoring, website) run on k3s |
|
||||
| Cloudflare Tunnel for remote access | No port forwarding needed, works behind any NAT |
|
||||
| app.yaml per stack | Separates deploy config from compose files, survives git pulls |
|
||||
| Password fields require explicit input | Prevents accidental empty-password deployments |
|
||||
| Health-aware state from Docker Status field | Docker's State says "running" even for unhealthy containers |
|
||||
| Memory limits via deploy.resources.limits | Prevents runaway containers; ~50% headroom over expected usage |
|
||||
| System info from /proc/meminfo + statfs | No external dependencies, cheap to read on each page load |
|
||||
| mem_request vs mem_limit (K8s-inspired) | Requests = expected usage (hard block), limits = peak (overcommit OK) |
|
||||
| 384MB reserved for system | Prevents deploying apps that would starve the OS/controller |
|
||||
| Logo SVG embedded as Go constant | Same approach as CSS/HTML — zero external file deps |
|
||||
| Git sync via os/exec git CLI | No Go git library needed, git is in the container image |
|
||||
| SHA-256 for content comparison | Only copy changed files, avoid unnecessary disk writes |
|
||||
| 30s debounce on manual sync | Prevents spamming the git server |
|
||||
| Orphan = deployed but not in catalog | Safe lifecycle: remove from catalog → mark orphaned → user deletes via UI |
|
||||
| FileBrowser as infra (not catalog) | Needed even after apps deleted (user browses HDD data); deployed by setup script |
|
||||
| Protected HDD paths | Safety net: never delete top-level HDD dirs (media, storage, Dokumentumok, appdata) |
|
||||
| Central scheduler (not ad-hoc goroutines) | Single place to register/monitor all periodic tasks, graceful shutdown, skip-if-running |
|
||||
| CPU sampling via background goroutine | /proc/stat delta needs two readings — collector runs every 5s, GetInfo() reads cached value |
|
||||
| Temperature from /host/sys (Docker mount) | Container can't read host /sys directly — mount /sys:/host/sys:ro, try /host/sys first |
|
||||
| Restic password auto-generated | No manual setup needed — generated on first backup run, stored in named volume |
|
||||
| DB discovery via docker inspect | No config needed — discovers postgres/mariadb containers by image name + env vars |
|
||||
| Backup orchestrator with running flag | Prevents concurrent backups, supports both scheduled and manual trigger |
|
||||
| modernc.org/sqlite (pure Go) | No CGO/gcc needed in Docker build stage — keeps `CGO_ENABLED=0` static binary |
|
||||
| AlertManager state-based refresh | Alerts regenerated every 5min from health report — no persistent storage needed, always reflects current state |
|
||||
| Notification relay via hub | Controller → hub → Resend → email. Hub acts as central relay: knows customer email, handles Resend API. Controller only needs hub URL + API key |
|
||||
| In-memory notification cooldowns | Per-event-type cooldown map (default 6h). Lost on restart = acceptable (better to re-notify than miss). No persistence needed |
|
||||
| Health status change detection | Only notify on degradation (ok→warn, ok→fail, warn→fail). Avoids spam on flapping. First run records baseline, doesn't notify |
|
||||
| Resend HTTP API (no SMTP) | Direct POST to api.resend.com — same pattern as website contact-mailer. Simpler than SMTP setup, good deliverability |
|
||||
| Preferences sync on save + startup | Controller pushes prefs to hub (not pull). Startup sync handles hub DB rebuild. Local save always succeeds even if sync fails |
|
||||
| Chart.js embedded locally | Customer hardware may not have internet — CDN not reliable for offline environments |
|
||||
| StackDataProvider interface | backup package needs stack data but can't import stacks (circular). Interface in backup, thin adapter in main.go |
|
||||
| Password sync to hub via report | Restic password in Docker named volume on SSD. Hub sync provides redundancy for disaster recovery |
|
||||
| App backup via HDD mounts only | Docker volumes at /var/lib/docker/volumes/ not mounted in controller. HDD data is the important user data; DB in volumes covered by nightly dump |
|
||||
| Restore uses running mutex | Prevents concurrent backup+restore on same restic repo. Reuses existing `m.running` flag |
|
||||
| Storage paths registry in settings.json | Multi-storage support: each app's HDD_PATH from app.yaml is authoritative. Auto-discovery on startup avoids manual config. Registry enables UI management + health monitoring per path |
|
||||
| /mnt:/mnt:rw mount in controller | Replaces per-path HDD_PATH mount. Enables multi-storage + restore writes. All customer HDD mounts are under /mnt/ by convention |
|
||||
| Per-app HDD_PATH resolution (app.yaml > global) | App's own env HDD_PATH is Priority 1, registered storage paths as fallback. Eliminates dependency on global controller.yaml hdd_path |
|
||||
| Mount-point detection via syscall.Stat_t.Dev | Compares device ID of path vs parent dir — reliable check that path is on separate filesystem. Prevents data writes to SSD |
|
||||
| Health severity: mount-point = warning | Non-mount-point is informational, not a service failure. FAIL reserved for genuinely broken things. Avoids false alarms on demo/test environments |
|
||||
| FS info via findmnt + sysfs | `findmnt -n -o SOURCE,FSTYPE --target <path>` for filesystem type/device. `/sys/block/<dev>/device/model` for disk model. Best-effort, returns nil on failure |
|
||||
| Query param flash messages | Stateless, no session store needed. Consistent with backup page pattern. `?storage_msg=success&storage_detail=...` |
|
||||
| StorageLabels map on stacks page | Separate map passed to template (not modifying Stack struct). Built from deployed apps' HDD_PATH → registered path label lookup |
|
||||
| Metrics downsampling via SQL | Bucket-based AVG in GROUP BY keeps Chart.js responsive with up to 30 days of data |
|
||||
| 60s metrics collection interval | Good balance of resolution vs. storage — ~44K rows/month for system metrics |
|
||||
| /etc/os-release mounted read-only | Container can't read host OS info directly — mount to /host/etc/os-release:ro |
|
||||
|
||||
## Key file locations on demo-felhom
|
||||
|
||||
```
|
||||
/opt/docker/felhom-controller/ # Controller compose + config
|
||||
├── controller.yaml # Customer config (domain, auth, paths)
|
||||
├── docker-compose.yml # Controller's own compose
|
||||
└── data/ # Controller persistent data (named volume)
|
||||
|
||||
/opt/docker/stacks/ # All app stacks
|
||||
├── traefik/ # Reverse proxy (protected)
|
||||
├── cloudflared/ # Tunnel (protected)
|
||||
├── paperless-ngx/ # First deployed app ✅
|
||||
│ ├── docker-compose.yml
|
||||
│ ├── .felhom.yml # App metadata
|
||||
│ └── app.yaml # Deploy config (env vars, locked fields)
|
||||
└── whoami/ # Test stack (not deployed)
|
||||
|
||||
/mnt/hdd_placeholder/storage/ # HDD storage for apps
|
||||
└── paperless/
|
||||
├── consume/ # Drop files here for OCR
|
||||
├── media/ # Processed documents
|
||||
└── export/ # Backup exports
|
||||
```
|
||||
|
||||
## Related repositories and their state
|
||||
|
||||
| Repository | Status | Notes |
|
||||
|------------|--------|-------|
|
||||
| felhom-controller | Active | This repo. Controller code + deploy scripts |
|
||||
| app-catalog-felhom.eu | Active | 10 app templates, all with .felhom.yml metadata + memory limits |
|
||||
| felhom.eu | Active | Website + hub/ subfolder (felhom-hub service) + k8s manifests |
|
||||
| homelab-manifests | Stable | k3s cluster running (dooplex.hu services) |
|
||||
| misc-scripts | Utility | collect-repo.sh, backup helpers |
|
||||
|
||||
## Gotchas & lessons learned
|
||||
|
||||
- `docker compose restart` ≠ `docker compose up -d` — restart doesn't pick up new images
|
||||
- Go maps have random iteration order — always sort slices before displaying
|
||||
- Docker `.State`="running" doesn't mean healthy — check `.Status` for "(health: starting)" / "(unhealthy)"
|
||||
- Paperless-ngx needs `PAPERLESS_OCR_LANGUAGES` (plural) to install language packs, `PAPERLESS_OCR_LANGUAGE` (singular) to select
|
||||
- In-memory Deployed flag must be set BEFORE `docker compose up -d` (not after) — compose can take 30-60s for image pulls, during which the UI would show a stale "Telepítés" button
|
||||
- Cloudflare Tunnel handles *.demo-felhom.eu → Traefik handles Host()-based routing to containers
|
||||
- BIOS "AC Power Recovery" must be enabled on N100 for auto-restart after power outage
|
||||
- `docker compose up -d` returns exit 0 even when containers immediately crash-loop — need post-start status check to detect this
|
||||
- When logging env vars for debugging, only log keys (not values) to avoid leaking secrets in log files
|
||||
- Mealie image (`ghcr.io/mealie-recipes/mealie`) doesn't include wget/curl — use Python TCP socket check for healthcheck
|
||||
- Mealie DB migrations on first start take ~40s (alembic) — use `start_period: 60s` to avoid premature unhealthy status
|
||||
- Alpine-based images (filebrowser, vaultwarden) have wget via BusyBox — healthchecks with `wget --spider` work fine
|
||||
- Deploy `sed` command to update image version must target only the `image:` line — naive `sed 's|name:OLD|name:NEW|'` also matches the service name line (e.g., `felhom-controller:` → `felhom-controller:0.2.12`), breaking YAML. Use `sudo sed -i 's|image:.*felhom-controller:[^ ]*|image: ...felhom-controller:NEW|'` or similar scoped pattern
|
||||
- Hungarian quotation marks `„"` in YAML: `„` (U+201E) is safe inside YAML double-quoted strings, but the closing `"` must NOT be ASCII `"` (0x22) — it terminates the YAML string. Use `\"` escape or Unicode `"` (U+201D). This caused a silent parse failure for the entire `.felhom.yml` file
|
||||
- Never silently swallow parse errors — always log them. Silent failures make debugging impossible (took a dedicated debug session to find a simple quoting issue)
|
||||
+11
-3
@@ -239,7 +239,7 @@ backups, monitoring and notifications. All Proxmox/disk operations are delegated
|
||||
| **Infra** | `internal/infra/` | Pure renderers (embedded `text/template`) for the base-infra stacks (traefik/cloudflared/filebrowser); **pinned image tags as the single source of truth** (web filebrowser sync delegates here) |
|
||||
| **Crypto** | `internal/crypto/` | AES-256-GCM encryption for sensitive app.yaml values (passwords, secrets), key management |
|
||||
| **Sync** | `internal/sync/` | Git-based app catalog sync (clone/pull, content-hash copy) |
|
||||
| **AppBackup** | `internal/appbackup/` | Self-contained app-data backup primitives: DB dump discovery/execution (`DiscoverDatabases`, `DumpOne`), Docker-volume/app-data discovery (`StackDataProvider`, `DiscoverAppData`), keep-side path helpers (`AppDBDumpPath`, `AppVolumeDumpPath`, `AppDataDir`). `DiscoverDatabases` takes the deployed-stack set so a DB container maps to the right stack even when a slug ends in a DB-role token (M19, v0.62.0). `ListDumpFiles` takes an optional `cached(name,size,mod)` lookup so an unchanged dump isn't re-validated (line-scan) every ~5-min cycle (M18, v0.62.0). No dependency on restic/cross-drive/drive-mount. Imported directly by `appexport` and `storage`. |
|
||||
| **AppBackup** | `internal/appbackup/` | Self-contained app-data backup primitives: DB dump discovery/execution (`DiscoverDatabases`, `DumpOne`), Docker-volume/app-data discovery (`StackDataProvider`, `DiscoverAppData`), keep-side path helpers (`AppDBDumpPath`, `AppVolumeDumpPath`, `AppDataDir`). `DiscoverDatabases` reads each container's `com.docker.compose.project` label and prefers it as the stack name (**v0.218.0, R-355**) — the label is the stack name by construction, since compose runs with `cmd.Dir` set to the stack directory and no `-p`; the deployed-stack set (M19, v0.62.0) remains the fallback for containers not started by compose, and an attribution that resolves to no known stack now WARNs instead of being returned silently. Before this, `paperless-ngx` (container `paperless-postgres`) was attributed to a non-existent stack `paperless`, so its 72-table PostgreSQL dump landed outside its recovery unit, never reached the off-site copy, and was never restored — and the same value reaching `writeSafetyDump` meant a destructive restore of that app took no undo copy at all. `ListDumpFiles` takes an optional `cached(name,size,mod)` lookup so an unchanged dump isn't re-validated (line-scan) every ~5-min cycle (M18, v0.62.0). No dependency on restic/cross-drive/drive-mount. Imported directly by `appexport` and `storage`. |
|
||||
| **Backup** | `internal/backup/` | Per-drive 3-layer backup: DB dumps → restic snapshots → cross-drive copies, restore. Re-exposes the `appbackup` primitives via aliases/forwarders (`appbackup_bridge.go`) for the disk/host-side code and the web/api/report consumers. |
|
||||
| **Storage** | `internal/storage/` | Disk scanning (`lsblk`), partitioning (`sfdisk`), formatting (`mkfs.ext4`), mounting, data migration (`rsync`) |
|
||||
| **System** | `internal/system/` | System info (`/proc`), CPU collector, mount points, disk usage, FS info |
|
||||
@@ -554,9 +554,17 @@ Each app can define rich metadata in `.felhom.yml`:
|
||||
- **Offsite reconstitution (v0.148.0, R-43 — `offbox_reconstitute.go`):** the leg that was missing.
|
||||
`ReconstituteFromOffsite` (`/backup/offbox/reconstitute`, „Teljes visszaállítás (fájlok +
|
||||
adatbázis)") makes the live app equal to the chosen snapshot: **safety dump → stop → files
|
||||
overwritten (`rsyncRestoreOverwrite`: no `--ignore-existing`, no `--delete`) → the DATABASE
|
||||
SERVICE ONLY started (`StartStackServices`, v0.153.0) → the snapshot's dump replayed
|
||||
overwritten (`rsyncRestoreOverwrite`: no `--ignore-existing`, no `--delete`) → **the snapshot's
|
||||
NAMED VOLUMES replayed (`restoreDockerVolumesFrom`, reading the SCRATCH unit — v0.218.0, R-354)**
|
||||
→ the DATABASE SERVICE ONLY started (`StartStackServices`, v0.153.0) → the snapshot's dump replayed
|
||||
(`reimportDBDumpsFrom`, reading the SCRATCH unit) → the full stack started → health wait**.
|
||||
**The volume leg did not exist before v0.218.0** — the archives live inside the recovery unit, whose
|
||||
placement is (correctly) skipped, so the off-site restore returned files and a database and silently
|
||||
nothing else. For the 40 of 53 catalogue apps that declare no data drive, that archive is the entire
|
||||
dataset. Volumes replay BEFORE the database, so a logical dump still wins over a volume-tar copy of
|
||||
the same database, and inside the stopped window because Docker will not replace a volume in use.
|
||||
`restoreDockerVolumesFrom` is the local restore path's own replay with an explicit directory — one
|
||||
implementation, two callers.
|
||||
Two invariants: nothing is ever deleted (post-snapshot files survive as extras), and the
|
||||
`pre-restore-` safety dump is verified on disk BEFORE anything is stopped or overwritten — if it
|
||||
cannot be taken the operation refuses with zero changes. Safety dumps appear in `ListDumpFiles`
|
||||
|
||||
Reference in New Issue
Block a user