# CONTEXT.md — Project Memory > This file serves as persistent project memory across Claude Code sessions. > It replaces the auto-generated "Memory" from the claude.ai Project. > **Update this file at the end of each working session** with current state, > recent decisions, and anything the next session needs to know. > > Ask Claude Code: "Please update CONTEXT.md with what we did today" Last updated: 2026-08-22 (v0.219.0 — R-356: the off-site restore refused every app that has no data drive) > **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.** > > **[DESIGN] The restore destination is resolved by the SAME rule as the capture destination.** The > drive if the app declares one, the system data path otherwise — `Manager.GetAppDrivePath`, one > expression, now used by `CaptureRecoveryUnit`, `ReconstituteFromOffsite` and `PlaceOffsiteRestore` > alike. Anything else and the restore aims somewhere the backup never came from, which surfaces as a > placement-mismatch prompt on a box where nothing actually moved. > > **The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351) > applies to apps that HAVE a drive to get wrong.** It used to be reached by `HDD_PATH == ""`, which > also stood in for "is this app installed?". Measured in the catalogue at `459766cb1639`: **53 > templates, 13 `needs_hdd: true`, 40 `false`** — so for 40 apps that test was permanently true and the > off-site restore refused them forever, while they were running, telling the customer to reinstall > them "in the same place", which those apps never offer. **An app with no drive is not misconfigured** > (`01-topology-and-trust.md` §8, `[DESIGN]`); it is the majority case, and between 19 and 22 August it > was called a defect four times. > > **Now:** *installed?* is asked of `ListDeployedStacks()` via `Manager.isStackDeployed`, which **fails > CLOSED on a nil provider** — "cannot tell" must not become "go ahead" when the next act is a write. > *Where?* is asked of `GetAppDrivePath`. A third refusal, with its own sentence and its own route, > covers installed-but-no-resolvable-data-root: widening `nincs telepítve` to cover that would send a > customer to reinstall a running app and hide the real fault. > > **FENCED, and not changed:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`). > Capture resolves an app's declared `userdata`/`import` file legs against that value; a system-data > fallback there would write a snapshot claiming to hold the customer's files and not holding them. The > fenced ACT is "introduce a fallback into capture-side path resolution" — reading the value elsewhere > is fine. > > **Proven live on `demo-hp`, 2026-08-22.** `privatebin` (driveless): planted through the app's own > HTTP API plus a direct file plant, off-sited, **deleted**, restored through > `POST /backup/offbox/reconstitute` — **15/15 files back byte for byte**, two Hungarian accented names > included, message „0 fájl és 1 adatkötet visszaállítva". The scratch was unit-only, exactly the shape > nobody had ever driven to completion before, and every downstream leg held. `calibre-web` (drive > app) walked the same way and did not move. Evidence: > `felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. > **2026-08-22 — v0.218.0 (R-354/R-355). The database nobody backed up, and the restore that > returned most apps nothing.** > > > Both fixes came out of the 2026-08-21 backup-truth drill. **R-355 went first because it is the only > place in the product where one customer action causes permanent total loss:** paperless-ngx's database > was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the > off-site copy or the restore — and the same misattribution meant a destructive restore of that app took > NO undo copy, then told the customer the app has no database. Fixed by reading the compose project > label, which is the stack name by construction. **One app of 53 affected**, established with a sweep > proven able to convict by planting a second mismatch. > > **R-354:** the off-site restore had no named-volume leg at all. The archives live inside the unit, whose > placement is correctly skipped, and the comment beside that skip said the dump is replayed from the > scratch "so nothing is lost" — true of the database, false of the volumes. `restoreDockerVolumesFrom` > now replays them from the scratch unit; `VolumesReplayed` reaches the message. > > **WHAT THIS DOES NOT FIX, and it is the blocking item for the apps that need it most:** the off-site > restore still REFUSES outright for the 40 of 53 apps that declare no data drive (**R-356**), saying a > running app „nincs telepítve". Those are exactly the apps whose entire dataset is a named volume, so > R-354's fix cannot reach them until R-356 is closed. Proven again on hardware 2026-08-21. The live > confirmation of R-354 was therefore done on `calibre-web` and `paperless-ngx`, which declare a drive > and can reach the restore. > > Also still open from the drill: the empty-restore success message (R-353), where the 40 apps' data > lives (R-352), and the remaining rows R-357..R-366. > > ## THE RESTORE'S OWN MEMORY (v0.217.0, 2026-08-21) — R-351 / R-352 / R-353 > > > **Both sides of a written fact must be checked, not just the writing side.** Every recovery unit > > has recorded `drive` and `namespace_root` since schema 1. **No non-test code in the repository ever > > read either back.** A restore into a different destination than the backup recorded therefore > > succeeded silently under a green message. This is the same shape as several defects closed this > > month, and the cheap test for it is one grep: *who reads this field?* > > > > **`IsRunning()` is still the wrong flag, in one more place than we knew.** v0.154.0 fixed the wizard > > and left a comment explaining why. The **seven handlers** were never moved over, so a second press > > genuinely started a second run and reported „…elindult". A comment explaining a trap does not fix > > the other call sites — grep for them. > > > > **A result nobody can see is the same defect as no result.** The banner gated its terminal state on > > a page-local `sawRunning`. The 8.666 s OpenGist restore finished before any poll saw it, so no > > screen said it had completed. Fixed with `RestoreOpStatus.LastRecent` — and the window now lives in > > `internal/backup` as ONE expression that both surfaces read. > > **State, and what is next.** > > - **Shipped:** placement comparison + named mismatch + `ack_placement`; not-installed refusal names > the recorded drive; deploy prefill from the app's own backup; `restoreOpBlocked()`; `LastRecent`; > off-site listing bounded-concurrent (measured 16.1 s → two waves). > - **R-352 partly closed.** 40 of 53 catalogue templates declare no data path, and > `GetDefaultStoragePath()` is read by nothing that places data — its comment `// new apps use this by > default` has never been true. **Only visibility shipped**; the deploy page now states where the data > will live. **No placement changed, nothing migrated.** Specification: > `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`. > - **R-353 is the next session's first item.** A restore whose unit carries no `db_dumps` and no > `volume_dumps` reports a bare completion. OpenGist's unit held configuration and nothing else, and > the restore said only that it had finished. Fix the outcome first; *then* prove the off-site > coverage of a named-volume app by running a dump cycle — do not close the first on the second. > - **Open and unproven:** whether the 40-class reaches the off-site tier at all has **not been > observed**. `runVolumeDumps` covers them on paper; every unit on the box read `volume_dumps: None` > because no nightly run had happened yet. > > --- > > ## THE TWO RULES THE RECOVERY JOURNEY LEANS ON (v0.203.0, 2026-08-06) > > > **1. A credential the hub stages is collected by the box, not waited for.** The reconcile that > > collects runs on a tick for exactly as long as the box's own declaration says it needs one — and > > stops the instant a target exists. It is driven from `OffboxReportStatus().State`, the same statement > > the hub acts on, so the two can never disagree about whether a retry is wanted. > > > > **2. A mount Felhom itself made is not "something else".** Enrolment mounts a drive twice — the > > managed path and a raw `/mnt/` on the host — and the host survives a guest rebuild while the > > guest's registry does not. The claimed check forgives a non-managed mount **only when corroborated** > > by the same device also being mounted under the managed path. **A genuinely foreign mount is still > > refused, and that fence has its own test.** > > **Why both are stated here rather than left in the code:** each was a dead end that kept the unaided > recovery journey failing, and each looked correct in isolation. R-218's declaration half shipped and > worked while nothing consumed what it asked for; R-220's check was right about foreign disks and wrong > about our own. **Neither is a bug in the thing it guards — both are about what runs, and when.** > > Two things that must not be "simplified" back: > - **The settle gate stays.** The retry goes through `ReconcileWhenSettled`, so the day-0 floor race is > unchanged. A retry that skipped it would trade one defect for another. > - **The R-220 exemption is corroborated, never a prefix.** Widening it to any `/mnt/*` path offers a > disk another system is using for formatting — the red-proof shows exactly that. > > ## THE UNLOCK PATH'S RULE (v0.202.0, 2026-08-06) — state it before changing anything there > > > **On the recovery unlock path the customer is blamed only after a real attempt REFUSED their code. > > Every other outcome — including one that cannot be classified — says something else.** > > This is the rule, and it outlives the bug that produced it. It was learned twice, because fixing it > once was not enough: > > - **v0.201.0** stopped an agent that is too OLD from being reported as a wrong code (R-216). > - **v0.202.0** found the same defect through a different door: an agent that is **stopped**, and a hub > that cannot be **reached**, still fell through to a message about the code. Measured with a > **correct** code at 0.0299 s and 0.0556 s, against ~1.0 s for a real unseal — the machine accused the > customer of something it had not tried (R-224). > - And the inverse: the one message that says *"check your ten words"* was unreachable on any box that > had re-escrowed, which is exactly the box a customer has just recovered (R-226). > > **How it is enforced.** `agentapi.ClassifyRecoveryFailure` maps the failure to one of five classes > **from the value, never the text**; the typing message is reachable from **one** of them > (`RecoveryAskedAndRefused`, i.e. HTTP 400, i.e. the bundle was fetched and `age` refused it); and the > zero value is `RecoveryUnknown`, which renders **neutral**. **The safe default is the load-bearing > part** — an unrecognised status must not fall into an accusation. > > **Two things that are deliberately NOT how it works, and must not be "fixed" into it:** > > 1. **Elapsed time is never a classifier.** It is what diagnosed this, it is logged for the operator, > and that is all. A duration guard would be a second thing that can be wrong. > 2. **The error's TEXT is never read.** A string match is a defect waiting for a rewording. When the > distinction was not available as a value, the **agent was changed to provide one** > (`escrow.ErrBundleFetch` → HTTP 502, agent v0.126.0, `MinAgent 0.126.0`) rather than parsed for. > > **The coupling degrades safely and silently:** an agent below 0.126.0 answers 400 for both causes, so > `FeatureRecoveryFailureClass` withholds the refusal reading and the 400 becomes neutral. The gate > blocks nothing; it only decides whether the customer may be told to check their typing. > > ## CAMPAIGN 11 — what changed in v0.201.0 (2026-08-05) > > **The off-site key recovery is a COUPLED feature and now declares it.** It needs agent **0.125.0** > (`POST /escrow/recover-offsite-password`). `FeatureOffsiteKeyRecovery` has a `featureProbes` row, a > `featureMinAgent` row and a `Supports` gate at the unlock entry point. > > ⚠ **That gate FAILS CLOSED — alone in that table.** The package default is fail-open, and that default > is what produced R-216: an agent that could not answer 404'd, the unlock was attempted anyway, and the > customer was told their correct recovery code was wrong. Anything but `SupportYes` now says *the > machine* cannot ask yet, and no attempt is made. Do not "fix" it back to the package default. > > **The box declares `needs_credential` until the TIER WORKS, not until a key exists.** The old > short-circuit on "a repository password is present" is deleted: installing one is the recovery > screen's whole job, so it made succeeding at recovery switch off the mechanism that delivers the > coordinates to use it. A disabled target still short-circuits at the first line (Scenario E). > > **The unlock finishes the job**: place the key → bring the tier up (`offsiteapply.Bridge.Reconcile`, > wired via `SetRecoveryTierUp`) → list. Without the middle step the promised listing can never render on > shape (a), because no key ⇒ no target ⇒ no inventory. > > **Four messages, not one.** Wrong code (the only one mentioning typing) · the machine cannot ask · > the store could not be read / the connection details have not arrived · the code belongs to a RETAINED > earlier package. The last one is driven by the ACK's `superseded_present`/`superseded_at` (hub > v0.97.0) and **promises nothing** — no read path for a superseded package exists. > > **Still open from the campaign:** R-214 (console pairing banner), R-220 (drives unenrollable after a > rebuild), R-221 (a rebuilt box cannot run the escrow ceremony), R-223 (the Day-0 manifest still vouches > agent 0.120.0 — operator decision). > > ## R-203 (v0.197.0) — the namespace-root contract, and what `ok` now means > > **The contract, in one line:** `appbackup`'s path helpers (`UserdataDir`, `PrimaryBackupPath`, > `RecoveryUnitPath`, `AppDataDir`) take a **NAMESPACE ROOT**. Anything that came out of `HDD_PATH` or a > `StoragePath` is a **DRIVE path** — put it through `appbackup.NamespaceRootFor(drive, systemDataPath)` > first. `UserdataDir(bareDrivePath)` still compiles and is still wrong; five callers proved it. > > **The rule now has ONE expression.** `NamespaceRootFor` / `IsEnrolledDrive` in `appbackup`; > `backup.Manager.namespaceRoot` and `stacks.Manager.inGuest` delegate. There were two copies before and > **they differed** — one compared without `filepath.Clean`, the other with it. > > **Why it was invisible:** on an enrolled drive the namespace root IS the drive path. The two diverge > only on the system-data fallback, which `paths.go:26` names as a supported arrangement. > > **`last_status` gains `incomplete`.** A run that could not capture a directory an app declares > MANDATORY is not a successful run. **Not `error`** — the rest of the run worked, so `SnapshotCount` > and `LastSuccess` still record what WAS captured. It reaches the operator via the existing > `backup_run_failures` digest (a new event type is a two-repo change; the hub drops unlisted types). > The Hungarian customer warning is unchanged; the page renders `! Hiányos`. > > **Still open, and NOT fixed here:** `resolveAbs` resolves `RootHDD` and `RootUserdata` against the same > root. Both callers now pass the namespace root so the export and the backup agree with each other, but > whether `${HDD_PATH}` should mean the namespace root on the system drive touches every deployed app's > binds and needs a decision, not a patch. > > **Blast radius, measured before changing anything:** exactly one app in the fleet had > `HDD_PATH == system_data_path` (`calibre-web` on demo-hp, the R-201 drill fixture). Its data was > migrated and its sentinel re-verified byte-identical. > > ## R-203 (2026-08-04) — a MANDATORY userdata directory can be absent from the off-site snapshot while the run says `ok` > > Found live on demo-hp while staging the R-201 drill, and it **halted that drill**. > > `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment > **when the drive IS the system data path** — `m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath != > m.systemDataPath)` (`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is > `/userdata` (`stacks/classify_binds.go:14`). > > With `system_data_path: /mnt/sys_drive` and `calibre-web` deployed at `HDD_PATH=/mnt/sys_drive`: > > live bind (files land here): /mnt/sys_drive/userdata/media/books ← exists > capture set looked for: /mnt/sys_drive/felhom-data/userdata/media/books ← does not > > **The same compose used BOTH roots** — `${IMPORT_PATH}` resolved *with* the segment, > `${USERDATA_PATH}` *without*. The run logged one `[WARN] mandatory data path missing on disk, skipped > from offsite`, then `0 mandatory path(s)` and **`backup OK: 3 app(s), 3 snapshot(s)`**, with > `last_status: ok`. Nothing customer-visible or hub-visible said the directory was dropped. > > **Not established:** whether `HDD_PATH == system_data_path` is a supported deploy. It was accepted > (HTTP 202) one call after the NAS path was correctly refused (R-108). **Either branch is a defect** — > broken resolution, or a missing refusal. > > **Two things a fix must do:** make the two roots one function, and make a skipped **MANDATORY** path a > customer/hub-visible signal rather than a container-log WARN. `opengist`/`privatebin` declare no > mandatory userdata paths and are unaffected. > > ## R-200 (v0.195.0) — the offsite key recovery diagnostic > > `--recover-offsite-check` is a `docker exec` escape hatch (the `--print-reset-code` shape), NOT a page > or an API a browser can reach. R comes from **STDIN** — never argv, never `ps`, never shell history, > never a transcript. It asks the agent (>= v0.125.0) to fetch this host's sealed bundle and open it, > then reports whether the recovered repository password matches the on-disk one **by sha256**. > > docker exec -i felhom-controller /usr/local/bin/felhom-controller --recover-offsite-check < /path/to/code > > **IT COMPARES AND NEVER INSTALLS.** `CheckOffsiteKeyRecoverable` must stay free of any write — if a > future change makes it place the recovered password, it stops being a diagnostic and needs the drill's > supervision (that is link 9, R-200's remaining half). Pinned by > `TestCheckOffsiteKeyRecoverable_WritesNothing`, whose red-proof is adding the install call. > > Exit codes are load-bearing: **0** match, **2** a clean MISMATCH, **1** a step failed. A mismatch is a > finding about the system; a failure is a finding about the run, and they must never share a status. > > **Proven live on demo-felhom 2026-08-04** — recovered sha256 == on-disk sha256 == the hub's stored > hash. Nothing customer-facing ships with it: no card, no form, no preview. > > ## About Viktor (project owner) > > - Works at Deutsche Telekom (Budapest), building Felhom.eu as a side business > - Felhom.eu: managed home-server service for Hungarian households > - Technical but prefers pragmatic solutions over over-engineering > - Runs all infrastructure on Gitea (gitea.dooplex.hu), k3s cluster for management > - Customer deployments use Docker Compose (not Kubernetes) for simplicity > > ### felhom-controller (this repo) > - **Version:** v0.16.1 > - **Phase 1:** ✅ COMPLETE — Stack Manager + Deploy Flow > - **Phase 2:** ✅ COMPLETE — Monitoring & Health (scheduler, CPU/temp, healthchecks.io pings) > - **Phase 3:** ✅ COMPLETE — Backups (DB dumps, restic integration, manual trigger, **dedicated backup page**) > - **Phase 4:** ✅ COMPLETE — Monitoring Page with Metrics Store (SQLite, Chart.js, system + container metrics) > - **Phase 5:** ✅ COMPLETE — Authentication, Persistence & Settings Page (settings.json, password change, session management) > - **Phase 6:** ✅ COMPLETE — Monitoring Warnings, Dashboard Alerts & Notification System > - **Phase 7:** ✅ COMPLETE — Storage Overview, Per-App Backup Toggles & Limited Restore > - **Phase A:** ✅ COMPLETE — Storage Paths Foundation (registry, auto-discovery, per-app HDD_PATH, deploy dropdown, health monitoring) > - **Phase B:** ✅ COMPLETE — Storage Management UI Polish & Health Severity Fix (flash messages, label editing, app details, FS info, deploy free space, backup context) > - **Phase C:** ✅ COMPLETE — Storage Init Wizard, Data Migration & Startup Fix (disk scan/format/mount wizard, rsync-based migration, startup pings) > - **v0.11.1 bugfix:** ✅ COMPLETE — Storage Scan: system disk detection via host fstab + blkid UUID resolution; FSType enrichment via `blkid -o export` > - **v0.11.2 bugfix:** ✅ COMPLETE — /host-dev mount for block device access; `HostDevicePath()` helper; all format/scan/safety ops use /host-dev > - **v0.11.3 bugfix:** ✅ COMPLETE — Added `fdisk` package to Dockerfile (provides `sfdisk`; not in `util-linux` on Debian bookworm) > - **v0.11.4 bugfix:** ✅ COMPLETE — FormatAndMount: fixed sfdisk (wipefs+force+`,,`), mount (explicit device path), mount propagation (rshared), ASCII label, smart partition skip, findmnt verification > - **v0.11.6:** ✅ COMPLETE — FileBrowser auto-mount sync (`syncFileBrowserMounts()`) + 3 UI fixes (badge color, progress bar, button text) > - **v0.11.7:** ✅ COMPLETE — Stale data cleanup + FileBrowser sync after migration + deploy page title fix > - **v0.11.8:** ✅ COMPLETE — Per-App Cross-Drive Backup (3-2-1 rule): rsync/restic to secondary drive, deploy page UI, backup page summary, scheduler jobs, API endpoints > - **v0.11.9:** ✅ COMPLETE — UI Polish Fixes: spacing, tooltip on "Módszer", status dot instead of disabled checkbox, progressive disclosure, emoji cleanup > - **First app deployed:** Paperless-ngx on demo-felhom.eu (2026-02-13) > - **Running on:** demo-felhom (N100 mini PC) at 192.168.0.162:8080, felhotest (Proxmox VM) at router.abonet.hu:33022 > - **All Phase 1-5 features working:** deploy, start/stop/restart/update, logs, health-aware states, auth, monitoring, backups, backup detail page, system monitoring page, settings page > > ## Architecture decisions > > | Decision | Rationale | > |----------|-----------| > | Go stdlib for web (no Gin/Echo) | Minimal dependencies, single binary, easy to embed templates | > | Templates as go:embed HTML/CSS files | Zero runtime file dependencies (compiled into binary), but each template is a separate editable file | > | Docker Compose for customers (not k8s) | Simpler troubleshooting, customers don't need k8s knowledge | > | k3s for management infra only | Viktor's own services (gitea, monitoring, website) run on k3s | > | Cloudflare Tunnel for remote access | No port forwarding needed, works behind any NAT | > | app.yaml per stack | Separates deploy config from compose files, survives git pulls | > | Password fields require explicit input | Prevents accidental empty-password deployments | > | Health-aware state from Docker Status field | Docker's State says "running" even for unhealthy containers | > | Memory limits via deploy.resources.limits | Prevents runaway containers; ~50% headroom over expected usage | > | System info from /proc/meminfo + statfs | No external dependencies, cheap to read on each page load | > | mem_request vs mem_limit (K8s-inspired) | Requests = expected usage (hard block), limits = peak (overcommit OK) | > | 384MB reserved for system | Prevents deploying apps that would starve the OS/controller | > | Logo SVG embedded as Go constant | Same approach as CSS/HTML — zero external file deps | > | Git sync via os/exec git CLI | No Go git library needed, git is in the container image | > | SHA-256 for content comparison | Only copy changed files, avoid unnecessary disk writes | > | 30s debounce on manual sync | Prevents spamming the git server | > | Orphan = deployed but not in catalog | Safe lifecycle: remove from catalog → mark orphaned → user deletes via UI | > | FileBrowser as infra (not catalog) | Needed even after apps deleted (user browses HDD data); deployed by setup script | > | Protected HDD paths | Safety net: never delete top-level HDD dirs (media, storage, Dokumentumok, appdata) | > | Central scheduler (not ad-hoc goroutines) | Single place to register/monitor all periodic tasks, graceful shutdown, skip-if-running | > | CPU sampling via background goroutine | /proc/stat delta needs two readings — collector runs every 5s, GetInfo() reads cached value | > | Temperature from /host/sys (Docker mount) | Container can't read host /sys directly — mount /sys:/host/sys:ro, try /host/sys first | > | Restic password auto-generated | No manual setup needed — generated on first backup run, stored in named volume | > | DB discovery via docker inspect | No config needed — discovers postgres/mariadb containers by image name + env vars | > | Backup orchestrator with running flag | Prevents concurrent backups, supports both scheduled and manual trigger | > | modernc.org/sqlite (pure Go) | No CGO/gcc needed in Docker build stage — keeps `CGO_ENABLED=0` static binary | > | AlertManager state-based refresh | Alerts regenerated every 5min from health report — no persistent storage needed, always reflects current state | > | Notification relay via hub | Controller → hub → Resend → email. Hub acts as central relay: knows customer email, handles Resend API. Controller only needs hub URL + API key | > | In-memory notification cooldowns | Per-event-type cooldown map (default 6h). Lost on restart = acceptable (better to re-notify than miss). No persistence needed | > | Health status change detection | Only notify on degradation (ok→warn, ok→fail, warn→fail). Avoids spam on flapping. First run records baseline, doesn't notify | > | Resend HTTP API (no SMTP) | Direct POST to api.resend.com — same pattern as website contact-mailer. Simpler than SMTP setup, good deliverability | > | Preferences sync on save + startup | Controller pushes prefs to hub (not pull). Startup sync handles hub DB rebuild. Local save always succeeds even if sync fails | > | Chart.js embedded locally | Customer hardware may not have internet — CDN not reliable for offline environments | > | StackDataProvider interface | backup package needs stack data but can't import stacks (circular). Interface in backup, thin adapter in main.go | > | Password sync to hub via report | Restic password in Docker named volume on SSD. Hub sync provides redundancy for disaster recovery | > | App backup via HDD mounts only | Docker volumes at /var/lib/docker/volumes/ not mounted in controller. HDD data is the important user data; DB in volumes covered by nightly dump | > | Restore uses running mutex | Prevents concurrent backup+restore on same restic repo. Reuses existing `m.running` flag | > | Storage paths registry in settings.json | Multi-storage support: each app's HDD_PATH from app.yaml is authoritative. Auto-discovery on startup avoids manual config. Registry enables UI management + health monitoring per path | > | /mnt:/mnt:rw mount in controller | Replaces per-path HDD_PATH mount. Enables multi-storage + restore writes. All customer HDD mounts are under /mnt/ by convention | > | Per-app HDD_PATH resolution (app.yaml > global) | App's own env HDD_PATH is Priority 1, registered storage paths as fallback. Eliminates dependency on global controller.yaml hdd_path | > | Mount-point detection via syscall.Stat_t.Dev | Compares device ID of path vs parent dir — reliable check that path is on separate filesystem. Prevents data writes to SSD | > | Health severity: mount-point = warning | Non-mount-point is informational, not a service failure. FAIL reserved for genuinely broken things. Avoids false alarms on demo/test environments | > | FS info via findmnt + sysfs | `findmnt -n -o SOURCE,FSTYPE --target ` for filesystem type/device. `/sys/block//device/model` for disk model. Best-effort, returns nil on failure | > | Query param flash messages | Stateless, no session store needed. Consistent with backup page pattern. `?storage_msg=success&storage_detail=...` | > | StorageLabels map on stacks page | Separate map passed to template (not modifying Stack struct). Built from deployed apps' HDD_PATH → registered path label lookup | > | Metrics downsampling via SQL | Bucket-based AVG in GROUP BY keeps Chart.js responsive with up to 30 days of data | > | 60s metrics collection interval | Good balance of resolution vs. storage — ~44K rows/month for system metrics | > | /etc/os-release mounted read-only | Container can't read host OS info directly — mount to /host/etc/os-release:ro | > > ## Key file locations on demo-felhom > > ``` > /opt/docker/felhom-controller/ # Controller compose + config > ├── controller.yaml # Customer config (domain, auth, paths) > ├── docker-compose.yml # Controller's own compose > └── data/ # Controller persistent data (named volume) > > /opt/docker/stacks/ # All app stacks > ├── traefik/ # Reverse proxy (protected) > ├── cloudflared/ # Tunnel (protected) > ├── paperless-ngx/ # First deployed app ✅ > │ ├── docker-compose.yml > │ ├── .felhom.yml # App metadata > │ └── app.yaml # Deploy config (env vars, locked fields) > └── whoami/ # Test stack (not deployed) > > /mnt/hdd_placeholder/storage/ # HDD storage for apps > └── paperless/ > ├── consume/ # Drop files here for OCR > ├── media/ # Processed documents > └── export/ # Backup exports > ``` > > ## Related repositories and their state > > | Repository | Status | Notes | > |------------|--------|-------| > | felhom-controller | Active | This repo. Controller code + deploy scripts | > | app-catalog-felhom.eu | Active | 10 app templates, all with .felhom.yml metadata + memory limits | > | felhom.eu | Active | Website + hub/ subfolder (felhom-hub service) + k8s manifests | > | homelab-manifests | Stable | k3s cluster running (dooplex.hu services) | > | misc-scripts | Utility | collect-repo.sh, backup helpers | > > ## Gotchas & lessons learned > > - `docker compose restart` ≠ `docker compose up -d` — restart doesn't pick up new images > - Go maps have random iteration order — always sort slices before displaying > - Docker `.State`="running" doesn't mean healthy — check `.Status` for "(health: starting)" / "(unhealthy)" > - Paperless-ngx needs `PAPERLESS_OCR_LANGUAGES` (plural) to install language packs, `PAPERLESS_OCR_LANGUAGE` (singular) to select > - In-memory Deployed flag must be set BEFORE `docker compose up -d` (not after) — compose can take 30-60s for image pulls, during which the UI would show a stale "Telepítés" button > - Cloudflare Tunnel handles *.demo-felhom.eu → Traefik handles Host()-based routing to containers > - BIOS "AC Power Recovery" must be enabled on N100 for auto-restart after power outage > - `docker compose up -d` returns exit 0 even when containers immediately crash-loop — need post-start status check to detect this > - When logging env vars for debugging, only log keys (not values) to avoid leaking secrets in log files > - Mealie image (`ghcr.io/mealie-recipes/mealie`) doesn't include wget/curl — use Python TCP socket check for healthcheck > - Mealie DB migrations on first start take ~40s (alembic) — use `start_period: 60s` to avoid premature unhealthy status > - Alpine-based images (filebrowser, vaultwarden) have wget via BusyBox — healthchecks with `wget --spider` work fine > - Deploy `sed` command to update image version must target only the `image:` line — naive `sed 's|name:OLD|name:NEW|'` also matches the service name line (e.g., `felhom-controller:` → `felhom-controller:0.2.12`), breaking YAML. Use `sudo sed -i 's|image:.*felhom-controller:[^ ]*|image: ...felhom-controller:NEW|'` or similar scoped pattern > - Hungarian quotation marks `„"` in YAML: `„` (U+201E) is safe inside YAML double-quoted strings, but the closing `"` must NOT be ASCII `"` (0x22) — it terminates the YAML string. Use `\"` escape or Unicode `"` (U+201D). This caused a silent parse failure for the entire `.felhom.yml` file > - Never silently swallow parse errors — always log them. Silent failures make debugging impossible (took a dedicated debug session to find a simple quoting issue) > **2026-08-14 — v0.215.0 (R-328..R-333). Disk health, phase 1: the alert that reached nobody.** > > ### A severity string is a WIRE CONTRACT with the hub, not a label we choose. > > The hub accepts exactly `{info, warning, error, critical}` and **silently coerces anything else to > `info`**, which `severityNotifies` then drops. `disk_health_degraded` shipped `"warn"` — one letter > short of the contract — so **every Figyelmeztetés-level disk alert this product ever produced was > emailed to nobody, on both legs.** Proven live side by side on 2026-08-14: `"warning"` → > `notification_log` status **`sent`**; `"warn"` → stored `info`, **no row at all**. > `app_start_failed` (`notifier.go` ~L546) carries the identical defect and was deliberately NOT > changed here — it needs its own decision on whether it should notify (**R-329**). > > ### A drive's own PASSED verdict cannot fail on bad sectors. Do not build on it. > > Attributes 187/197/198 all carry `thresh: 0`; a normalized SMART value floors at 1 and can never drop > to or below the threshold. The real drive (ST3000VX010, S/N Z6A07P2G) read `PASSED` at **352** pending > sectors and **1001** reported-uncorrectable reads. Felhom already read the raw counters, which is the > only reason it would have noticed at all. Evidence + fixtures: > `felhom.eu/documentation/audits/DIAG-smart-passed-trap-2026-08-14.md`. > > **DECISIONS MADE, so they are not re-litigated:** > > - **Predicted failure is labelled „Hiba" — there is no fourth verdict word.** A fourth Hungarian word > sharing a root with „Figyelmeztetés" would make the MORE severe state read as the milder one. Four > labels, final: Rendben / Figyelmeztetés / Hiba / Nincs adat. > - **Sustain is the primary rule; the count is the backstop.** Truth-table row 6 (unreadable sectors > present again at the next check) sits ABOVE row 8 (count >= 64) because on the real drive sustain > fires 12 Aug and the count not until 13 Aug. Row 8 exists only for a box powered off across the > sustain window. > - **The numbers and where they come from.** 64: the benign excursion peaked at 16 and cleared inside > an hour; the terminal run passed 64 at 13 Aug 11:28 and never returned. **It is a judgement from ONE > drive** — a static backstop, expected to be replaced by growth-rate detection in Phase 3. 55/60 °C: > the operator's existing Prometheus bands on DooPlex, adopted unchanged so the two systems cannot > disagree about the same drive. **These bands are SPINNING-DISK bands and are questionable for NVMe** > — demo-hp's healthy Toshiba NVMe idles at **53 °C**, 2 °C below Figyelmeztetés (**R-333**). > - **Phase 2 owns the new SMART attributes (187 Reported_Uncorrect, 199, 188).** They are a declared > wire change, so under the G-1 gate the hub must model them in the same session. Putting them here > would have turned a one-word severity fix into a three-repo change (**R-330**). Everything v0.215.0 > needs was already on the wire. > - **Phase 1 state is one small record per disk, NOT a sample series.** `metrics.MetricsStore` is the > right home for Phase 2/3 history; using it now would have put a schema migration on the critical > path of the severity fix. > > **The trap this change nearly shipped, caught by a test and not by review:** the card and the check > share one verdict function so the chip and the email can never disagree — but the check CONSUMES the > prior and then overwrites it, so a card rendering afterwards read its own check's write and showed one > level MORE severe than the alert. Fixed by `diskRecord.PriorSawUncorrectable`, which replays the prior > that produced the stored verdict. The guarantee was previously asserted in a comment only. > > **Cadence is hourly, and it was MEASURED:** demo-hp `/disks` costs median 0.821s (min 0.805 / max > 0.841, 10 calls, 3 physical rows) — 6x under the 5s bar. Open question deliberately NOT acted on: the > agent runs bare `smartctl -a -j` with **no `-n standby`**, so an hourly poll would wake a spun-down > HDD. demo-hp is all-flash so the measurement could not show it (**R-333**). > > **NOT live-validated:** the Fail-from-counters path has never fired on real hardware — only against > the fixture's values in unit tests (**R-332**). > **2026-08-08 — v0.208.0 (R-254). THE RULE, stated so it outlives this session:** > > ### A secret is never in a page's response body. It is fetched by an explicit act, and the act is recorded. > > Three instances of one pattern shipped in two days, each found by hand: the retrieval passphrase > (R-249), an app's real first-login password (R-254 site one), and an already-deployed app's generated > secret field (R-254 site two). Every one was "hidden" with `display:none`, `hidden`, or > `type="password"` — **instructions a browser honours when DRAWING and nothing else.** The plaintext > was in the bytes; a `curl` returned it; caches, history, saved pages and screen-shares had it. > > **The shape of the fix, now used three times:** the page carries a BOOLEAN; the value comes from a > **POST** (so CSRF covers it and it is not re-fetchable from history) with **`Cache-Control: > no-store`**; the reveal is **LOGGED as an act** — reading a value off markup left no trace anywhere, > which is why nobody can say whether any of these was ever read. **Per-secret endpoints, never one > generic "reveal any named secret"** — that would turn three narrow exposures into one lever. > > **And the test must assert the RAW RESPONSE BODY.** Every test that asked what the customer *sees* > passed while the bytes carried the secret. That is precisely how this survived three times. > > **What is NOT this defect:** a form must carry what it submits. The pre-deploy hidden input round-trips > a generated secret deliberately (README §318) so the saved value is the one the customer wrote down. > The defect there was the neighbouring READONLY input on an already-deployed app, where nothing is > submitted at all. > > **The gate:** `scripts/secret_in_markup_gate.py`. Name-based, all 36 templates, **blind to a secret > arriving under a neutral page-data key** — measured, not assumed. The complementary runtime > body-assertion covers 4 of 27 page templates; the other 23 are **R-255**. > > **A correction to v0.207.0's report:** it said HTML comments ship in the response body. They do not > here — `html/template` strips them (`text/template` does not). Measured. > **2026-08-08 — v0.207.0 (R-249, R-252, R-253). Three things the fifth walk exposed BY PASSING.** > The walk closed R-201 (both halves) on 2026-08-07; none of the below touches the recovery path it > proved. > > **R-249 — a secret was living in the page source.** `settings_security.html` rendered the retrieval > passphrase into a `display:none` span behind a „Megjelenít" button. That toggle stops a browser > DRAWING it and nothing else: the plaintext was in the response body of every render. Found by doing > exactly that — it landed in a session transcript while driving the documented rebuild path. > **THE RULE, which the codebase already stated for R and this page did not follow:** a secret is > revealed by an XHR, never templated server-side into HTML (`escrow_handlers.go`). The page now > carries only `HasRetrievalPassword`; the value comes from `POST /settings/retrieval-password/reveal` > — CSRF-covered, `no-store`, and **logged as an act**, which reading it off the markup never was. > **The test asserts the RAW RESPONSE BODY** — every test that asked what the customer *sees* passed > while the bytes carried the secret, and that is why it survived. > **The census found two more instances** (`app_info.html`, a real per-install app password in a > `hidden` span; `deploy.html`, a generated secret in a `value=`) — **filed as R-254, not fixed.** > > **R-252 / R-253 — the two obstacles, and the rule they share.** A rebuilt box keeps its drives but > loses their REGISTRATION, so every restore refused with a sentence naming no next step; and the > restore list promised „a visszaállítás előbb újratelepíti" three lines above a refusal that fired > *because* the app was not installed. **The promise was the wrong half:** reconstitution writes to > the app's own `GetStackHDDPath`, which exists only once the CUSTOMER has chosen a drive at deploy > time — an automatic reinstall would mean the product making that choice for them, which is the one > decision this recovery path exists to leave with them. Both now name a reason and route to the step > that clears it, and both notices are conditional (a healthy box is byte-identical, pinned by a test > that fails if either becomes unconditional). > > **The page and the resolver ask ONE question:** `HasRestoreDestination()` reads the same > `GetSchedulableStoragePaths()` the scratch resolver reads. A second copy of that predicate is > exactly how a page ends up promising what the handler refuses — which is R-253 itself. > **2026-08-07 — v0.206.0 (R-241). THE RULING, and it reversed the fix: this was a MINTING defect, > not a screen-predicate defect.** The recovery screen was telling the truth — there genuinely was > nothing recoverable under the key the box held, because **the box minted that key itself over the > top of a sealed package it already knew the hub was holding**. Fixing the predicate would have > papered over a machine quietly making its own backups unopenable. > > **THE RULE: a box does not create a repository key while the hub holds a sealed package for it.** > The guard is a conjunction (package held AND no key), so a first-time box is untouched, and the > refusal is a HOLDING state rather than a failure — the transport is still configured so the > recovery screen can bring the tier up the moment the key arrives. > > **THE SECOND RULE: the fact that answers a question must be kept where the question is asked.** The > hub-vs-local key comparison had been computed on every ACK since SLICE 3 and persisted nowhere; on > the venue it logged the right answer thirty-five minutes before the customer looked at a screen > that could not see it. It is now persisted and drives shape (c) of the offer. > > **THE THIRD RULE (the operator's, and it generalises): fix the state, do not remember that it is > wrong.** Abandoning the old history now starts a 14-day countdown that removes the set-aside store > and its sealed package TOGETHER, after which the offer falls silent on its own because there is > nothing left to compare — rather than a "they decided" flag suppressing a screen over a state that > is still wrong. The recovery offer stays reachable for the whole grace; a grace in which recovery > is impossible is decorative. > > **Surface:** the full page appears once per ENTRY into the offered state, not once ever — a box > rebuilt months later is a new situation. Three dismissal levers with three scopes, and none of them > removes the entry point on the backups page. > > **Needs hub v0.98.0** for the superseded-package purge. `felhom-agent` untouched. > > **Two real bugs were caught by tests rather than by review** — a missing `t.Enabled` (an existing > test) and a missing falling-edge sync that reintroduced the very defect the epoch exists to fix. > > **NOT built, deliberately:** the automatic 30-day abandonment (R-245, with the operator's reasoning > recorded), and R-242's release-to-golden gate. > **2026-08-06 — v0.205.0 (R-234).** THE RULE: **a run that skipped an app the customer selected is > not a successful run.** The R-203 verdict block already said *"a warning beside a success is read > as a success"* and applied it to one of the two shapes it describes — a missing declared FOLDER > made the run `incomplete`, an app skipped ENTIRELY did not. Now both do. A selected-but-UNDEPLOYED > app is named with what to do but does NOT move the verdict, because a box left permanently amber by > an app somebody removed is a status nobody reads. > > **§7.3, MEASURED rather than assumed — and the answer was "already done".** `CaptureRecoveryUnit` > writes compose config + a manifest (a few KB), only ENUMERATES dumps rather than creating them, is > idempotent, and does NOT stop the app; the off-site run already calls it for every deployed stack in > its own pre-dump phase, through `admitApp`. So there is no wait to remove for a deployed app, and > **nothing was built**. Proven on demo-hp: a unit moved aside was RECREATED by the run. > > **AND THE FILED MECHANISM WAS NOT THE MEASURED CAUSE.** R-234 was filed as "the first run after a > toggle finds no bundle and skips the app". That cannot happen for a deployed app (above). What did > happen on 2026-08-06: the manual run was dropped by the **single-flight** while an earlier run was > still going; `runOffboxBackup` returned nil; the handler had already said „elindult”; and the card > then showed the PREVIOUS run's „✓ Rendben”. Fixed by taking that decision synchronously in the > handler. **The nightly path deliberately still returns nil** — nobody asked, and it retries. > **2026-08-05 — v0.200.0 (R-193 CLOSED).** The customer-facing recovery screen. Until now a customer > whose machine was rebuilt had everything needed to get their data back and no way to find out — the > only route was a command line. > > **IT UNLOCKS AND ONLY UNLOCKS** (operator ruling). Explains, takes the recovery code, opens the > repository, lists what is in it (apps, dates, sizes). **Restores nothing** — restore is per-app and > lives in the backups area; the put-back is **R-213** and its stated requirement is a > live-versus-backup comparison. > > **ONE CORE, TWO CALLERS.** `backup.RecoverInstallCore` is the only fetch→unseal→compare→install path. > `RecoverAndInstall` is now a thin CLI wrapper — exit codes and printed lines byte-identical, every > pre-existing CLI test passed unchanged — and the handler calls the same function. Asserted from > source by AST on BOTH sides, plus a test that the routes and the landing-page interception exist. > > **THE TRIGGER HAS TWO SHAPES and the second is the one that matters.** `OffsiteRecoveryOffer` = the > hub holds a package AND (no repository password OR the tier is orphaned). The literal "no repository > password" alone is a window that CLOSES BY ITSELF — `WriteOffboxSecrets` auto-generates one on > re-apply (R-193's own orphaning mechanism) and hub v0.96.0's self-heal re-applies within ~15–30 min. > Shape (b) is also what the shipped move-aside requires, which is why the discard choice can reach it. > > **CLAIMED is part of the predicate** — a legacy-open box passes through `RequireAuth`, so without an > explicit `authEnabled()` check the interception fired for an unauthenticated visitor. A test caught it. > > **„Most nem" suppresses the FULL PAGE ONLY.** The backups-area entry point is bound to > `recoveryOffer`, never to the postpone flag. > > **The code:** POST body only, never logged/persisted/echoed, cleared on every path, `no-store`, > `autocomplete=off`. **No lockout** — a ten-word phrase is not guessable and locking a customer out of > their own data for a typo is worse; failures are logged locally without the code, and NO operator > alert is raised (reasoning in REPORT.md §4). > > *Live:* demo-felhom is genuinely in shape (b), so validation needed no arrangement — `/launcher` → > 302 `/recovery`, both mandatory sentences rendered, three wrong codes refused with the `offbox/` > listing byte-identical and no lockout, and the code found in no file, log or ring **with a > planted-copy positive control that first exposed a mis-aimed sweep**. **NOT proven live: a CORRECT > code** — none was kept for demo-felhom's orphaned history and demo-hp's is operator-held. > **2026-08-05 — v0.199.0 (R-204 item 4 / R-193).** The last of the four manual interventions the > 2026-08-04 drill needed. **Operator ruling: automate it, and the trigger is a state the BOX > DECLARES.** From the hub an absent off-site object has FOUR meanings — never configured, > mid-restart, a transient config read failure, rebuilt-and-stranded — and the hub cannot tell them > apart. The box can. > > **The declaration needs BOTH halves** (`backup.needsOffsiteCredential`): a fresh data area (no > repository password) AND a hub-held recovery package (the ACK's `identity_blob_present`). Freshness > alone is a box that never had off-site backups; dropping that condition makes the whole fleet ask > for credentials, which is what `TestOffsiteDeclare_NeverHadOffsiteSaysNothing` catches. A merely > DISABLED target is the customer's own choice and never declares. > > **The ACK field stopped being discarded.** `EscrowAutoConfirmer.Reconcile` returns early when the box > is neither pending nor escrowed — exactly a rebuilt box — so the fact was thrown away every cycle. It > is recorded FIRST, before every gate, via `RecordPresence`, wired in main.go and asserted by > `TestMainWiresRecordPresence` (AST, comments dropped). Last-write-wins, not set-only, so a customer > RESET turns the declaration back off; a nil ACK escrow records nothing. > > **Inert to every existing reader:** `enabled:false` + zero sizes, so the hub's `isStale` and > `fillBand` both short-circuit; an unknown `state` string is ignored by encoding/json. **A configured > box's report JSON is byte-identical to v0.198.0's.** The one reader that would have misread it is the > hub's `reportHasOffsite`, tightened in hub v0.96.0 to require `enabled:true`. > > *Live:* both demo boxes now record `hub_escrow_identity_present=true` in settings.json (the recorder > working on a HEALTHY box). demo-felhom 9201, arranged reversibly into the stranded shape, produced > report id=16743 carrying `{enabled:false, state:needs_credential, quota_gb:0, repo_size_bytes:0}`; > the single declaration was absorbed by the hub's debounce (no self-heal event) and the box was > restored the same minute. **The hub half is felhom.eu v0.96.0.** > > *Rider:* `.githooks/pre-push` in all four repos now refuses a push from a clone outside > `/mnt/5_hdd/felhom.eu`. Proven both ways against a scratch clone. > **2026-08-05 — v0.198.0 (R-204 items 1 & 3).** The 2026-08-04 drill (R-201) passed only because a > person was there; four manual interventions stood between a recovered key and a restored file. Two > of the three defects are in this repo. > > **Item 1 — the reset code needed a restart.** `--print-reset-code` is a SEPARATE process; it > persisted a new code while the running server kept the old one cached, so the code the customer was > told to type was refused until the controller restarted, and nothing said so. `effectiveClaimCode` > now calls `settings.ReloadClaimCode()` first. **The settings-vs-config precedence is unchanged** — > the defect was freshness, not precedence. **Read-through, not a TTL, and that is the point:** a TTL > makes the new code visible AND leaves a window in which the superseded one still works, which is > worse than the bug. That is the mutation `TestClaimCode_SupersededByASecondMint_RefusedImmediately` > exists to kill, and its red-proof produced exactly *"the SUPERSEDED code was accepted"*. The > function now returns an error and **every caller fails closed**; an absent settings file is NOT an > error. `ClaimConsumedGeneration` is deliberately NOT re-read — this process is its only writer and > re-reading could move it BACKWARDS if a save had failed, resurrecting a consumed code. > > **Item 3 — the restore's default returned the wrong thing silently.** `mode=unit` restores the > recovery unit (definition + config + DB dumps) and not the customer's files. `restoreScratchOutcomeMsg` > now names what came back, what did not, and the next step; the wizard's intent card states its scope > before the choice. **The size gate is untouched** and pinned unchanged by > `TestOffboxRestore_FullPathUnchanged`. **The default stays `unit`** — all three wizard forms set > `mode` explicitly, so a change would alter nothing visible while silently changing a mode-less POST. > > *Live-validated endpoint-level (no browser on DooPlex):* on demo-felhom 9201 with `restarts=0` > across both mints, a superseded code returned „Hibás vagy lejárt kód" and the current one was > accepted first time; on demo-hp 9201 a `privatebin` unit restore produced the scoped Hungarian > outcome and `mode=full` without confirm revealed `full_size=6.8+KB` without restoring anything. > demo-hp's drill scratch (`calibre-web`) was not touched. > > **Item 2 is the hub's** (felhom.eu v0.95.0, R-196). **Item 4 — a rebuilt box cannot obtain an > off-site credential unaided — remains OPEN (R-193)** and was deliberately not begun. > **2026-08-02 — v0.190.0 (R-157 mechanism A · R-170 · R-171).** Three items, one live validation > cycle, because all three are boot behaviour and all three are proven by hard-resetting the box. > > **DIAGNOSE BEFORE THEORISING — and the first diagnosis was a FALSE NEGATIVE.** A hole was reasoned > out of the v0.189.0 diff (a drive-gate-stopped app has zero containers and `desired_state: running`, > so it now reads as a boot orphan) and confirmed on hardware BEFORE any fix was written. **Attempt 1 > produced `no boot-orphaned apps` and would have been reported as a disproof.** It was a race: > unmounting only the parent bind is healed by the agent within ~60 s, so the drive gate's startup > reconcile restarted the apps **one second before** the sweep looked. Holding the drive genuinely > absent reproduced the defect immediately. **"It didn't happen this time" is not a mechanism.** > > **The confirmation moved the severity in BOTH directions.** The write hazard did not materialise — > compose failed `mkdir …/userdata: permission denied` because the unbound mountpoint is > host-root-owned and the guest is unprivileged. **That protection is ACCIDENTAL**: no code chose it, > no test pinned it, and it is one `chown` or one privileged guest away from gone. But the harm that > DID occur was not in the hypothesis and is real on every box: two wasted attempts and a **false > dead-app alarm for an app the drive gate is deliberately holding**. > > **The fix already existed one path over.** `startGatedByMissingDrive` (the API) refuses a customer's > start on an absent drive; the sweep bypassed it by calling `Manager.StartStack` directly. > **`StartStack` HAS NO GATE OF ITS OWN** — carry this: every caller that is not the customer must > decide for itself whether the app may run. New consumer-side `bootrecon.StartGate`, fail-safe > (cannot determine ⇒ do not start). > > **Widening a window makes previously-unreachable overlaps reachable — a design input, not an > afterthought.** The old T+5 s sweep never met a quiesce or an in-flight app-data operation; a 50 s > window can. All three holders answer ONE seam because they differ only in the reason string. > > **A TEST REJECTED MY FIRST CONSTANT, and the comment says so.** `settle + budget + one retry` must > fit inside `deadAppBootGrace`; 60 s gave 95 s against 90 s. The budget is 50 s **because a test said > so** — recorded in the code rather than presented as taste. Widening the grace was rejected: it > hides a late recovery instead of reporting one (`recordLateRecovery`). > > **THE FIX HAD ITS OWN DEFECT, FOUND LIVE AND NOT BY REVIEW.** The window sampled `GetStacks()` — the > Manager's map, refreshed by the scheduler every **10 s** — every 5 s, so two identical samples could > mean *the cache did not update*. Observed: a container removed ~5 s before the window closed was > still in the sampled fleet and the sweep logged `no boot-orphaned apps` for an app that had none. > `sampleBootFleet` now refreshes first. **Generalise: a settle detector is only as good as the > freshness of what it samples — if the source is cached, refresh it, or you are watching the cache > settle rather than the system.** > > **R-170:** `shouldRecreateOnBoot` reads intent with the identical three-way table; absent keeps the > old `hasContainers` behaviour exactly; `presentStable` untouched and still load-bearing. Its comment > argued at length FOR the count and was rewritten. Agreement pinned from BOTH sides against one > fixture table (an import cycle prevents testing the two gates together). > > **Live: 6/6 hard resets** (every app back; the customer-stopped app down all six), settle times > 10/40/10/10/15/15 s. Sharpest evidence: same app, same box — missed at 18:08:35, recovered at > 18:18:50. R-170 proven in one reboot (calibre-web recreated, immich left stopped). 27/27 packages; > 7 red-proofs. Detail: `REPORT.md`. Last updated: 2026-08-02 (v0.189.0 — R-166 / D-b: the box stops guessing what the customer wanted) > **2026-08-02 — v0.189.0 (R-166, operator decision D-b).** When an app was not running the box had > to work out *why*, and it did so **by counting containers**: zero meant "the customer stopped it", > some meant "something broke". A **power cut mid-compose** and an **interrupted deploy** also leave > zero containers, so both were read as deliberate stops and stranded **silently** (R-157 mechanism > B) — and a backup that stopped an app and died left it stopped with **nothing on disk** recording > that it was owed a restart. The settling fact — what the customer asked for — **was written down > nowhere**: `app.yaml` recorded *installed*, never *meant to be running*. > > **DECISION — one owner: the customer's action, and nothing else.** A census found **14 callers of > `StartStack`/`StopStack`, of which exactly 2 are the customer**; the rest are quiesce, the volume > dump, offbox reconstitution, app export/restore, the storage gate, migration and the boot > reconciler. So the primitives are deliberately **not** writers — intent there would make a nightly > backup indistinguishable from the customer pressing Stop. Writers: the API action switch, > `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, the `.fab` import. Intent is written > **BEFORE** the act and a failed write **REFUSES** the act. > > **DECISION — absent means UNKNOWN, never "running", and this is the whole safety property.** Every > `app.yaml` on every box predates the field, so absent is what the fleet reads on upgrade; reading > it as running would start every deliberately-stopped app on the first boot after the upgrade. The > legacy branch of `isBootOrphan` keeps the old container-count rule **byte-for-byte**, and its test > asserts BOTH legacy rows together because the safety property is the pair. Backfill is > **running-only** — "zero containers ⇒ stopped" IS the defect, so an ambiguous app stays ambiguous. > > **Part 2 — `backup.AppStopGuard`**, a persisted marker over every stop→work→start window (volume > dump, offbox reconstitute, `.fab` export), in its **own** file (one file, one writer). Written > before the stop, cleared only after a restart that **succeeded**, kept when one fails. `Recover` > **returns** its outcome instead of using a notifier seam, because it must complete before the boot > reconciler (`main.go` ~236) while the notifier is not built until ~307 — a seam wired after the > fact is a seam that never fires. > > **THE TEST LESSON, and it is the one worth carrying:** Scenario E's first version called > `appStop.Begin` itself, and **survived the red-proof that deleted the production call**. It proved > the marker type, not that `DumpAppVolumesSafe` uses it. Rewritten to drive the real function with a > simulated hard abort (an unwind that skips the restart statement, since a `defer` is not > crash-safety — Campaign 8 fault 10). **A test that constructs the thing it is meant to prove the > caller constructs is hollow, and its red-proof will say so if you run it.** > > **FOUND EN ROUTE — `SaveAppConfig` rebuilt `AppConfig` field-by-field**, the R-100 shape (v0.181.0 > shipped two live instances). The literal named five fields, so `desired_state` would have been > dropped on **every** save across nine call sites — a customer's Stop erased by the next unrelated > `app.yaml` write. Copy-and-overlay (`saveCfg := *cfg`) is safe by construction. **Generalise it: > treat any field-by-field struct rebuild in a save path as a defect on sight.** Measured, not > assumed: `app.yaml` does NOT round-trip YAML keys the struct does not model (pinned by test). > > **R-157: mechanism B closed, mechanism A untouched** (the sweep observes ~5 s after start and never > re-checks) — and B's fix makes A cost more, since the sweep now has more it could recover. > **NEW R-170:** `shouldRecreateOnBoot` (`internal/web/intermediary.go:131`) still infers a Stop from > `hasContainers` — the same defect one gate over, for drive-backed apps. Left deliberately. > > **Live on 9201, three flows** (stop survives a restart; a zero-container `running` app recovered by > name; a legacy app.yaml skipped and never inferred stopped). The **interrupted-operation half is > IMPLEMENTED, not PROVEN-LIVE** — nobody killed the controller mid-backup on metal. 27/27 packages; > 7 red-proofs observed FAIL then restored. Detail: `REPORT.md`. Last updated: 2026-07-28 (v0.182.0 — R-101 + F-DIAG: the restore dialog names the last SUCCESSFUL copy) > **2026-07-28 — v0.182.0 (R-101 + F-DIAG).** `Tier2LastRun` is the ATTEMPT clock (written on failure) > and was rendered as „Legutóbbi másolat" in the **restore confirm dialog** — misinformation at a > decision point: the restore fills in MISSING files, so a customer with a failing Tier-2 restored and > silently got OLDER files. New `CrossDriveBackup.LastSuccess` + **`SuccessTracked`**; the marker is > load-bearing because **all 7 fleet rows were pre-anchor at deploy** — without it every customer sees > „Még nincs sikeres másolat" at once. Legacy rows migrate on first touch (`ok` adopts its time, > `error` seeds nothing). **PART 2 — the three `record*` helpers rebuilt the WHOLE struct with only 2 > fields carried over; the naive fix would have had `recordTier2Failure` CLEAR the anchor.** Replaced > by `tier2Update` (copy-and-overlay = safe by construction). New `fmtTimeStr` → Budapest-local dates > in the dialog instead of raw UTC RFC3339. **F-DIAG:** 6 classes incl. an honest `unknown`, and the > notification no longer passes `err.Error()` through raw — **LESSON: my first sanitiser was regex-only > and leaked a bare hostname; its own test caught it. Redact KNOWN values, don't guess at shapes.** > Live on demo-hp: rendered dialog read in the failed, healthy AND legacy states. F-OPS documented at > `felhom.eu/documentation/runbooks/RUNBOOK-manual-guest-restore.md`. > **2026-07-28 — v0.181.0 (R-100).** `OffboxTarget.LastSuccess` + wire field `last_success`; the hub > (v0.80.0) anchors offsite staleness on it. **`LastRun` is written unconditionally on every run > INCLUDING failures** — it records an ATTEMPT — so the hub's "how long since LastRun" verdict read a > nightly-failing tier as perfectly fresh forever. The rule is the pure `offboxAnchorAfterRun(prev, at, > runErr)`: a failure neither ADVANCES nor CLEARS the anchor (both are distinct bugs; clearing it would > make one bad night look like never-succeeded). `LastStatus == "error" ⇒ stale` was rejected — it pages > on every blip, the F-A1 noise mode. **TWO SILENT-WIPE SITES CLOSED** (`offboxConfigHandler` and > `ApplyOffsiteTarget` both rebuild the target and copy runtime status field-by-field — omitting > LastSuccess would erase the anchor on any settings save or hub re-apply). **LESSON: my first test > modelled the rule in a local closure and stayed GREEN when production was mutated — hollow; the > extraction to a pure function is what made the red-proof bite.** Live on demo-hp: failing run advanced > `last_run` to 11:25:48Z while `last_success` HELD at 11:24:20Z; demo-felhom healthy → advanced. The > settings-save preservation was proven live too. Detail: `REPORT.md` + `felhom.eu/REPORT-r100.md`. > **2026-07-28 — v0.180.0 (F-OBS).** Source: `audits/CAMPAIGN-8-backup-restore-2026-07-27.md`. > On a default `logging.level: info` box there was **no positive observable that `deadapp-check` had > run**: its per-cycle line goes through `Scheduler.dbg()`, gated on `level==debug`, so on a default > box it was never *produced* and could not even reach the always-DEBUG ring. "No alarms" was > therefore indistinguishable from "the detector never ran" — standing rule 3's exact fallacy, and it > undermines F-CRIT-1's fix, which is a fix to **this same detector**. > `noteDeadAppScan()` now emits an INFO line every **20th** scan (10 min at the 30 s cadence) carrying > scans-since-boot / evaluated / currently-down. It reports **what it saw**, not that it ran, and it > summarises rather than floods — one line per run is 2880/day, which is what made silence attractive > in the first place. Both bounds are pinned by test in the direction that would break them. > **The same shape then turned up in the agent's brand-new guest-power watchdog** (v0.107.0, shipped > hours earlier): it logged only at startup and when it acted. Fixed in agent v0.109.0 with the same > pattern. The anti-pattern reproduces itself — which is the argument for not having dropped this part. > Live on demo-hp at INFO on a default-level box; deployed on both boxes. Detail: `REPORT.md`. > **2026-07-26 — v0.173.0 (R-77).** Source: `audits/DIAG-agent-channel-2026-07-26.md`. > > **UNRESOLVED AND DELIBERATELY DEFERRED — which file is authoritative for `local_api`?** R-77 ships > DETECTION ONLY. `controller.yaml` and `bootstrap.json` can disagree; the controller dials > `controller.yaml`. The obvious "fix" — reconcile from `bootstrap.json` on every boot — has a failure > mode **as severe as the bug it fixes**: on a guest whose `controller.yaml` is correct and whose > `bootstrap.json` is stale (a re-provision that half-completed, a hand-repaired guest, a > setup-wizard box), auto-reconcile would clobber a WORKING channel on the next restart — fleet-wide, > silently, at the moment of a routine deploy. R-77's position is that **naming the drift is enough**: > it would have converted the 17.5 h outage into a specific alert on the first health cycle. The > authority ruling is **R-78** and needs its own spike — do not resolve it opportunistically. > > Corollary for anyone editing `bootstrap.MaybeIngest`/`ensureLocalAPI`: `ensureLocalAPI` is the ONLY > writer, it fires only when the endpoint is EMPTY, and `DetectEndpointDrift` must stay write-free. > Scenario A's test asserts `controller.yaml` is byte-identical after the check, and its red-proof > covers the auto-correcting variant precisely because that is the tempting wrong turn. > > **Also settled here:** the samba protected-set must mirror EVERY early return in > `reconcileSambaAt` (currently two: `!smb.Enabled`, `!smb.UserSet`). A third would need the same > mirror, and the doc comment above `EffectiveProtected` must be updated with it. > **2026-07-26 — v0.172.0 (R-75).** Spike `felhom.eu/documentation/audits/SPIKE-catalog-data-paths-2026-07-26.md`; > feature doc `felhom.eu/documentation/controller/import-and-data-paths.md`. > > **RULING — the import root is CANONICAL on the system drive, overriding the spike's Fork-1 > recommendation of per-drive roots.** The spike weighed sidebar clutter and per-app link ambiguity and > concluded per-drive; the operator overruled it on an argument the spike missed: each drop-zone app has > exactly ONE ingest bind, so on a two-drive box every import folder except the app's own would look like > a drop-zone and silently do nothing — and because `import/*` is `class: excluded`, files stranded there > are never backed up either. A canonical root is the only shape with no dead drop-zone. Recorded as a > deliberate deviation, not an oversight. > > **Phase-0 probe changed the shape of Part 6.** The system drive is NOT a registered `StoragePath` on > either demo box (`/mnt/felhom-drives/hdd_1` on demo-felhom; `nvme-1tb` + `Felhom-Share` on demo-hp), > so `sharingResolvePath` REFUSES `/userdata/import` — verified against the real guard with a > passing control. Registering the drive was rejected (it would make the 50 GB volume holding the > recovery units a customer-visible drive, deploy target and wipe candidate, and `SharingDeniedRoots` > would then deny the namespace-consistent shape anyway). **Chosen: leave it unregistered and have the > controller write the `beolvasas` share directly** — the picker guard validates CUSTOMER-supplied paths, > a controller-generated constant is a different trust class. No guard was weakened. > > Also note: `withUserdataPath` computes `USERDATA_PATH` as `/userdata`, NOT > `NamespaceRoot(hdd)/userdata`. For an app on the system drive those disagree > (`/mnt/sys_drive/userdata` vs the `felhom-data` namespace). Latent — no app with a userdata bind has > ever been deployed there — but it is a real inconsistency, left untouched here. > > The other three forks followed the spike unchanged: all-apps skeleton / deployed-only in the UI; > unknown role fails OPEN while a malformed path whole-block rejects; drop-zone copy driven by the > derived backup class. > **2026-07-24 — v0.169.0 (disk-health card + degradation alert).** Consumes the agent's new `smart` > field (agent v0.94.0; MinAgent floor unchanged — feature-detect by presence). **Rulings:** (1) ONE > pure verdict fn `agentapi.DiskVerdictFor` is the shared truth for the card chip AND the 6h check — they > can never disagree. Thresholds: FAILING→Hiba; PASSED + any(reallocated>0/pending>0/offline_unc>0/ > critical_warning>0/media_errors>0/percentage_used **≥90**)→Figyelmeztetés; PASSED clean→Rendben; > nil/UNKNOWN→Nincs adat (never alarms). (2) **No global alert banner** — the card + email carry disk > health; banner fatigue is a real cost, so this is deliberately NOT wired into the dead-app/alert-banner > machinery. (3) Degradation-only notification with an in-memory baseline: first run baselines silently, > recovery never notifies, **UNKNOWN excluded both directions** (a transient blip neither fires nor erases > history). (4) **Controller restart re-baselines silently** (in-memory baseline lost on restart) — an > accepted trade consistent with the health-change pattern (a real post-restart degradation still fires on > the following 6h check once a baseline exists). (5) A **60s TTL cache** wraps the card's /disks call so > dashboard refresh-spam can't smartctl-storm the host; the 6h check fetches FRESH (cache-independent). > Pairs with hub +1 (allowlist `disk_health_degraded`). No new smartctl load — serialization only. Last updated: 2026-07-24 (v0.168.0 — customer-configurable backup window "Mentési időablak") > **2026-07-24 — v0.168.0 (customer-configurable backup window).** ONE customer setting — the window > start W ("Mentési időablak kezdete") — drives every nightly leg at FIXED, never-stored offsets so > misordering is impossible: DB dump at W, tier-2 at W+60m, off-box at W+105m (wrap-safe). **Design > rulings:** offsets are DERIVED and computed everywhere, never persisted and never exposed in the UI; > precedence is settings > controller.yaml `db_dump_schedule` > "02:30" (mirrors PasswordHash); a change > applies WITHOUT restart via the new scheduler seam `UpdateDaily` (per-daily-job buffered `resched` > chan + a select case in `runDailyJob`). New pure package `internal/backupwindow` holds all the time > math (ParseHHMM/FmtHHMM/LegTimes/GateWindow/EffectiveWindow). **Disk-tier (whole-guest PBS/vzdump) > gate:** the quiesce loop's SCHEDULED cycles run only inside [W+2h, W+6h) (wall-clock Europe/Budapest), > with a safety valve — last successful backup older than cadence+24h (or none) runs regardless, so a > box only ever on outside its window never starves. **Manual "Mentés most"/TriggerNow is NEVER gated** > (bypasses runOnce). The `quiesce.Backend.Due` seam now also returns the backup age (from the agent's > own `/backup/due`); the agent, its cadence, and `/backup/due` are untouched. Window read fresh each > poll (WindowStartFn) so runtime changes take effect. Cadence defaults to 24h controller-side (the > response carries no cadence). Backup page gets a "Mentési időablak" card (time input + derived rows + > the "kb. W+2h–W+6h között" rendszermentés line); POST /backups/window (RequireAuth+CsrfProtect). > **2026-07-24 — v0.167.0 (outlined logo + favicon — Part 4 unblocked).** Viktor pushed the > text-outlined `logo.svg` to felhom.eu `main` (`be9edb4`); the wordmark is now 17 real `` > glyphs. `FelhomLogoSVG` swapped to it; Inkscape's leftover **empty `` shells + font-* leftovers > on the paths** were stripped via an lxml DOM pass (glyphs untouched — CC did NOT do text-to-path), > editor `` dropped. `FelhomFaviconSVG` vestigial `` removed. Both constants: > **0 ` Inkscape "Object→Path" leaves empty `` shells AND copies `style="…font-family:…"` onto the > resulting ``s — a search for `svg:text` misses them (elements are ``, no prefix); grep > ` first** (`assetsSyncer.Resolve`) and fall back to the constant only if none is on disk — on 9201 the > constant is what's live (verified). Still open (separate follow-up): website + hub serve their own > non-outlined logo copies; login.html stylesheet link still unversioned. > **2026-07-24 — v0.166.0 (mobile nav = off-canvas drawer; sidebar cleanup; ?v= on logo/favicon).** > Mobile nav was broken: the ≤768px block predated the v0.146.0 accordion and flattened `.nav-links` > into a horizontal `overflow-x` strip, clipping the accordion's nested sub-lists (they share the > `.nav-links` class). **Decision: mobile nav = a sticky top bar + off-canvas left drawer that REUSES > the vertical sidebar (Option A).** The accordion handler is untouched and works inside the drawer; > a `no-js` html-class fallback renders the sidebar static inline so nothing dead-ends without JS. > Options B (separate mobile menu) and C (exclude nested lists from the strip) were rejected. z-index > ladder topbar 800 < backdrop 900 < drawer 950 < modal 1000; `100dvh`; reduced-motion disables the > slide; focus-trap deliberately omitted (navigations reset state). **Sidebar customer-name removed** > (logo only); `{{.CustomerName}}` stays in base data + login subtitle. **Logo policy decision: the > wordmark must be OUTLINED paths, never live ``** — under `` secure static mode only > locally-installed fonts resolve, so `font-family` in the SVG renders a fallback font everywhere. > **Part 4 (swap `FelhomLogoSVG`/`FelhomFaviconSVG` to the outlined master) is GATED OUT** — §3a check > against live felhom.eu `main` (`be9edb44`) found `website/assets/logo.svg` still has ``/ > `font-family`; the outlined master is Viktor's manual Inkscape push, still pending. Only the `?v=` > cache-bust (logo/favicon/login-logo, Cloudflare 4h edge-cache — the 0.126.1 failure mode) shipped > from the logo work. Follow-up: when Viktor pushes the outlined asset, ship Part 4 (swap constants + > clean the favicon's vestigial `` nodes). Separately, the website + hub still serve their own > non-outlined logo copies — propagation is a distinct follow-up. > **2026-07-24 — v0.165.1 (native "Megosztás…" in the share modal, Web Share API).** The share modal > gains a feature-detected `navigator.share` button (OS share sheet → Messenger/WhatsApp/email), > sending **title + text + URL only**. Hidden unless supported; "Link másolása" stays the universal > fallback (and catches the non-cancel rejection); `AbortError` (user cancel) is silent. **Ruling: the > QR is NOT attached** (no Web Share Level-2 `files:`) — file-share support is narrow and several > targets drop the URL when handed file+URL, leaving an unscannable QR picture in a chat; the QR's job > (physical cross-device scanning) is already served by the modal image (mobile long-press). Template > JS + tests only; the OS sheet interaction is an operator manual check (not endpoint-testable). > **2026-07-24 — v0.165.0 (Indítópult megosztása — guest launcher via capability URL).** The admin > launcher gets an "Indítópult megosztása" button that mints a **capability URL** > (`https:///s/`, 160-bit `crypto/rand` token) serving a standalone, read-only guest > launcher — same tiles, opens apps in new tabs — with **no account and no admin session**. **Security > ruling: the link grants INFORMATION ONLY, ZERO CONTROL** — app names + public URLs; every privilege > stays behind each app's own auth and the controller admin password. The token IS the secret (160-bit > entropy is the whole defence for the GET — never rate-limited, never logged, `subtle.ConstantTimeCompare` > only; an empty stored token = sharing OFF, matches nothing, so a wrong/disabled token is byte-identical > to the mux default 404). Optional per-share password is a SEPARATE credential (own bcrypt hash, own > attempt map — NEVER the admin ones); one pass mints a cookie = HMAC(`token|passwordHash`) keyed with > the persisted `web.session_secret`, so rotate-token OR change-password invalidates all cookies for free. > **Part-2 secret decision: REUSED `web.session_secret`** (persisted + box-scoped + stable — the SAME > secret the claim pre-auth CSRF already trusts; not per-boot, not claim-generation-scoped → the reuse > branch), so no `ShareCookieSecret` field was added. **Design rulings recorded:** member accounts are > **superseded** by this capability-URL model; **per-member tile visibility is PARKED under the SSO arc.** > Guest state labels ride the v0.164.0 invariants: `StateStopped` ⇒ "A tulajdonos leállította"; any > other non-clickable state ⇒ "Átmenetileg nem elérhető" (guests never see stopped/exited/degraded/ > unhealthy). Accepted residuals (documented, no code action): link-preview crawlers fetch once and see > app names (noindex prevents indexing); reverse-proxy/CF access logs may hold the path (ops-tier); the > modal link carries the request Host, so a LAN-IP admin session yields a LAN-IP link. New dep: > `github.com/skip2/go-qrcode`. Tests: Groups A–G (14 tests) + 3 red-proofs verified red. > **2026-07-24 — v0.164.0 (stopped ≠ fault).** Operator finding on 9201: a UI stop (Leállítás) raised > the global "Telepített alkalmazás nem fut: … (stopped)" banner on every page AND fired the > `app_start_failed` email. RULING: **a deliberate user action must not alarm anywhere.** One-line > filter at the single fix-3 derivation point — `scanDeployedAppRunStates`'s pure core extracted to > `classifyRunStates([]stacks.Stack)`, down predicate now > `stacks.IsDownState(st.State) && st.State != stacks.StateStopped`. `StateStopped` is dropped from > BOTH the banner dead-list and the notifier Down-set (⇒ no banner, no event, clean tracker). Rests on > **two invariants that MUST both hold for this suppression to be correct:** **I1** — the UI stop path > `Manager.StopStack` runs `docker compose down` → containers removed → a deployed stack with zero > containers aggregates to `StateStopped` (refreshStatusLocked). **I2** — the P2 restart-policy census > (2026-07-21, 53 templates / 78 services) found every catalog service on `unless-stopped`, so a crash > never rests at `stopped` — faults surface as `exited`/`degraded`/`restarting`/`unhealthy`. **If > either invariant changes, revisit this suppression.** `IsDownState` UNCHANGED (other callers rely on > stopped=down). Out-of-band `docker compose stop` (containers remain → `StateExited`) still alerts — > correct, tampering is reportable. The `stopped_by_user` intent flag was considered and PARKED (only > adds value against out-of-band stops, which should keep alerting). Tests +4 (notify 3→4, main 4→7), > both red-proofs verified. No template/funcmap/notifier/counter/copy change. > **2026-07-24 — v0.163.1 (launcher polish).** Two v0.163.0 live findings fixed. RULE recorded: > **every app-logo surface ends in a visible placeholder** (`SVG → PNG → /static/app-placeholder.svg`, > infra rows → `infra-logo.svg`) — the four sibling `onerror` chains (`backups_apps`, `stacks`, > `app_info` hero, `deploy`) now match `app_row.html`; `app_info` screenshots deliberately still > vanish on error. And the **launcher monogram is launcher-only AND failure-only**: hidden by default, > revealed when the tile's img chain fails (`onerror` adds `.launch-tile--noimg`) — it was bleeding > through every transparent white glyph. Template/CSS only; no handler/funcmap change. 5 tests + 2 > red-proofs. [[launcher-v0163-2026-07-24]] > **2026-07-24 — v0.163.0 (Indítópult app launcher + universal placeholder icon).** New > customer-facing `/launcher` page: the FIRST sidebar item (above Vezérlőpult), a grid of large > tappable tiles for openable deployed apps. `/` stays the Vezérlőpult — the launcher is ADDITIVE. > Design rulings recorded here: > - **(a) The felhom brand mark is NEVER an app placeholder** — brand = platform identity only. The > logo-less fallback everywhere is the new generic `AppPlaceholderSVG` (a 2×2 app-grid glyph, > `/static/app-placeholder.svg`), now the DEFAULT `FallbackIcon` on `app_list_row` (was > `visibility:hidden`). On the launcher tile the fallback is the **monogram**, not the placeholder. > - **(b) A launcher tile exists ⟺ a „Megnyitás" button would** — subdomain presence (env `SUBDOMAIN` > > `.felhom.yml` subdomain > `protectedStackSubdomains`) is the single openability criterion. The > controller stack is excluded by name. The subdomain assembly was extracted to > `Server.subdomainMap` (3 callers: dashboard, Alkalmazások, launcher; priority byte-unchanged). > - **(c) Colored-tile + mono-glyph design.** `tileColor` = validated `.felhom.yml` `brand_color` > (`#rgb`/`#rrggbb`, new `Metadata.BrandColor`, omitempty) OR a deterministic FNV-1a-of-slug HSL > (fixed S/L, hue per app). Invalid `brand_color` silently falls back to the hash color (the one > §8 exception to no-silent-failure — cosmetic). `tileColor` returns `template.CSS` (we > validate/compute in Go; html/template's CSS filter mangles a legit `hsl()` from a func pipeline). > - **(d) `/` remains the Vezérlőpult.** No role/auth gating — member-role gating is a future arc > (ROADMAP: member role → launcher becomes the member landing page). No catalog app sets > `brand_color` yet (curation parked). > No agent coupling; MinAgent unchanged. 10 new test functions + 4 red-proofs (all observed FAIL then > restored). Gates green (app_row_dedup / template_id / emoji). > **2026-07-24 — v0.162.0 (R-71a), SHIPPED + deployed BOTH boxes (demo-felhom 9201 + demo-hp 9201 > via G1 break-glass), clean+healthy, settle-gate GO line captured on both.** B′ live note: both > above-floor boxes GOed correctly but NOT literally first-poll — the floor is in-memory (not > persisted), unknown at t=0, so the gate logged `awaiting floor knowledge` then GOed ~10 s later the > instant the report ACK landed (report-ACK latency = exactly what the 90 s sub-bound is sized to; > zero-wait-when-floor-known is unit-proven, test E). The gate correctly did NOT burn the one-time > password before the update picture was clear. > The structural fix for the F10 day-0 race (DIAG-f10): the apply-bridge no longer consumes the > single-use offsite password while a managed floor-update is in flight or imminent (below floor). > New seam `offsiteapply.SettleProvider.SettleState()` + `SettleFunc` adapter over the updater's own > `GetFloor()`/`IsUpdateRunning()` (no second floor path); `Bridge.AwaitSettle` polls 10 s BEFORE the > 3-min Reconcile ctx (deferral never eats the reconcile budget), bounds 90 s floor sub-bound / 5 min > overall (both GO+WARN — the "hub that can't serve a floor can't serve a consume → no burn" argument, > R-71c is the belt). At/above floor → GO first poll, zero wait (B′). Bridge goroutine MOVED after the > updater in main.go; wired only when an updater exists. **Ordering-only** — consume/persist/404 > contract untouched; R-71(b) rejected-by-design. **FINDING:** the floor is in-memory > (report-ACK-derived ~5–10 s), NOT persisted → unknown on any restart until the first ACK (sized the > 90 s sub-bound to that). 5 test scenarios (A–E) + nil-provider + cancelled-gate; **4 red-proofs all > observed FAIL then restored** (gate/updateRunning/sub-bound/overall-bound). Deferral paths NOT > live-fired (precondition now structurally prevented by the v1.25.0 build gate). **Layering: gate > prevents, (a) defers, (c) heals.** ROADMAP R-71 → SHIPPED (a)+(c). Live leg = the B′ first-poll GO > line on both above-floor boxes. > **2026-07-23 — v0.161.0 (R-70 controller leg), SHIPPED + deployed BOTH boxes.** When > `offsite.enabled` is in controller.yaml but no `offbox` target exists (pre-apply window / burned > credential — the F10 shape), Távoli mentés now shows „Felhom offsite tárhely kiépítve — a > beállítás automatikus, folyamatban…" on BOTH empty surfaces (status card + target line) instead > of „igényelhető" / „Még nincs beállítva". Data key `OffsiteHubEnabled` (from `Server.cfg`, no new > wiring); render tests per gate branch; banner leg is unit-proven/live-pending (no healthy box > occupies the window; next fresh onboarding is the natural live leg). Hub sibling v0.72.0 carries > the detector + `offsite_delivery_stuck` + the R-71c self-heal. Origin + rulings: > `felhom.eu/documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`. > **2026-07-22 — v0.160.0 (R-67), SHIPPED + deployed BOTH boxes, full live leg on demo-hp.** > Network shares now bind their share ROOT into FileBrowser (`…/:/srv/:rslave`) — no > skeleton/userdata toward the NAS, ever. Pure assembly = `buildFileBrowserPaths` + `fbPathDeps` > (handlers.go), returning mounts AND config sources together so they can't disagree. > > **DECISION — two classes, two gates:** drives keep the drive-absent gate (byte-identical, > tested + observed live: demo-felhom logged a no-op sync); network shares use the STUB classifier > gate instead (stub ⇒ excluded from both lists + WARN — an exposed stub swallows uploads the real > mount later shadows; idle autofs is HEALTHY and included; unknown fails open). Never force-wake > in the sync (doctrine). > > **Phase-0 probe = GO:** in-container access through an rslave bind WAKES an idle autofs trigger > (proved on demo-hp against the real Felhom-Share). Live leg: upload from demo-hp's filebrowser > container (uid 1000) landed on demo-felhom's share dir and deleted clean; dead-NAS gave > `Host is down` in seconds (no hang) and recovered unaided after samba restart. RESIDUAL for the > operator: the FileBrowser HTTP click-through — its admin credential is customer-held (CC got 401 > on admin/admin and the demo password; by design). ROADMAP R-67 SHIPPED (coupled to R-64). > **2026-07-22 — v0.159.0 (R-66), SHIPPED + deployed to BOTH boxes.** Three legs: „Hálózat" card on > Beállítások → Rendszer (Helyi cím / Hálózati név only-while-Megosztás / Átjáró; „—" fallback), > `network` section in the Debug dump (best-effort per item), and the NetBIOS trap named on the NAS > add form (Szerver helper text + a purely lexical hint on `unreachable` for single-label non-IP > names). > > **DECISION (the load-bearing one): all guest-net reads go through the samba netns door.** The > controller is bridge-netns'd, so `/proc/net/route`/resolv.conf/net.Interfaces in-process answer > for the CONTAINER (172.x / 127.0.0.11) — the S-2 trap. `internal/stacks/guestnet.go` docker-execs > into host-networked felhom-samba (one `guestNetExecFn` seam); Megosztás off ⇒ door closed ⇒ „—" / > in-place error strings, never a plausible-wrong substitute (S-5). Nothing stored anywhere. > > Deploy: 0.159.0 on demo-felhom 9201 (open-door path live: .104/.1/\\FELHOM) AND demo-hp 9201 via > G1 break-glass (closed-door path live: dashes, no name row, in-place dump errors; secret shredded). > demo-hp gotcha worth keeping: the controller 404s on direct container-IP probes without the > customer-domain Host header (`felhom.enkisfelhom.hu` there). Red-proofs A2 + C2 run and recorded. > ROADMAP: R-66 SHIPPED; R-64 (pairing blessed, drill = evidence leg) + R-65 (buddy-box replication, > post-alpha spike-first) minted. NAS doc gained the naming-caveat paragraph. > **2026-07-21 — v0.155.0.** v0.154.0's wizard sourced "is an op running" from `Manager.IsRunning()` > — the CONCURRENCY single-flight, acquired inside the goroutine, and **`RestoreOffboxScratch` never > acquires it**. So the execution step was unreachable for „Ellenőrzés" and the full-restore > preparation: live buttons while a restore downloaded, with the progress banner contradicting the > phase strip on the same screen. Found by the operator on the first live click-through. > > **DECISION: display reads `RestoreStatus()` (the `opRunning` flag), never `IsRunning()`**, through > the named `restoreOpInFlight` seam, and the handler reads the status ONCE per render so the strip, > the suppression and the running-op name cannot diverge. The lesson generalises: `opstatus.go` is the > DISPLAY surface and says so in its own header — the concurrency flag is not a substitute. > > **The test lesson:** a table test over a pure function proves the function, not the caller. Scenario > E passed throughout because it injected `OpRunning=true` directly. The new test drives a real > `Manager` through `BeginRestoreOp` and asserts the render. > > **DECISION: „Eredmény" earns its place.** The strip's highlight is now `Phase`, derived separately > from `Step`: a finished restore is back on the intent step while the strip reads „Eredmény" and an > outcome card shows the result — window-bounded (10 min) and app-bound. > **2026-07-21 — v0.154.0 (R-48).** Collapses the offsite restore controls to a single > „Visszaállítás…" entry per app row plus a per-app wizard at `GET /backups/restore/app?name=`. > The defect it closes is the CAUSE of the round-2 incident: the list rendered up to five inline > forms per row, two of which — the missing-only merge and the true reconstitution — were sibling > buttons whose difference is whether the data comes back. The rule it establishes: *two adjacent > controls whose difference is "your data comes back" vs "your data cannot come back" must never be > distinguishable only by layout.* > > **DECISION: the wizard is server-rendered on the EXISTING endpoints.** No new mutation endpoint, > no JSON state API, no client router. Every card is a real form POST to > `/backup/offbox/{restore,place,reconstitute}` with the same field names and gates, and the server > renders the next step — so it works with JavaScript disabled. `TestRestoreWizard_NoNewMutationEndpoints` > makes that structural: adding a form that posts somewhere new fails the suite by design. > > **DECISION: R-45 stays its own item.** The wizard polls the two existing status surfaces as-is; the > generalized job registry (and with it a real per-phase progress feed) is not built here. > > **DECISION: the step is derived, never requested.** `deriveWizardStep` is pure over (op running, > size-gate flash, scratch ready). Precedence is load-bearing — a running op outranks a stale > `?full_prep=` in the URL, or a commit button reappears mid-restore. While ANY op runs every > mutation form is suppressed server-side rather than offered and then refused with a 409. > > Latent bug found and fixed on the way: `offboxRedirectTo` hardcoded `"?"` when appending its flash, > which would have buried the flash inside `?name=`. **No agent coupling — MinAgent stays > 0.90.0.** 9 new tests + the Group-B red-proof; full suite green. > > **NOT live-validated at commit time by design:** v0.154.0 is published but deliberately NOT > hand-deployed — the operator's hub floor save (0.153.0 → 0.154.0) pulls it via the self-update > path, and that swap IS the R-23(a) single-fire validation (STOP-1). > **2026-07-20 — v0.153.0 (R-47).** Closes the H4 race on **BOTH** restore paths. The replay needs a > running DB container, so both paths started the WHOLE stack first — giving the application a window > to rebuild the schema objects the dump was about to create. Measured at 8 s on 2026-07-19 > (`DIAG-immich-restore-round2-2026-07-19`): immich-server rebuilt `clip_index` two seconds before > the dump's `CREATE INDEX`, the replay aborted `already exists` under `ON_ERROR_STOP=1`, and immich > then reported schema drift. The photos came back **by accident** — `pg_dump` emits COPY before > CREATE INDEX, so the abort landed after the rows; a collision earlier in the script would have left > a genuinely half-restored database, reported identically. > > **DECISION: the DB-only bring-up is done by compose SERVICE scoping**, not by container tricks — > `StartStackServices(name, []string{svc})` → `compose up -d `. Every catalog template's > dependency direction is app→db, so naming the DB starts the DB and nothing else. `docker start > ` was never an option: `StopStack` is `compose down`, so the containers no longer exist. > `RestartStack`/`RedeployFromEnv` are traps here — both end in a full `up -d`. > > **DECISION: fail-closed.** A `.sql` dump with no identifiable DB service refuses BEFORE the first > mutation, on both paths (one Hungarian string, shared). The alternative would be to start everything > and replay into the race. It should be structurally unreachable — `dbTypeForImage` is now shared by > `DiscoverDatabases` and `DBServiceNames`, and a dump can only exist because discovery matched the > container's image, which IS the compose `image:` value — so this is the belt for template drift. > > Enablers: `RedeployFromEnv` split into `PersistUnitRedeployConfig` (persist, starts nothing) + the > unchanged tail; `StackDataProvider.RecreateStackFromUnit` renamed to > `RecreateStackDefinitionFromUnit` because the old name promised less than the method did — the > hidden `up -d` inside it is what carried the defect on the local path. `StartStackServices` REFUSES > an empty list (argument-less `up -d` is a full start). **No agent coupling — MinAgent stays 0.90.0.** > 19 new tests, 3 red-proofs, 23/23 green. **NOT live-validated yet:** STOP-1 supervised reconstitute, > golden 0.153.0 bake (P3 registry-reachability probe from the vacation site is load-bearing), Viktor's > two hub saves, and his C6 customer-restore UI run. > **2026-07-20 — v0.152.0 + felhom-samba 1.1.0 (Megosztás on a Mac).** Closes **S-3**. **A capture > on the box overturned the earlier guess:** macOS DOES send a correct NBNS query for `<20>` and > nmbd DOES answer it correctly in 140 µs (flags `0x8580`, RCODE=0, right address) — macOS simply > never acts on it. NetBIOS there feeds legacy browsing, not `smb://` URL resolution, so **the bare > `smb://` can never work from a Mac** and nmbd was never the broken part (it is what serves > Windows). felhom-samba 1.1.0 adds **avahi + dbus**, templating `avahi-daemon.conf` and the > `_smb._tcp` service file from `FELHOM_SERVER_NAME` so a rename re-advertises; both daemons are > non-fatal on failure. v0.151.0's card had offered `smb://` for Mac — the one dead form — now > `smb://.local`; Windows keeps flat `\\`. Spiked live by hand and confirmed from the > operator's Mac BEFORE publishing the image (the operator's call, and it chose the design too). > **STILL OPEN: Finder-sidebar discovery is NOT shipped** — the record is published and answers > browse queries, but was never observed working; likely a Finder Settings → Sidebar toggle, but > unverified. **Windows was not retested.** Two test bugs fixed en route, neither a production > defect: `TestRenderSambaCompose` pinned a literal image tag, and `TestFabUpload_GCAndIdleTimeout` > asserted an async unlink synchronously (it passed alone, failed in the full package once the new > render tests made `web` heavier). 23/23 green twice; 2 red-proofs. > **2026-07-20 — v0.151.0 (Megosztás).** Closes **S-1/S-2/S-4-core/S-5** of > `felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md`; **S-3 (no mDNS/Bonjour) stays OPEN**, > awaiting Viktor's `smbutil lookup FELHOM` + `dns-sd -B _smb._tcp` from the Mac. **The `/sharing` > page had been reload-looping at ~1.2 s for every customer with sharing enabled since v0.147.0** — > `/sharing/status` coerced `idle`→`running` on the JOB phase channel, and the client answers a > terminal `running` with a one-shot `location.reload()`, so the first poll of every steady-state > page load re-armed it. The rule this leaves behind, now recorded against R-45 too: **a phase a > client answers with a one-shot action is an EDGE — never synthesise it from a level, and serve it > exactly once.** Both halves are server-side; `sharing.html`'s `