# CONTEXT.md — Project Memory > This file serves as persistent project memory across Claude Code sessions. > It replaces the auto-generated "Memory" from the claude.ai Project. > **Update this file at the end of each working session** with current state, > recent decisions, and anything the next session needs to know. > > Ask Claude Code: "Please update CONTEXT.md with what we did today" Last updated: 2026-09-18 (v0.255.0 — the globe on the sign-in-flow pages: styled, and inside the card) > **2026-09-18 — v0.255.0 (fixes the v0.254.0 globe).** The sign-in-flow pages requested `style.css` with **no `?v=`**, so a browser holding a pre-0.254.0 copy served CSS with no `.lang-globe` rules and the globe rendered as a bare unstyled `
` — **invisible to every test, because they all read markup and the fault was in which CSS file the browser fetched**. It was FIVE shells, not three (both guest share pages too), and `.Version` was missing from three of their data maps — now set once in `executeTemplateLang`. The globe also moved INSIDE the card, centred under the footer, menu opening upward via the shared rule. 15 shell fixtures re-captured, **zero dashboard ones**. R-579. > **2026-09-18 — FLEET FLOOR RAISED to 0.254.0** (operator asked), `min_agent` 0.131.0 declared — above the vouched golden 0.246.0, so the declaration carries it (R-472). Hub: `managed floor SERVED … from declared`; demo-felhom went 0.253.0 -> 0.254.0 **by itself in ~40 s**, healthy, four other containers up, own log `settle-gate: GO — at/above floor 0.254.0`, and its sign-in page shows **one globe, zero old text links**. **Peti Proxmox (DOWN 65 d, 0.115.0) and Tester 1 (DOWN 1 d, 0.245.0) did NOT get it** and take it unattended when they return — untested on this version. Evidence: `felhom.eu/documentation/audits/i18n-slice2-2026-09-18/floor-raise-0.254.0.md`. > **2026-09-18 — v0.254.0 (R-557 slice 2 release C — SLICE 2 CLOSED).** Saved notes follow the box language at WRITE time (§16 option 1): ~70 producers, and `EndRestoreOp` no longer takes a Hungarian literal from anywhere. The language switch is a **globe** (`
`, no script) in the sidebar footer and on the sign-in/claim/recovery pages; a visitor's choice lives in a display-only `felhom_lang` cookie that `langFor` reads ONLY when there is no session, and a claim carries it into the household setting on success. `POST /lang` is CSRF-exempt for a narrow reason written at the exemption; `safeBackPath` refuses `//evil.example` too. **Two parity exceptions, measured with a real diff: exactly two change shapes across 106 fixtures, 5 byte-identical.** **A DEADLOCK was introduced and caught by the suite hanging** — a note rendered inside `UpdateOffboxStatus`'s callback takes the settings read lock while the write lock is held; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` guards it now. 629 literals left, 0 errors, 0 saved notes. New row R-577. > **2026-09-18 — FLEET FLOOR RAISED to 0.253.0** (operator asked), `min_agent` 0.131.0 declared — the floor is above the vouched golden 0.246.0, so the declaration is what carries it (R-472). Hub: `managed floor SERVED … from declared`; demo-felhom went 0.250.0 -> 0.253.0 **by itself in ~20 s**, healthy, its other four containers up, and its own log reads `settle-gate: GO — at/above floor 0.253.0`. Both live boxes run agent 0.132.0 (above the requirement). **Peti Proxmox (DOWN 65 d, 0.115.0) and Tester 1 (DOWN 1 d, 0.245.0) did NOT get it** and will take it unattended when they return — untested on this version. Evidence: `felhom.eu/documentation/audits/i18n-slice2-2026-09-18/floor-raise-0.253.0.md`. > **2026-09-18 — v0.253.0 (R-557 slice 2 release B).** All **179** Hungarian error messages carry a key (`util.MsgError`): `Error()` is still the Hungarian byte for byte, `errors.Is` answers for the kind AND the wrapped cause, and an inner error argument renders recursively. 76 display sites go through `errText`, pinned by `TestNoErrErrorInPageOutput`. `memoryVerdict` returns an error, `UpdateRefusal` gained a `Cause`. Plurals: a key with `.one`/`.other` takes its count FIRST (bundle rule, not a call-site flag) — the guard caught a real key collision. **Two tooling defects found and recorded:** the bulk converter dropped multi-line concatenations (7 producers; the parity gate could not see it, two behaviour tests could), and my counter was case-sensitive and under-reported. 705 literals left, 0 of them errors. > **2026-09-18 — v0.252.0 (R-557 slice 2 release A).** The sentences the program BUILDS now follow the language: 226 Go literals converted (flash lines, page data, API answers, alert banners, 237 country names, the four app-named page titles — R-566 closed). Flash travels as a KEY in the redirect URL with `fa` parameters; an old link's prose is still shown verbatim. New gate `i18n_go_parity.py` refuses any key whose Hungarian is not byte-identical to the base commit (7 467 literals frozen; 3 decoys). Wire goldens freeze the report warnings and the event messages the hub MAILS — those stay Hungarian until slice 3 (R-558). Remaining: 894 literals, 176 of them `fmt.Errorf` (release B); persisted text (release C, §16 option 1). New rows R-572/R-573/R-574. > **2026-09-17 night — v0.251.0 (R-553 + R-563 CLOSED).** Five decisions that read their own Hungarian words now read a kind: deploy status sentinels (`internal/stacks/deploy_errors.go`), `backup.ErrOffsiteQuota`, `monitor.WarningKinds`, `settings.LastWarningKind` (legacy text fallback until R-570), and `data-status` on the remote-backup status. `util.KindErrorf` keeps message bytes identical. Slice 2 (R-557) is unblocked except the one producer named in R-570. > **2026-09-17 night — v0.250.0 (R-556 release C, slice 1 CLOSED).** All templates converted; the language switch is on every dashboard page (decision 6 superseded); login/claim/guest/catch-all through `executeTemplateLang`; Go-side messages and three app-name titles stay Hungarian (slice 2). > **2026-09-17 evening — v0.248.0 + v0.249.0 (R-556 releases A, B).** Apps/settings and backup pages > converted; `TestI18nParityCoversEveryMarker` (every marker rendered by a case); fixtures only written > when missing (`-update-i18n-golden`, `-i18n-golden-only`); `executeTemplateLang` for pages outside the > dashboard chrome (recovery now). Words a page COMPARES stay unconverted (R-563). Split Hungarian verbs > escape the retrieval gate's Hungarian stems (R-564). > **2026-09-17 — v0.247.0 (localisation starter; design `felhom.eu/documentation/architecture/10-localisation.md`).** > Message bundles `internal/i18n/locales/{hu,en}.json`; templates carry `{{T "key"}}`, EXPANDED > before parsing — one template set per language (`loadTemplates` / `parseTemplateSet`); `s.tmpl` is > the Hungarian set. Converted: launcher, `/backups`, `/apps/`, `layout.html`. **Release gate for > every future slice: `TestI18nParity`** — Hungarian render vs fixtures captured from UNCONVERTED > templates; never regenerate fixtures to make a conversion pass. `settings.json` `language`, > `POST /settings/language`, `?lang=` override; switch shown only on non-Hungarian pages (decided by CC, > operator may reverse). Report carries `"language"`. Copy gates now read templates expanded > (`scripts/i18n_bundle.py`) — they had gone blind/stale when copy moved. New `i18n_missing_gate.py` > (ratchets: English gap 0, formal „ön" forms 6). Converting a page: `scripts/i18n_extract.py`, then > REVIEW what it left (ASCII-only Hungarian slips past it). Plan rows R-553..R-562 (R-553 first). > **2026-09-13 night — v0.242.0 (R-487 / R-491 / R-490 / R-489 / R-476 / R-456).** The local backup > lists are keyed on the DRIVES, not on what is deployed (`ListRemovedAppUnits`, the R-237 rule one > tier down): a removed app whose unit was kept is listed with its restore, the picker and > `/api/backup/snapshots` answer for it, and the restore opens the unit where it sits > (`primaryUnitDirFor`). A removal clears an update hold. `/api/system/info` reaches the API router. > `volumes_removed` is a before/after difference — **but a volume recreated by a unit restore has no > compose label and is missed (R-489 open, residual: list by `_` prefix too)**. Tier-2 copy > date = newest dump (`UnitDataDate`) unless the leg was preserved. **The scratch guest 9202 on demo-hp > is where releases are validated now** (hand-set image; hub/tunnel off) — `felhom.eu/CONTEXT.md`. > **2026-09-13 — v0.237.0 (slice 4: R-448, R-443, R-439).** Update is a guarded 202 job: > refusals (hold, busy, memory, disk, no restorable Tier-2 copy) → backup-first if the proven copy is > older than `update.backup_max_age` (24h) → safety dump → pin → pull (failure: pin BACK) → up → > health (`.felhom.yml` check or 60 s settle, `update.health_timeout` 5m; failure: stop + HOLD, > `RestoreHold.Reason=update_failed`, pin stays). `UpdateStack` deleted. **The copy is aged by the last > successful Tier-2 copy, NOT the manifest `created_at`** — measured on demo-hp that `created_at` moves > only on definition changes. Journal `update-journal.json`; `RecoverUpdates` before the boot sweep. > Three unattended paths that ignored a hold now honour it (drive-return gate ×2, nightly volume dump); > capture and Tier-2 skip held apps. **No automatic rollback** — measured per-app; route back = restore. > **2026-09-13 — v0.236.0 (R-442).** Removal resolves the drive from the app's OWN `app.yaml` > `HDD_PATH` (the `07` ~L437 rule), never the global `cfg.Paths.HDDPath` (set on no box); a data > removal it cannot resolve is REFUSED (409, exact Hungarian sentence, typed `RemoveRefusedError`) > before anything is touched and the app is kept. SSD app → `hdd_paths_removed: []`, never `null`. > Backup-path refusals reach the response. Proven live on demo-hp — see `REPORT.md`. > **2026-09-06 — v0.235.0. THE RULING, AND THE TRAP IT SET.** > > **1. OPERATOR RULING: freeze the version, keep the fixes flowing.** R-447 sat `BLOCKED` because > R-438 established that `RestartStack`'s `up -d` was a CHOSEN behaviour with its reason in its own > comment. Option 1 was taken: an app's version is frozen to what the customer has and only a > deliberate Update moves it, while health-check fixes, memory limits and self-healing keep arriving > on the 15-minute cycle. **Both halves of the old behaviour were examined; only the version change > was unwanted.** > > **2. THE MECHANISM IS A RENDER, NOT A GATE — and that distinction is the whole design.** Nothing was > added to any of the thirteen `compose up -d` call sites. Most of them are REPAIRS (boot reconciler, > drive-return gate, app-stop guard), and a repair that refuses to repair leaves a customer's app > down. They are made safe by removing the reason: the file they act on no longer changes version. > > **3. `pinned_images` IS INTENT; `installed_images` IS AN OBSERVATION. NEVER FEED ONE FROM THE > OTHER.** Letting a reading become a deployment is the R-166 category error one field over. They will > normally agree; when they disagree that is a signal. > > **4. THE TRAP THIS RELEASE SET FOR ITSELF, and it would have shipped silently.** `Stack.TemplateImages` > is read from the app's LIVE compose file — which is now the RENDERED one. On a frozen app that file > names the OLD version, so the update badge would have found installed == template and answered > **„Naprakész" on exactly the apps that are behind**, with every test still green, because the new > field has the same type and shape. The badge now reads `Stack.CatalogImages`, from the syncer's own > clone. **A feature that silently inverts a previous feature is the failure mode to look for whenever > a file changes meaning.** > **2026-09-03 — v0.234.0. ONE GAP CLOSED, ONE TEST DEFECT OF MY OWN.** > > **1. A RECORD THAT ONLY THE BRING-UP PATHS WRITE NEVER REACHES A QUIET BOX.** v0.233.0 shipped with > "the record appears after the next lifecycle action" written down as a known limitation. One day > later the operator looked at demo-felhom and saw OpenGist — up 15 hours, running exactly the catalog > pin, **no badge at all**. The limitation WAS the feature not working. `BackfillInstalledImages` now > seeds the absences at startup by READING containers. **The general lesson: a feature that only fills > itself in on an event nobody triggers is, on the quiet installations, not shipped.** > > **2. THE BACKFILL IS STRICTER THAN THE BRING-UP PATHS, AND THE ASYMMETRY IS THE DESIGN.** The badge > reads a service-count mismatch as BEHIND. The bring-up recorder runs right after a SUCCESSFUL > `up -d`, where a missing container is real news; a backfill meets a box in whatever state it is in, > so a partial seed would render „Frissítés elérhető" over an app that is perfectly current. It > therefore refuses to seed anything it cannot observe COMPLETELY. **Same data, two writers, two > different admission rules — do not "make them consistent".** > > **3. A TEST THAT HARDCODES A DATE AND ASSERTS AN AGE IS GREEN ONLY ON THE DAY IT IS WRITTEN.** > `TestGroupD_BadgeRendersOnBothSurfaces` pinned `catalog_since: "2026-07-18"` and the string > "46 napja". The pure tests inject a clock; the RENDER path goes through the funcmap and reads > `time.Now()`. It passed on 2026-09-02 and was red on 2026-09-03. Now derived from the same clock the > code reads. **R-457** names six other test files that mix a literal date with `time.Now()` — as > unchecked candidates, not accusations. > **2026-09-02 — v0.233.0. TWO DECISIONS, AND ONE LIMITATION THAT IS NOT A DEFECT.** > > **1. `installed_images` is an OBSERVATION, so a failed write NEVER refuses the action — deliberately > the opposite of `desired_state`.** `SetDesiredState` refuses the act when the record fails, because > intent that could not be recorded recreates the exact ambiguity R-166 closed. `recordInstalledImages` > does the reverse: refusing to start a customer's app because we could not write down which version > it is trades a real outage for a bookkeeping gap. **The rule is "refuse on intent, log on > observation", and the reason is in both code comments** so neither gets "made consistent" later. > > **2. THE RECORD READS THE CONTAINER, NEVER THE COMPOSE FILE — and the file is not a lesser source, > it is a WRONG one.** The catalog syncer overwrites a deployed app's `docker-compose.yml` on a > 15-minute cycle with no deployed check at all; the spike measured the file saying `v2.8.5` while the > container ran `v2.8.6` for 25 minutes. Anything derived from that file answers "what will happen > next time something runs `up -d`", which is a different question from "what is running". > > **3. THE LIMITATION, STATED: 23 catalog pins FLOAT, so „Naprakész" CAN BE FALSE.** The comparison is > reference-to-reference and queries no registry (a box must not need eight upstream registries to > render a page). For `postgres:16-alpine`, `mariadb:11.6` and 21 others the ref can be identical while > the image behind it has moved — the spike caught `mariadb:11.4` and `mariadb:12.3` already moved. > **This is a known gap with a register row, not an oversight.** Digest-level comparison needs a > registry query and is deferred. > > **Nothing about updating changed.** No behaviour, no new endpoint, no auto-update, and the three > lifecycle buttons are byte-identical (`TestScenarioE_TheUpdateButtonIsUntouched`). The behaviour work > is slice 3 and needs an operator ruling. The whole arc's reasoning now has a home: > `felhom.eu/documentation/architecture/09-update-architecture.md` (R-438 — its absence was a finding). > **2026-09-01 — v0.232.0. THREE RULINGS.** > > **1. `restic stats` TAKES A REPOSITORY LOCK, and that is the fact the whole R-411 chain rested on.** > Nobody had it. Clean-room measured on demo-hp 2026-08-31: nothing else running, four invocations, > the sampler reads `locks=1`. A customer FULL restore shells `stats` in its size probe, so it holds a > lock — and until v0.232.0 it held no single-writer flag, so the integrity check was not blocked, ran, > met that lock, and `resticStep` removed it with `unlock --remove-all` while logging *"a stale > exclusive lock left by a previous crash"*. There was no crash. Also measured, and recorded so the > next reader does not re-derive it: `restic check` takes a lock; `restic snapshots` and `restic list` > do **not**. > > **2. THE INVARIANT IS PINNED BY A WALK, NOT BY A COMMENT — and the walk is the deliverable, not the > acquire.** `offbox_integrity.go:28` asserted *"Every off-site operation takes `acquireRunning`"* from > v0.227.0 and it was false for months, which is the ninth instance of this project's most-repeated > class. `TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered` is an AST pass over > `internal/backup` — deliberately not `strings.Contains`, because a commented-out call still contains > the string. **On its first run it found three entry points nobody had named**: > `OffboxRestorePrepareFull` (the request the customer's UI reaches FIRST, and the one that shells > `stats`), `RestoreSharesScratch` (R-411's exact shape on the shares tier, with a live caller) and > `RestoreOffbox` (no caller today). All four now take the flag. `OffsiteInventoryList` is registered > EXEMPT with its reason — `snapshots` only, measured not to lock, and flagging it would make a page > refuse to load during a backup for no safety gain. **Adding a line to `offsiteExempt` is a deliberate > act and belongs in the commit that adds it.** > > **3. R-414's BRANCH, AND THE EVIDENCE FOR IT.** The question was whether `offboxRestoreScratchDir` > was MISSED by R-356's unification or EXCLUDED on purpose. **Neither label fits: it was consciously > OUT OF SCOPE.** R-356's own commit (`08eb1a6`) says so in its test comments — *"the prepared scratch > still resolves to the registered storage path … only the DESTINATION moves, which is precisely what > this change is about"* — and every one of its fixtures assumed a registered storage path exists. > `demo-felhom`, with `storage_paths: []`, is the case it never had. It was **never ruled out on > state-only grounds**: the one comment about a `systemDataPath` fallback belonged to > `PlaceOffsiteRestore`, concerned bulk **userdata**, and R-356 deleted it deliberately. This > function's own documented exclusion is `cfg.Paths.DataDir` — the **rootfs** — a different filesystem. > > **So §6.3's `[DESIGN]` rule applies and now has a FOURTH consumer.** But the fallback is **SCOPED**, > because the two callers ask different questions and one predicate answering both is the R-356 defect > itself: **unit-only** may fall back (§7 records as `[FACT]` that a driveless app's unit already lives > on `systemDataPath` indefinitely and that the same-device placement is *"intended, not a defect"*); > **full** keeps the R-252 refusal, because it pulls bulk userdata onto a state-only tier (§2.2). > > **AND THE SILENCE ENDS EITHER WAY.** `ProofResultCannotRun` is recorded through > `RecordProofVerdict`, so `last_proof_result` is never ABSENT — absent already means *"controller > older than v0.231.0"*, and giving one field two meanings is the `StatsKnown` trap one level up. It > does **not** advance per-snapshot due-ness: nothing was proved, and marking one proved would stop the > app being retried once a drive is finally registered. > > **A MISTAKE OF MINE, RECORDED BECAUSE LIVE VALIDATION IS WHAT CAUGHT IT.** The fallback resolved a > scratch that `removeProofScratch` then refused to delete — its accepted-roots list is built from > REGISTERED drives, and a driveless box has none. Observed on `demo-felhom`: *"refusing to remove … > it is not inside a proof root"*, with the copy still on disk. Every nightly proof would have left one > behind, on exactly the boxes the fallback exists for. **The unit tests all registered a drive, so > none of them could see it.** Fixed, and pinned by a pair — one that the copy IS removed on a > driveless box, one that a path outside every proof root is still REFUSED, so the fix is not a > widening into uselessness. > **2026-08-31 — v0.231.0. FOUR RULINGS, recorded so none is re-litigated.** > > **1. THE ACCEPTANCE RULE IS NOT "EVERY DECLARED FILE IS PRESENT".** That is the spike's own one-line > summary and taken literally it is worthless: a hollow unit declares nothing, so everything it > declares is present, and the check passes on exactly the shape it exists to catch. The rule has TWO > parts and needs both — (1) everything declared is present, AND (2) the manifest declares what the app > is SUPPOSED to have. **Part 2 is the whole value; part 1 alone is the trap.** > `TestR87_HollowUnitForAnAppWithADatabaseFAILS` is the fence and its red-proof models the naive rule. > > **2. THE EXPECTATION COMES FROM INSIDE THE UNIT, NEVER FROM THE LIVE BOX.** R-403's guard could tell > hollow from legitimately-empty because it had TWO copies to compare; this has ONE. The snapshot may > predate the app's current shape, and the point is to judge the snapshot on its own terms — so the > source is the unit's own `compose/docker-compose.yml`, and `GetDockerVolumes` (live Docker, > `backup.go`) is explicitly NOT it. Database half: `DBServiceNames`, the same discriminator > `RestoreFromRecoveryUnit` already uses, so this cannot disagree with the restore path about what an > app is. Volume half: `ParseComposeNamedVolumes`. > > **THE VOLUME HALF IS AN EXISTENCE CHECK AND NOT A NAME MATCH, and that half-rule is deliberate.** > Volume tars are `_.tar`; `ResolveDockerVolumeNames` derives the project from > `filepath.Base(filepath.Dir(composePath))`, which inside a unit is the literal string `compose`, not > the stack. Measured 2026-08-31 on all eight real units on demo-hp: the counts match exactly > (bookstack 2/2, docmost 3/3, kimai 2/2, opengist 1/1, privatebin 1/1, calibre-web 1/1, paperless-ngx > 3/3, romm 3/3) and `_.tar` held in every case. **"Held on eight" is not "derivable"** > — R-355's standing rule is that a claim about the app must never be inferred from a counter. Half a > rule that is true beats a whole rule that is invented. Name-level matching is available the moment > the capture records the project, and is not worth inventing before then. > > **3. THE PROOF'S RESTORE IS A READ-ONLY VARIANT, NOT A CHANGE TO THE CUSTOMER'S PATH.** > `RestoreOffboxScratch` is the customer's restore, it is proven, and a customer restore taking a lock > is correct — so it was NOT changed. What was shared instead of forked: the scratch-dir resolver is > now parameterised on its ROOT builder (`offboxScratchDirIn`), so the drive-preference rules, the > network-storage refusal and the R-252 wording have exactly one implementation; and the unit-only > headroom gate is extracted to `unitOnlyHeadroom` so both paths refuse at the same floor with the same > Hungarian sentence. The proof's own three differences are the ones that must differ: `--no-lock`, no > `unlockStale`, and `m.runner()` instead of `resticStep` so the `unlock --remove-all` escalation is > unreachable rather than unlikely (REUSE.md's rule: replacing `resticStep` would hide the escalation > from the assertion that must see it). > > **THE PROOF SCRATCH IS A SEPARATE ROOT (`backups/offsite-proof`) AND THAT IS A SAFETY DECISION, NOT > TIDINESS.** The proof deletes its copy on every path including failure. Sharing > `backups/offsite-restore/` would mean a nightly background job deleting the verification copy a > CUSTOMER made and is looking at — a poorer actor destroying a richer one, R-403's shape in different > clothes. The separate root also keeps the proof copy invisible to `DeleteOffsiteRestoreCopy`, the > copy listing and `OffboxFullScratchReady`, so it can never be offered for placement into a live app. > > **4. THE ALARM IS A NEW EVENT TYPE, AND REUSING `backup_integrity_failed` WOULD HAVE BEEN WRONG.** > That type is the nearest existing one and it carries a hub-side Hungarian template saying the > integrity check found an error — i.e. **the store is damaged**. Here the store is sound and the > CONTENT is missing: a different fact, a different cause, a different customer action, and telling > someone their backups are damaged when they are not is the more expensive mistake (the same asymmetry > `looksLikeRepositoryDamage` is shaped around). So `offsite_proof_empty` was minted, severity `error`, > with NO `customerMessages` entry so the controller's dynamic Hungarian survives, and **the hub half — > `allowedEventTypes` plus `operatorOnlyEvents` — ships in the same commit**, because an unallowlisted > type is answered 400 and vanishes, and a missing customerMessages entry is not a routing block. > **This widened the task's stated scope to `felhom.eu/hub/`** and the reason is recorded here rather > than left as an unexplained diff. > > **DELIBERATELY NOT DONE:** `07` §8 matrix row 4 was NOT moved. This proves the snapshot CONTAINS a > recoverable unit; it does not prove a restore puts data back into a running app. R-408 (the missing > `acquireRunning` on `RestoreOffboxScratch`) was NOT fixed — the job takes the flag itself and the row > stays open. > **2026-08-31 — v0.230.0. THREE RULINGS, recorded so none is re-litigated.** > > **1. HOLLOWNESS IS A MANIFEST QUESTION, NEVER A SIZE QUESTION.** `unitCarriesData` asks whether the > unit's manifest lists any database dump or any volume tar, and nothing else. `dirSizeBytes` lives > two files away and is the obvious wrong answer: a unit with a fat compose capture and no dumps is > exactly the shape that deleted 120 MB on `demo-hp`, and a 360-byte unit belonging to a tiny app is > perfectly healthy. Size answers *how big*; the question is *is there anything to recover*. Absent or > unparseable manifest ⇒ hollow, fail closed: a unit whose contents cannot be vouched for must never > authorise a delete of one whose contents can. `TestR403_SizeIsNeverConsulted` is the fence. > > **2. THE GUARD FENCES ONE SHAPE, NOT SHRINKING — because the derived-copy rebuild is a DESIGN > DECISION.** `07-backup-architecture.md` §8 row 5 records that the secondary is a derived copy, > rebuilt on the next run, and `tier2.go`'s own header records that a classified app's copy > legitimately shrinks as `export` drops out of its class set. `rsync -a --delete` stays, the data legs > are untouched, and complete→hollow and hollow→hollow both still mirror. The ONLY refusal is a source > carrying no data over a destination that carries some. Widening this to "the secondary never shrinks" > would be calling a decision a defect; `TestR403_DataLegShrinkIsUnaffected` is the guard on the guard. > > **3. THE REHYDRATE HAPPENS INSIDE THE RESTORE, BECAUSE A FOLLOW-UP JOB RACES THE CAPTURE.** The > hollow manifest was written **two seconds** after a Tier-2 unit restore, by the 5-minute > `backup-cache` job. A goroutine, a scheduled refresh or a "do it on the next run" would each lose > that race some of the time, and the failure mode is silent. `RestoreTier2Unit` refills the primary > before it returns, and the test asserts ORDERING rather than sleeping. > **And the capture is deliberately NOT guarded:** a capture that describes an empty drive as empty is > CORRECT. With the primary refilled there is no hollow state left to describe. Guarding the capture > would have made the manifest lie, which is the opposite of every other fix this week. > > **A defect the LIVE run caught and the unit tests did not, worth remembering as a shape.** The first > draft of `UnitRestoreDate` also flagged "the package is older than the run" by comparing their dates > — and a unit is ALWAYS captured shortly before the run that mirrors it, so it was true for every > healthy app on the box. Four apps would have been told their package was stale. **A warning that > fires on everything costs the same as the comforting lie it replaces.** The flag is now > `UnitLegPreserved` and nothing else. > > **The credential rider.** `felhom.eu/scripts/read_credential.py` is now the one place that value is > read. Three occurrences (2026-07-20, and twice on 2026-08-31, the third of which rewrote a live box's > password hash) happened while the project already had a memory file, a worked recipe and a session > report describing the mistake. **A note is read by whoever thinks to look; a check runs whether or > not anyone remembers.** > **2026-08-31 — v0.229.0. THREE RULINGS, recorded so none is re-litigated.** > > **1. The source moves; the destination does not.** `RestoreFromRecoveryUnitAt(stack, unitDir)` takes > the recovery-unit DIRECTORY, so the same restore reads a unit from the primary drive or from the > Tier-2 mirror on the second drive. What it must NEVER take is a destination: data still lands in the > live Docker volumes and the live database container, and the definition in the guest, resolved by > `GetAppDrivePath` exactly as the capture is. A restore that also relocated an app's data would be a > migration wearing a restore's label, and the customer pressed a button that said neither. > > **2. Two predicates, never one wider one — and it is the SECOND time this is written down.** > `Tier2Coverage.CanRestore()` answers *"can the additive file restore run?"* and nothing else; > `CanRestoreUnit()` answers *"can the unit restore open this copy?"*. `HasUnit` keeps its third, > distinct meaning: *"is there captured data the file restore is not looking at?"* — true even for a > half-copied mirror the unit restore refuses, because the disclosure is still owed. The temptation is > always to widen the predicate already there. **R-356 is what that costs:** one predicate meaning both > *"has this app a drive?"* and *"is this app installed?"* refused 40 running apps for months, while > they were running, with a message telling their owners to reinstall them somewhere those apps never > offer. > > **3. A destructive operation reached from a non-destructive surface must carry the difference in the > CONFIRM, not in the label.** „Teljes visszaállítás a másolatból" sits beside „Fájlok > visszaállítása" on the same row; one overwrites the app's database and internal volumes, the other > only adds files that are missing and never overwrites anything. The confirm says exactly that, names > the copy's date, and says so DIFFERENTLY when that date is only an attempt clock (R-101). It is built > from named Go constants (`tier2UnitConfirmBase` / `…DateFmt` / `…DateUnprovenFmt` / `…Contrast`) and > asserted verbatim, because a sentence assembled inside an HTML attribute cannot be pinned and R-364 > makes grepping accented Hungarian out of rendered markup unreliable on top of that. > > **The count is settled and must not be re-derived.** At catalogue `459766cb1639`, by the production > rule: **A = 7 · B = 45 · C = 1**. The C9-F1 Phase-0 count (9/43/1) was wrong by two — **radarr and > sonarr**, whose `${USERDATA_PATH}` binds are WRITABLE (so the `:ro` default rule Phase 0 applied does > not catch them) and are excluded by an explicit `class: excluded` entry instead. C is **bentopdf**. > > **R-403, filed and NOT fixed.** Two seconds after a restore that ran with the primary unit absent, the > 5-minute status refresh (`captureAllRecoveryUnits`) rewrote the primary unit from a drive with no > dumps, yielding `"db_dumps": []` / `"volume_dumps": null`. Measured on demo-hp 2026-08-31. The > dangerous half — that the next Tier-2 run would mirror that hollow unit over the good secondary copy, > `rsyncMirror` carrying `--delete` — **was not tested and is recorded as unverified.** > **2026-08-31 — v0.228.0. TWO RULINGS, recorded so neither is re-litigated.** > > **1. The off-site integrity check ships at FULL depth — `--read-data-subset=100%` — and `off` is the > way back.** Viktor's ruling, 2026-08-31, taken on a measurement rather than a claim: on `demo-hp`, > 2026-08-30, a pack damaged **without changing its size** made plain `restic check` report > `no errors were found` and exit clean; every read-data form caught it. Cost on that store > (140 829 678 B / 2 651 blobs / 67 snapshots): structure 35.0 s, 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. > > Three consequences that are decided, not open: > - **Empty means "not configured", therefore the default.** It does NOT mean off. `off` (any case) is > the off token, and it exists because a setting with no off switch is not a setting. > - **A malformed value falls back to the DEFAULT, never to structure.** Falling back to structure > would silently remove the protection on a typo — R-357's shape, a guard that opens quietly. > - **The default lives in `internal/backup/offbox_integrity.go`, NOT in `config.applyDefaults`.** Both > integrity defaults are resolved in one accessor each, beside the argument that justifies them; > symmetry with the other `Monitoring` defaults is worth less than that. > > **The thing a future session will get wrong: there is exactly ONE data point, on a 134 MB store.** > The cost curves are governed by different quantities — structure tracks the index, read-data tracks > the data — so nothing here extrapolates. That is why v0.228.0 ships a *notice* (a WARN over 5 minutes > naming R-401) and NOT a rotation schedule, a size threshold or a bandwidth budget. Every one of those > would be a number invented from one measurement, which is the shape of the four production designs > this project has already specced against nothing. **R-401's trigger is that WARN firing on any box, > not a calendar date.** > > **2. Implement or delete FIRST, register the gate SECOND.** R-400 found seven debug-page controls > with no handler. The gate that makes that impossible (`controller/scripts/debug_route_gate.py`) was > written and registered only after all seven were resolved — a registered-but-failing gate refuses > every push, exactly as `instructions_gate` established. The gate is deliberately ten lines: two lists > and a difference, in both directions, because a cleverer gate needs maintaining and an unmaintained > gate is how the class hides in the first place. **Keep `handleDebugAPI`'s exact-match switch with its > `NotFound` default** — a prefix match would have made the original defect invisible instead of merely > silent. > > **The shape worth remembering is worse than "seven dead buttons":** three of the seven fetched on > page LOAD, so those panels were permanently blank on the page an operator opens when something is > already wrong. > **2026-08-30 — v0.227.0/v0.227.1. THREE RULINGS, recorded so none is re-litigated.** > > **1. The integrity check TAKES the single-writer flag and SKIPS rather than waits.** `resticStep` > self-heals a crash lock by running `unlock --remove-all` and retrying, and its own comment records > why that is safe: every caller holds the in-process single-flight mutex, so any lock it meets is > stale. A check that did not take the flag could meet a **live** `forget --prune`'s lock from this > same box, remove it, and retry over the top of it. It skips rather than waits because waiting would > pin the nightly backup behind a check, and a skip costs nothing — due-ness makes tomorrow try again. > > **2. Due-ness, not a weekday.** A daily job asking "is the last successful check older than 7 days?" > catches up after downtime; a Sunday-gated job silently skips a week every time the box is off on a > Sunday. **R-341 is that failure**, and no `Weekly` primitive was added to the scheduler. > > **3. The result is published on `OffboxReportStatus`, NOT on `report.BackupReport`'s `IntegrityOK` / > `LastIntegrityCheck`.** R-331 retired those the day before, because the hub card rendering them read > `Integrity Unknown` for every customer forever. Giving them a live value would resurrect a card that > was deliberately removed and break `TestBackupReport_DeadFieldsStayZero`. > > **AND THE MEASUREMENT THAT MATTERS MOST, because it is the thing a future session will assume > wrongly: the structure check that ships ON does NOT catch silent corruption.** Measured on demo-hp — > a pack corrupted without changing its size returned `no errors were found`, exit 0. Only > `--read-data*` caught it. The structure check does catch missing packs, broken indexes and unreadable > snapshots, which are real; it does not re-hash pack contents. **R-399 is therefore not merely a > bandwidth question.** The cost curve is measured and small at today's store size (100% costs +12% > wall-clock over structure-only, 39.2 s vs 35.0 s on 134.3 MB) but does NOT extrapolate — the > structure check's cost tracks the index, read-data's tracks the data. > **2026-08-30 — v0.226.0. TWO RULINGS THIS SESSION MAKES, recorded so neither is re-litigated.** > > **1. A local restore's outcome is a claim about THE BACKUP, never about the app.** This is R-355 > extended from the off-site path to the Tier-1 path. On the off-site path `SafetyDump` is an honest > discriminator for "does this app have a database" — a path is returned only when a live database was > found AND dumped. **The local unit-restore path has no such discriminator at all**, so no claim about > the app is available to it. „ennek az alkalmazásnak nincs adata" and every variant is forbidden in > `unitRestoreOutcomeMsg`. It is not merely unproven but unprovable from a manifest: > 07-backup-architecture §6.3 records that an absent dump has causes that say nothing about the app — > R-361 destroyed apps' canonical `.sql` files for four months, and a restore in that window would have > "proven" a database-bearing app had none. > > **2. The destructive reconstitute uses NO headroom margin**, matching `PlaceOffsiteRestore` and > deliberately NOT `OffboxRestorePrepareFull`'s ×1.1. The ×1.1 exists because that gate is *predicting* > the size of a download it has not made. The reconstitute copies a tree that already exists on disk, > so its size is measured, not estimated, and a margin over a measured value is a refusal with no fault > behind it. Stated in a comment at the gate as well, because the two neighbouring gates disagreeing > looks like an oversight to a reader who does not know which is predicting. > > Also this session: three earlier-day releases stacked in front of this one — v0.224.0 (R-330, the > nightly backup alarming about the apps it was holding down) and v0.225.0 (R-331, `stats_known` on the > wire). **The task specifying v0.226.0's work was written against v0.223.0 and targeted v0.224.0**; the > drift was re-confirmed against live Gitea before the first edit rather than assumed. > **2026-08-23 — v0.223.0 (R-329 + R-386), and a defect that only became visible once another was fixed.** > > **[RULING] The severity a controller sends is the HUB's vocabulary: exactly > `{info, warning, error, critical}`.** Anything else is **coerced to `info` at ingest, silently**, and > `severityNotifies` drops `info` **before both** delivery legs. `app_start_failed` emitted `"warn"`. > **Measured on the live hub DB: 91 such events stored all-time, ZERO notification rows ever.** > > **[FACT] This was the SECOND occurrence, and the first one's comment had recorded the lesson.** > `DiskAlertKind.Severity` emitted `"warn"` until v0.215.0. **A comment is not a guard** — the guard is > now an AST walk over the whole controller. **grep cannot do this job:** `"warn"` is a legitimate > *healthcheck status* in `internal/monitor` and `internal/selftest`; the sweep hit nine such strings > and exactly one defect. The walk cannot follow a variable, so the **six** dynamic call sites are > registered by name with the values each can take — **an unlisted limit is not a limit, it is a hole**. > The guard found two of those six that the hand sweep had missed. > > **[FACT] It hid because another defect hid it.** R-384's ordering bug meant `app_start_failed` could > not fire at all, so a broken severity had nothing to break. **Fixing one defect made another > reachable** — and the same shape appeared again downstream: the operator cooldown key carries no app > identifier, so **only the first app-down per hour now e-mails the operator** (R-182's shape, newly > load-bearing, filed not fixed). > > **[RULING] `app_start_failed`: operator always, customer OFF by default.** `processOperator` never > consults customer preferences, so one word fixed the operator leg and left the customer leg where the > ruling wanted it. **Deliberately NOT in `operatorOnlyEvents`** — that would make the new toggle > visible, flickable and structurally incapable of delivering. > > **[RULING, R-386] "The customer stopped this" is a RECORD, never an inference.** `aggregateState` > folds `StateExited` into the stopped counter, so an out-of-band stop and a customer's Stop are > byte-identical on the Docker side — no state test can separate them. Ask `DesiredState`, which has > exactly one writer. `Stopped` → no alarm; `Running` → **alarm**; **absent → UNKNOWN, keep the old > behaviour AND announce it**, because reading absent as "nobody asked" would e-mail about every app > anyone ever stopped, fleet-wide, on the first cycle after upgrade. **The backfill cannot help — it > seeds `Running` only from an observed-UP reading.** > > **[MECHANISM] `IntentUnknown` + an INFO line naming the apps.** A rule without a mechanism is a wish. > Measured on `demo-hp`: **0 of 8** deployed apps carry an absent intent. > > **[FENCE] Adding a `DesiredState` WRITER is the fenced act; reading is fine.** And `failedRestart` > must still lift a `Stopped` intent, or F-CRIT-1 re-opens. > > **[TRAP, cost three attempts] An HTTP 200 can be a REFUSAL.** The settings save answers 200 while > rendering the empty-email wipe-guard error. Scenario G's before/after hashes matched twice because > **nothing was saved**, not because nothing changed. And the email `` spans three lines, so a > single-line grep reads it empty. **Assert the refusal banner is ABSENT before believing a save.** > > **[TRAP] A red-proof that passes may mean an INERT mutation.** `if next <= prev` → `if next < prev` > in fillwatch changes nothing, because an earlier `if next == prev { continue }` already removed the > equal case. Check the mutation applied before believing either verdict. > > **[RULING] The compound-toggle split's risk was the MIGRATION, not the split.** A save whose event > SET is unchanged now stores the existing slice **verbatim**, so byte-identity is by construction — > without that guard the defaults case reorders, and the red-proof caught it. > **2026-08-23 — v0.222.0 (R-384 + R-383), and a bigger hole found by a measurement that was told not to fix it.** > > **[DECISION] A dead SUPERVISED member is asked about BEFORE a failing healthcheck, because they are > different questions and the second was answering the first.** `aggregateState` returned > `StateUnhealthy` the moment `unhealthy > 0`, and the R-51 mixed-case block that asks "is a supervised > member dead?" sat below it. A two-container app whose database exits goes `unhealthy` seconds later > *because it cannot reach that database* — so **the symptom the fault causes was what suppressed the > alarm for it.** The supervised test is now hoisted above the unhealthy/starting/restarting returns. > > **[DECISION] "Some members are up" means ANY member not in the down bucket** — running, unhealthy, > starting or restarting. The old guard was `running > 0` counting `StateRunning` alone, which made the > R-51 block **unreachable in precisely the case it was written for**: an unhealthy survivor beside a > dead database counted as nothing being up. Either half alone leaves the defect standing, and the two > red-proofs convict independently. > > **[FENCE, unchanged] `IsDownState` is byte-identical and `unhealthy` stays excluded.** An unhealthy > container is RUNNING; folding it in reintroduces the flapping that exclusion exists to stop. **No new > state was minted** — `StateDegraded` already meant this and every consumer already handled it. The > fix is an ORDER, not a widening. > > **[FACT] The register's own suggested remedy was wrong.** R-384's row proposed a sustained-`unhealthy` > threshold on the `crashLoopAfter` model. The defect needed no threshold at all. **A register's > "recommended fix" is a hypothesis written before the diagnosis, and must be re-derived from source.** > > **[TRAP] Three existing subtests pinned the DEFECT as settled behaviour.** > `TestAggregateState_UnchangedBranches` asserted an unhealthy/starting/restarting member beat an > `exited` peer that was on `unless-stopped`. They were amended (down member given a benign policy, > which is the only case where that sentence was ever true) and the change is reported, not buried. > **A green suite can be green about the wrong thing.** > > **[FINDING — R-386, OPEN, NOT FIXED] A single-container app stopped out of band raises NO alarm, and > a comment states the opposite.** `aggregateState` folds `StateExited` into the `stopped` counter, so > an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**; > `classifyRunStates` then whitelists `stopped` as a deliberate user stop. The comment at > `cmd/controller/main.go` claiming an out-of-band `docker compose stop` "still alerts" is **measured > false** — `privatebin`, 9 scans, 0 events, 0 banner. **Case #10 of "a comment asserting an invariant > the code does not provide".** The task asked for this as a MEASUREMENT and forbade a fix; it is filed. > > **[DECISION, R-383] A message may not assert a file exists without asking the disk.** The > double-failure sentence named the undo copy as present, built from the returned path — and a missing > file is one of the two ways that rollback fails. `undoCopyPhrase` now reads from disk; a zero-length > dump counts as MISSING; and the absent case still names WHERE the file should have been, because > R-351's lesson is that a refusal naming nothing forces someone to remember what the product knows. > > **[DECISION] The alarm ladder now has an owning document** — > `felhom.eu/documentation/architecture/08-alarm-ladder.md`. Until 2026-08-23 no document owned it; the > rules lived as comments in four packages, each locally correct, with the ordering between them legible > only by reading one function top to bottom. **That absence is why R-384 survived review.** > > **[GOTCHA] `app_start_failed` still ships severity `warn` (R-329)**, which is not in the hub's > vocabulary and coerces silently to `info`, e-mailing nobody, while the POST returns 200. R-384 moved > this event from unreachable to load-bearing, so the severity bug now matters. > **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.** > > **[DECISION] `db_dumps` lists the app's OWN dumps, not the `pre-restore-*` undo copies.** They are > local material for a restore that went wrong, not part of the app's recovery set. **Every consumer > of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`** (the > declaration, the enumeration, the change-detection compare); none reads it for recovery, and no hub > or agent consumer exists. Three copies per app were being pushed off-site permanently for no > recovery value. **The files are neither deleted nor hidden** — their visibility is a separate > recorded decision and it stands. > > **[TRAP, and it bit within minutes] A stable `db_dumps` lets `CaptureRecoveryUnit`'s already-current > early return fire.** Anything that must happen on EVERY capture — bounding the undo copies — has to > sit ABOVE that check. It did not, and the cap silently stopped applying: four copies against a cap > of three, counted on the box. Fixed in v0.221.1. **One change made another unreachable, and only > counting files on a real machine showed it.** > > **[FACT] The comment was the defect.** `writeSafetyDump` called `DumpOne` into the app's own unit > and renamed afterwards; `DumpOne` writes the canonical `-.sql`, so every safety dump > destroyed the app's real backup. The comment said the rename meant it "can never overwrite the app's > real dump" — false as written, for four months. The fix is a DESTINATION (`DumpOneTo`), not a > rename, and the `.tmp` derives from the final path so a nightly dump beside it cannot collide. > **`DumpOne`'s signature did not move.** > > **[NEGATIVE — do not re-derive this] A HELD app does NOT raise the dead-app alarm.** It was read > from source that it would, because it keeps its database container and so is not `StateStopped`. > Measured on the shipped v0.220.2: it aggregates to `unhealthy`, `aggregateState` checks > `unhealthy > 0` before the mixed-case degraded branch, and `IsDownState` excludes `unhealthy`. > Heartbeat read `0 currently down` throughout. **No suppression was built.** The same measurement > exposed **R-384**: an app whose database has died is `unhealthy` too, and is likewise silent. > > **Proven live on `demo-hp`:** the canonical dump's sha256 unchanged across a restore on both engines > — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`. Evidence: > `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`. > **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).** > > **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the > customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A > running app on a half-written database lets the customer type into it, and that turns a recoverable > state into a permanent one. The alternative — start it and mark it — was put to the operator and > declined. If that judgement is ever revisited, this is the sentence to revisit. > > **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored > database; the only difference was whether it looked broken (Postgres emptied and crash-looping, > MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not > transactional. Putting the customer's own copy back is what removes the half state, and it is the > same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps. > > **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such > files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older > state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a > two-database app would have restored one and left the other half-written. > > **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found > by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as > `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec` > against the dead id timed out — so the app was held for an infrastructure reason while its data was > recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and > never look at container identity.** > > **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure > into it makes our fault indistinguishable from their choice. It is not the app-stop marker either — > that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start > the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared > `driveStartGate` **above** its driveless early return, because the apps this exists for have no > drive. > > **The way out is `--clear-restore-hold `, and it REQUIRES A CONTROLLER RESTART** — it runs as a > second process and the running controller keeps its in-memory settings. v0.220.2 makes the command > say so. Clearing through the running controller is the right shape later; it needs an operator tier > the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own > hold is what the hold prevents). > > **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state > (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was > measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`. > **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.** > > **[DESIGN] The restore destination is resolved by the SAME rule as the capture destination.** The > drive if the app declares one, the system data path otherwise — `Manager.GetAppDrivePath`, one > expression, now used by `CaptureRecoveryUnit`, `ReconstituteFromOffsite` and `PlaceOffsiteRestore` > alike. Anything else and the restore aims somewhere the backup never came from, which surfaces as a > placement-mismatch prompt on a box where nothing actually moved. > > **The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351) > applies to apps that HAVE a drive to get wrong.** It used to be reached by `HDD_PATH == ""`, which > also stood in for "is this app installed?". Measured in the catalogue at `459766cb1639`: **53 > templates, 13 `needs_hdd: true`, 40 `false`** — so for 40 apps that test was permanently true and the > off-site restore refused them forever, while they were running, telling the customer to reinstall > them "in the same place", which those apps never offer. **An app with no drive is not misconfigured** > (`01-topology-and-trust.md` §8, `[DESIGN]`); it is the majority case, and between 19 and 22 August it > was called a defect four times. > > **Now:** *installed?* is asked of `ListDeployedStacks()` via `Manager.isStackDeployed`, which **fails > CLOSED on a nil provider** — "cannot tell" must not become "go ahead" when the next act is a write. > *Where?* is asked of `GetAppDrivePath`. A third refusal, with its own sentence and its own route, > covers installed-but-no-resolvable-data-root: widening `nincs telepítve` to cover that would send a > customer to reinstall a running app and hide the real fault. > > **FENCED, and not changed:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`). > Capture resolves an app's declared `userdata`/`import` file legs against that value; a system-data > fallback there would write a snapshot claiming to hold the customer's files and not holding them. The > fenced ACT is "introduce a fallback into capture-side path resolution" — reading the value elsewhere > is fine. > > **Proven live on `demo-hp`, 2026-08-22.** `privatebin` (driveless): planted through the app's own > HTTP API plus a direct file plant, off-sited, **deleted**, restored through > `POST /backup/offbox/reconstitute` — **15/15 files back byte for byte**, two Hungarian accented names > included, message „0 fájl és 1 adatkötet visszaállítva". The scratch was unit-only, exactly the shape > nobody had ever driven to completion before, and every downstream leg held. `calibre-web` (drive > app) walked the same way and did not move. Evidence: > `felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. > **2026-08-22 — v0.218.0 (R-354/R-355). The database nobody backed up, and the restore that > returned most apps nothing.** > > > Both fixes came out of the 2026-08-21 backup-truth drill. **R-355 went first because it is the only > place in the product where one customer action causes permanent total loss:** paperless-ngx's database > was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the > off-site copy or the restore — and the same misattribution meant a destructive restore of that app took > NO undo copy, then told the customer the app has no database. Fixed by reading the compose project > label, which is the stack name by construction. **One app of 53 affected**, established with a sweep > proven able to convict by planting a second mismatch. > > **R-354:** the off-site restore had no named-volume leg at all. The archives live inside the unit, whose > placement is correctly skipped, and the comment beside that skip said the dump is replayed from the > scratch "so nothing is lost" — true of the database, false of the volumes. `restoreDockerVolumesFrom` > now replays them from the scratch unit; `VolumesReplayed` reaches the message. > > **WHAT THIS DOES NOT FIX, and it is the blocking item for the apps that need it most:** the off-site > restore still REFUSES outright for the 40 of 53 apps that declare no data drive (**R-356**), saying a > running app „nincs telepítve". Those are exactly the apps whose entire dataset is a named volume, so > R-354's fix cannot reach them until R-356 is closed. Proven again on hardware 2026-08-21. The live > confirmation of R-354 was therefore done on `calibre-web` and `paperless-ngx`, which declare a drive > and can reach the restore. > > Also still open from the drill: the empty-restore success message (R-353), where the 40 apps' data > lives (R-352), and the remaining rows R-357..R-366. > > ## THE RESTORE'S OWN MEMORY (v0.217.0, 2026-08-21) — R-351 / R-352 / R-353 > > > **Both sides of a written fact must be checked, not just the writing side.** Every recovery unit > > has recorded `drive` and `namespace_root` since schema 1. **No non-test code in the repository ever > > read either back.** A restore into a different destination than the backup recorded therefore > > succeeded silently under a green message. This is the same shape as several defects closed this > > month, and the cheap test for it is one grep: *who reads this field?* > > > > **`IsRunning()` is still the wrong flag, in one more place than we knew.** v0.154.0 fixed the wizard > > and left a comment explaining why. The **seven handlers** were never moved over, so a second press > > genuinely started a second run and reported „…elindult". A comment explaining a trap does not fix > > the other call sites — grep for them. > > > > **A result nobody can see is the same defect as no result.** The banner gated its terminal state on > > a page-local `sawRunning`. The 8.666 s OpenGist restore finished before any poll saw it, so no > > screen said it had completed. Fixed with `RestoreOpStatus.LastRecent` — and the window now lives in > > `internal/backup` as ONE expression that both surfaces read. > > **State, and what is next.** > > - **Shipped:** placement comparison + named mismatch + `ack_placement`; not-installed refusal names > the recorded drive; deploy prefill from the app's own backup; `restoreOpBlocked()`; `LastRecent`; > off-site listing bounded-concurrent (measured 16.1 s → two waves). > - **R-352 partly closed.** 40 of 53 catalogue templates declare no data path, and > `GetDefaultStoragePath()` is read by nothing that places data — its comment `// new apps use this by > default` has never been true. **Only visibility shipped**; the deploy page now states where the data > will live. **No placement changed, nothing migrated.** Specification: > `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`. > - **R-353 is the next session's first item.** A restore whose unit carries no `db_dumps` and no > `volume_dumps` reports a bare completion. OpenGist's unit held configuration and nothing else, and > the restore said only that it had finished. Fix the outcome first; *then* prove the off-site > coverage of a named-volume app by running a dump cycle — do not close the first on the second. > - **Open and unproven:** whether the 40-class reaches the off-site tier at all has **not been > observed**. `runVolumeDumps` covers them on paper; every unit on the box read `volume_dumps: None` > because no nightly run had happened yet. > > --- > > ## THE TWO RULES THE RECOVERY JOURNEY LEANS ON (v0.203.0, 2026-08-06) > > > **1. A credential the hub stages is collected by the box, not waited for.** The reconcile that > > collects runs on a tick for exactly as long as the box's own declaration says it needs one — and > > stops the instant a target exists. It is driven from `OffboxReportStatus().State`, the same statement > > the hub acts on, so the two can never disagree about whether a retry is wanted. > > > > **2. A mount Felhom itself made is not "something else".** Enrolment mounts a drive twice — the > > managed path and a raw `/mnt/` on the host — and the host survives a guest rebuild while the > > guest's registry does not. The claimed check forgives a non-managed mount **only when corroborated** > > by the same device also being mounted under the managed path. **A genuinely foreign mount is still > > refused, and that fence has its own test.** > > **Why both are stated here rather than left in the code:** each was a dead end that kept the unaided > recovery journey failing, and each looked correct in isolation. R-218's declaration half shipped and > worked while nothing consumed what it asked for; R-220's check was right about foreign disks and wrong > about our own. **Neither is a bug in the thing it guards — both are about what runs, and when.** > > Two things that must not be "simplified" back: > - **The settle gate stays.** The retry goes through `ReconcileWhenSettled`, so the day-0 floor race is > unchanged. A retry that skipped it would trade one defect for another. > - **The R-220 exemption is corroborated, never a prefix.** Widening it to any `/mnt/*` path offers a > disk another system is using for formatting — the red-proof shows exactly that. > > ## THE UNLOCK PATH'S RULE (v0.202.0, 2026-08-06) — state it before changing anything there > > > **On the recovery unlock path the customer is blamed only after a real attempt REFUSED their code. > > Every other outcome — including one that cannot be classified — says something else.** > > This is the rule, and it outlives the bug that produced it. It was learned twice, because fixing it > once was not enough: > > - **v0.201.0** stopped an agent that is too OLD from being reported as a wrong code (R-216). > - **v0.202.0** found the same defect through a different door: an agent that is **stopped**, and a hub > that cannot be **reached**, still fell through to a message about the code. Measured with a > **correct** code at 0.0299 s and 0.0556 s, against ~1.0 s for a real unseal — the machine accused the > customer of something it had not tried (R-224). > - And the inverse: the one message that says *"check your ten words"* was unreachable on any box that > had re-escrowed, which is exactly the box a customer has just recovered (R-226). > > **How it is enforced.** `agentapi.ClassifyRecoveryFailure` maps the failure to one of five classes > **from the value, never the text**; the typing message is reachable from **one** of them > (`RecoveryAskedAndRefused`, i.e. HTTP 400, i.e. the bundle was fetched and `age` refused it); and the > zero value is `RecoveryUnknown`, which renders **neutral**. **The safe default is the load-bearing > part** — an unrecognised status must not fall into an accusation. > > **Two things that are deliberately NOT how it works, and must not be "fixed" into it:** > > 1. **Elapsed time is never a classifier.** It is what diagnosed this, it is logged for the operator, > and that is all. A duration guard would be a second thing that can be wrong. > 2. **The error's TEXT is never read.** A string match is a defect waiting for a rewording. When the > distinction was not available as a value, the **agent was changed to provide one** > (`escrow.ErrBundleFetch` → HTTP 502, agent v0.126.0, `MinAgent 0.126.0`) rather than parsed for. > > **The coupling degrades safely and silently:** an agent below 0.126.0 answers 400 for both causes, so > `FeatureRecoveryFailureClass` withholds the refusal reading and the 400 becomes neutral. The gate > blocks nothing; it only decides whether the customer may be told to check their typing. > > ## CAMPAIGN 11 — what changed in v0.201.0 (2026-08-05) > > **The off-site key recovery is a COUPLED feature and now declares it.** It needs agent **0.125.0** > (`POST /escrow/recover-offsite-password`). `FeatureOffsiteKeyRecovery` has a `featureProbes` row, a > `featureMinAgent` row and a `Supports` gate at the unlock entry point. > > ⚠ **That gate FAILS CLOSED — alone in that table.** The package default is fail-open, and that default > is what produced R-216: an agent that could not answer 404'd, the unlock was attempted anyway, and the > customer was told their correct recovery code was wrong. Anything but `SupportYes` now says *the > machine* cannot ask yet, and no attempt is made. Do not "fix" it back to the package default. > > **The box declares `needs_credential` until the TIER WORKS, not until a key exists.** The old > short-circuit on "a repository password is present" is deleted: installing one is the recovery > screen's whole job, so it made succeeding at recovery switch off the mechanism that delivers the > coordinates to use it. A disabled target still short-circuits at the first line (Scenario E). > > **The unlock finishes the job**: place the key → bring the tier up (`offsiteapply.Bridge.Reconcile`, > wired via `SetRecoveryTierUp`) → list. Without the middle step the promised listing can never render on > shape (a), because no key ⇒ no target ⇒ no inventory. > > **Four messages, not one.** Wrong code (the only one mentioning typing) · the machine cannot ask · > the store could not be read / the connection details have not arrived · the code belongs to a RETAINED > earlier package. The last one is driven by the ACK's `superseded_present`/`superseded_at` (hub > v0.97.0) and **promises nothing** — no read path for a superseded package exists. > > **Still open from the campaign:** R-214 (console pairing banner), R-220 (drives unenrollable after a > rebuild), R-221 (a rebuilt box cannot run the escrow ceremony), R-223 (the Day-0 manifest still vouches > agent 0.120.0 — operator decision). > > ## R-203 (v0.197.0) — the namespace-root contract, and what `ok` now means > > **The contract, in one line:** `appbackup`'s path helpers (`UserdataDir`, `PrimaryBackupPath`, > `RecoveryUnitPath`, `AppDataDir`) take a **NAMESPACE ROOT**. Anything that came out of `HDD_PATH` or a > `StoragePath` is a **DRIVE path** — put it through `appbackup.NamespaceRootFor(drive, systemDataPath)` > first. `UserdataDir(bareDrivePath)` still compiles and is still wrong; five callers proved it. > > **The rule now has ONE expression.** `NamespaceRootFor` / `IsEnrolledDrive` in `appbackup`; > `backup.Manager.namespaceRoot` and `stacks.Manager.inGuest` delegate. There were two copies before and > **they differed** — one compared without `filepath.Clean`, the other with it. > > **Why it was invisible:** on an enrolled drive the namespace root IS the drive path. The two diverge > only on the system-data fallback, which `paths.go:26` names as a supported arrangement. > > **`last_status` gains `incomplete`.** A run that could not capture a directory an app declares > MANDATORY is not a successful run. **Not `error`** — the rest of the run worked, so `SnapshotCount` > and `LastSuccess` still record what WAS captured. It reaches the operator via the existing > `backup_run_failures` digest (a new event type is a two-repo change; the hub drops unlisted types). > The Hungarian customer warning is unchanged; the page renders `! Hiányos`. > > **Still open, and NOT fixed here:** `resolveAbs` resolves `RootHDD` and `RootUserdata` against the same > root. Both callers now pass the namespace root so the export and the backup agree with each other, but > whether `${HDD_PATH}` should mean the namespace root on the system drive touches every deployed app's > binds and needs a decision, not a patch. > > **Blast radius, measured before changing anything:** exactly one app in the fleet had > `HDD_PATH == system_data_path` (`calibre-web` on demo-hp, the R-201 drill fixture). Its data was > migrated and its sentinel re-verified byte-identical. > > ## R-203 (2026-08-04) — a MANDATORY userdata directory can be absent from the off-site snapshot while the run says `ok` > > Found live on demo-hp while staging the R-201 drill, and it **halted that drill**. > > `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment > **when the drive IS the system data path** — `m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath != > m.systemDataPath)` (`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is > `/userdata` (`stacks/classify_binds.go:14`). > > With `system_data_path: /mnt/sys_drive` and `calibre-web` deployed at `HDD_PATH=/mnt/sys_drive`: > > live bind (files land here): /mnt/sys_drive/userdata/media/books ← exists > capture set looked for: /mnt/sys_drive/felhom-data/userdata/media/books ← does not > > **The same compose used BOTH roots** — `${IMPORT_PATH}` resolved *with* the segment, > `${USERDATA_PATH}` *without*. The run logged one `[WARN] mandatory data path missing on disk, skipped > from offsite`, then `0 mandatory path(s)` and **`backup OK: 3 app(s), 3 snapshot(s)`**, with > `last_status: ok`. Nothing customer-visible or hub-visible said the directory was dropped. > > **Not established:** whether `HDD_PATH == system_data_path` is a supported deploy. It was accepted > (HTTP 202) one call after the NAS path was correctly refused (R-108). **Either branch is a defect** — > broken resolution, or a missing refusal. > > **Two things a fix must do:** make the two roots one function, and make a skipped **MANDATORY** path a > customer/hub-visible signal rather than a container-log WARN. `opengist`/`privatebin` declare no > mandatory userdata paths and are unaffected. > > ## R-200 (v0.195.0) — the offsite key recovery diagnostic > > `--recover-offsite-check` is a `docker exec` escape hatch (the `--print-reset-code` shape), NOT a page > or an API a browser can reach. R comes from **STDIN** — never argv, never `ps`, never shell history, > never a transcript. It asks the agent (>= v0.125.0) to fetch this host's sealed bundle and open it, > then reports whether the recovered repository password matches the on-disk one **by sha256**. > > docker exec -i felhom-controller /usr/local/bin/felhom-controller --recover-offsite-check < /path/to/code > > **IT COMPARES AND NEVER INSTALLS.** `CheckOffsiteKeyRecoverable` must stay free of any write — if a > future change makes it place the recovered password, it stops being a diagnostic and needs the drill's > supervision (that is link 9, R-200's remaining half). Pinned by > `TestCheckOffsiteKeyRecoverable_WritesNothing`, whose red-proof is adding the install call. > > Exit codes are load-bearing: **0** match, **2** a clean MISMATCH, **1** a step failed. A mismatch is a > finding about the system; a failure is a finding about the run, and they must never share a status. > > **Proven live on demo-felhom 2026-08-04** — recovered sha256 == on-disk sha256 == the hub's stored > hash. Nothing customer-facing ships with it: no card, no form, no preview. > > ## About Viktor (project owner) > > - Works at Deutsche Telekom (Budapest), building Felhom.eu as a side business > - Felhom.eu: managed home-server service for Hungarian households > - Technical but prefers pragmatic solutions over over-engineering > - Runs all infrastructure on Gitea (gitea.dooplex.hu), k3s cluster for management > - Customer deployments use Docker Compose (not Kubernetes) for simplicity > > ### felhom-controller (this repo) > - **Version:** v0.16.1 > - **Phase 1:** ✅ COMPLETE — Stack Manager + Deploy Flow > - **Phase 2:** ✅ COMPLETE — Monitoring & Health (scheduler, CPU/temp, healthchecks.io pings) > - **Phase 3:** ✅ COMPLETE — Backups (DB dumps, restic integration, manual trigger, **dedicated backup page**) > - **Phase 4:** ✅ COMPLETE — Monitoring Page with Metrics Store (SQLite, Chart.js, system + container metrics) > - **Phase 5:** ✅ COMPLETE — Authentication, Persistence & Settings Page (settings.json, password change, session management) > - **Phase 6:** ✅ COMPLETE — Monitoring Warnings, Dashboard Alerts & Notification System > - **Phase 7:** ✅ COMPLETE — Storage Overview, Per-App Backup Toggles & Limited Restore > - **Phase A:** ✅ COMPLETE — Storage Paths Foundation (registry, auto-discovery, per-app HDD_PATH, deploy dropdown, health monitoring) > - **Phase B:** ✅ COMPLETE — Storage Management UI Polish & Health Severity Fix (flash messages, label editing, app details, FS info, deploy free space, backup context) > - **Phase C:** ✅ COMPLETE — Storage Init Wizard, Data Migration & Startup Fix (disk scan/format/mount wizard, rsync-based migration, startup pings) > - **v0.11.1 bugfix:** ✅ COMPLETE — Storage Scan: system disk detection via host fstab + blkid UUID resolution; FSType enrichment via `blkid -o export` > - **v0.11.2 bugfix:** ✅ COMPLETE — /host-dev mount for block device access; `HostDevicePath()` helper; all format/scan/safety ops use /host-dev > - **v0.11.3 bugfix:** ✅ COMPLETE — Added `fdisk` package to Dockerfile (provides `sfdisk`; not in `util-linux` on Debian bookworm) > - **v0.11.4 bugfix:** ✅ COMPLETE — FormatAndMount: fixed sfdisk (wipefs+force+`,,`), mount (explicit device path), mount propagation (rshared), ASCII label, smart partition skip, findmnt verification > - **v0.11.6:** ✅ COMPLETE — FileBrowser auto-mount sync (`syncFileBrowserMounts()`) + 3 UI fixes (badge color, progress bar, button text) > - **v0.11.7:** ✅ COMPLETE — Stale data cleanup + FileBrowser sync after migration + deploy page title fix > - **v0.11.8:** ✅ COMPLETE — Per-App Cross-Drive Backup (3-2-1 rule): rsync/restic to secondary drive, deploy page UI, backup page summary, scheduler jobs, API endpoints > - **v0.11.9:** ✅ COMPLETE — UI Polish Fixes: spacing, tooltip on "Módszer", status dot instead of disabled checkbox, progressive disclosure, emoji cleanup > - **First app deployed:** Paperless-ngx on demo-felhom.eu (2026-02-13) > - **Running on:** demo-felhom (N100 mini PC) at 192.168.0.162:8080, felhotest (Proxmox VM) at router.abonet.hu:33022 > - **All Phase 1-5 features working:** deploy, start/stop/restart/update, logs, health-aware states, auth, monitoring, backups, backup detail page, system monitoring page, settings page > > ## Architecture decisions > > | Decision | Rationale | > |----------|-----------| > | Go stdlib for web (no Gin/Echo) | Minimal dependencies, single binary, easy to embed templates | > | Templates as go:embed HTML/CSS files | Zero runtime file dependencies (compiled into binary), but each template is a separate editable file | > | Docker Compose for customers (not k8s) | Simpler troubleshooting, customers don't need k8s knowledge | > | k3s for management infra only | Viktor's own services (gitea, monitoring, website) run on k3s | > | Cloudflare Tunnel for remote access | No port forwarding needed, works behind any NAT | > | app.yaml per stack | Separates deploy config from compose files, survives git pulls | > | Password fields require explicit input | Prevents accidental empty-password deployments | > | Health-aware state from Docker Status field | Docker's State says "running" even for unhealthy containers | > | Memory limits via deploy.resources.limits | Prevents runaway containers; ~50% headroom over expected usage | > | System info from /proc/meminfo + statfs | No external dependencies, cheap to read on each page load | > | mem_request vs mem_limit (K8s-inspired) | Requests = expected usage (hard block), limits = peak (overcommit OK) | > | 384MB reserved for system | Prevents deploying apps that would starve the OS/controller | > | Logo SVG embedded as Go constant | Same approach as CSS/HTML — zero external file deps | > | Git sync via os/exec git CLI | No Go git library needed, git is in the container image | > | SHA-256 for content comparison | Only copy changed files, avoid unnecessary disk writes | > | 30s debounce on manual sync | Prevents spamming the git server | > | Orphan = deployed but not in catalog | Safe lifecycle: remove from catalog → mark orphaned → user deletes via UI | > | FileBrowser as infra (not catalog) | Needed even after apps deleted (user browses HDD data); deployed by setup script | > | Protected HDD paths | Safety net: never delete top-level HDD dirs (media, storage, Dokumentumok, appdata) | > | Central scheduler (not ad-hoc goroutines) | Single place to register/monitor all periodic tasks, graceful shutdown, skip-if-running | > | CPU sampling via background goroutine | /proc/stat delta needs two readings — collector runs every 5s, GetInfo() reads cached value | > | Temperature from /host/sys (Docker mount) | Container can't read host /sys directly — mount /sys:/host/sys:ro, try /host/sys first | > | Restic password auto-generated | No manual setup needed — generated on first backup run, stored in named volume | > | DB discovery via docker inspect | No config needed — discovers postgres/mariadb containers by image name + env vars | > | Backup orchestrator with running flag | Prevents concurrent backups, supports both scheduled and manual trigger | > | modernc.org/sqlite (pure Go) | No CGO/gcc needed in Docker build stage — keeps `CGO_ENABLED=0` static binary | > | AlertManager state-based refresh | Alerts regenerated every 5min from health report — no persistent storage needed, always reflects current state | > | Notification relay via hub | Controller → hub → Resend → email. Hub acts as central relay: knows customer email, handles Resend API. Controller only needs hub URL + API key | > | In-memory notification cooldowns | Per-event-type cooldown map (default 6h). Lost on restart = acceptable (better to re-notify than miss). No persistence needed | > | Health status change detection | Only notify on degradation (ok→warn, ok→fail, warn→fail). Avoids spam on flapping. First run records baseline, doesn't notify | > | Resend HTTP API (no SMTP) | Direct POST to api.resend.com — same pattern as website contact-mailer. Simpler than SMTP setup, good deliverability | > | Preferences sync on save + startup | Controller pushes prefs to hub (not pull). Startup sync handles hub DB rebuild. Local save always succeeds even if sync fails | > | Chart.js embedded locally | Customer hardware may not have internet — CDN not reliable for offline environments | > | StackDataProvider interface | backup package needs stack data but can't import stacks (circular). Interface in backup, thin adapter in main.go | > | Password sync to hub via report | Restic password in Docker named volume on SSD. Hub sync provides redundancy for disaster recovery | > | App backup via HDD mounts only | Docker volumes at /var/lib/docker/volumes/ not mounted in controller. HDD data is the important user data; DB in volumes covered by nightly dump | > | Restore uses running mutex | Prevents concurrent backup+restore on same restic repo. Reuses existing `m.running` flag | > | Storage paths registry in settings.json | Multi-storage support: each app's HDD_PATH from app.yaml is authoritative. Auto-discovery on startup avoids manual config. Registry enables UI management + health monitoring per path | > | /mnt:/mnt:rw mount in controller | Replaces per-path HDD_PATH mount. Enables multi-storage + restore writes. All customer HDD mounts are under /mnt/ by convention | > | Per-app HDD_PATH resolution (app.yaml > global) | App's own env HDD_PATH is Priority 1, registered storage paths as fallback. Eliminates dependency on global controller.yaml hdd_path | > | Mount-point detection via syscall.Stat_t.Dev | Compares device ID of path vs parent dir — reliable check that path is on separate filesystem. Prevents data writes to SSD | > | Health severity: mount-point = warning | Non-mount-point is informational, not a service failure. FAIL reserved for genuinely broken things. Avoids false alarms on demo/test environments | > | FS info via findmnt + sysfs | `findmnt -n -o SOURCE,FSTYPE --target ` for filesystem type/device. `/sys/block//device/model` for disk model. Best-effort, returns nil on failure | > | Query param flash messages | Stateless, no session store needed. Consistent with backup page pattern. `?storage_msg=success&storage_detail=...` | > | StorageLabels map on stacks page | Separate map passed to template (not modifying Stack struct). Built from deployed apps' HDD_PATH → registered path label lookup | > | Metrics downsampling via SQL | Bucket-based AVG in GROUP BY keeps Chart.js responsive with up to 30 days of data | > | 60s metrics collection interval | Good balance of resolution vs. storage — ~44K rows/month for system metrics | > | /etc/os-release mounted read-only | Container can't read host OS info directly — mount to /host/etc/os-release:ro | > > ## Key file locations on demo-felhom > > ``` > /opt/docker/felhom-controller/ # Controller compose + config > ├── controller.yaml # Customer config (domain, auth, paths) > ├── docker-compose.yml # Controller's own compose > └── data/ # Controller persistent data (named volume) > > /opt/docker/stacks/ # All app stacks > ├── traefik/ # Reverse proxy (protected) > ├── cloudflared/ # Tunnel (protected) > ├── paperless-ngx/ # First deployed app ✅ > │ ├── docker-compose.yml > │ ├── .felhom.yml # App metadata > │ └── app.yaml # Deploy config (env vars, locked fields) > └── whoami/ # Test stack (not deployed) > > /mnt/hdd_placeholder/storage/ # HDD storage for apps > └── paperless/ > ├── consume/ # Drop files here for OCR > ├── media/ # Processed documents > └── export/ # Backup exports > ``` > > ## Related repositories and their state > > | Repository | Status | Notes | > |------------|--------|-------| > | felhom-controller | Active | This repo. Controller code + deploy scripts | > | app-catalog-felhom.eu | Active | 10 app templates, all with .felhom.yml metadata + memory limits | > | felhom.eu | Active | Website + hub/ subfolder (felhom-hub service) + k8s manifests | > | homelab-manifests | Stable | k3s cluster running (dooplex.hu services) | > | misc-scripts | Utility | collect-repo.sh, backup helpers | > > ## Gotchas & lessons learned > > - `docker compose restart` ≠ `docker compose up -d` — restart doesn't pick up new images > - Go maps have random iteration order — always sort slices before displaying > - Docker `.State`="running" doesn't mean healthy — check `.Status` for "(health: starting)" / "(unhealthy)" > - Paperless-ngx needs `PAPERLESS_OCR_LANGUAGES` (plural) to install language packs, `PAPERLESS_OCR_LANGUAGE` (singular) to select > - In-memory Deployed flag must be set BEFORE `docker compose up -d` (not after) — compose can take 30-60s for image pulls, during which the UI would show a stale "Telepítés" button > - Cloudflare Tunnel handles *.demo-felhom.eu → Traefik handles Host()-based routing to containers > - BIOS "AC Power Recovery" must be enabled on N100 for auto-restart after power outage > - `docker compose up -d` returns exit 0 even when containers immediately crash-loop — need post-start status check to detect this > - When logging env vars for debugging, only log keys (not values) to avoid leaking secrets in log files > - Mealie image (`ghcr.io/mealie-recipes/mealie`) doesn't include wget/curl — use Python TCP socket check for healthcheck > - Mealie DB migrations on first start take ~40s (alembic) — use `start_period: 60s` to avoid premature unhealthy status > - Alpine-based images (filebrowser, vaultwarden) have wget via BusyBox — healthchecks with `wget --spider` work fine > - Deploy `sed` command to update image version must target only the `image:` line — naive `sed 's|name:OLD|name:NEW|'` also matches the service name line (e.g., `felhom-controller:` → `felhom-controller:0.2.12`), breaking YAML. Use `sudo sed -i 's|image:.*felhom-controller:[^ ]*|image: ...felhom-controller:NEW|'` or similar scoped pattern > - Hungarian quotation marks `„"` in YAML: `„` (U+201E) is safe inside YAML double-quoted strings, but the closing `"` must NOT be ASCII `"` (0x22) — it terminates the YAML string. Use `\"` escape or Unicode `"` (U+201D). This caused a silent parse failure for the entire `.felhom.yml` file > - Never silently swallow parse errors — always log them. Silent failures make debugging impossible (took a dedicated debug session to find a simple quoting issue) > **2026-08-14 — v0.215.0 (R-328..R-333). Disk health, phase 1: the alert that reached nobody.** > > ### A severity string is a WIRE CONTRACT with the hub, not a label we choose. > > The hub accepts exactly `{info, warning, error, critical}` and **silently coerces anything else to > `info`**, which `severityNotifies` then drops. `disk_health_degraded` shipped `"warn"` — one letter > short of the contract — so **every Figyelmeztetés-level disk alert this product ever produced was > emailed to nobody, on both legs.** Proven live side by side on 2026-08-14: `"warning"` → > `notification_log` status **`sent`**; `"warn"` → stored `info`, **no row at all**. > `app_start_failed` (`notifier.go` ~L546) carries the identical defect and was deliberately NOT > changed here — it needs its own decision on whether it should notify (**R-329**). > > ### A drive's own PASSED verdict cannot fail on bad sectors. Do not build on it. > > Attributes 187/197/198 all carry `thresh: 0`; a normalized SMART value floors at 1 and can never drop > to or below the threshold. The real drive (ST3000VX010, S/N Z6A07P2G) read `PASSED` at **352** pending > sectors and **1001** reported-uncorrectable reads. Felhom already read the raw counters, which is the > only reason it would have noticed at all. Evidence + fixtures: > `felhom.eu/documentation/audits/DIAG-smart-passed-trap-2026-08-14.md`. > > **DECISIONS MADE, so they are not re-litigated:** > > - **Predicted failure is labelled „Hiba" — there is no fourth verdict word.** A fourth Hungarian word > sharing a root with „Figyelmeztetés" would make the MORE severe state read as the milder one. Four > labels, final: Rendben / Figyelmeztetés / Hiba / Nincs adat. > - **Sustain is the primary rule; the count is the backstop.** Truth-table row 6 (unreadable sectors > present again at the next check) sits ABOVE row 8 (count >= 64) because on the real drive sustain > fires 12 Aug and the count not until 13 Aug. Row 8 exists only for a box powered off across the > sustain window. > - **The numbers and where they come from.** 64: the benign excursion peaked at 16 and cleared inside > an hour; the terminal run passed 64 at 13 Aug 11:28 and never returned. **It is a judgement from ONE > drive** — a static backstop, expected to be replaced by growth-rate detection in Phase 3. 55/60 °C: > the operator's existing Prometheus bands on DooPlex, adopted unchanged so the two systems cannot > disagree about the same drive. **These bands are SPINNING-DISK bands and are questionable for NVMe** > — demo-hp's healthy Toshiba NVMe idles at **53 °C**, 2 °C below Figyelmeztetés (**R-333**). > - **Phase 2 owns the new SMART attributes (187 Reported_Uncorrect, 199, 188).** They are a declared > wire change, so under the G-1 gate the hub must model them in the same session. Putting them here > would have turned a one-word severity fix into a three-repo change (**R-330**). Everything v0.215.0 > needs was already on the wire. > - **Phase 1 state is one small record per disk, NOT a sample series.** `metrics.MetricsStore` is the > right home for Phase 2/3 history; using it now would have put a schema migration on the critical > path of the severity fix. > > **The trap this change nearly shipped, caught by a test and not by review:** the card and the check > share one verdict function so the chip and the email can never disagree — but the check CONSUMES the > prior and then overwrites it, so a card rendering afterwards read its own check's write and showed one > level MORE severe than the alert. Fixed by `diskRecord.PriorSawUncorrectable`, which replays the prior > that produced the stored verdict. The guarantee was previously asserted in a comment only. > > **Cadence is hourly, and it was MEASURED:** demo-hp `/disks` costs median 0.821s (min 0.805 / max > 0.841, 10 calls, 3 physical rows) — 6x under the 5s bar. Open question deliberately NOT acted on: the > agent runs bare `smartctl -a -j` with **no `-n standby`**, so an hourly poll would wake a spun-down > HDD. demo-hp is all-flash so the measurement could not show it (**R-333**). > > **NOT live-validated:** the Fail-from-counters path has never fired on real hardware — only against > the fixture's values in unit tests (**R-332**). > **2026-08-08 — v0.208.0 (R-254). THE RULE, stated so it outlives this session:** > > ### A secret is never in a page's response body. It is fetched by an explicit act, and the act is recorded. > > Three instances of one pattern shipped in two days, each found by hand: the retrieval passphrase > (R-249), an app's real first-login password (R-254 site one), and an already-deployed app's generated > secret field (R-254 site two). Every one was "hidden" with `display:none`, `hidden`, or > `type="password"` — **instructions a browser honours when DRAWING and nothing else.** The plaintext > was in the bytes; a `curl` returned it; caches, history, saved pages and screen-shares had it. > > **The shape of the fix, now used three times:** the page carries a BOOLEAN; the value comes from a > **POST** (so CSRF covers it and it is not re-fetchable from history) with **`Cache-Control: > no-store`**; the reveal is **LOGGED as an act** — reading a value off markup left no trace anywhere, > which is why nobody can say whether any of these was ever read. **Per-secret endpoints, never one > generic "reveal any named secret"** — that would turn three narrow exposures into one lever. > > **And the test must assert the RAW RESPONSE BODY.** Every test that asked what the customer *sees* > passed while the bytes carried the secret. That is precisely how this survived three times. > > **What is NOT this defect:** a form must carry what it submits. The pre-deploy hidden input round-trips > a generated secret deliberately (README §318) so the saved value is the one the customer wrote down. > The defect there was the neighbouring READONLY input on an already-deployed app, where nothing is > submitted at all. > > **The gate:** `scripts/secret_in_markup_gate.py`. Name-based, all 36 templates, **blind to a secret > arriving under a neutral page-data key** — measured, not assumed. The complementary runtime > body-assertion covers 4 of 27 page templates; the other 23 are **R-255**. > > **A correction to v0.207.0's report:** it said HTML comments ship in the response body. They do not > here — `html/template` strips them (`text/template` does not). Measured. > **2026-08-08 — v0.207.0 (R-249, R-252, R-253). Three things the fifth walk exposed BY PASSING.** > The walk closed R-201 (both halves) on 2026-08-07; none of the below touches the recovery path it > proved. > > **R-249 — a secret was living in the page source.** `settings_security.html` rendered the retrieval > passphrase into a `display:none` span behind a „Megjelenít" button. That toggle stops a browser > DRAWING it and nothing else: the plaintext was in the response body of every render. Found by doing > exactly that — it landed in a session transcript while driving the documented rebuild path. > **THE RULE, which the codebase already stated for R and this page did not follow:** a secret is > revealed by an XHR, never templated server-side into HTML (`escrow_handlers.go`). The page now > carries only `HasRetrievalPassword`; the value comes from `POST /settings/retrieval-password/reveal` > — CSRF-covered, `no-store`, and **logged as an act**, which reading it off the markup never was. > **The test asserts the RAW RESPONSE BODY** — every test that asked what the customer *sees* passed > while the bytes carried the secret, and that is why it survived. > **The census found two more instances** (`app_info.html`, a real per-install app password in a > `hidden` span; `deploy.html`, a generated secret in a `value=`) — **filed as R-254, not fixed.** > > **R-252 / R-253 — the two obstacles, and the rule they share.** A rebuilt box keeps its drives but > loses their REGISTRATION, so every restore refused with a sentence naming no next step; and the > restore list promised „a visszaállítás előbb újratelepíti" three lines above a refusal that fired > *because* the app was not installed. **The promise was the wrong half:** reconstitution writes to > the app's own `GetStackHDDPath`, which exists only once the CUSTOMER has chosen a drive at deploy > time — an automatic reinstall would mean the product making that choice for them, which is the one > decision this recovery path exists to leave with them. Both now name a reason and route to the step > that clears it, and both notices are conditional (a healthy box is byte-identical, pinned by a test > that fails if either becomes unconditional). > > **The page and the resolver ask ONE question:** `HasRestoreDestination()` reads the same > `GetSchedulableStoragePaths()` the scratch resolver reads. A second copy of that predicate is > exactly how a page ends up promising what the handler refuses — which is R-253 itself. > **2026-08-07 — v0.206.0 (R-241). THE RULING, and it reversed the fix: this was a MINTING defect, > not a screen-predicate defect.** The recovery screen was telling the truth — there genuinely was > nothing recoverable under the key the box held, because **the box minted that key itself over the > top of a sealed package it already knew the hub was holding**. Fixing the predicate would have > papered over a machine quietly making its own backups unopenable. > > **THE RULE: a box does not create a repository key while the hub holds a sealed package for it.** > The guard is a conjunction (package held AND no key), so a first-time box is untouched, and the > refusal is a HOLDING state rather than a failure — the transport is still configured so the > recovery screen can bring the tier up the moment the key arrives. > > **THE SECOND RULE: the fact that answers a question must be kept where the question is asked.** The > hub-vs-local key comparison had been computed on every ACK since SLICE 3 and persisted nowhere; on > the venue it logged the right answer thirty-five minutes before the customer looked at a screen > that could not see it. It is now persisted and drives shape (c) of the offer. > > **THE THIRD RULE (the operator's, and it generalises): fix the state, do not remember that it is > wrong.** Abandoning the old history now starts a 14-day countdown that removes the set-aside store > and its sealed package TOGETHER, after which the offer falls silent on its own because there is > nothing left to compare — rather than a "they decided" flag suppressing a screen over a state that > is still wrong. The recovery offer stays reachable for the whole grace; a grace in which recovery > is impossible is decorative. > > **Surface:** the full page appears once per ENTRY into the offered state, not once ever — a box > rebuilt months later is a new situation. Three dismissal levers with three scopes, and none of them > removes the entry point on the backups page. > > **Needs hub v0.98.0** for the superseded-package purge. `felhom-agent` untouched. > > **Two real bugs were caught by tests rather than by review** — a missing `t.Enabled` (an existing > test) and a missing falling-edge sync that reintroduced the very defect the epoch exists to fix. > > **NOT built, deliberately:** the automatic 30-day abandonment (R-245, with the operator's reasoning > recorded), and R-242's release-to-golden gate. > **2026-08-06 — v0.205.0 (R-234).** THE RULE: **a run that skipped an app the customer selected is > not a successful run.** The R-203 verdict block already said *"a warning beside a success is read > as a success"* and applied it to one of the two shapes it describes — a missing declared FOLDER > made the run `incomplete`, an app skipped ENTIRELY did not. Now both do. A selected-but-UNDEPLOYED > app is named with what to do but does NOT move the verdict, because a box left permanently amber by > an app somebody removed is a status nobody reads. > > **§7.3, MEASURED rather than assumed — and the answer was "already done".** `CaptureRecoveryUnit` > writes compose config + a manifest (a few KB), only ENUMERATES dumps rather than creating them, is > idempotent, and does NOT stop the app; the off-site run already calls it for every deployed stack in > its own pre-dump phase, through `admitApp`. So there is no wait to remove for a deployed app, and > **nothing was built**. Proven on demo-hp: a unit moved aside was RECREATED by the run. > > **AND THE FILED MECHANISM WAS NOT THE MEASURED CAUSE.** R-234 was filed as "the first run after a > toggle finds no bundle and skips the app". That cannot happen for a deployed app (above). What did > happen on 2026-08-06: the manual run was dropped by the **single-flight** while an earlier run was > still going; `runOffboxBackup` returned nil; the handler had already said „elindult”; and the card > then showed the PREVIOUS run's „✓ Rendben”. Fixed by taking that decision synchronously in the > handler. **The nightly path deliberately still returns nil** — nobody asked, and it retries. > **2026-08-05 — v0.200.0 (R-193 CLOSED).** The customer-facing recovery screen. Until now a customer > whose machine was rebuilt had everything needed to get their data back and no way to find out — the > only route was a command line. > > **IT UNLOCKS AND ONLY UNLOCKS** (operator ruling). Explains, takes the recovery code, opens the > repository, lists what is in it (apps, dates, sizes). **Restores nothing** — restore is per-app and > lives in the backups area; the put-back is **R-213** and its stated requirement is a > live-versus-backup comparison. > > **ONE CORE, TWO CALLERS.** `backup.RecoverInstallCore` is the only fetch→unseal→compare→install path. > `RecoverAndInstall` is now a thin CLI wrapper — exit codes and printed lines byte-identical, every > pre-existing CLI test passed unchanged — and the handler calls the same function. Asserted from > source by AST on BOTH sides, plus a test that the routes and the landing-page interception exist. > > **THE TRIGGER HAS TWO SHAPES and the second is the one that matters.** `OffsiteRecoveryOffer` = the > hub holds a package AND (no repository password OR the tier is orphaned). The literal "no repository > password" alone is a window that CLOSES BY ITSELF — `WriteOffboxSecrets` auto-generates one on > re-apply (R-193's own orphaning mechanism) and hub v0.96.0's self-heal re-applies within ~15–30 min. > Shape (b) is also what the shipped move-aside requires, which is why the discard choice can reach it. > > **CLAIMED is part of the predicate** — a legacy-open box passes through `RequireAuth`, so without an > explicit `authEnabled()` check the interception fired for an unauthenticated visitor. A test caught it. > > **„Most nem" suppresses the FULL PAGE ONLY.** The backups-area entry point is bound to > `recoveryOffer`, never to the postpone flag. > > **The code:** POST body only, never logged/persisted/echoed, cleared on every path, `no-store`, > `autocomplete=off`. **No lockout** — a ten-word phrase is not guessable and locking a customer out of > their own data for a typo is worse; failures are logged locally without the code, and NO operator > alert is raised (reasoning in REPORT.md §4). > > *Live:* demo-felhom is genuinely in shape (b), so validation needed no arrangement — `/launcher` → > 302 `/recovery`, both mandatory sentences rendered, three wrong codes refused with the `offbox/` > listing byte-identical and no lockout, and the code found in no file, log or ring **with a > planted-copy positive control that first exposed a mis-aimed sweep**. **NOT proven live: a CORRECT > code** — none was kept for demo-felhom's orphaned history and demo-hp's is operator-held. > **2026-08-05 — v0.199.0 (R-204 item 4 / R-193).** The last of the four manual interventions the > 2026-08-04 drill needed. **Operator ruling: automate it, and the trigger is a state the BOX > DECLARES.** From the hub an absent off-site object has FOUR meanings — never configured, > mid-restart, a transient config read failure, rebuilt-and-stranded — and the hub cannot tell them > apart. The box can. > > **The declaration needs BOTH halves** (`backup.needsOffsiteCredential`): a fresh data area (no > repository password) AND a hub-held recovery package (the ACK's `identity_blob_present`). Freshness > alone is a box that never had off-site backups; dropping that condition makes the whole fleet ask > for credentials, which is what `TestOffsiteDeclare_NeverHadOffsiteSaysNothing` catches. A merely > DISABLED target is the customer's own choice and never declares. > > **The ACK field stopped being discarded.** `EscrowAutoConfirmer.Reconcile` returns early when the box > is neither pending nor escrowed — exactly a rebuilt box — so the fact was thrown away every cycle. It > is recorded FIRST, before every gate, via `RecordPresence`, wired in main.go and asserted by > `TestMainWiresRecordPresence` (AST, comments dropped). Last-write-wins, not set-only, so a customer > RESET turns the declaration back off; a nil ACK escrow records nothing. > > **Inert to every existing reader:** `enabled:false` + zero sizes, so the hub's `isStale` and > `fillBand` both short-circuit; an unknown `state` string is ignored by encoding/json. **A configured > box's report JSON is byte-identical to v0.198.0's.** The one reader that would have misread it is the > hub's `reportHasOffsite`, tightened in hub v0.96.0 to require `enabled:true`. > > *Live:* both demo boxes now record `hub_escrow_identity_present=true` in settings.json (the recorder > working on a HEALTHY box). demo-felhom 9201, arranged reversibly into the stranded shape, produced > report id=16743 carrying `{enabled:false, state:needs_credential, quota_gb:0, repo_size_bytes:0}`; > the single declaration was absorbed by the hub's debounce (no self-heal event) and the box was > restored the same minute. **The hub half is felhom.eu v0.96.0.** > > *Rider:* `.githooks/pre-push` in all four repos now refuses a push from a clone outside > `/mnt/5_hdd/felhom.eu`. Proven both ways against a scratch clone. > **2026-08-05 — v0.198.0 (R-204 items 1 & 3).** The 2026-08-04 drill (R-201) passed only because a > person was there; four manual interventions stood between a recovered key and a restored file. Two > of the three defects are in this repo. > > **Item 1 — the reset code needed a restart.** `--print-reset-code` is a SEPARATE process; it > persisted a new code while the running server kept the old one cached, so the code the customer was > told to type was refused until the controller restarted, and nothing said so. `effectiveClaimCode` > now calls `settings.ReloadClaimCode()` first. **The settings-vs-config precedence is unchanged** — > the defect was freshness, not precedence. **Read-through, not a TTL, and that is the point:** a TTL > makes the new code visible AND leaves a window in which the superseded one still works, which is > worse than the bug. That is the mutation `TestClaimCode_SupersededByASecondMint_RefusedImmediately` > exists to kill, and its red-proof produced exactly *"the SUPERSEDED code was accepted"*. The > function now returns an error and **every caller fails closed**; an absent settings file is NOT an > error. `ClaimConsumedGeneration` is deliberately NOT re-read — this process is its only writer and > re-reading could move it BACKWARDS if a save had failed, resurrecting a consumed code. > > **Item 3 — the restore's default returned the wrong thing silently.** `mode=unit` restores the > recovery unit (definition + config + DB dumps) and not the customer's files. `restoreScratchOutcomeMsg` > now names what came back, what did not, and the next step; the wizard's intent card states its scope > before the choice. **The size gate is untouched** and pinned unchanged by > `TestOffboxRestore_FullPathUnchanged`. **The default stays `unit`** — all three wizard forms set > `mode` explicitly, so a change would alter nothing visible while silently changing a mode-less POST. > > *Live-validated endpoint-level (no browser on DooPlex):* on demo-felhom 9201 with `restarts=0` > across both mints, a superseded code returned „Hibás vagy lejárt kód" and the current one was > accepted first time; on demo-hp 9201 a `privatebin` unit restore produced the scoped Hungarian > outcome and `mode=full` without confirm revealed `full_size=6.8+KB` without restoring anything. > demo-hp's drill scratch (`calibre-web`) was not touched. > > **Item 2 is the hub's** (felhom.eu v0.95.0, R-196). **Item 4 — a rebuilt box cannot obtain an > off-site credential unaided — remains OPEN (R-193)** and was deliberately not begun. > **2026-08-02 — v0.190.0 (R-157 mechanism A · R-170 · R-171).** Three items, one live validation > cycle, because all three are boot behaviour and all three are proven by hard-resetting the box. > > **DIAGNOSE BEFORE THEORISING — and the first diagnosis was a FALSE NEGATIVE.** A hole was reasoned > out of the v0.189.0 diff (a drive-gate-stopped app has zero containers and `desired_state: running`, > so it now reads as a boot orphan) and confirmed on hardware BEFORE any fix was written. **Attempt 1 > produced `no boot-orphaned apps` and would have been reported as a disproof.** It was a race: > unmounting only the parent bind is healed by the agent within ~60 s, so the drive gate's startup > reconcile restarted the apps **one second before** the sweep looked. Holding the drive genuinely > absent reproduced the defect immediately. **"It didn't happen this time" is not a mechanism.** > > **The confirmation moved the severity in BOTH directions.** The write hazard did not materialise — > compose failed `mkdir …/userdata: permission denied` because the unbound mountpoint is > host-root-owned and the guest is unprivileged. **That protection is ACCIDENTAL**: no code chose it, > no test pinned it, and it is one `chown` or one privileged guest away from gone. But the harm that > DID occur was not in the hypothesis and is real on every box: two wasted attempts and a **false > dead-app alarm for an app the drive gate is deliberately holding**. > > **The fix already existed one path over.** `startGatedByMissingDrive` (the API) refuses a customer's > start on an absent drive; the sweep bypassed it by calling `Manager.StartStack` directly. > **`StartStack` HAS NO GATE OF ITS OWN** — carry this: every caller that is not the customer must > decide for itself whether the app may run. New consumer-side `bootrecon.StartGate`, fail-safe > (cannot determine ⇒ do not start). > > **Widening a window makes previously-unreachable overlaps reachable — a design input, not an > afterthought.** The old T+5 s sweep never met a quiesce or an in-flight app-data operation; a 50 s > window can. All three holders answer ONE seam because they differ only in the reason string. > > **A TEST REJECTED MY FIRST CONSTANT, and the comment says so.** `settle + budget + one retry` must > fit inside `deadAppBootGrace`; 60 s gave 95 s against 90 s. The budget is 50 s **because a test said > so** — recorded in the code rather than presented as taste. Widening the grace was rejected: it > hides a late recovery instead of reporting one (`recordLateRecovery`). > > **THE FIX HAD ITS OWN DEFECT, FOUND LIVE AND NOT BY REVIEW.** The window sampled `GetStacks()` — the > Manager's map, refreshed by the scheduler every **10 s** — every 5 s, so two identical samples could > mean *the cache did not update*. Observed: a container removed ~5 s before the window closed was > still in the sampled fleet and the sweep logged `no boot-orphaned apps` for an app that had none. > `sampleBootFleet` now refreshes first. **Generalise: a settle detector is only as good as the > freshness of what it samples — if the source is cached, refresh it, or you are watching the cache > settle rather than the system.** > > **R-170:** `shouldRecreateOnBoot` reads intent with the identical three-way table; absent keeps the > old `hasContainers` behaviour exactly; `presentStable` untouched and still load-bearing. Its comment > argued at length FOR the count and was rewritten. Agreement pinned from BOTH sides against one > fixture table (an import cycle prevents testing the two gates together). > > **Live: 6/6 hard resets** (every app back; the customer-stopped app down all six), settle times > 10/40/10/10/15/15 s. Sharpest evidence: same app, same box — missed at 18:08:35, recovered at > 18:18:50. R-170 proven in one reboot (calibre-web recreated, immich left stopped). 27/27 packages; > 7 red-proofs. Detail: `REPORT.md`. Last updated: 2026-08-02 (v0.189.0 — R-166 / D-b: the box stops guessing what the customer wanted) > **2026-08-02 — v0.189.0 (R-166, operator decision D-b).** When an app was not running the box had > to work out *why*, and it did so **by counting containers**: zero meant "the customer stopped it", > some meant "something broke". A **power cut mid-compose** and an **interrupted deploy** also leave > zero containers, so both were read as deliberate stops and stranded **silently** (R-157 mechanism > B) — and a backup that stopped an app and died left it stopped with **nothing on disk** recording > that it was owed a restart. The settling fact — what the customer asked for — **was written down > nowhere**: `app.yaml` recorded *installed*, never *meant to be running*. > > **DECISION — one owner: the customer's action, and nothing else.** A census found **14 callers of > `StartStack`/`StopStack`, of which exactly 2 are the customer**; the rest are quiesce, the volume > dump, offbox reconstitution, app export/restore, the storage gate, migration and the boot > reconciler. So the primitives are deliberately **not** writers — intent there would make a nightly > backup indistinguishable from the customer pressing Stop. Writers: the API action switch, > `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, the `.fab` import. Intent is written > **BEFORE** the act and a failed write **REFUSES** the act. > > **DECISION — absent means UNKNOWN, never "running", and this is the whole safety property.** Every > `app.yaml` on every box predates the field, so absent is what the fleet reads on upgrade; reading > it as running would start every deliberately-stopped app on the first boot after the upgrade. The > legacy branch of `isBootOrphan` keeps the old container-count rule **byte-for-byte**, and its test > asserts BOTH legacy rows together because the safety property is the pair. Backfill is > **running-only** — "zero containers ⇒ stopped" IS the defect, so an ambiguous app stays ambiguous. > > **Part 2 — `backup.AppStopGuard`**, a persisted marker over every stop→work→start window (volume > dump, offbox reconstitute, `.fab` export), in its **own** file (one file, one writer). Written > before the stop, cleared only after a restart that **succeeded**, kept when one fails. `Recover` > **returns** its outcome instead of using a notifier seam, because it must complete before the boot > reconciler (`main.go` ~236) while the notifier is not built until ~307 — a seam wired after the > fact is a seam that never fires. > > **THE TEST LESSON, and it is the one worth carrying:** Scenario E's first version called > `appStop.Begin` itself, and **survived the red-proof that deleted the production call**. It proved > the marker type, not that `DumpAppVolumesSafe` uses it. Rewritten to drive the real function with a > simulated hard abort (an unwind that skips the restart statement, since a `defer` is not > crash-safety — Campaign 8 fault 10). **A test that constructs the thing it is meant to prove the > caller constructs is hollow, and its red-proof will say so if you run it.** > > **FOUND EN ROUTE — `SaveAppConfig` rebuilt `AppConfig` field-by-field**, the R-100 shape (v0.181.0 > shipped two live instances). The literal named five fields, so `desired_state` would have been > dropped on **every** save across nine call sites — a customer's Stop erased by the next unrelated > `app.yaml` write. Copy-and-overlay (`saveCfg := *cfg`) is safe by construction. **Generalise it: > treat any field-by-field struct rebuild in a save path as a defect on sight.** Measured, not > assumed: `app.yaml` does NOT round-trip YAML keys the struct does not model (pinned by test). > > **R-157: mechanism B closed, mechanism A untouched** (the sweep observes ~5 s after start and never > re-checks) — and B's fix makes A cost more, since the sweep now has more it could recover. > **NEW R-170:** `shouldRecreateOnBoot` (`internal/web/intermediary.go:131`) still infers a Stop from > `hasContainers` — the same defect one gate over, for drive-backed apps. Left deliberately. > > **Live on 9201, three flows** (stop survives a restart; a zero-container `running` app recovered by > name; a legacy app.yaml skipped and never inferred stopped). The **interrupted-operation half is > IMPLEMENTED, not PROVEN-LIVE** — nobody killed the controller mid-backup on metal. 27/27 packages; > 7 red-proofs observed FAIL then restored. Detail: `REPORT.md`. Last updated: 2026-07-28 (v0.182.0 — R-101 + F-DIAG: the restore dialog names the last SUCCESSFUL copy) > **2026-07-28 — v0.182.0 (R-101 + F-DIAG).** `Tier2LastRun` is the ATTEMPT clock (written on failure) > and was rendered as „Legutóbbi másolat" in the **restore confirm dialog** — misinformation at a > decision point: the restore fills in MISSING files, so a customer with a failing Tier-2 restored and > silently got OLDER files. New `CrossDriveBackup.LastSuccess` + **`SuccessTracked`**; the marker is > load-bearing because **all 7 fleet rows were pre-anchor at deploy** — without it every customer sees > „Még nincs sikeres másolat" at once. Legacy rows migrate on first touch (`ok` adopts its time, > `error` seeds nothing). **PART 2 — the three `record*` helpers rebuilt the WHOLE struct with only 2 > fields carried over; the naive fix would have had `recordTier2Failure` CLEAR the anchor.** Replaced > by `tier2Update` (copy-and-overlay = safe by construction). New `fmtTimeStr` → Budapest-local dates > in the dialog instead of raw UTC RFC3339. **F-DIAG:** 6 classes incl. an honest `unknown`, and the > notification no longer passes `err.Error()` through raw — **LESSON: my first sanitiser was regex-only > and leaked a bare hostname; its own test caught it. Redact KNOWN values, don't guess at shapes.** > Live on demo-hp: rendered dialog read in the failed, healthy AND legacy states. F-OPS documented at > `felhom.eu/documentation/runbooks/RUNBOOK-manual-guest-restore.md`. > **2026-07-28 — v0.181.0 (R-100).** `OffboxTarget.LastSuccess` + wire field `last_success`; the hub > (v0.80.0) anchors offsite staleness on it. **`LastRun` is written unconditionally on every run > INCLUDING failures** — it records an ATTEMPT — so the hub's "how long since LastRun" verdict read a > nightly-failing tier as perfectly fresh forever. The rule is the pure `offboxAnchorAfterRun(prev, at, > runErr)`: a failure neither ADVANCES nor CLEARS the anchor (both are distinct bugs; clearing it would > make one bad night look like never-succeeded). `LastStatus == "error" ⇒ stale` was rejected — it pages > on every blip, the F-A1 noise mode. **TWO SILENT-WIPE SITES CLOSED** (`offboxConfigHandler` and > `ApplyOffsiteTarget` both rebuild the target and copy runtime status field-by-field — omitting > LastSuccess would erase the anchor on any settings save or hub re-apply). **LESSON: my first test > modelled the rule in a local closure and stayed GREEN when production was mutated — hollow; the > extraction to a pure function is what made the red-proof bite.** Live on demo-hp: failing run advanced > `last_run` to 11:25:48Z while `last_success` HELD at 11:24:20Z; demo-felhom healthy → advanced. The > settings-save preservation was proven live too. Detail: `REPORT.md` + `felhom.eu/REPORT-r100.md`. > **2026-07-28 — v0.180.0 (F-OBS).** Source: `audits/CAMPAIGN-8-backup-restore-2026-07-27.md`. > On a default `logging.level: info` box there was **no positive observable that `deadapp-check` had > run**: its per-cycle line goes through `Scheduler.dbg()`, gated on `level==debug`, so on a default > box it was never *produced* and could not even reach the always-DEBUG ring. "No alarms" was > therefore indistinguishable from "the detector never ran" — standing rule 3's exact fallacy, and it > undermines F-CRIT-1's fix, which is a fix to **this same detector**. > `noteDeadAppScan()` now emits an INFO line every **20th** scan (10 min at the 30 s cadence) carrying > scans-since-boot / evaluated / currently-down. It reports **what it saw**, not that it ran, and it > summarises rather than floods — one line per run is 2880/day, which is what made silence attractive > in the first place. Both bounds are pinned by test in the direction that would break them. > **The same shape then turned up in the agent's brand-new guest-power watchdog** (v0.107.0, shipped > hours earlier): it logged only at startup and when it acted. Fixed in agent v0.109.0 with the same > pattern. The anti-pattern reproduces itself — which is the argument for not having dropped this part. > Live on demo-hp at INFO on a default-level box; deployed on both boxes. Detail: `REPORT.md`. > **2026-07-26 — v0.173.0 (R-77).** Source: `audits/DIAG-agent-channel-2026-07-26.md`. > > **UNRESOLVED AND DELIBERATELY DEFERRED — which file is authoritative for `local_api`?** R-77 ships > DETECTION ONLY. `controller.yaml` and `bootstrap.json` can disagree; the controller dials > `controller.yaml`. The obvious "fix" — reconcile from `bootstrap.json` on every boot — has a failure > mode **as severe as the bug it fixes**: on a guest whose `controller.yaml` is correct and whose > `bootstrap.json` is stale (a re-provision that half-completed, a hand-repaired guest, a > setup-wizard box), auto-reconcile would clobber a WORKING channel on the next restart — fleet-wide, > silently, at the moment of a routine deploy. R-77's position is that **naming the drift is enough**: > it would have converted the 17.5 h outage into a specific alert on the first health cycle. The > authority ruling is **R-78** and needs its own spike — do not resolve it opportunistically. > > Corollary for anyone editing `bootstrap.MaybeIngest`/`ensureLocalAPI`: `ensureLocalAPI` is the ONLY > writer, it fires only when the endpoint is EMPTY, and `DetectEndpointDrift` must stay write-free. > Scenario A's test asserts `controller.yaml` is byte-identical after the check, and its red-proof > covers the auto-correcting variant precisely because that is the tempting wrong turn. > > **Also settled here:** the samba protected-set must mirror EVERY early return in > `reconcileSambaAt` (currently two: `!smb.Enabled`, `!smb.UserSet`). A third would need the same > mirror, and the doc comment above `EffectiveProtected` must be updated with it. > **2026-07-26 — v0.172.0 (R-75).** Spike `felhom.eu/documentation/audits/SPIKE-catalog-data-paths-2026-07-26.md`; > feature doc `felhom.eu/documentation/controller/import-and-data-paths.md`. > > **RULING — the import root is CANONICAL on the system drive, overriding the spike's Fork-1 > recommendation of per-drive roots.** The spike weighed sidebar clutter and per-app link ambiguity and > concluded per-drive; the operator overruled it on an argument the spike missed: each drop-zone app has > exactly ONE ingest bind, so on a two-drive box every import folder except the app's own would look like > a drop-zone and silently do nothing — and because `import/*` is `class: excluded`, files stranded there > are never backed up either. A canonical root is the only shape with no dead drop-zone. Recorded as a > deliberate deviation, not an oversight. > > **Phase-0 probe changed the shape of Part 6.** The system drive is NOT a registered `StoragePath` on > either demo box (`/mnt/felhom-drives/hdd_1` on demo-felhom; `nvme-1tb` + `Felhom-Share` on demo-hp), > so `sharingResolvePath` REFUSES `/userdata/import` — verified against the real guard with a > passing control. Registering the drive was rejected (it would make the 50 GB volume holding the > recovery units a customer-visible drive, deploy target and wipe candidate, and `SharingDeniedRoots` > would then deny the namespace-consistent shape anyway). **Chosen: leave it unregistered and have the > controller write the `beolvasas` share directly** — the picker guard validates CUSTOMER-supplied paths, > a controller-generated constant is a different trust class. No guard was weakened. > > Also note: `withUserdataPath` computes `USERDATA_PATH` as `/userdata`, NOT > `NamespaceRoot(hdd)/userdata`. For an app on the system drive those disagree > (`/mnt/sys_drive/userdata` vs the `felhom-data` namespace). Latent — no app with a userdata bind has > ever been deployed there — but it is a real inconsistency, left untouched here. > > The other three forks followed the spike unchanged: all-apps skeleton / deployed-only in the UI; > unknown role fails OPEN while a malformed path whole-block rejects; drop-zone copy driven by the > derived backup class. > **2026-07-24 — v0.169.0 (disk-health card + degradation alert).** Consumes the agent's new `smart` > field (agent v0.94.0; MinAgent floor unchanged — feature-detect by presence). **Rulings:** (1) ONE > pure verdict fn `agentapi.DiskVerdictFor` is the shared truth for the card chip AND the 6h check — they > can never disagree. Thresholds: FAILING→Hiba; PASSED + any(reallocated>0/pending>0/offline_unc>0/ > critical_warning>0/media_errors>0/percentage_used **≥90**)→Figyelmeztetés; PASSED clean→Rendben; > nil/UNKNOWN→Nincs adat (never alarms). (2) **No global alert banner** — the card + email carry disk > health; banner fatigue is a real cost, so this is deliberately NOT wired into the dead-app/alert-banner > machinery. (3) Degradation-only notification with an in-memory baseline: first run baselines silently, > recovery never notifies, **UNKNOWN excluded both directions** (a transient blip neither fires nor erases > history). (4) **Controller restart re-baselines silently** (in-memory baseline lost on restart) — an > accepted trade consistent with the health-change pattern (a real post-restart degradation still fires on > the following 6h check once a baseline exists). (5) A **60s TTL cache** wraps the card's /disks call so > dashboard refresh-spam can't smartctl-storm the host; the 6h check fetches FRESH (cache-independent). > Pairs with hub +1 (allowlist `disk_health_degraded`). No new smartctl load — serialization only. Last updated: 2026-07-24 (v0.168.0 — customer-configurable backup window "Mentési időablak") > **2026-07-24 — v0.168.0 (customer-configurable backup window).** ONE customer setting — the window > start W ("Mentési időablak kezdete") — drives every nightly leg at FIXED, never-stored offsets so > misordering is impossible: DB dump at W, tier-2 at W+60m, off-box at W+105m (wrap-safe). **Design > rulings:** offsets are DERIVED and computed everywhere, never persisted and never exposed in the UI; > precedence is settings > controller.yaml `db_dump_schedule` > "02:30" (mirrors PasswordHash); a change > applies WITHOUT restart via the new scheduler seam `UpdateDaily` (per-daily-job buffered `resched` > chan + a select case in `runDailyJob`). New pure package `internal/backupwindow` holds all the time > math (ParseHHMM/FmtHHMM/LegTimes/GateWindow/EffectiveWindow). **Disk-tier (whole-guest PBS/vzdump) > gate:** the quiesce loop's SCHEDULED cycles run only inside [W+2h, W+6h) (wall-clock Europe/Budapest), > with a safety valve — last successful backup older than cadence+24h (or none) runs regardless, so a > box only ever on outside its window never starves. **Manual "Mentés most"/TriggerNow is NEVER gated** > (bypasses runOnce). The `quiesce.Backend.Due` seam now also returns the backup age (from the agent's > own `/backup/due`); the agent, its cadence, and `/backup/due` are untouched. Window read fresh each > poll (WindowStartFn) so runtime changes take effect. Cadence defaults to 24h controller-side (the > response carries no cadence). Backup page gets a "Mentési időablak" card (time input + derived rows + > the "kb. W+2h–W+6h között" rendszermentés line); POST /backups/window (RequireAuth+CsrfProtect). > **2026-07-24 — v0.167.0 (outlined logo + favicon — Part 4 unblocked).** Viktor pushed the > text-outlined `logo.svg` to felhom.eu `main` (`be9edb4`); the wordmark is now 17 real `` > glyphs. `FelhomLogoSVG` swapped to it; Inkscape's leftover **empty `` shells + font-* leftovers > on the paths** were stripped via an lxml DOM pass (glyphs untouched — CC did NOT do text-to-path), > editor `` dropped. `FelhomFaviconSVG` vestigial `` removed. Both constants: > **0 ` Inkscape "Object→Path" leaves empty `` shells AND copies `style="…font-family:…"` onto the > resulting ``s — a search for `svg:text` misses them (elements are ``, no prefix); grep > ` first** (`assetsSyncer.Resolve`) and fall back to the constant only if none is on disk — on 9201 the > constant is what's live (verified). Still open (separate follow-up): website + hub serve their own > non-outlined logo copies; login.html stylesheet link still unversioned. > **2026-07-24 — v0.166.0 (mobile nav = off-canvas drawer; sidebar cleanup; ?v= on logo/favicon).** > Mobile nav was broken: the ≤768px block predated the v0.146.0 accordion and flattened `.nav-links` > into a horizontal `overflow-x` strip, clipping the accordion's nested sub-lists (they share the > `.nav-links` class). **Decision: mobile nav = a sticky top bar + off-canvas left drawer that REUSES > the vertical sidebar (Option A).** The accordion handler is untouched and works inside the drawer; > a `no-js` html-class fallback renders the sidebar static inline so nothing dead-ends without JS. > Options B (separate mobile menu) and C (exclude nested lists from the strip) were rejected. z-index > ladder topbar 800 < backdrop 900 < drawer 950 < modal 1000; `100dvh`; reduced-motion disables the > slide; focus-trap deliberately omitted (navigations reset state). **Sidebar customer-name removed** > (logo only); `{{.CustomerName}}` stays in base data + login subtitle. **Logo policy decision: the > wordmark must be OUTLINED paths, never live ``** — under `` secure static mode only > locally-installed fonts resolve, so `font-family` in the SVG renders a fallback font everywhere. > **Part 4 (swap `FelhomLogoSVG`/`FelhomFaviconSVG` to the outlined master) is GATED OUT** — §3a check > against live felhom.eu `main` (`be9edb44`) found `website/assets/logo.svg` still has ``/ > `font-family`; the outlined master is Viktor's manual Inkscape push, still pending. Only the `?v=` > cache-bust (logo/favicon/login-logo, Cloudflare 4h edge-cache — the 0.126.1 failure mode) shipped > from the logo work. Follow-up: when Viktor pushes the outlined asset, ship Part 4 (swap constants + > clean the favicon's vestigial `` nodes). Separately, the website + hub still serve their own > non-outlined logo copies — propagation is a distinct follow-up. > **2026-07-24 — v0.165.1 (native "Megosztás…" in the share modal, Web Share API).** The share modal > gains a feature-detected `navigator.share` button (OS share sheet → Messenger/WhatsApp/email), > sending **title + text + URL only**. Hidden unless supported; "Link másolása" stays the universal > fallback (and catches the non-cancel rejection); `AbortError` (user cancel) is silent. **Ruling: the > QR is NOT attached** (no Web Share Level-2 `files:`) — file-share support is narrow and several > targets drop the URL when handed file+URL, leaving an unscannable QR picture in a chat; the QR's job > (physical cross-device scanning) is already served by the modal image (mobile long-press). Template > JS + tests only; the OS sheet interaction is an operator manual check (not endpoint-testable). > **2026-07-24 — v0.165.0 (Indítópult megosztása — guest launcher via capability URL).** The admin > launcher gets an "Indítópult megosztása" button that mints a **capability URL** > (`https:///s/`, 160-bit `crypto/rand` token) serving a standalone, read-only guest > launcher — same tiles, opens apps in new tabs — with **no account and no admin session**. **Security > ruling: the link grants INFORMATION ONLY, ZERO CONTROL** — app names + public URLs; every privilege > stays behind each app's own auth and the controller admin password. The token IS the secret (160-bit > entropy is the whole defence for the GET — never rate-limited, never logged, `subtle.ConstantTimeCompare` > only; an empty stored token = sharing OFF, matches nothing, so a wrong/disabled token is byte-identical > to the mux default 404). Optional per-share password is a SEPARATE credential (own bcrypt hash, own > attempt map — NEVER the admin ones); one pass mints a cookie = HMAC(`token|passwordHash`) keyed with > the persisted `web.session_secret`, so rotate-token OR change-password invalidates all cookies for free. > **Part-2 secret decision: REUSED `web.session_secret`** (persisted + box-scoped + stable — the SAME > secret the claim pre-auth CSRF already trusts; not per-boot, not claim-generation-scoped → the reuse > branch), so no `ShareCookieSecret` field was added. **Design rulings recorded:** member accounts are > **superseded** by this capability-URL model; **per-member tile visibility is PARKED under the SSO arc.** > Guest state labels ride the v0.164.0 invariants: `StateStopped` ⇒ "A tulajdonos leállította"; any > other non-clickable state ⇒ "Átmenetileg nem elérhető" (guests never see stopped/exited/degraded/ > unhealthy). Accepted residuals (documented, no code action): link-preview crawlers fetch once and see > app names (noindex prevents indexing); reverse-proxy/CF access logs may hold the path (ops-tier); the > modal link carries the request Host, so a LAN-IP admin session yields a LAN-IP link. New dep: > `github.com/skip2/go-qrcode`. Tests: Groups A–G (14 tests) + 3 red-proofs verified red. > **2026-07-24 — v0.164.0 (stopped ≠ fault).** Operator finding on 9201: a UI stop (Leállítás) raised > the global "Telepített alkalmazás nem fut: … (stopped)" banner on every page AND fired the > `app_start_failed` email. RULING: **a deliberate user action must not alarm anywhere.** One-line > filter at the single fix-3 derivation point — `scanDeployedAppRunStates`'s pure core extracted to > `classifyRunStates([]stacks.Stack)`, down predicate now > `stacks.IsDownState(st.State) && st.State != stacks.StateStopped`. `StateStopped` is dropped from > BOTH the banner dead-list and the notifier Down-set (⇒ no banner, no event, clean tracker). Rests on > **two invariants that MUST both hold for this suppression to be correct:** **I1** — the UI stop path > `Manager.StopStack` runs `docker compose down` → containers removed → a deployed stack with zero > containers aggregates to `StateStopped` (refreshStatusLocked). **I2** — the P2 restart-policy census > (2026-07-21, 53 templates / 78 services) found every catalog service on `unless-stopped`, so a crash > never rests at `stopped` — faults surface as `exited`/`degraded`/`restarting`/`unhealthy`. **If > either invariant changes, revisit this suppression.** `IsDownState` UNCHANGED (other callers rely on > stopped=down). Out-of-band `docker compose stop` (containers remain → `StateExited`) still alerts — > correct, tampering is reportable. The `stopped_by_user` intent flag was considered and PARKED (only > adds value against out-of-band stops, which should keep alerting). Tests +4 (notify 3→4, main 4→7), > both red-proofs verified. No template/funcmap/notifier/counter/copy change. > **2026-07-24 — v0.163.1 (launcher polish).** Two v0.163.0 live findings fixed. RULE recorded: > **every app-logo surface ends in a visible placeholder** (`SVG → PNG → /static/app-placeholder.svg`, > infra rows → `infra-logo.svg`) — the four sibling `onerror` chains (`backups_apps`, `stacks`, > `app_info` hero, `deploy`) now match `app_row.html`; `app_info` screenshots deliberately still > vanish on error. And the **launcher monogram is launcher-only AND failure-only**: hidden by default, > revealed when the tile's img chain fails (`onerror` adds `.launch-tile--noimg`) — it was bleeding > through every transparent white glyph. Template/CSS only; no handler/funcmap change. 5 tests + 2 > red-proofs. [[launcher-v0163-2026-07-24]] > **2026-07-24 — v0.163.0 (Indítópult app launcher + universal placeholder icon).** New > customer-facing `/launcher` page: the FIRST sidebar item (above Vezérlőpult), a grid of large > tappable tiles for openable deployed apps. `/` stays the Vezérlőpult — the launcher is ADDITIVE. > Design rulings recorded here: > - **(a) The felhom brand mark is NEVER an app placeholder** — brand = platform identity only. The > logo-less fallback everywhere is the new generic `AppPlaceholderSVG` (a 2×2 app-grid glyph, > `/static/app-placeholder.svg`), now the DEFAULT `FallbackIcon` on `app_list_row` (was > `visibility:hidden`). On the launcher tile the fallback is the **monogram**, not the placeholder. > - **(b) A launcher tile exists ⟺ a „Megnyitás" button would** — subdomain presence (env `SUBDOMAIN` > > `.felhom.yml` subdomain > `protectedStackSubdomains`) is the single openability criterion. The > controller stack is excluded by name. The subdomain assembly was extracted to > `Server.subdomainMap` (3 callers: dashboard, Alkalmazások, launcher; priority byte-unchanged). > - **(c) Colored-tile + mono-glyph design.** `tileColor` = validated `.felhom.yml` `brand_color` > (`#rgb`/`#rrggbb`, new `Metadata.BrandColor`, omitempty) OR a deterministic FNV-1a-of-slug HSL > (fixed S/L, hue per app). Invalid `brand_color` silently falls back to the hash color (the one > §8 exception to no-silent-failure — cosmetic). `tileColor` returns `template.CSS` (we > validate/compute in Go; html/template's CSS filter mangles a legit `hsl()` from a func pipeline). > - **(d) `/` remains the Vezérlőpult.** No role/auth gating — member-role gating is a future arc > (ROADMAP: member role → launcher becomes the member landing page). No catalog app sets > `brand_color` yet (curation parked). > No agent coupling; MinAgent unchanged. 10 new test functions + 4 red-proofs (all observed FAIL then > restored). Gates green (app_row_dedup / template_id / emoji). > **2026-07-24 — v0.162.0 (R-71a), SHIPPED + deployed BOTH boxes (demo-felhom 9201 + demo-hp 9201 > via G1 break-glass), clean+healthy, settle-gate GO line captured on both.** B′ live note: both > above-floor boxes GOed correctly but NOT literally first-poll — the floor is in-memory (not > persisted), unknown at t=0, so the gate logged `awaiting floor knowledge` then GOed ~10 s later the > instant the report ACK landed (report-ACK latency = exactly what the 90 s sub-bound is sized to; > zero-wait-when-floor-known is unit-proven, test E). The gate correctly did NOT burn the one-time > password before the update picture was clear. > The structural fix for the F10 day-0 race (DIAG-f10): the apply-bridge no longer consumes the > single-use offsite password while a managed floor-update is in flight or imminent (below floor). > New seam `offsiteapply.SettleProvider.SettleState()` + `SettleFunc` adapter over the updater's own > `GetFloor()`/`IsUpdateRunning()` (no second floor path); `Bridge.AwaitSettle` polls 10 s BEFORE the > 3-min Reconcile ctx (deferral never eats the reconcile budget), bounds 90 s floor sub-bound / 5 min > overall (both GO+WARN — the "hub that can't serve a floor can't serve a consume → no burn" argument, > R-71c is the belt). At/above floor → GO first poll, zero wait (B′). Bridge goroutine MOVED after the > updater in main.go; wired only when an updater exists. **Ordering-only** — consume/persist/404 > contract untouched; R-71(b) rejected-by-design. **FINDING:** the floor is in-memory > (report-ACK-derived ~5–10 s), NOT persisted → unknown on any restart until the first ACK (sized the > 90 s sub-bound to that). 5 test scenarios (A–E) + nil-provider + cancelled-gate; **4 red-proofs all > observed FAIL then restored** (gate/updateRunning/sub-bound/overall-bound). Deferral paths NOT > live-fired (precondition now structurally prevented by the v1.25.0 build gate). **Layering: gate > prevents, (a) defers, (c) heals.** ROADMAP R-71 → SHIPPED (a)+(c). Live leg = the B′ first-poll GO > line on both above-floor boxes. > **2026-07-23 — v0.161.0 (R-70 controller leg), SHIPPED + deployed BOTH boxes.** When > `offsite.enabled` is in controller.yaml but no `offbox` target exists (pre-apply window / burned > credential — the F10 shape), Távoli mentés now shows „Felhom offsite tárhely kiépítve — a > beállítás automatikus, folyamatban…" on BOTH empty surfaces (status card + target line) instead > of „igényelhető" / „Még nincs beállítva". Data key `OffsiteHubEnabled` (from `Server.cfg`, no new > wiring); render tests per gate branch; banner leg is unit-proven/live-pending (no healthy box > occupies the window; next fresh onboarding is the natural live leg). Hub sibling v0.72.0 carries > the detector + `offsite_delivery_stuck` + the R-71c self-heal. Origin + rulings: > `felhom.eu/documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`. > **2026-07-22 — v0.160.0 (R-67), SHIPPED + deployed BOTH boxes, full live leg on demo-hp.** > Network shares now bind their share ROOT into FileBrowser (`…/:/srv/:rslave`) — no > skeleton/userdata toward the NAS, ever. Pure assembly = `buildFileBrowserPaths` + `fbPathDeps` > (handlers.go), returning mounts AND config sources together so they can't disagree. > > **DECISION — two classes, two gates:** drives keep the drive-absent gate (byte-identical, > tested + observed live: demo-felhom logged a no-op sync); network shares use the STUB classifier > gate instead (stub ⇒ excluded from both lists + WARN — an exposed stub swallows uploads the real > mount later shadows; idle autofs is HEALTHY and included; unknown fails open). Never force-wake > in the sync (doctrine). > > **Phase-0 probe = GO:** in-container access through an rslave bind WAKES an idle autofs trigger > (proved on demo-hp against the real Felhom-Share). Live leg: upload from demo-hp's filebrowser > container (uid 1000) landed on demo-felhom's share dir and deleted clean; dead-NAS gave > `Host is down` in seconds (no hang) and recovered unaided after samba restart. RESIDUAL for the > operator: the FileBrowser HTTP click-through — its admin credential is customer-held (CC got 401 > on admin/admin and the demo password; by design). ROADMAP R-67 SHIPPED (coupled to R-64). > **2026-07-22 — v0.159.0 (R-66), SHIPPED + deployed to BOTH boxes.** Three legs: „Hálózat" card on > Beállítások → Rendszer (Helyi cím / Hálózati név only-while-Megosztás / Átjáró; „—" fallback), > `network` section in the Debug dump (best-effort per item), and the NetBIOS trap named on the NAS > add form (Szerver helper text + a purely lexical hint on `unreachable` for single-label non-IP > names). > > **DECISION (the load-bearing one): all guest-net reads go through the samba netns door.** The > controller is bridge-netns'd, so `/proc/net/route`/resolv.conf/net.Interfaces in-process answer > for the CONTAINER (172.x / 127.0.0.11) — the S-2 trap. `internal/stacks/guestnet.go` docker-execs > into host-networked felhom-samba (one `guestNetExecFn` seam); Megosztás off ⇒ door closed ⇒ „—" / > in-place error strings, never a plausible-wrong substitute (S-5). Nothing stored anywhere. > > Deploy: 0.159.0 on demo-felhom 9201 (open-door path live: .104/.1/\\FELHOM) AND demo-hp 9201 via > G1 break-glass (closed-door path live: dashes, no name row, in-place dump errors; secret shredded). > demo-hp gotcha worth keeping: the controller 404s on direct container-IP probes without the > customer-domain Host header (`felhom.enkisfelhom.hu` there). Red-proofs A2 + C2 run and recorded. > ROADMAP: R-66 SHIPPED; R-64 (pairing blessed, drill = evidence leg) + R-65 (buddy-box replication, > post-alpha spike-first) minted. NAS doc gained the naming-caveat paragraph. > **2026-07-21 — v0.155.0.** v0.154.0's wizard sourced "is an op running" from `Manager.IsRunning()` > — the CONCURRENCY single-flight, acquired inside the goroutine, and **`RestoreOffboxScratch` never > acquires it**. So the execution step was unreachable for „Ellenőrzés" and the full-restore > preparation: live buttons while a restore downloaded, with the progress banner contradicting the > phase strip on the same screen. Found by the operator on the first live click-through. > > **DECISION: display reads `RestoreStatus()` (the `opRunning` flag), never `IsRunning()`**, through > the named `restoreOpInFlight` seam, and the handler reads the status ONCE per render so the strip, > the suppression and the running-op name cannot diverge. The lesson generalises: `opstatus.go` is the > DISPLAY surface and says so in its own header — the concurrency flag is not a substitute. > > **The test lesson:** a table test over a pure function proves the function, not the caller. Scenario > E passed throughout because it injected `OpRunning=true` directly. The new test drives a real > `Manager` through `BeginRestoreOp` and asserts the render. > > **DECISION: „Eredmény" earns its place.** The strip's highlight is now `Phase`, derived separately > from `Step`: a finished restore is back on the intent step while the strip reads „Eredmény" and an > outcome card shows the result — window-bounded (10 min) and app-bound. > **2026-07-21 — v0.154.0 (R-48).** Collapses the offsite restore controls to a single > „Visszaállítás…" entry per app row plus a per-app wizard at `GET /backups/restore/app?name=`. > The defect it closes is the CAUSE of the round-2 incident: the list rendered up to five inline > forms per row, two of which — the missing-only merge and the true reconstitution — were sibling > buttons whose difference is whether the data comes back. The rule it establishes: *two adjacent > controls whose difference is "your data comes back" vs "your data cannot come back" must never be > distinguishable only by layout.* > > **DECISION: the wizard is server-rendered on the EXISTING endpoints.** No new mutation endpoint, > no JSON state API, no client router. Every card is a real form POST to > `/backup/offbox/{restore,place,reconstitute}` with the same field names and gates, and the server > renders the next step — so it works with JavaScript disabled. `TestRestoreWizard_NoNewMutationEndpoints` > makes that structural: adding a form that posts somewhere new fails the suite by design. > > **DECISION: R-45 stays its own item.** The wizard polls the two existing status surfaces as-is; the > generalized job registry (and with it a real per-phase progress feed) is not built here. > > **DECISION: the step is derived, never requested.** `deriveWizardStep` is pure over (op running, > size-gate flash, scratch ready). Precedence is load-bearing — a running op outranks a stale > `?full_prep=` in the URL, or a commit button reappears mid-restore. While ANY op runs every > mutation form is suppressed server-side rather than offered and then refused with a 409. > > Latent bug found and fixed on the way: `offboxRedirectTo` hardcoded `"?"` when appending its flash, > which would have buried the flash inside `?name=`. **No agent coupling — MinAgent stays > 0.90.0.** 9 new tests + the Group-B red-proof; full suite green. > > **NOT live-validated at commit time by design:** v0.154.0 is published but deliberately NOT > hand-deployed — the operator's hub floor save (0.153.0 → 0.154.0) pulls it via the self-update > path, and that swap IS the R-23(a) single-fire validation (STOP-1). > **2026-07-20 — v0.153.0 (R-47).** Closes the H4 race on **BOTH** restore paths. The replay needs a > running DB container, so both paths started the WHOLE stack first — giving the application a window > to rebuild the schema objects the dump was about to create. Measured at 8 s on 2026-07-19 > (`DIAG-immich-restore-round2-2026-07-19`): immich-server rebuilt `clip_index` two seconds before > the dump's `CREATE INDEX`, the replay aborted `already exists` under `ON_ERROR_STOP=1`, and immich > then reported schema drift. The photos came back **by accident** — `pg_dump` emits COPY before > CREATE INDEX, so the abort landed after the rows; a collision earlier in the script would have left > a genuinely half-restored database, reported identically. > > **DECISION: the DB-only bring-up is done by compose SERVICE scoping**, not by container tricks — > `StartStackServices(name, []string{svc})` → `compose up -d `. Every catalog template's > dependency direction is app→db, so naming the DB starts the DB and nothing else. `docker start > ` was never an option: `StopStack` is `compose down`, so the containers no longer exist. > `RestartStack`/`RedeployFromEnv` are traps here — both end in a full `up -d`. > > **DECISION: fail-closed.** A `.sql` dump with no identifiable DB service refuses BEFORE the first > mutation, on both paths (one Hungarian string, shared). The alternative would be to start everything > and replay into the race. It should be structurally unreachable — `dbTypeForImage` is now shared by > `DiscoverDatabases` and `DBServiceNames`, and a dump can only exist because discovery matched the > container's image, which IS the compose `image:` value — so this is the belt for template drift. > > Enablers: `RedeployFromEnv` split into `PersistUnitRedeployConfig` (persist, starts nothing) + the > unchanged tail; `StackDataProvider.RecreateStackFromUnit` renamed to > `RecreateStackDefinitionFromUnit` because the old name promised less than the method did — the > hidden `up -d` inside it is what carried the defect on the local path. `StartStackServices` REFUSES > an empty list (argument-less `up -d` is a full start). **No agent coupling — MinAgent stays 0.90.0.** > 19 new tests, 3 red-proofs, 23/23 green. **NOT live-validated yet:** STOP-1 supervised reconstitute, > golden 0.153.0 bake (P3 registry-reachability probe from the vacation site is load-bearing), Viktor's > two hub saves, and his C6 customer-restore UI run. > **2026-07-20 — v0.152.0 + felhom-samba 1.1.0 (Megosztás on a Mac).** Closes **S-3**. **A capture > on the box overturned the earlier guess:** macOS DOES send a correct NBNS query for `<20>` and > nmbd DOES answer it correctly in 140 µs (flags `0x8580`, RCODE=0, right address) — macOS simply > never acts on it. NetBIOS there feeds legacy browsing, not `smb://` URL resolution, so **the bare > `smb://` can never work from a Mac** and nmbd was never the broken part (it is what serves > Windows). felhom-samba 1.1.0 adds **avahi + dbus**, templating `avahi-daemon.conf` and the > `_smb._tcp` service file from `FELHOM_SERVER_NAME` so a rename re-advertises; both daemons are > non-fatal on failure. v0.151.0's card had offered `smb://` for Mac — the one dead form — now > `smb://.local`; Windows keeps flat `\\`. Spiked live by hand and confirmed from the > operator's Mac BEFORE publishing the image (the operator's call, and it chose the design too). > **STILL OPEN: Finder-sidebar discovery is NOT shipped** — the record is published and answers > browse queries, but was never observed working; likely a Finder Settings → Sidebar toggle, but > unverified. **Windows was not retested.** Two test bugs fixed en route, neither a production > defect: `TestRenderSambaCompose` pinned a literal image tag, and `TestFabUpload_GCAndIdleTimeout` > asserted an async unlink synchronously (it passed alone, failed in the full package once the new > render tests made `web` heavier). 23/23 green twice; 2 red-proofs. > **2026-07-20 — v0.151.0 (Megosztás).** Closes **S-1/S-2/S-4-core/S-5** of > `felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md`; **S-3 (no mDNS/Bonjour) stays OPEN**, > awaiting Viktor's `smbutil lookup FELHOM` + `dns-sd -B _smb._tcp` from the Mac. **The `/sharing` > page had been reload-looping at ~1.2 s for every customer with sharing enabled since v0.147.0** — > `/sharing/status` coerced `idle`→`running` on the JOB phase channel, and the client answers a > terminal `running` with a one-shot `location.reload()`, so the first poll of every steady-state > page load re-armed it. The rule this leaves behind, now recorded against R-45 too: **a phase a > client answers with a one-shot action is an EDGE — never synthesise it from a level, and serve it > exactly once.** Both halves are server-side; `sharing.html`'s `