# CONTEXT.md — Project Memory
> This file serves as persistent project memory across Claude Code sessions.
> It replaces the auto-generated "Memory" from the claude.ai Project.
> **Update this file at the end of each working session** with current state,
> recent decisions, and anything the next session needs to know.
>
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-08-23 (v0.223.0 — R-329: the alarm reached nobody; R-386: ask the field that knows)
> **2026-08-23 — v0.223.0 (R-329 + R-386), and a defect that only became visible once another was fixed.**
>
> **[RULING] The severity a controller sends is the HUB's vocabulary: exactly
> `{info, warning, error, critical}`.** Anything else is **coerced to `info` at ingest, silently**, and
> `severityNotifies` drops `info` **before both** delivery legs. `app_start_failed` emitted `"warn"`.
> **Measured on the live hub DB: 91 such events stored all-time, ZERO notification rows ever.**
>
> **[FACT] This was the SECOND occurrence, and the first one's comment had recorded the lesson.**
> `DiskAlertKind.Severity` emitted `"warn"` until v0.215.0. **A comment is not a guard** — the guard is
> now an AST walk over the whole controller. **grep cannot do this job:** `"warn"` is a legitimate
> *healthcheck status* in `internal/monitor` and `internal/selftest`; the sweep hit nine such strings
> and exactly one defect. The walk cannot follow a variable, so the **six** dynamic call sites are
> registered by name with the values each can take — **an unlisted limit is not a limit, it is a hole**.
> The guard found two of those six that the hand sweep had missed.
>
> **[FACT] It hid because another defect hid it.** R-384's ordering bug meant `app_start_failed` could
> not fire at all, so a broken severity had nothing to break. **Fixing one defect made another
> reachable** — and the same shape appeared again downstream: the operator cooldown key carries no app
> identifier, so **only the first app-down per hour now e-mails the operator** (R-182's shape, newly
> load-bearing, filed not fixed).
>
> **[RULING] `app_start_failed`: operator always, customer OFF by default.** `processOperator` never
> consults customer preferences, so one word fixed the operator leg and left the customer leg where the
> ruling wanted it. **Deliberately NOT in `operatorOnlyEvents`** — that would make the new toggle
> visible, flickable and structurally incapable of delivering.
>
> **[RULING, R-386] "The customer stopped this" is a RECORD, never an inference.** `aggregateState`
> folds `StateExited` into the stopped counter, so an out-of-band stop and a customer's Stop are
> byte-identical on the Docker side — no state test can separate them. Ask `DesiredState`, which has
> exactly one writer. `Stopped` → no alarm; `Running` → **alarm**; **absent → UNKNOWN, keep the old
> behaviour AND announce it**, because reading absent as "nobody asked" would e-mail about every app
> anyone ever stopped, fleet-wide, on the first cycle after upgrade. **The backfill cannot help — it
> seeds `Running` only from an observed-UP reading.**
>
> **[MECHANISM] `IntentUnknown` + an INFO line naming the apps.** A rule without a mechanism is a wish.
> Measured on `demo-hp`: **0 of 8** deployed apps carry an absent intent.
>
> **[FENCE] Adding a `DesiredState` WRITER is the fenced act; reading is fine.** And `failedRestart`
> must still lift a `Stopped` intent, or F-CRIT-1 re-opens.
>
> **[TRAP, cost three attempts] An HTTP 200 can be a REFUSAL.** The settings save answers 200 while
> rendering the empty-email wipe-guard error. Scenario G's before/after hashes matched twice because
> **nothing was saved**, not because nothing changed. And the email `` spans three lines, so a
> single-line grep reads it empty. **Assert the refusal banner is ABSENT before believing a save.**
>
> **[TRAP] A red-proof that passes may mean an INERT mutation.** `if next <= prev` → `if next < prev`
> in fillwatch changes nothing, because an earlier `if next == prev { continue }` already removed the
> equal case. Check the mutation applied before believing either verdict.
>
> **[RULING] The compound-toggle split's risk was the MIGRATION, not the split.** A save whose event
> SET is unchanged now stores the existing slice **verbatim**, so byte-identity is by construction —
> without that guard the defaults case reorders, and the red-proof caught it.
> **2026-08-23 — v0.222.0 (R-384 + R-383), and a bigger hole found by a measurement that was told not to fix it.**
>
> **[DECISION] A dead SUPERVISED member is asked about BEFORE a failing healthcheck, because they are
> different questions and the second was answering the first.** `aggregateState` returned
> `StateUnhealthy` the moment `unhealthy > 0`, and the R-51 mixed-case block that asks "is a supervised
> member dead?" sat below it. A two-container app whose database exits goes `unhealthy` seconds later
> *because it cannot reach that database* — so **the symptom the fault causes was what suppressed the
> alarm for it.** The supervised test is now hoisted above the unhealthy/starting/restarting returns.
>
> **[DECISION] "Some members are up" means ANY member not in the down bucket** — running, unhealthy,
> starting or restarting. The old guard was `running > 0` counting `StateRunning` alone, which made the
> R-51 block **unreachable in precisely the case it was written for**: an unhealthy survivor beside a
> dead database counted as nothing being up. Either half alone leaves the defect standing, and the two
> red-proofs convict independently.
>
> **[FENCE, unchanged] `IsDownState` is byte-identical and `unhealthy` stays excluded.** An unhealthy
> container is RUNNING; folding it in reintroduces the flapping that exclusion exists to stop. **No new
> state was minted** — `StateDegraded` already meant this and every consumer already handled it. The
> fix is an ORDER, not a widening.
>
> **[FACT] The register's own suggested remedy was wrong.** R-384's row proposed a sustained-`unhealthy`
> threshold on the `crashLoopAfter` model. The defect needed no threshold at all. **A register's
> "recommended fix" is a hypothesis written before the diagnosis, and must be re-derived from source.**
>
> **[TRAP] Three existing subtests pinned the DEFECT as settled behaviour.**
> `TestAggregateState_UnchangedBranches` asserted an unhealthy/starting/restarting member beat an
> `exited` peer that was on `unless-stopped`. They were amended (down member given a benign policy,
> which is the only case where that sentence was ever true) and the change is reported, not buried.
> **A green suite can be green about the wrong thing.**
>
> **[FINDING — R-386, OPEN, NOT FIXED] A single-container app stopped out of band raises NO alarm, and
> a comment states the opposite.** `aggregateState` folds `StateExited` into the `stopped` counter, so
> an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**;
> `classifyRunStates` then whitelists `stopped` as a deliberate user stop. The comment at
> `cmd/controller/main.go` claiming an out-of-band `docker compose stop` "still alerts" is **measured
> false** — `privatebin`, 9 scans, 0 events, 0 banner. **Case #10 of "a comment asserting an invariant
> the code does not provide".** The task asked for this as a MEASUREMENT and forbade a fix; it is filed.
>
> **[DECISION, R-383] A message may not assert a file exists without asking the disk.** The
> double-failure sentence named the undo copy as present, built from the returned path — and a missing
> file is one of the two ways that rollback fails. `undoCopyPhrase` now reads from disk; a zero-length
> dump counts as MISSING; and the absent case still names WHERE the file should have been, because
> R-351's lesson is that a refusal naming nothing forces someone to remember what the product knows.
>
> **[DECISION] The alarm ladder now has an owning document** —
> `felhom.eu/documentation/architecture/08-alarm-ladder.md`. Until 2026-08-23 no document owned it; the
> rules lived as comments in four packages, each locally correct, with the ordering between them legible
> only by reading one function top to bottom. **That absence is why R-384 survived review.**
>
> **[GOTCHA] `app_start_failed` still ships severity `warn` (R-329)**, which is not in the hub's
> vocabulary and coerces silently to `info`, e-mailing nobody, while the POST returns 200. R-384 moved
> this event from unreachable to load-bearing, so the severity bug now matters.
> **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.**
>
> **[DECISION] `db_dumps` lists the app's OWN dumps, not the `pre-restore-*` undo copies.** They are
> local material for a restore that went wrong, not part of the app's recovery set. **Every consumer
> of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`** (the
> declaration, the enumeration, the change-detection compare); none reads it for recovery, and no hub
> or agent consumer exists. Three copies per app were being pushed off-site permanently for no
> recovery value. **The files are neither deleted nor hidden** — their visibility is a separate
> recorded decision and it stands.
>
> **[TRAP, and it bit within minutes] A stable `db_dumps` lets `CaptureRecoveryUnit`'s already-current
> early return fire.** Anything that must happen on EVERY capture — bounding the undo copies — has to
> sit ABOVE that check. It did not, and the cap silently stopped applying: four copies against a cap
> of three, counted on the box. Fixed in v0.221.1. **One change made another unreachable, and only
> counting files on a real machine showed it.**
>
> **[FACT] The comment was the defect.** `writeSafetyDump` called `DumpOne` into the app's own unit
> and renamed afterwards; `DumpOne` writes the canonical `-.sql`, so every safety dump
> destroyed the app's real backup. The comment said the rename meant it "can never overwrite the app's
> real dump" — false as written, for four months. The fix is a DESTINATION (`DumpOneTo`), not a
> rename, and the `.tmp` derives from the final path so a nightly dump beside it cannot collide.
> **`DumpOne`'s signature did not move.**
>
> **[NEGATIVE — do not re-derive this] A HELD app does NOT raise the dead-app alarm.** It was read
> from source that it would, because it keeps its database container and so is not `StateStopped`.
> Measured on the shipped v0.220.2: it aggregates to `unhealthy`, `aggregateState` checks
> `unhealthy > 0` before the mixed-case degraded branch, and `IsDownState` excludes `unhealthy`.
> Heartbeat read `0 currently down` throughout. **No suppression was built.** The same measurement
> exposed **R-384**: an app whose database has died is `unhealthy` too, and is likewise silent.
>
> **Proven live on `demo-hp`:** the canonical dump's sha256 unchanged across a restore on both engines
> — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`. Evidence:
> `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
>
> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the
> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A
> running app on a half-written database lets the customer type into it, and that turns a recoverable
> state into a permanent one. The alternative — start it and mark it — was put to the operator and
> declined. If that judgement is ever revisited, this is the sentence to revisit.
>
> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored
> database; the only difference was whether it looked broken (Postgres emptied and crash-looping,
> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not
> transactional. Putting the customer's own copy back is what removes the half state, and it is the
> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps.
>
> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such
> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older
> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a
> two-database app would have restored one and left the other half-written.
>
> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found
> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as
> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec`
> against the dead id timed out — so the app was held for an infrastructure reason while its data was
> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and
> never look at container identity.**
>
> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure
> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either —
> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start
> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared
> `driveStartGate` **above** its driveless early return, because the apps this exists for have no
> drive.
>
> **The way out is `--clear-restore-hold `, and it REQUIRES A CONTROLLER RESTART** — it runs as a
> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command
> say so. Clearing through the running controller is the right shape later; it needs an operator tier
> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own
> hold is what the hold prevents).
>
> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state
> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was
> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`.
> **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.**
>
> **[DESIGN] The restore destination is resolved by the SAME rule as the capture destination.** The
> drive if the app declares one, the system data path otherwise — `Manager.GetAppDrivePath`, one
> expression, now used by `CaptureRecoveryUnit`, `ReconstituteFromOffsite` and `PlaceOffsiteRestore`
> alike. Anything else and the restore aims somewhere the backup never came from, which surfaces as a
> placement-mismatch prompt on a box where nothing actually moved.
>
> **The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351)
> applies to apps that HAVE a drive to get wrong.** It used to be reached by `HDD_PATH == ""`, which
> also stood in for "is this app installed?". Measured in the catalogue at `459766cb1639`: **53
> templates, 13 `needs_hdd: true`, 40 `false`** — so for 40 apps that test was permanently true and the
> off-site restore refused them forever, while they were running, telling the customer to reinstall
> them "in the same place", which those apps never offer. **An app with no drive is not misconfigured**
> (`01-topology-and-trust.md` §8, `[DESIGN]`); it is the majority case, and between 19 and 22 August it
> was called a defect four times.
>
> **Now:** *installed?* is asked of `ListDeployedStacks()` via `Manager.isStackDeployed`, which **fails
> CLOSED on a nil provider** — "cannot tell" must not become "go ahead" when the next act is a write.
> *Where?* is asked of `GetAppDrivePath`. A third refusal, with its own sentence and its own route,
> covers installed-but-no-resolvable-data-root: widening `nincs telepítve` to cover that would send a
> customer to reinstall a running app and hide the real fault.
>
> **FENCED, and not changed:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`).
> Capture resolves an app's declared `userdata`/`import` file legs against that value; a system-data
> fallback there would write a snapshot claiming to hold the customer's files and not holding them. The
> fenced ACT is "introduce a fallback into capture-side path resolution" — reading the value elsewhere
> is fine.
>
> **Proven live on `demo-hp`, 2026-08-22.** `privatebin` (driveless): planted through the app's own
> HTTP API plus a direct file plant, off-sited, **deleted**, restored through
> `POST /backup/offbox/reconstitute` — **15/15 files back byte for byte**, two Hungarian accented names
> included, message „0 fájl és 1 adatkötet visszaállítva". The scratch was unit-only, exactly the shape
> nobody had ever driven to completion before, and every downstream leg held. `calibre-web` (drive
> app) walked the same way and did not move. Evidence:
> `felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`.
> **2026-08-22 — v0.218.0 (R-354/R-355). The database nobody backed up, and the restore that
> returned most apps nothing.**
>
>
> Both fixes came out of the 2026-08-21 backup-truth drill. **R-355 went first because it is the only
> place in the product where one customer action causes permanent total loss:** paperless-ngx's database
> was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the
> off-site copy or the restore — and the same misattribution meant a destructive restore of that app took
> NO undo copy, then told the customer the app has no database. Fixed by reading the compose project
> label, which is the stack name by construction. **One app of 53 affected**, established with a sweep
> proven able to convict by planting a second mismatch.
>
> **R-354:** the off-site restore had no named-volume leg at all. The archives live inside the unit, whose
> placement is correctly skipped, and the comment beside that skip said the dump is replayed from the
> scratch "so nothing is lost" — true of the database, false of the volumes. `restoreDockerVolumesFrom`
> now replays them from the scratch unit; `VolumesReplayed` reaches the message.
>
> **WHAT THIS DOES NOT FIX, and it is the blocking item for the apps that need it most:** the off-site
> restore still REFUSES outright for the 40 of 53 apps that declare no data drive (**R-356**), saying a
> running app „nincs telepítve". Those are exactly the apps whose entire dataset is a named volume, so
> R-354's fix cannot reach them until R-356 is closed. Proven again on hardware 2026-08-21. The live
> confirmation of R-354 was therefore done on `calibre-web` and `paperless-ngx`, which declare a drive
> and can reach the restore.
>
> Also still open from the drill: the empty-restore success message (R-353), where the 40 apps' data
> lives (R-352), and the remaining rows R-357..R-366.
>
> ## THE RESTORE'S OWN MEMORY (v0.217.0, 2026-08-21) — R-351 / R-352 / R-353
>
> > **Both sides of a written fact must be checked, not just the writing side.** Every recovery unit
> > has recorded `drive` and `namespace_root` since schema 1. **No non-test code in the repository ever
> > read either back.** A restore into a different destination than the backup recorded therefore
> > succeeded silently under a green message. This is the same shape as several defects closed this
> > month, and the cheap test for it is one grep: *who reads this field?*
> >
> > **`IsRunning()` is still the wrong flag, in one more place than we knew.** v0.154.0 fixed the wizard
> > and left a comment explaining why. The **seven handlers** were never moved over, so a second press
> > genuinely started a second run and reported „…elindult". A comment explaining a trap does not fix
> > the other call sites — grep for them.
> >
> > **A result nobody can see is the same defect as no result.** The banner gated its terminal state on
> > a page-local `sawRunning`. The 8.666 s OpenGist restore finished before any poll saw it, so no
> > screen said it had completed. Fixed with `RestoreOpStatus.LastRecent` — and the window now lives in
> > `internal/backup` as ONE expression that both surfaces read.
>
> **State, and what is next.**
>
> - **Shipped:** placement comparison + named mismatch + `ack_placement`; not-installed refusal names
> the recorded drive; deploy prefill from the app's own backup; `restoreOpBlocked()`; `LastRecent`;
> off-site listing bounded-concurrent (measured 16.1 s → two waves).
> - **R-352 partly closed.** 40 of 53 catalogue templates declare no data path, and
> `GetDefaultStoragePath()` is read by nothing that places data — its comment `// new apps use this by
> default` has never been true. **Only visibility shipped**; the deploy page now states where the data
> will live. **No placement changed, nothing migrated.** Specification:
> `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
> - **R-353 is the next session's first item.** A restore whose unit carries no `db_dumps` and no
> `volume_dumps` reports a bare completion. OpenGist's unit held configuration and nothing else, and
> the restore said only that it had finished. Fix the outcome first; *then* prove the off-site
> coverage of a named-volume app by running a dump cycle — do not close the first on the second.
> - **Open and unproven:** whether the 40-class reaches the off-site tier at all has **not been
> observed**. `runVolumeDumps` covers them on paper; every unit on the box read `volume_dumps: None`
> because no nightly run had happened yet.
>
> ---
>
> ## THE TWO RULES THE RECOVERY JOURNEY LEANS ON (v0.203.0, 2026-08-06)
>
> > **1. A credential the hub stages is collected by the box, not waited for.** The reconcile that
> > collects runs on a tick for exactly as long as the box's own declaration says it needs one — and
> > stops the instant a target exists. It is driven from `OffboxReportStatus().State`, the same statement
> > the hub acts on, so the two can never disagree about whether a retry is wanted.
> >
> > **2. A mount Felhom itself made is not "something else".** Enrolment mounts a drive twice — the
> > managed path and a raw `/mnt/` on the host — and the host survives a guest rebuild while the
> > guest's registry does not. The claimed check forgives a non-managed mount **only when corroborated**
> > by the same device also being mounted under the managed path. **A genuinely foreign mount is still
> > refused, and that fence has its own test.**
>
> **Why both are stated here rather than left in the code:** each was a dead end that kept the unaided
> recovery journey failing, and each looked correct in isolation. R-218's declaration half shipped and
> worked while nothing consumed what it asked for; R-220's check was right about foreign disks and wrong
> about our own. **Neither is a bug in the thing it guards — both are about what runs, and when.**
>
> Two things that must not be "simplified" back:
> - **The settle gate stays.** The retry goes through `ReconcileWhenSettled`, so the day-0 floor race is
> unchanged. A retry that skipped it would trade one defect for another.
> - **The R-220 exemption is corroborated, never a prefix.** Widening it to any `/mnt/*` path offers a
> disk another system is using for formatting — the red-proof shows exactly that.
>
> ## THE UNLOCK PATH'S RULE (v0.202.0, 2026-08-06) — state it before changing anything there
>
> > **On the recovery unlock path the customer is blamed only after a real attempt REFUSED their code.
> > Every other outcome — including one that cannot be classified — says something else.**
>
> This is the rule, and it outlives the bug that produced it. It was learned twice, because fixing it
> once was not enough:
>
> - **v0.201.0** stopped an agent that is too OLD from being reported as a wrong code (R-216).
> - **v0.202.0** found the same defect through a different door: an agent that is **stopped**, and a hub
> that cannot be **reached**, still fell through to a message about the code. Measured with a
> **correct** code at 0.0299 s and 0.0556 s, against ~1.0 s for a real unseal — the machine accused the
> customer of something it had not tried (R-224).
> - And the inverse: the one message that says *"check your ten words"* was unreachable on any box that
> had re-escrowed, which is exactly the box a customer has just recovered (R-226).
>
> **How it is enforced.** `agentapi.ClassifyRecoveryFailure` maps the failure to one of five classes
> **from the value, never the text**; the typing message is reachable from **one** of them
> (`RecoveryAskedAndRefused`, i.e. HTTP 400, i.e. the bundle was fetched and `age` refused it); and the
> zero value is `RecoveryUnknown`, which renders **neutral**. **The safe default is the load-bearing
> part** — an unrecognised status must not fall into an accusation.
>
> **Two things that are deliberately NOT how it works, and must not be "fixed" into it:**
>
> 1. **Elapsed time is never a classifier.** It is what diagnosed this, it is logged for the operator,
> and that is all. A duration guard would be a second thing that can be wrong.
> 2. **The error's TEXT is never read.** A string match is a defect waiting for a rewording. When the
> distinction was not available as a value, the **agent was changed to provide one**
> (`escrow.ErrBundleFetch` → HTTP 502, agent v0.126.0, `MinAgent 0.126.0`) rather than parsed for.
>
> **The coupling degrades safely and silently:** an agent below 0.126.0 answers 400 for both causes, so
> `FeatureRecoveryFailureClass` withholds the refusal reading and the 400 becomes neutral. The gate
> blocks nothing; it only decides whether the customer may be told to check their typing.
>
> ## CAMPAIGN 11 — what changed in v0.201.0 (2026-08-05)
>
> **The off-site key recovery is a COUPLED feature and now declares it.** It needs agent **0.125.0**
> (`POST /escrow/recover-offsite-password`). `FeatureOffsiteKeyRecovery` has a `featureProbes` row, a
> `featureMinAgent` row and a `Supports` gate at the unlock entry point.
>
> ⚠ **That gate FAILS CLOSED — alone in that table.** The package default is fail-open, and that default
> is what produced R-216: an agent that could not answer 404'd, the unlock was attempted anyway, and the
> customer was told their correct recovery code was wrong. Anything but `SupportYes` now says *the
> machine* cannot ask yet, and no attempt is made. Do not "fix" it back to the package default.
>
> **The box declares `needs_credential` until the TIER WORKS, not until a key exists.** The old
> short-circuit on "a repository password is present" is deleted: installing one is the recovery
> screen's whole job, so it made succeeding at recovery switch off the mechanism that delivers the
> coordinates to use it. A disabled target still short-circuits at the first line (Scenario E).
>
> **The unlock finishes the job**: place the key → bring the tier up (`offsiteapply.Bridge.Reconcile`,
> wired via `SetRecoveryTierUp`) → list. Without the middle step the promised listing can never render on
> shape (a), because no key ⇒ no target ⇒ no inventory.
>
> **Four messages, not one.** Wrong code (the only one mentioning typing) · the machine cannot ask ·
> the store could not be read / the connection details have not arrived · the code belongs to a RETAINED
> earlier package. The last one is driven by the ACK's `superseded_present`/`superseded_at` (hub
> v0.97.0) and **promises nothing** — no read path for a superseded package exists.
>
> **Still open from the campaign:** R-214 (console pairing banner), R-220 (drives unenrollable after a
> rebuild), R-221 (a rebuilt box cannot run the escrow ceremony), R-223 (the Day-0 manifest still vouches
> agent 0.120.0 — operator decision).
>
> ## R-203 (v0.197.0) — the namespace-root contract, and what `ok` now means
>
> **The contract, in one line:** `appbackup`'s path helpers (`UserdataDir`, `PrimaryBackupPath`,
> `RecoveryUnitPath`, `AppDataDir`) take a **NAMESPACE ROOT**. Anything that came out of `HDD_PATH` or a
> `StoragePath` is a **DRIVE path** — put it through `appbackup.NamespaceRootFor(drive, systemDataPath)`
> first. `UserdataDir(bareDrivePath)` still compiles and is still wrong; five callers proved it.
>
> **The rule now has ONE expression.** `NamespaceRootFor` / `IsEnrolledDrive` in `appbackup`;
> `backup.Manager.namespaceRoot` and `stacks.Manager.inGuest` delegate. There were two copies before and
> **they differed** — one compared without `filepath.Clean`, the other with it.
>
> **Why it was invisible:** on an enrolled drive the namespace root IS the drive path. The two diverge
> only on the system-data fallback, which `paths.go:26` names as a supported arrangement.
>
> **`last_status` gains `incomplete`.** A run that could not capture a directory an app declares
> MANDATORY is not a successful run. **Not `error`** — the rest of the run worked, so `SnapshotCount`
> and `LastSuccess` still record what WAS captured. It reaches the operator via the existing
> `backup_run_failures` digest (a new event type is a two-repo change; the hub drops unlisted types).
> The Hungarian customer warning is unchanged; the page renders `! Hiányos`.
>
> **Still open, and NOT fixed here:** `resolveAbs` resolves `RootHDD` and `RootUserdata` against the same
> root. Both callers now pass the namespace root so the export and the backup agree with each other, but
> whether `${HDD_PATH}` should mean the namespace root on the system drive touches every deployed app's
> binds and needs a decision, not a patch.
>
> **Blast radius, measured before changing anything:** exactly one app in the fleet had
> `HDD_PATH == system_data_path` (`calibre-web` on demo-hp, the R-201 drill fixture). Its data was
> migrated and its sentinel re-verified byte-identical.
>
> ## R-203 (2026-08-04) — a MANDATORY userdata directory can be absent from the off-site snapshot while the run says `ok`
>
> Found live on demo-hp while staging the R-201 drill, and it **halted that drill**.
>
> `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment
> **when the drive IS the system data path** — `m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath !=
> m.systemDataPath)` (`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is
> `/userdata` (`stacks/classify_binds.go:14`).
>
> With `system_data_path: /mnt/sys_drive` and `calibre-web` deployed at `HDD_PATH=/mnt/sys_drive`:
>
> live bind (files land here): /mnt/sys_drive/userdata/media/books ← exists
> capture set looked for: /mnt/sys_drive/felhom-data/userdata/media/books ← does not
>
> **The same compose used BOTH roots** — `${IMPORT_PATH}` resolved *with* the segment,
> `${USERDATA_PATH}` *without*. The run logged one `[WARN] mandatory data path missing on disk, skipped
> from offsite`, then `0 mandatory path(s)` and **`backup OK: 3 app(s), 3 snapshot(s)`**, with
> `last_status: ok`. Nothing customer-visible or hub-visible said the directory was dropped.
>
> **Not established:** whether `HDD_PATH == system_data_path` is a supported deploy. It was accepted
> (HTTP 202) one call after the NAS path was correctly refused (R-108). **Either branch is a defect** —
> broken resolution, or a missing refusal.
>
> **Two things a fix must do:** make the two roots one function, and make a skipped **MANDATORY** path a
> customer/hub-visible signal rather than a container-log WARN. `opengist`/`privatebin` declare no
> mandatory userdata paths and are unaffected.
>
> ## R-200 (v0.195.0) — the offsite key recovery diagnostic
>
> `--recover-offsite-check` is a `docker exec` escape hatch (the `--print-reset-code` shape), NOT a page
> or an API a browser can reach. R comes from **STDIN** — never argv, never `ps`, never shell history,
> never a transcript. It asks the agent (>= v0.125.0) to fetch this host's sealed bundle and open it,
> then reports whether the recovered repository password matches the on-disk one **by sha256**.
>
> docker exec -i felhom-controller /usr/local/bin/felhom-controller --recover-offsite-check < /path/to/code
>
> **IT COMPARES AND NEVER INSTALLS.** `CheckOffsiteKeyRecoverable` must stay free of any write — if a
> future change makes it place the recovered password, it stops being a diagnostic and needs the drill's
> supervision (that is link 9, R-200's remaining half). Pinned by
> `TestCheckOffsiteKeyRecoverable_WritesNothing`, whose red-proof is adding the install call.
>
> Exit codes are load-bearing: **0** match, **2** a clean MISMATCH, **1** a step failed. A mismatch is a
> finding about the system; a failure is a finding about the run, and they must never share a status.
>
> **Proven live on demo-felhom 2026-08-04** — recovered sha256 == on-disk sha256 == the hub's stored
> hash. Nothing customer-facing ships with it: no card, no form, no preview.
>
> ## About Viktor (project owner)
>
> - Works at Deutsche Telekom (Budapest), building Felhom.eu as a side business
> - Felhom.eu: managed home-server service for Hungarian households
> - Technical but prefers pragmatic solutions over over-engineering
> - Runs all infrastructure on Gitea (gitea.dooplex.hu), k3s cluster for management
> - Customer deployments use Docker Compose (not Kubernetes) for simplicity
>
> ### felhom-controller (this repo)
> - **Version:** v0.16.1
> - **Phase 1:** ✅ COMPLETE — Stack Manager + Deploy Flow
> - **Phase 2:** ✅ COMPLETE — Monitoring & Health (scheduler, CPU/temp, healthchecks.io pings)
> - **Phase 3:** ✅ COMPLETE — Backups (DB dumps, restic integration, manual trigger, **dedicated backup page**)
> - **Phase 4:** ✅ COMPLETE — Monitoring Page with Metrics Store (SQLite, Chart.js, system + container metrics)
> - **Phase 5:** ✅ COMPLETE — Authentication, Persistence & Settings Page (settings.json, password change, session management)
> - **Phase 6:** ✅ COMPLETE — Monitoring Warnings, Dashboard Alerts & Notification System
> - **Phase 7:** ✅ COMPLETE — Storage Overview, Per-App Backup Toggles & Limited Restore
> - **Phase A:** ✅ COMPLETE — Storage Paths Foundation (registry, auto-discovery, per-app HDD_PATH, deploy dropdown, health monitoring)
> - **Phase B:** ✅ COMPLETE — Storage Management UI Polish & Health Severity Fix (flash messages, label editing, app details, FS info, deploy free space, backup context)
> - **Phase C:** ✅ COMPLETE — Storage Init Wizard, Data Migration & Startup Fix (disk scan/format/mount wizard, rsync-based migration, startup pings)
> - **v0.11.1 bugfix:** ✅ COMPLETE — Storage Scan: system disk detection via host fstab + blkid UUID resolution; FSType enrichment via `blkid -o export`
> - **v0.11.2 bugfix:** ✅ COMPLETE — /host-dev mount for block device access; `HostDevicePath()` helper; all format/scan/safety ops use /host-dev
> - **v0.11.3 bugfix:** ✅ COMPLETE — Added `fdisk` package to Dockerfile (provides `sfdisk`; not in `util-linux` on Debian bookworm)
> - **v0.11.4 bugfix:** ✅ COMPLETE — FormatAndMount: fixed sfdisk (wipefs+force+`,,`), mount (explicit device path), mount propagation (rshared), ASCII label, smart partition skip, findmnt verification
> - **v0.11.6:** ✅ COMPLETE — FileBrowser auto-mount sync (`syncFileBrowserMounts()`) + 3 UI fixes (badge color, progress bar, button text)
> - **v0.11.7:** ✅ COMPLETE — Stale data cleanup + FileBrowser sync after migration + deploy page title fix
> - **v0.11.8:** ✅ COMPLETE — Per-App Cross-Drive Backup (3-2-1 rule): rsync/restic to secondary drive, deploy page UI, backup page summary, scheduler jobs, API endpoints
> - **v0.11.9:** ✅ COMPLETE — UI Polish Fixes: spacing, tooltip on "Módszer", status dot instead of disabled checkbox, progressive disclosure, emoji cleanup
> - **First app deployed:** Paperless-ngx on demo-felhom.eu (2026-02-13)
> - **Running on:** demo-felhom (N100 mini PC) at 192.168.0.162:8080, felhotest (Proxmox VM) at router.abonet.hu:33022
> - **All Phase 1-5 features working:** deploy, start/stop/restart/update, logs, health-aware states, auth, monitoring, backups, backup detail page, system monitoring page, settings page
>
> ## Architecture decisions
>
> | Decision | Rationale |
> |----------|-----------|
> | Go stdlib for web (no Gin/Echo) | Minimal dependencies, single binary, easy to embed templates |
> | Templates as go:embed HTML/CSS files | Zero runtime file dependencies (compiled into binary), but each template is a separate editable file |
> | Docker Compose for customers (not k8s) | Simpler troubleshooting, customers don't need k8s knowledge |
> | k3s for management infra only | Viktor's own services (gitea, monitoring, website) run on k3s |
> | Cloudflare Tunnel for remote access | No port forwarding needed, works behind any NAT |
> | app.yaml per stack | Separates deploy config from compose files, survives git pulls |
> | Password fields require explicit input | Prevents accidental empty-password deployments |
> | Health-aware state from Docker Status field | Docker's State says "running" even for unhealthy containers |
> | Memory limits via deploy.resources.limits | Prevents runaway containers; ~50% headroom over expected usage |
> | System info from /proc/meminfo + statfs | No external dependencies, cheap to read on each page load |
> | mem_request vs mem_limit (K8s-inspired) | Requests = expected usage (hard block), limits = peak (overcommit OK) |
> | 384MB reserved for system | Prevents deploying apps that would starve the OS/controller |
> | Logo SVG embedded as Go constant | Same approach as CSS/HTML — zero external file deps |
> | Git sync via os/exec git CLI | No Go git library needed, git is in the container image |
> | SHA-256 for content comparison | Only copy changed files, avoid unnecessary disk writes |
> | 30s debounce on manual sync | Prevents spamming the git server |
> | Orphan = deployed but not in catalog | Safe lifecycle: remove from catalog → mark orphaned → user deletes via UI |
> | FileBrowser as infra (not catalog) | Needed even after apps deleted (user browses HDD data); deployed by setup script |
> | Protected HDD paths | Safety net: never delete top-level HDD dirs (media, storage, Dokumentumok, appdata) |
> | Central scheduler (not ad-hoc goroutines) | Single place to register/monitor all periodic tasks, graceful shutdown, skip-if-running |
> | CPU sampling via background goroutine | /proc/stat delta needs two readings — collector runs every 5s, GetInfo() reads cached value |
> | Temperature from /host/sys (Docker mount) | Container can't read host /sys directly — mount /sys:/host/sys:ro, try /host/sys first |
> | Restic password auto-generated | No manual setup needed — generated on first backup run, stored in named volume |
> | DB discovery via docker inspect | No config needed — discovers postgres/mariadb containers by image name + env vars |
> | Backup orchestrator with running flag | Prevents concurrent backups, supports both scheduled and manual trigger |
> | modernc.org/sqlite (pure Go) | No CGO/gcc needed in Docker build stage — keeps `CGO_ENABLED=0` static binary |
> | AlertManager state-based refresh | Alerts regenerated every 5min from health report — no persistent storage needed, always reflects current state |
> | Notification relay via hub | Controller → hub → Resend → email. Hub acts as central relay: knows customer email, handles Resend API. Controller only needs hub URL + API key |
> | In-memory notification cooldowns | Per-event-type cooldown map (default 6h). Lost on restart = acceptable (better to re-notify than miss). No persistence needed |
> | Health status change detection | Only notify on degradation (ok→warn, ok→fail, warn→fail). Avoids spam on flapping. First run records baseline, doesn't notify |
> | Resend HTTP API (no SMTP) | Direct POST to api.resend.com — same pattern as website contact-mailer. Simpler than SMTP setup, good deliverability |
> | Preferences sync on save + startup | Controller pushes prefs to hub (not pull). Startup sync handles hub DB rebuild. Local save always succeeds even if sync fails |
> | Chart.js embedded locally | Customer hardware may not have internet — CDN not reliable for offline environments |
> | StackDataProvider interface | backup package needs stack data but can't import stacks (circular). Interface in backup, thin adapter in main.go |
> | Password sync to hub via report | Restic password in Docker named volume on SSD. Hub sync provides redundancy for disaster recovery |
> | App backup via HDD mounts only | Docker volumes at /var/lib/docker/volumes/ not mounted in controller. HDD data is the important user data; DB in volumes covered by nightly dump |
> | Restore uses running mutex | Prevents concurrent backup+restore on same restic repo. Reuses existing `m.running` flag |
> | Storage paths registry in settings.json | Multi-storage support: each app's HDD_PATH from app.yaml is authoritative. Auto-discovery on startup avoids manual config. Registry enables UI management + health monitoring per path |
> | /mnt:/mnt:rw mount in controller | Replaces per-path HDD_PATH mount. Enables multi-storage + restore writes. All customer HDD mounts are under /mnt/ by convention |
> | Per-app HDD_PATH resolution (app.yaml > global) | App's own env HDD_PATH is Priority 1, registered storage paths as fallback. Eliminates dependency on global controller.yaml hdd_path |
> | Mount-point detection via syscall.Stat_t.Dev | Compares device ID of path vs parent dir — reliable check that path is on separate filesystem. Prevents data writes to SSD |
> | Health severity: mount-point = warning | Non-mount-point is informational, not a service failure. FAIL reserved for genuinely broken things. Avoids false alarms on demo/test environments |
> | FS info via findmnt + sysfs | `findmnt -n -o SOURCE,FSTYPE --target ` for filesystem type/device. `/sys/block//device/model` for disk model. Best-effort, returns nil on failure |
> | Query param flash messages | Stateless, no session store needed. Consistent with backup page pattern. `?storage_msg=success&storage_detail=...` |
> | StorageLabels map on stacks page | Separate map passed to template (not modifying Stack struct). Built from deployed apps' HDD_PATH → registered path label lookup |
> | Metrics downsampling via SQL | Bucket-based AVG in GROUP BY keeps Chart.js responsive with up to 30 days of data |
> | 60s metrics collection interval | Good balance of resolution vs. storage — ~44K rows/month for system metrics |
> | /etc/os-release mounted read-only | Container can't read host OS info directly — mount to /host/etc/os-release:ro |
>
> ## Key file locations on demo-felhom
>
> ```
> /opt/docker/felhom-controller/ # Controller compose + config
> ├── controller.yaml # Customer config (domain, auth, paths)
> ├── docker-compose.yml # Controller's own compose
> └── data/ # Controller persistent data (named volume)
>
> /opt/docker/stacks/ # All app stacks
> ├── traefik/ # Reverse proxy (protected)
> ├── cloudflared/ # Tunnel (protected)
> ├── paperless-ngx/ # First deployed app ✅
> │ ├── docker-compose.yml
> │ ├── .felhom.yml # App metadata
> │ └── app.yaml # Deploy config (env vars, locked fields)
> └── whoami/ # Test stack (not deployed)
>
> /mnt/hdd_placeholder/storage/ # HDD storage for apps
> └── paperless/
> ├── consume/ # Drop files here for OCR
> ├── media/ # Processed documents
> └── export/ # Backup exports
> ```
>
> ## Related repositories and their state
>
> | Repository | Status | Notes |
> |------------|--------|-------|
> | felhom-controller | Active | This repo. Controller code + deploy scripts |
> | app-catalog-felhom.eu | Active | 10 app templates, all with .felhom.yml metadata + memory limits |
> | felhom.eu | Active | Website + hub/ subfolder (felhom-hub service) + k8s manifests |
> | homelab-manifests | Stable | k3s cluster running (dooplex.hu services) |
> | misc-scripts | Utility | collect-repo.sh, backup helpers |
>
> ## Gotchas & lessons learned
>
> - `docker compose restart` ≠ `docker compose up -d` — restart doesn't pick up new images
> - Go maps have random iteration order — always sort slices before displaying
> - Docker `.State`="running" doesn't mean healthy — check `.Status` for "(health: starting)" / "(unhealthy)"
> - Paperless-ngx needs `PAPERLESS_OCR_LANGUAGES` (plural) to install language packs, `PAPERLESS_OCR_LANGUAGE` (singular) to select
> - In-memory Deployed flag must be set BEFORE `docker compose up -d` (not after) — compose can take 30-60s for image pulls, during which the UI would show a stale "Telepítés" button
> - Cloudflare Tunnel handles *.demo-felhom.eu → Traefik handles Host()-based routing to containers
> - BIOS "AC Power Recovery" must be enabled on N100 for auto-restart after power outage
> - `docker compose up -d` returns exit 0 even when containers immediately crash-loop — need post-start status check to detect this
> - When logging env vars for debugging, only log keys (not values) to avoid leaking secrets in log files
> - Mealie image (`ghcr.io/mealie-recipes/mealie`) doesn't include wget/curl — use Python TCP socket check for healthcheck
> - Mealie DB migrations on first start take ~40s (alembic) — use `start_period: 60s` to avoid premature unhealthy status
> - Alpine-based images (filebrowser, vaultwarden) have wget via BusyBox — healthchecks with `wget --spider` work fine
> - Deploy `sed` command to update image version must target only the `image:` line — naive `sed 's|name:OLD|name:NEW|'` also matches the service name line (e.g., `felhom-controller:` → `felhom-controller:0.2.12`), breaking YAML. Use `sudo sed -i 's|image:.*felhom-controller:[^ ]*|image: ...felhom-controller:NEW|'` or similar scoped pattern
> - Hungarian quotation marks `„"` in YAML: `„` (U+201E) is safe inside YAML double-quoted strings, but the closing `"` must NOT be ASCII `"` (0x22) — it terminates the YAML string. Use `\"` escape or Unicode `"` (U+201D). This caused a silent parse failure for the entire `.felhom.yml` file
> - Never silently swallow parse errors — always log them. Silent failures make debugging impossible (took a dedicated debug session to find a simple quoting issue)
> **2026-08-14 — v0.215.0 (R-328..R-333). Disk health, phase 1: the alert that reached nobody.**
>
> ### A severity string is a WIRE CONTRACT with the hub, not a label we choose.
>
> The hub accepts exactly `{info, warning, error, critical}` and **silently coerces anything else to
> `info`**, which `severityNotifies` then drops. `disk_health_degraded` shipped `"warn"` — one letter
> short of the contract — so **every Figyelmeztetés-level disk alert this product ever produced was
> emailed to nobody, on both legs.** Proven live side by side on 2026-08-14: `"warning"` →
> `notification_log` status **`sent`**; `"warn"` → stored `info`, **no row at all**.
> `app_start_failed` (`notifier.go` ~L546) carries the identical defect and was deliberately NOT
> changed here — it needs its own decision on whether it should notify (**R-329**).
>
> ### A drive's own PASSED verdict cannot fail on bad sectors. Do not build on it.
>
> Attributes 187/197/198 all carry `thresh: 0`; a normalized SMART value floors at 1 and can never drop
> to or below the threshold. The real drive (ST3000VX010, S/N Z6A07P2G) read `PASSED` at **352** pending
> sectors and **1001** reported-uncorrectable reads. Felhom already read the raw counters, which is the
> only reason it would have noticed at all. Evidence + fixtures:
> `felhom.eu/documentation/audits/DIAG-smart-passed-trap-2026-08-14.md`.
>
> **DECISIONS MADE, so they are not re-litigated:**
>
> - **Predicted failure is labelled „Hiba" — there is no fourth verdict word.** A fourth Hungarian word
> sharing a root with „Figyelmeztetés" would make the MORE severe state read as the milder one. Four
> labels, final: Rendben / Figyelmeztetés / Hiba / Nincs adat.
> - **Sustain is the primary rule; the count is the backstop.** Truth-table row 6 (unreadable sectors
> present again at the next check) sits ABOVE row 8 (count >= 64) because on the real drive sustain
> fires 12 Aug and the count not until 13 Aug. Row 8 exists only for a box powered off across the
> sustain window.
> - **The numbers and where they come from.** 64: the benign excursion peaked at 16 and cleared inside
> an hour; the terminal run passed 64 at 13 Aug 11:28 and never returned. **It is a judgement from ONE
> drive** — a static backstop, expected to be replaced by growth-rate detection in Phase 3. 55/60 °C:
> the operator's existing Prometheus bands on DooPlex, adopted unchanged so the two systems cannot
> disagree about the same drive. **These bands are SPINNING-DISK bands and are questionable for NVMe**
> — demo-hp's healthy Toshiba NVMe idles at **53 °C**, 2 °C below Figyelmeztetés (**R-333**).
> - **Phase 2 owns the new SMART attributes (187 Reported_Uncorrect, 199, 188).** They are a declared
> wire change, so under the G-1 gate the hub must model them in the same session. Putting them here
> would have turned a one-word severity fix into a three-repo change (**R-330**). Everything v0.215.0
> needs was already on the wire.
> - **Phase 1 state is one small record per disk, NOT a sample series.** `metrics.MetricsStore` is the
> right home for Phase 2/3 history; using it now would have put a schema migration on the critical
> path of the severity fix.
>
> **The trap this change nearly shipped, caught by a test and not by review:** the card and the check
> share one verdict function so the chip and the email can never disagree — but the check CONSUMES the
> prior and then overwrites it, so a card rendering afterwards read its own check's write and showed one
> level MORE severe than the alert. Fixed by `diskRecord.PriorSawUncorrectable`, which replays the prior
> that produced the stored verdict. The guarantee was previously asserted in a comment only.
>
> **Cadence is hourly, and it was MEASURED:** demo-hp `/disks` costs median 0.821s (min 0.805 / max
> 0.841, 10 calls, 3 physical rows) — 6x under the 5s bar. Open question deliberately NOT acted on: the
> agent runs bare `smartctl -a -j` with **no `-n standby`**, so an hourly poll would wake a spun-down
> HDD. demo-hp is all-flash so the measurement could not show it (**R-333**).
>
> **NOT live-validated:** the Fail-from-counters path has never fired on real hardware — only against
> the fixture's values in unit tests (**R-332**).
> **2026-08-08 — v0.208.0 (R-254). THE RULE, stated so it outlives this session:**
>
> ### A secret is never in a page's response body. It is fetched by an explicit act, and the act is recorded.
>
> Three instances of one pattern shipped in two days, each found by hand: the retrieval passphrase
> (R-249), an app's real first-login password (R-254 site one), and an already-deployed app's generated
> secret field (R-254 site two). Every one was "hidden" with `display:none`, `hidden`, or
> `type="password"` — **instructions a browser honours when DRAWING and nothing else.** The plaintext
> was in the bytes; a `curl` returned it; caches, history, saved pages and screen-shares had it.
>
> **The shape of the fix, now used three times:** the page carries a BOOLEAN; the value comes from a
> **POST** (so CSRF covers it and it is not re-fetchable from history) with **`Cache-Control:
> no-store`**; the reveal is **LOGGED as an act** — reading a value off markup left no trace anywhere,
> which is why nobody can say whether any of these was ever read. **Per-secret endpoints, never one
> generic "reveal any named secret"** — that would turn three narrow exposures into one lever.
>
> **And the test must assert the RAW RESPONSE BODY.** Every test that asked what the customer *sees*
> passed while the bytes carried the secret. That is precisely how this survived three times.
>
> **What is NOT this defect:** a form must carry what it submits. The pre-deploy hidden input round-trips
> a generated secret deliberately (README §318) so the saved value is the one the customer wrote down.
> The defect there was the neighbouring READONLY input on an already-deployed app, where nothing is
> submitted at all.
>
> **The gate:** `scripts/secret_in_markup_gate.py`. Name-based, all 36 templates, **blind to a secret
> arriving under a neutral page-data key** — measured, not assumed. The complementary runtime
> body-assertion covers 4 of 27 page templates; the other 23 are **R-255**.
>
> **A correction to v0.207.0's report:** it said HTML comments ship in the response body. They do not
> here — `html/template` strips them (`text/template` does not). Measured.
> **2026-08-08 — v0.207.0 (R-249, R-252, R-253). Three things the fifth walk exposed BY PASSING.**
> The walk closed R-201 (both halves) on 2026-08-07; none of the below touches the recovery path it
> proved.
>
> **R-249 — a secret was living in the page source.** `settings_security.html` rendered the retrieval
> passphrase into a `display:none` span behind a „Megjelenít" button. That toggle stops a browser
> DRAWING it and nothing else: the plaintext was in the response body of every render. Found by doing
> exactly that — it landed in a session transcript while driving the documented rebuild path.
> **THE RULE, which the codebase already stated for R and this page did not follow:** a secret is
> revealed by an XHR, never templated server-side into HTML (`escrow_handlers.go`). The page now
> carries only `HasRetrievalPassword`; the value comes from `POST /settings/retrieval-password/reveal`
> — CSRF-covered, `no-store`, and **logged as an act**, which reading it off the markup never was.
> **The test asserts the RAW RESPONSE BODY** — every test that asked what the customer *sees* passed
> while the bytes carried the secret, and that is why it survived.
> **The census found two more instances** (`app_info.html`, a real per-install app password in a
> `hidden` span; `deploy.html`, a generated secret in a `value=`) — **filed as R-254, not fixed.**
>
> **R-252 / R-253 — the two obstacles, and the rule they share.** A rebuilt box keeps its drives but
> loses their REGISTRATION, so every restore refused with a sentence naming no next step; and the
> restore list promised „a visszaállítás előbb újratelepíti" three lines above a refusal that fired
> *because* the app was not installed. **The promise was the wrong half:** reconstitution writes to
> the app's own `GetStackHDDPath`, which exists only once the CUSTOMER has chosen a drive at deploy
> time — an automatic reinstall would mean the product making that choice for them, which is the one
> decision this recovery path exists to leave with them. Both now name a reason and route to the step
> that clears it, and both notices are conditional (a healthy box is byte-identical, pinned by a test
> that fails if either becomes unconditional).
>
> **The page and the resolver ask ONE question:** `HasRestoreDestination()` reads the same
> `GetSchedulableStoragePaths()` the scratch resolver reads. A second copy of that predicate is
> exactly how a page ends up promising what the handler refuses — which is R-253 itself.
> **2026-08-07 — v0.206.0 (R-241). THE RULING, and it reversed the fix: this was a MINTING defect,
> not a screen-predicate defect.** The recovery screen was telling the truth — there genuinely was
> nothing recoverable under the key the box held, because **the box minted that key itself over the
> top of a sealed package it already knew the hub was holding**. Fixing the predicate would have
> papered over a machine quietly making its own backups unopenable.
>
> **THE RULE: a box does not create a repository key while the hub holds a sealed package for it.**
> The guard is a conjunction (package held AND no key), so a first-time box is untouched, and the
> refusal is a HOLDING state rather than a failure — the transport is still configured so the
> recovery screen can bring the tier up the moment the key arrives.
>
> **THE SECOND RULE: the fact that answers a question must be kept where the question is asked.** The
> hub-vs-local key comparison had been computed on every ACK since SLICE 3 and persisted nowhere; on
> the venue it logged the right answer thirty-five minutes before the customer looked at a screen
> that could not see it. It is now persisted and drives shape (c) of the offer.
>
> **THE THIRD RULE (the operator's, and it generalises): fix the state, do not remember that it is
> wrong.** Abandoning the old history now starts a 14-day countdown that removes the set-aside store
> and its sealed package TOGETHER, after which the offer falls silent on its own because there is
> nothing left to compare — rather than a "they decided" flag suppressing a screen over a state that
> is still wrong. The recovery offer stays reachable for the whole grace; a grace in which recovery
> is impossible is decorative.
>
> **Surface:** the full page appears once per ENTRY into the offered state, not once ever — a box
> rebuilt months later is a new situation. Three dismissal levers with three scopes, and none of them
> removes the entry point on the backups page.
>
> **Needs hub v0.98.0** for the superseded-package purge. `felhom-agent` untouched.
>
> **Two real bugs were caught by tests rather than by review** — a missing `t.Enabled` (an existing
> test) and a missing falling-edge sync that reintroduced the very defect the epoch exists to fix.
>
> **NOT built, deliberately:** the automatic 30-day abandonment (R-245, with the operator's reasoning
> recorded), and R-242's release-to-golden gate.
> **2026-08-06 — v0.205.0 (R-234).** THE RULE: **a run that skipped an app the customer selected is
> not a successful run.** The R-203 verdict block already said *"a warning beside a success is read
> as a success"* and applied it to one of the two shapes it describes — a missing declared FOLDER
> made the run `incomplete`, an app skipped ENTIRELY did not. Now both do. A selected-but-UNDEPLOYED
> app is named with what to do but does NOT move the verdict, because a box left permanently amber by
> an app somebody removed is a status nobody reads.
>
> **§7.3, MEASURED rather than assumed — and the answer was "already done".** `CaptureRecoveryUnit`
> writes compose config + a manifest (a few KB), only ENUMERATES dumps rather than creating them, is
> idempotent, and does NOT stop the app; the off-site run already calls it for every deployed stack in
> its own pre-dump phase, through `admitApp`. So there is no wait to remove for a deployed app, and
> **nothing was built**. Proven on demo-hp: a unit moved aside was RECREATED by the run.
>
> **AND THE FILED MECHANISM WAS NOT THE MEASURED CAUSE.** R-234 was filed as "the first run after a
> toggle finds no bundle and skips the app". That cannot happen for a deployed app (above). What did
> happen on 2026-08-06: the manual run was dropped by the **single-flight** while an earlier run was
> still going; `runOffboxBackup` returned nil; the handler had already said „elindult”; and the card
> then showed the PREVIOUS run's „✓ Rendben”. Fixed by taking that decision synchronously in the
> handler. **The nightly path deliberately still returns nil** — nobody asked, and it retries.
> **2026-08-05 — v0.200.0 (R-193 CLOSED).** The customer-facing recovery screen. Until now a customer
> whose machine was rebuilt had everything needed to get their data back and no way to find out — the
> only route was a command line.
>
> **IT UNLOCKS AND ONLY UNLOCKS** (operator ruling). Explains, takes the recovery code, opens the
> repository, lists what is in it (apps, dates, sizes). **Restores nothing** — restore is per-app and
> lives in the backups area; the put-back is **R-213** and its stated requirement is a
> live-versus-backup comparison.
>
> **ONE CORE, TWO CALLERS.** `backup.RecoverInstallCore` is the only fetch→unseal→compare→install path.
> `RecoverAndInstall` is now a thin CLI wrapper — exit codes and printed lines byte-identical, every
> pre-existing CLI test passed unchanged — and the handler calls the same function. Asserted from
> source by AST on BOTH sides, plus a test that the routes and the landing-page interception exist.
>
> **THE TRIGGER HAS TWO SHAPES and the second is the one that matters.** `OffsiteRecoveryOffer` = the
> hub holds a package AND (no repository password OR the tier is orphaned). The literal "no repository
> password" alone is a window that CLOSES BY ITSELF — `WriteOffboxSecrets` auto-generates one on
> re-apply (R-193's own orphaning mechanism) and hub v0.96.0's self-heal re-applies within ~15–30 min.
> Shape (b) is also what the shipped move-aside requires, which is why the discard choice can reach it.
>
> **CLAIMED is part of the predicate** — a legacy-open box passes through `RequireAuth`, so without an
> explicit `authEnabled()` check the interception fired for an unauthenticated visitor. A test caught it.
>
> **„Most nem" suppresses the FULL PAGE ONLY.** The backups-area entry point is bound to
> `recoveryOffer`, never to the postpone flag.
>
> **The code:** POST body only, never logged/persisted/echoed, cleared on every path, `no-store`,
> `autocomplete=off`. **No lockout** — a ten-word phrase is not guessable and locking a customer out of
> their own data for a typo is worse; failures are logged locally without the code, and NO operator
> alert is raised (reasoning in REPORT.md §4).
>
> *Live:* demo-felhom is genuinely in shape (b), so validation needed no arrangement — `/launcher` →
> 302 `/recovery`, both mandatory sentences rendered, three wrong codes refused with the `offbox/`
> listing byte-identical and no lockout, and the code found in no file, log or ring **with a
> planted-copy positive control that first exposed a mis-aimed sweep**. **NOT proven live: a CORRECT
> code** — none was kept for demo-felhom's orphaned history and demo-hp's is operator-held.
> **2026-08-05 — v0.199.0 (R-204 item 4 / R-193).** The last of the four manual interventions the
> 2026-08-04 drill needed. **Operator ruling: automate it, and the trigger is a state the BOX
> DECLARES.** From the hub an absent off-site object has FOUR meanings — never configured,
> mid-restart, a transient config read failure, rebuilt-and-stranded — and the hub cannot tell them
> apart. The box can.
>
> **The declaration needs BOTH halves** (`backup.needsOffsiteCredential`): a fresh data area (no
> repository password) AND a hub-held recovery package (the ACK's `identity_blob_present`). Freshness
> alone is a box that never had off-site backups; dropping that condition makes the whole fleet ask
> for credentials, which is what `TestOffsiteDeclare_NeverHadOffsiteSaysNothing` catches. A merely
> DISABLED target is the customer's own choice and never declares.
>
> **The ACK field stopped being discarded.** `EscrowAutoConfirmer.Reconcile` returns early when the box
> is neither pending nor escrowed — exactly a rebuilt box — so the fact was thrown away every cycle. It
> is recorded FIRST, before every gate, via `RecordPresence`, wired in main.go and asserted by
> `TestMainWiresRecordPresence` (AST, comments dropped). Last-write-wins, not set-only, so a customer
> RESET turns the declaration back off; a nil ACK escrow records nothing.
>
> **Inert to every existing reader:** `enabled:false` + zero sizes, so the hub's `isStale` and
> `fillBand` both short-circuit; an unknown `state` string is ignored by encoding/json. **A configured
> box's report JSON is byte-identical to v0.198.0's.** The one reader that would have misread it is the
> hub's `reportHasOffsite`, tightened in hub v0.96.0 to require `enabled:true`.
>
> *Live:* both demo boxes now record `hub_escrow_identity_present=true` in settings.json (the recorder
> working on a HEALTHY box). demo-felhom 9201, arranged reversibly into the stranded shape, produced
> report id=16743 carrying `{enabled:false, state:needs_credential, quota_gb:0, repo_size_bytes:0}`;
> the single declaration was absorbed by the hub's debounce (no self-heal event) and the box was
> restored the same minute. **The hub half is felhom.eu v0.96.0.**
>
> *Rider:* `.githooks/pre-push` in all four repos now refuses a push from a clone outside
> `/mnt/5_hdd/felhom.eu`. Proven both ways against a scratch clone.
> **2026-08-05 — v0.198.0 (R-204 items 1 & 3).** The 2026-08-04 drill (R-201) passed only because a
> person was there; four manual interventions stood between a recovered key and a restored file. Two
> of the three defects are in this repo.
>
> **Item 1 — the reset code needed a restart.** `--print-reset-code` is a SEPARATE process; it
> persisted a new code while the running server kept the old one cached, so the code the customer was
> told to type was refused until the controller restarted, and nothing said so. `effectiveClaimCode`
> now calls `settings.ReloadClaimCode()` first. **The settings-vs-config precedence is unchanged** —
> the defect was freshness, not precedence. **Read-through, not a TTL, and that is the point:** a TTL
> makes the new code visible AND leaves a window in which the superseded one still works, which is
> worse than the bug. That is the mutation `TestClaimCode_SupersededByASecondMint_RefusedImmediately`
> exists to kill, and its red-proof produced exactly *"the SUPERSEDED code was accepted"*. The
> function now returns an error and **every caller fails closed**; an absent settings file is NOT an
> error. `ClaimConsumedGeneration` is deliberately NOT re-read — this process is its only writer and
> re-reading could move it BACKWARDS if a save had failed, resurrecting a consumed code.
>
> **Item 3 — the restore's default returned the wrong thing silently.** `mode=unit` restores the
> recovery unit (definition + config + DB dumps) and not the customer's files. `restoreScratchOutcomeMsg`
> now names what came back, what did not, and the next step; the wizard's intent card states its scope
> before the choice. **The size gate is untouched** and pinned unchanged by
> `TestOffboxRestore_FullPathUnchanged`. **The default stays `unit`** — all three wizard forms set
> `mode` explicitly, so a change would alter nothing visible while silently changing a mode-less POST.
>
> *Live-validated endpoint-level (no browser on DooPlex):* on demo-felhom 9201 with `restarts=0`
> across both mints, a superseded code returned „Hibás vagy lejárt kód" and the current one was
> accepted first time; on demo-hp 9201 a `privatebin` unit restore produced the scoped Hungarian
> outcome and `mode=full` without confirm revealed `full_size=6.8+KB` without restoring anything.
> demo-hp's drill scratch (`calibre-web`) was not touched.
>
> **Item 2 is the hub's** (felhom.eu v0.95.0, R-196). **Item 4 — a rebuilt box cannot obtain an
> off-site credential unaided — remains OPEN (R-193)** and was deliberately not begun.
> **2026-08-02 — v0.190.0 (R-157 mechanism A · R-170 · R-171).** Three items, one live validation
> cycle, because all three are boot behaviour and all three are proven by hard-resetting the box.
>
> **DIAGNOSE BEFORE THEORISING — and the first diagnosis was a FALSE NEGATIVE.** A hole was reasoned
> out of the v0.189.0 diff (a drive-gate-stopped app has zero containers and `desired_state: running`,
> so it now reads as a boot orphan) and confirmed on hardware BEFORE any fix was written. **Attempt 1
> produced `no boot-orphaned apps` and would have been reported as a disproof.** It was a race:
> unmounting only the parent bind is healed by the agent within ~60 s, so the drive gate's startup
> reconcile restarted the apps **one second before** the sweep looked. Holding the drive genuinely
> absent reproduced the defect immediately. **"It didn't happen this time" is not a mechanism.**
>
> **The confirmation moved the severity in BOTH directions.** The write hazard did not materialise —
> compose failed `mkdir …/userdata: permission denied` because the unbound mountpoint is
> host-root-owned and the guest is unprivileged. **That protection is ACCIDENTAL**: no code chose it,
> no test pinned it, and it is one `chown` or one privileged guest away from gone. But the harm that
> DID occur was not in the hypothesis and is real on every box: two wasted attempts and a **false
> dead-app alarm for an app the drive gate is deliberately holding**.
>
> **The fix already existed one path over.** `startGatedByMissingDrive` (the API) refuses a customer's
> start on an absent drive; the sweep bypassed it by calling `Manager.StartStack` directly.
> **`StartStack` HAS NO GATE OF ITS OWN** — carry this: every caller that is not the customer must
> decide for itself whether the app may run. New consumer-side `bootrecon.StartGate`, fail-safe
> (cannot determine ⇒ do not start).
>
> **Widening a window makes previously-unreachable overlaps reachable — a design input, not an
> afterthought.** The old T+5 s sweep never met a quiesce or an in-flight app-data operation; a 50 s
> window can. All three holders answer ONE seam because they differ only in the reason string.
>
> **A TEST REJECTED MY FIRST CONSTANT, and the comment says so.** `settle + budget + one retry` must
> fit inside `deadAppBootGrace`; 60 s gave 95 s against 90 s. The budget is 50 s **because a test said
> so** — recorded in the code rather than presented as taste. Widening the grace was rejected: it
> hides a late recovery instead of reporting one (`recordLateRecovery`).
>
> **THE FIX HAD ITS OWN DEFECT, FOUND LIVE AND NOT BY REVIEW.** The window sampled `GetStacks()` — the
> Manager's map, refreshed by the scheduler every **10 s** — every 5 s, so two identical samples could
> mean *the cache did not update*. Observed: a container removed ~5 s before the window closed was
> still in the sampled fleet and the sweep logged `no boot-orphaned apps` for an app that had none.
> `sampleBootFleet` now refreshes first. **Generalise: a settle detector is only as good as the
> freshness of what it samples — if the source is cached, refresh it, or you are watching the cache
> settle rather than the system.**
>
> **R-170:** `shouldRecreateOnBoot` reads intent with the identical three-way table; absent keeps the
> old `hasContainers` behaviour exactly; `presentStable` untouched and still load-bearing. Its comment
> argued at length FOR the count and was rewritten. Agreement pinned from BOTH sides against one
> fixture table (an import cycle prevents testing the two gates together).
>
> **Live: 6/6 hard resets** (every app back; the customer-stopped app down all six), settle times
> 10/40/10/10/15/15 s. Sharpest evidence: same app, same box — missed at 18:08:35, recovered at
> 18:18:50. R-170 proven in one reboot (calibre-web recreated, immich left stopped). 27/27 packages;
> 7 red-proofs. Detail: `REPORT.md`.
Last updated: 2026-08-02 (v0.189.0 — R-166 / D-b: the box stops guessing what the customer wanted)
> **2026-08-02 — v0.189.0 (R-166, operator decision D-b).** When an app was not running the box had
> to work out *why*, and it did so **by counting containers**: zero meant "the customer stopped it",
> some meant "something broke". A **power cut mid-compose** and an **interrupted deploy** also leave
> zero containers, so both were read as deliberate stops and stranded **silently** (R-157 mechanism
> B) — and a backup that stopped an app and died left it stopped with **nothing on disk** recording
> that it was owed a restart. The settling fact — what the customer asked for — **was written down
> nowhere**: `app.yaml` recorded *installed*, never *meant to be running*.
>
> **DECISION — one owner: the customer's action, and nothing else.** A census found **14 callers of
> `StartStack`/`StopStack`, of which exactly 2 are the customer**; the rest are quiesce, the volume
> dump, offbox reconstitution, app export/restore, the storage gate, migration and the boot
> reconciler. So the primitives are deliberately **not** writers — intent there would make a nightly
> backup indistinguishable from the customer pressing Stop. Writers: the API action switch,
> `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, the `.fab` import. Intent is written
> **BEFORE** the act and a failed write **REFUSES** the act.
>
> **DECISION — absent means UNKNOWN, never "running", and this is the whole safety property.** Every
> `app.yaml` on every box predates the field, so absent is what the fleet reads on upgrade; reading
> it as running would start every deliberately-stopped app on the first boot after the upgrade. The
> legacy branch of `isBootOrphan` keeps the old container-count rule **byte-for-byte**, and its test
> asserts BOTH legacy rows together because the safety property is the pair. Backfill is
> **running-only** — "zero containers ⇒ stopped" IS the defect, so an ambiguous app stays ambiguous.
>
> **Part 2 — `backup.AppStopGuard`**, a persisted marker over every stop→work→start window (volume
> dump, offbox reconstitute, `.fab` export), in its **own** file (one file, one writer). Written
> before the stop, cleared only after a restart that **succeeded**, kept when one fails. `Recover`
> **returns** its outcome instead of using a notifier seam, because it must complete before the boot
> reconciler (`main.go` ~236) while the notifier is not built until ~307 — a seam wired after the
> fact is a seam that never fires.
>
> **THE TEST LESSON, and it is the one worth carrying:** Scenario E's first version called
> `appStop.Begin` itself, and **survived the red-proof that deleted the production call**. It proved
> the marker type, not that `DumpAppVolumesSafe` uses it. Rewritten to drive the real function with a
> simulated hard abort (an unwind that skips the restart statement, since a `defer` is not
> crash-safety — Campaign 8 fault 10). **A test that constructs the thing it is meant to prove the
> caller constructs is hollow, and its red-proof will say so if you run it.**
>
> **FOUND EN ROUTE — `SaveAppConfig` rebuilt `AppConfig` field-by-field**, the R-100 shape (v0.181.0
> shipped two live instances). The literal named five fields, so `desired_state` would have been
> dropped on **every** save across nine call sites — a customer's Stop erased by the next unrelated
> `app.yaml` write. Copy-and-overlay (`saveCfg := *cfg`) is safe by construction. **Generalise it:
> treat any field-by-field struct rebuild in a save path as a defect on sight.** Measured, not
> assumed: `app.yaml` does NOT round-trip YAML keys the struct does not model (pinned by test).
>
> **R-157: mechanism B closed, mechanism A untouched** (the sweep observes ~5 s after start and never
> re-checks) — and B's fix makes A cost more, since the sweep now has more it could recover.
> **NEW R-170:** `shouldRecreateOnBoot` (`internal/web/intermediary.go:131`) still infers a Stop from
> `hasContainers` — the same defect one gate over, for drive-backed apps. Left deliberately.
>
> **Live on 9201, three flows** (stop survives a restart; a zero-container `running` app recovered by
> name; a legacy app.yaml skipped and never inferred stopped). The **interrupted-operation half is
> IMPLEMENTED, not PROVEN-LIVE** — nobody killed the controller mid-backup on metal. 27/27 packages;
> 7 red-proofs observed FAIL then restored. Detail: `REPORT.md`.
Last updated: 2026-07-28 (v0.182.0 — R-101 + F-DIAG: the restore dialog names the last SUCCESSFUL copy)
> **2026-07-28 — v0.182.0 (R-101 + F-DIAG).** `Tier2LastRun` is the ATTEMPT clock (written on failure)
> and was rendered as „Legutóbbi másolat" in the **restore confirm dialog** — misinformation at a
> decision point: the restore fills in MISSING files, so a customer with a failing Tier-2 restored and
> silently got OLDER files. New `CrossDriveBackup.LastSuccess` + **`SuccessTracked`**; the marker is
> load-bearing because **all 7 fleet rows were pre-anchor at deploy** — without it every customer sees
> „Még nincs sikeres másolat" at once. Legacy rows migrate on first touch (`ok` adopts its time,
> `error` seeds nothing). **PART 2 — the three `record*` helpers rebuilt the WHOLE struct with only 2
> fields carried over; the naive fix would have had `recordTier2Failure` CLEAR the anchor.** Replaced
> by `tier2Update` (copy-and-overlay = safe by construction). New `fmtTimeStr` → Budapest-local dates
> in the dialog instead of raw UTC RFC3339. **F-DIAG:** 6 classes incl. an honest `unknown`, and the
> notification no longer passes `err.Error()` through raw — **LESSON: my first sanitiser was regex-only
> and leaked a bare hostname; its own test caught it. Redact KNOWN values, don't guess at shapes.**
> Live on demo-hp: rendered dialog read in the failed, healthy AND legacy states. F-OPS documented at
> `felhom.eu/documentation/runbooks/RUNBOOK-manual-guest-restore.md`.
> **2026-07-28 — v0.181.0 (R-100).** `OffboxTarget.LastSuccess` + wire field `last_success`; the hub
> (v0.80.0) anchors offsite staleness on it. **`LastRun` is written unconditionally on every run
> INCLUDING failures** — it records an ATTEMPT — so the hub's "how long since LastRun" verdict read a
> nightly-failing tier as perfectly fresh forever. The rule is the pure `offboxAnchorAfterRun(prev, at,
> runErr)`: a failure neither ADVANCES nor CLEARS the anchor (both are distinct bugs; clearing it would
> make one bad night look like never-succeeded). `LastStatus == "error" ⇒ stale` was rejected — it pages
> on every blip, the F-A1 noise mode. **TWO SILENT-WIPE SITES CLOSED** (`offboxConfigHandler` and
> `ApplyOffsiteTarget` both rebuild the target and copy runtime status field-by-field — omitting
> LastSuccess would erase the anchor on any settings save or hub re-apply). **LESSON: my first test
> modelled the rule in a local closure and stayed GREEN when production was mutated — hollow; the
> extraction to a pure function is what made the red-proof bite.** Live on demo-hp: failing run advanced
> `last_run` to 11:25:48Z while `last_success` HELD at 11:24:20Z; demo-felhom healthy → advanced. The
> settings-save preservation was proven live too. Detail: `REPORT.md` + `felhom.eu/REPORT-r100.md`.
> **2026-07-28 — v0.180.0 (F-OBS).** Source: `audits/CAMPAIGN-8-backup-restore-2026-07-27.md`.
> On a default `logging.level: info` box there was **no positive observable that `deadapp-check` had
> run**: its per-cycle line goes through `Scheduler.dbg()`, gated on `level==debug`, so on a default
> box it was never *produced* and could not even reach the always-DEBUG ring. "No alarms" was
> therefore indistinguishable from "the detector never ran" — standing rule 3's exact fallacy, and it
> undermines F-CRIT-1's fix, which is a fix to **this same detector**.
> `noteDeadAppScan()` now emits an INFO line every **20th** scan (10 min at the 30 s cadence) carrying
> scans-since-boot / evaluated / currently-down. It reports **what it saw**, not that it ran, and it
> summarises rather than floods — one line per run is 2880/day, which is what made silence attractive
> in the first place. Both bounds are pinned by test in the direction that would break them.
> **The same shape then turned up in the agent's brand-new guest-power watchdog** (v0.107.0, shipped
> hours earlier): it logged only at startup and when it acted. Fixed in agent v0.109.0 with the same
> pattern. The anti-pattern reproduces itself — which is the argument for not having dropped this part.
> Live on demo-hp at INFO on a default-level box; deployed on both boxes. Detail: `REPORT.md`.
> **2026-07-26 — v0.173.0 (R-77).** Source: `audits/DIAG-agent-channel-2026-07-26.md`.
>
> **UNRESOLVED AND DELIBERATELY DEFERRED — which file is authoritative for `local_api`?** R-77 ships
> DETECTION ONLY. `controller.yaml` and `bootstrap.json` can disagree; the controller dials
> `controller.yaml`. The obvious "fix" — reconcile from `bootstrap.json` on every boot — has a failure
> mode **as severe as the bug it fixes**: on a guest whose `controller.yaml` is correct and whose
> `bootstrap.json` is stale (a re-provision that half-completed, a hand-repaired guest, a
> setup-wizard box), auto-reconcile would clobber a WORKING channel on the next restart — fleet-wide,
> silently, at the moment of a routine deploy. R-77's position is that **naming the drift is enough**:
> it would have converted the 17.5 h outage into a specific alert on the first health cycle. The
> authority ruling is **R-78** and needs its own spike — do not resolve it opportunistically.
>
> Corollary for anyone editing `bootstrap.MaybeIngest`/`ensureLocalAPI`: `ensureLocalAPI` is the ONLY
> writer, it fires only when the endpoint is EMPTY, and `DetectEndpointDrift` must stay write-free.
> Scenario A's test asserts `controller.yaml` is byte-identical after the check, and its red-proof
> covers the auto-correcting variant precisely because that is the tempting wrong turn.
>
> **Also settled here:** the samba protected-set must mirror EVERY early return in
> `reconcileSambaAt` (currently two: `!smb.Enabled`, `!smb.UserSet`). A third would need the same
> mirror, and the doc comment above `EffectiveProtected` must be updated with it.
> **2026-07-26 — v0.172.0 (R-75).** Spike `felhom.eu/documentation/audits/SPIKE-catalog-data-paths-2026-07-26.md`;
> feature doc `felhom.eu/documentation/controller/import-and-data-paths.md`.
>
> **RULING — the import root is CANONICAL on the system drive, overriding the spike's Fork-1
> recommendation of per-drive roots.** The spike weighed sidebar clutter and per-app link ambiguity and
> concluded per-drive; the operator overruled it on an argument the spike missed: each drop-zone app has
> exactly ONE ingest bind, so on a two-drive box every import folder except the app's own would look like
> a drop-zone and silently do nothing — and because `import/*` is `class: excluded`, files stranded there
> are never backed up either. A canonical root is the only shape with no dead drop-zone. Recorded as a
> deliberate deviation, not an oversight.
>
> **Phase-0 probe changed the shape of Part 6.** The system drive is NOT a registered `StoragePath` on
> either demo box (`/mnt/felhom-drives/hdd_1` on demo-felhom; `nvme-1tb` + `Felhom-Share` on demo-hp),
> so `sharingResolvePath` REFUSES `/userdata/import` — verified against the real guard with a
> passing control. Registering the drive was rejected (it would make the 50 GB volume holding the
> recovery units a customer-visible drive, deploy target and wipe candidate, and `SharingDeniedRoots`
> would then deny the namespace-consistent shape anyway). **Chosen: leave it unregistered and have the
> controller write the `beolvasas` share directly** — the picker guard validates CUSTOMER-supplied paths,
> a controller-generated constant is a different trust class. No guard was weakened.
>
> Also note: `withUserdataPath` computes `USERDATA_PATH` as `/userdata`, NOT
> `NamespaceRoot(hdd)/userdata`. For an app on the system drive those disagree
> (`/mnt/sys_drive/userdata` vs the `felhom-data` namespace). Latent — no app with a userdata bind has
> ever been deployed there — but it is a real inconsistency, left untouched here.
>
> The other three forks followed the spike unchanged: all-apps skeleton / deployed-only in the UI;
> unknown role fails OPEN while a malformed path whole-block rejects; drop-zone copy driven by the
> derived backup class.
> **2026-07-24 — v0.169.0 (disk-health card + degradation alert).** Consumes the agent's new `smart`
> field (agent v0.94.0; MinAgent floor unchanged — feature-detect by presence). **Rulings:** (1) ONE
> pure verdict fn `agentapi.DiskVerdictFor` is the shared truth for the card chip AND the 6h check — they
> can never disagree. Thresholds: FAILING→Hiba; PASSED + any(reallocated>0/pending>0/offline_unc>0/
> critical_warning>0/media_errors>0/percentage_used **≥90**)→Figyelmeztetés; PASSED clean→Rendben;
> nil/UNKNOWN→Nincs adat (never alarms). (2) **No global alert banner** — the card + email carry disk
> health; banner fatigue is a real cost, so this is deliberately NOT wired into the dead-app/alert-banner
> machinery. (3) Degradation-only notification with an in-memory baseline: first run baselines silently,
> recovery never notifies, **UNKNOWN excluded both directions** (a transient blip neither fires nor erases
> history). (4) **Controller restart re-baselines silently** (in-memory baseline lost on restart) — an
> accepted trade consistent with the health-change pattern (a real post-restart degradation still fires on
> the following 6h check once a baseline exists). (5) A **60s TTL cache** wraps the card's /disks call so
> dashboard refresh-spam can't smartctl-storm the host; the 6h check fetches FRESH (cache-independent).
> Pairs with hub +1 (allowlist `disk_health_degraded`). No new smartctl load — serialization only.
Last updated: 2026-07-24 (v0.168.0 — customer-configurable backup window "Mentési időablak")
> **2026-07-24 — v0.168.0 (customer-configurable backup window).** ONE customer setting — the window
> start W ("Mentési időablak kezdete") — drives every nightly leg at FIXED, never-stored offsets so
> misordering is impossible: DB dump at W, tier-2 at W+60m, off-box at W+105m (wrap-safe). **Design
> rulings:** offsets are DERIVED and computed everywhere, never persisted and never exposed in the UI;
> precedence is settings > controller.yaml `db_dump_schedule` > "02:30" (mirrors PasswordHash); a change
> applies WITHOUT restart via the new scheduler seam `UpdateDaily` (per-daily-job buffered `resched`
> chan + a select case in `runDailyJob`). New pure package `internal/backupwindow` holds all the time
> math (ParseHHMM/FmtHHMM/LegTimes/GateWindow/EffectiveWindow). **Disk-tier (whole-guest PBS/vzdump)
> gate:** the quiesce loop's SCHEDULED cycles run only inside [W+2h, W+6h) (wall-clock Europe/Budapest),
> with a safety valve — last successful backup older than cadence+24h (or none) runs regardless, so a
> box only ever on outside its window never starves. **Manual "Mentés most"/TriggerNow is NEVER gated**
> (bypasses runOnce). The `quiesce.Backend.Due` seam now also returns the backup age (from the agent's
> own `/backup/due`); the agent, its cadence, and `/backup/due` are untouched. Window read fresh each
> poll (WindowStartFn) so runtime changes take effect. Cadence defaults to 24h controller-side (the
> response carries no cadence). Backup page gets a "Mentési időablak" card (time input + derived rows +
> the "kb. W+2h–W+6h között" rendszermentés line); POST /backups/window (RequireAuth+CsrfProtect).
> **2026-07-24 — v0.167.0 (outlined logo + favicon — Part 4 unblocked).** Viktor pushed the
> text-outlined `logo.svg` to felhom.eu `main` (`be9edb4`); the wordmark is now 17 real ``
> glyphs. `FelhomLogoSVG` swapped to it; Inkscape's leftover **empty `` shells + font-* leftovers
> on the paths** were stripped via an lxml DOM pass (glyphs untouched — CC did NOT do text-to-path),
> editor `` dropped. `FelhomFaviconSVG` vestigial `` removed. Both constants:
> **0 ` Inkscape "Object→Path" leaves empty `` shells AND copies `style="…font-family:…"` onto the
> resulting ``s — a search for `svg:text` misses them (elements are ``, no prefix); grep
> ` first** (`assetsSyncer.Resolve`) and fall back to the constant only if none is on disk — on 9201 the
> constant is what's live (verified). Still open (separate follow-up): website + hub serve their own
> non-outlined logo copies; login.html stylesheet link still unversioned.
> **2026-07-24 — v0.166.0 (mobile nav = off-canvas drawer; sidebar cleanup; ?v= on logo/favicon).**
> Mobile nav was broken: the ≤768px block predated the v0.146.0 accordion and flattened `.nav-links`
> into a horizontal `overflow-x` strip, clipping the accordion's nested sub-lists (they share the
> `.nav-links` class). **Decision: mobile nav = a sticky top bar + off-canvas left drawer that REUSES
> the vertical sidebar (Option A).** The accordion handler is untouched and works inside the drawer;
> a `no-js` html-class fallback renders the sidebar static inline so nothing dead-ends without JS.
> Options B (separate mobile menu) and C (exclude nested lists from the strip) were rejected. z-index
> ladder topbar 800 < backdrop 900 < drawer 950 < modal 1000; `100dvh`; reduced-motion disables the
> slide; focus-trap deliberately omitted (navigations reset state). **Sidebar customer-name removed**
> (logo only); `{{.CustomerName}}` stays in base data + login subtitle. **Logo policy decision: the
> wordmark must be OUTLINED paths, never live ``** — under `` secure static mode only
> locally-installed fonts resolve, so `font-family` in the SVG renders a fallback font everywhere.
> **Part 4 (swap `FelhomLogoSVG`/`FelhomFaviconSVG` to the outlined master) is GATED OUT** — §3a check
> against live felhom.eu `main` (`be9edb44`) found `website/assets/logo.svg` still has ``/
> `font-family`; the outlined master is Viktor's manual Inkscape push, still pending. Only the `?v=`
> cache-bust (logo/favicon/login-logo, Cloudflare 4h edge-cache — the 0.126.1 failure mode) shipped
> from the logo work. Follow-up: when Viktor pushes the outlined asset, ship Part 4 (swap constants +
> clean the favicon's vestigial `` nodes). Separately, the website + hub still serve their own
> non-outlined logo copies — propagation is a distinct follow-up.
> **2026-07-24 — v0.165.1 (native "Megosztás…" in the share modal, Web Share API).** The share modal
> gains a feature-detected `navigator.share` button (OS share sheet → Messenger/WhatsApp/email),
> sending **title + text + URL only**. Hidden unless supported; "Link másolása" stays the universal
> fallback (and catches the non-cancel rejection); `AbortError` (user cancel) is silent. **Ruling: the
> QR is NOT attached** (no Web Share Level-2 `files:`) — file-share support is narrow and several
> targets drop the URL when handed file+URL, leaving an unscannable QR picture in a chat; the QR's job
> (physical cross-device scanning) is already served by the modal image (mobile long-press). Template
> JS + tests only; the OS sheet interaction is an operator manual check (not endpoint-testable).
> **2026-07-24 — v0.165.0 (Indítópult megosztása — guest launcher via capability URL).** The admin
> launcher gets an "Indítópult megosztása" button that mints a **capability URL**
> (`https:///s/`, 160-bit `crypto/rand` token) serving a standalone, read-only guest
> launcher — same tiles, opens apps in new tabs — with **no account and no admin session**. **Security
> ruling: the link grants INFORMATION ONLY, ZERO CONTROL** — app names + public URLs; every privilege
> stays behind each app's own auth and the controller admin password. The token IS the secret (160-bit
> entropy is the whole defence for the GET — never rate-limited, never logged, `subtle.ConstantTimeCompare`
> only; an empty stored token = sharing OFF, matches nothing, so a wrong/disabled token is byte-identical
> to the mux default 404). Optional per-share password is a SEPARATE credential (own bcrypt hash, own
> attempt map — NEVER the admin ones); one pass mints a cookie = HMAC(`token|passwordHash`) keyed with
> the persisted `web.session_secret`, so rotate-token OR change-password invalidates all cookies for free.
> **Part-2 secret decision: REUSED `web.session_secret`** (persisted + box-scoped + stable — the SAME
> secret the claim pre-auth CSRF already trusts; not per-boot, not claim-generation-scoped → the reuse
> branch), so no `ShareCookieSecret` field was added. **Design rulings recorded:** member accounts are
> **superseded** by this capability-URL model; **per-member tile visibility is PARKED under the SSO arc.**
> Guest state labels ride the v0.164.0 invariants: `StateStopped` ⇒ "A tulajdonos leállította"; any
> other non-clickable state ⇒ "Átmenetileg nem elérhető" (guests never see stopped/exited/degraded/
> unhealthy). Accepted residuals (documented, no code action): link-preview crawlers fetch once and see
> app names (noindex prevents indexing); reverse-proxy/CF access logs may hold the path (ops-tier); the
> modal link carries the request Host, so a LAN-IP admin session yields a LAN-IP link. New dep:
> `github.com/skip2/go-qrcode`. Tests: Groups A–G (14 tests) + 3 red-proofs verified red.
> **2026-07-24 — v0.164.0 (stopped ≠ fault).** Operator finding on 9201: a UI stop (Leállítás) raised
> the global "Telepített alkalmazás nem fut: … (stopped)" banner on every page AND fired the
> `app_start_failed` email. RULING: **a deliberate user action must not alarm anywhere.** One-line
> filter at the single fix-3 derivation point — `scanDeployedAppRunStates`'s pure core extracted to
> `classifyRunStates([]stacks.Stack)`, down predicate now
> `stacks.IsDownState(st.State) && st.State != stacks.StateStopped`. `StateStopped` is dropped from
> BOTH the banner dead-list and the notifier Down-set (⇒ no banner, no event, clean tracker). Rests on
> **two invariants that MUST both hold for this suppression to be correct:** **I1** — the UI stop path
> `Manager.StopStack` runs `docker compose down` → containers removed → a deployed stack with zero
> containers aggregates to `StateStopped` (refreshStatusLocked). **I2** — the P2 restart-policy census
> (2026-07-21, 53 templates / 78 services) found every catalog service on `unless-stopped`, so a crash
> never rests at `stopped` — faults surface as `exited`/`degraded`/`restarting`/`unhealthy`. **If
> either invariant changes, revisit this suppression.** `IsDownState` UNCHANGED (other callers rely on
> stopped=down). Out-of-band `docker compose stop` (containers remain → `StateExited`) still alerts —
> correct, tampering is reportable. The `stopped_by_user` intent flag was considered and PARKED (only
> adds value against out-of-band stops, which should keep alerting). Tests +4 (notify 3→4, main 4→7),
> both red-proofs verified. No template/funcmap/notifier/counter/copy change.
> **2026-07-24 — v0.163.1 (launcher polish).** Two v0.163.0 live findings fixed. RULE recorded:
> **every app-logo surface ends in a visible placeholder** (`SVG → PNG → /static/app-placeholder.svg`,
> infra rows → `infra-logo.svg`) — the four sibling `onerror` chains (`backups_apps`, `stacks`,
> `app_info` hero, `deploy`) now match `app_row.html`; `app_info` screenshots deliberately still
> vanish on error. And the **launcher monogram is launcher-only AND failure-only**: hidden by default,
> revealed when the tile's img chain fails (`onerror` adds `.launch-tile--noimg`) — it was bleeding
> through every transparent white glyph. Template/CSS only; no handler/funcmap change. 5 tests + 2
> red-proofs. [[launcher-v0163-2026-07-24]]
> **2026-07-24 — v0.163.0 (Indítópult app launcher + universal placeholder icon).** New
> customer-facing `/launcher` page: the FIRST sidebar item (above Vezérlőpult), a grid of large
> tappable tiles for openable deployed apps. `/` stays the Vezérlőpult — the launcher is ADDITIVE.
> Design rulings recorded here:
> - **(a) The felhom brand mark is NEVER an app placeholder** — brand = platform identity only. The
> logo-less fallback everywhere is the new generic `AppPlaceholderSVG` (a 2×2 app-grid glyph,
> `/static/app-placeholder.svg`), now the DEFAULT `FallbackIcon` on `app_list_row` (was
> `visibility:hidden`). On the launcher tile the fallback is the **monogram**, not the placeholder.
> - **(b) A launcher tile exists ⟺ a „Megnyitás" button would** — subdomain presence (env `SUBDOMAIN`
> > `.felhom.yml` subdomain > `protectedStackSubdomains`) is the single openability criterion. The
> controller stack is excluded by name. The subdomain assembly was extracted to
> `Server.subdomainMap` (3 callers: dashboard, Alkalmazások, launcher; priority byte-unchanged).
> - **(c) Colored-tile + mono-glyph design.** `tileColor` = validated `.felhom.yml` `brand_color`
> (`#rgb`/`#rrggbb`, new `Metadata.BrandColor`, omitempty) OR a deterministic FNV-1a-of-slug HSL
> (fixed S/L, hue per app). Invalid `brand_color` silently falls back to the hash color (the one
> §8 exception to no-silent-failure — cosmetic). `tileColor` returns `template.CSS` (we
> validate/compute in Go; html/template's CSS filter mangles a legit `hsl()` from a func pipeline).
> - **(d) `/` remains the Vezérlőpult.** No role/auth gating — member-role gating is a future arc
> (ROADMAP: member role → launcher becomes the member landing page). No catalog app sets
> `brand_color` yet (curation parked).
> No agent coupling; MinAgent unchanged. 10 new test functions + 4 red-proofs (all observed FAIL then
> restored). Gates green (app_row_dedup / template_id / emoji).
> **2026-07-24 — v0.162.0 (R-71a), SHIPPED + deployed BOTH boxes (demo-felhom 9201 + demo-hp 9201
> via G1 break-glass), clean+healthy, settle-gate GO line captured on both.** B′ live note: both
> above-floor boxes GOed correctly but NOT literally first-poll — the floor is in-memory (not
> persisted), unknown at t=0, so the gate logged `awaiting floor knowledge` then GOed ~10 s later the
> instant the report ACK landed (report-ACK latency = exactly what the 90 s sub-bound is sized to;
> zero-wait-when-floor-known is unit-proven, test E). The gate correctly did NOT burn the one-time
> password before the update picture was clear.
> The structural fix for the F10 day-0 race (DIAG-f10): the apply-bridge no longer consumes the
> single-use offsite password while a managed floor-update is in flight or imminent (below floor).
> New seam `offsiteapply.SettleProvider.SettleState()` + `SettleFunc` adapter over the updater's own
> `GetFloor()`/`IsUpdateRunning()` (no second floor path); `Bridge.AwaitSettle` polls 10 s BEFORE the
> 3-min Reconcile ctx (deferral never eats the reconcile budget), bounds 90 s floor sub-bound / 5 min
> overall (both GO+WARN — the "hub that can't serve a floor can't serve a consume → no burn" argument,
> R-71c is the belt). At/above floor → GO first poll, zero wait (B′). Bridge goroutine MOVED after the
> updater in main.go; wired only when an updater exists. **Ordering-only** — consume/persist/404
> contract untouched; R-71(b) rejected-by-design. **FINDING:** the floor is in-memory
> (report-ACK-derived ~5–10 s), NOT persisted → unknown on any restart until the first ACK (sized the
> 90 s sub-bound to that). 5 test scenarios (A–E) + nil-provider + cancelled-gate; **4 red-proofs all
> observed FAIL then restored** (gate/updateRunning/sub-bound/overall-bound). Deferral paths NOT
> live-fired (precondition now structurally prevented by the v1.25.0 build gate). **Layering: gate
> prevents, (a) defers, (c) heals.** ROADMAP R-71 → SHIPPED (a)+(c). Live leg = the B′ first-poll GO
> line on both above-floor boxes.
> **2026-07-23 — v0.161.0 (R-70 controller leg), SHIPPED + deployed BOTH boxes.** When
> `offsite.enabled` is in controller.yaml but no `offbox` target exists (pre-apply window / burned
> credential — the F10 shape), Távoli mentés now shows „Felhom offsite tárhely kiépítve — a
> beállítás automatikus, folyamatban…" on BOTH empty surfaces (status card + target line) instead
> of „igényelhető" / „Még nincs beállítva". Data key `OffsiteHubEnabled` (from `Server.cfg`, no new
> wiring); render tests per gate branch; banner leg is unit-proven/live-pending (no healthy box
> occupies the window; next fresh onboarding is the natural live leg). Hub sibling v0.72.0 carries
> the detector + `offsite_delivery_stuck` + the R-71c self-heal. Origin + rulings:
> `felhom.eu/documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`.
> **2026-07-22 — v0.160.0 (R-67), SHIPPED + deployed BOTH boxes, full live leg on demo-hp.**
> Network shares now bind their share ROOT into FileBrowser (`…/:/srv/:rslave`) — no
> skeleton/userdata toward the NAS, ever. Pure assembly = `buildFileBrowserPaths` + `fbPathDeps`
> (handlers.go), returning mounts AND config sources together so they can't disagree.
>
> **DECISION — two classes, two gates:** drives keep the drive-absent gate (byte-identical,
> tested + observed live: demo-felhom logged a no-op sync); network shares use the STUB classifier
> gate instead (stub ⇒ excluded from both lists + WARN — an exposed stub swallows uploads the real
> mount later shadows; idle autofs is HEALTHY and included; unknown fails open). Never force-wake
> in the sync (doctrine).
>
> **Phase-0 probe = GO:** in-container access through an rslave bind WAKES an idle autofs trigger
> (proved on demo-hp against the real Felhom-Share). Live leg: upload from demo-hp's filebrowser
> container (uid 1000) landed on demo-felhom's share dir and deleted clean; dead-NAS gave
> `Host is down` in seconds (no hang) and recovered unaided after samba restart. RESIDUAL for the
> operator: the FileBrowser HTTP click-through — its admin credential is customer-held (CC got 401
> on admin/admin and the demo password; by design). ROADMAP R-67 SHIPPED (coupled to R-64).
> **2026-07-22 — v0.159.0 (R-66), SHIPPED + deployed to BOTH boxes.** Three legs: „Hálózat" card on
> Beállítások → Rendszer (Helyi cím / Hálózati név only-while-Megosztás / Átjáró; „—" fallback),
> `network` section in the Debug dump (best-effort per item), and the NetBIOS trap named on the NAS
> add form (Szerver helper text + a purely lexical hint on `unreachable` for single-label non-IP
> names).
>
> **DECISION (the load-bearing one): all guest-net reads go through the samba netns door.** The
> controller is bridge-netns'd, so `/proc/net/route`/resolv.conf/net.Interfaces in-process answer
> for the CONTAINER (172.x / 127.0.0.11) — the S-2 trap. `internal/stacks/guestnet.go` docker-execs
> into host-networked felhom-samba (one `guestNetExecFn` seam); Megosztás off ⇒ door closed ⇒ „—" /
> in-place error strings, never a plausible-wrong substitute (S-5). Nothing stored anywhere.
>
> Deploy: 0.159.0 on demo-felhom 9201 (open-door path live: .104/.1/\\FELHOM) AND demo-hp 9201 via
> G1 break-glass (closed-door path live: dashes, no name row, in-place dump errors; secret shredded).
> demo-hp gotcha worth keeping: the controller 404s on direct container-IP probes without the
> customer-domain Host header (`felhom.enkisfelhom.hu` there). Red-proofs A2 + C2 run and recorded.
> ROADMAP: R-66 SHIPPED; R-64 (pairing blessed, drill = evidence leg) + R-65 (buddy-box replication,
> post-alpha spike-first) minted. NAS doc gained the naming-caveat paragraph.
> **2026-07-21 — v0.155.0.** v0.154.0's wizard sourced "is an op running" from `Manager.IsRunning()`
> — the CONCURRENCY single-flight, acquired inside the goroutine, and **`RestoreOffboxScratch` never
> acquires it**. So the execution step was unreachable for „Ellenőrzés" and the full-restore
> preparation: live buttons while a restore downloaded, with the progress banner contradicting the
> phase strip on the same screen. Found by the operator on the first live click-through.
>
> **DECISION: display reads `RestoreStatus()` (the `opRunning` flag), never `IsRunning()`**, through
> the named `restoreOpInFlight` seam, and the handler reads the status ONCE per render so the strip,
> the suppression and the running-op name cannot diverge. The lesson generalises: `opstatus.go` is the
> DISPLAY surface and says so in its own header — the concurrency flag is not a substitute.
>
> **The test lesson:** a table test over a pure function proves the function, not the caller. Scenario
> E passed throughout because it injected `OpRunning=true` directly. The new test drives a real
> `Manager` through `BeginRestoreOp` and asserts the render.
>
> **DECISION: „Eredmény" earns its place.** The strip's highlight is now `Phase`, derived separately
> from `Step`: a finished restore is back on the intent step while the strip reads „Eredmény" and an
> outcome card shows the result — window-bounded (10 min) and app-bound.
> **2026-07-21 — v0.154.0 (R-48).** Collapses the offsite restore controls to a single
> „Visszaállítás…" entry per app row plus a per-app wizard at `GET /backups/restore/app?name=`.
> The defect it closes is the CAUSE of the round-2 incident: the list rendered up to five inline
> forms per row, two of which — the missing-only merge and the true reconstitution — were sibling
> buttons whose difference is whether the data comes back. The rule it establishes: *two adjacent
> controls whose difference is "your data comes back" vs "your data cannot come back" must never be
> distinguishable only by layout.*
>
> **DECISION: the wizard is server-rendered on the EXISTING endpoints.** No new mutation endpoint,
> no JSON state API, no client router. Every card is a real form POST to
> `/backup/offbox/{restore,place,reconstitute}` with the same field names and gates, and the server
> renders the next step — so it works with JavaScript disabled. `TestRestoreWizard_NoNewMutationEndpoints`
> makes that structural: adding a form that posts somewhere new fails the suite by design.
>
> **DECISION: R-45 stays its own item.** The wizard polls the two existing status surfaces as-is; the
> generalized job registry (and with it a real per-phase progress feed) is not built here.
>
> **DECISION: the step is derived, never requested.** `deriveWizardStep` is pure over (op running,
> size-gate flash, scratch ready). Precedence is load-bearing — a running op outranks a stale
> `?full_prep=` in the URL, or a commit button reappears mid-restore. While ANY op runs every
> mutation form is suppressed server-side rather than offered and then refused with a 409.
>
> Latent bug found and fixed on the way: `offboxRedirectTo` hardcoded `"?"` when appending its flash,
> which would have buried the flash inside `?name=`. **No agent coupling — MinAgent stays
> 0.90.0.** 9 new tests + the Group-B red-proof; full suite green.
>
> **NOT live-validated at commit time by design:** v0.154.0 is published but deliberately NOT
> hand-deployed — the operator's hub floor save (0.153.0 → 0.154.0) pulls it via the self-update
> path, and that swap IS the R-23(a) single-fire validation (STOP-1).
> **2026-07-20 — v0.153.0 (R-47).** Closes the H4 race on **BOTH** restore paths. The replay needs a
> running DB container, so both paths started the WHOLE stack first — giving the application a window
> to rebuild the schema objects the dump was about to create. Measured at 8 s on 2026-07-19
> (`DIAG-immich-restore-round2-2026-07-19`): immich-server rebuilt `clip_index` two seconds before
> the dump's `CREATE INDEX`, the replay aborted `already exists` under `ON_ERROR_STOP=1`, and immich
> then reported schema drift. The photos came back **by accident** — `pg_dump` emits COPY before
> CREATE INDEX, so the abort landed after the rows; a collision earlier in the script would have left
> a genuinely half-restored database, reported identically.
>
> **DECISION: the DB-only bring-up is done by compose SERVICE scoping**, not by container tricks —
> `StartStackServices(name, []string{svc})` → `compose up -d `. Every catalog template's
> dependency direction is app→db, so naming the DB starts the DB and nothing else. `docker start
> ` was never an option: `StopStack` is `compose down`, so the containers no longer exist.
> `RestartStack`/`RedeployFromEnv` are traps here — both end in a full `up -d`.
>
> **DECISION: fail-closed.** A `.sql` dump with no identifiable DB service refuses BEFORE the first
> mutation, on both paths (one Hungarian string, shared). The alternative would be to start everything
> and replay into the race. It should be structurally unreachable — `dbTypeForImage` is now shared by
> `DiscoverDatabases` and `DBServiceNames`, and a dump can only exist because discovery matched the
> container's image, which IS the compose `image:` value — so this is the belt for template drift.
>
> Enablers: `RedeployFromEnv` split into `PersistUnitRedeployConfig` (persist, starts nothing) + the
> unchanged tail; `StackDataProvider.RecreateStackFromUnit` renamed to
> `RecreateStackDefinitionFromUnit` because the old name promised less than the method did — the
> hidden `up -d` inside it is what carried the defect on the local path. `StartStackServices` REFUSES
> an empty list (argument-less `up -d` is a full start). **No agent coupling — MinAgent stays 0.90.0.**
> 19 new tests, 3 red-proofs, 23/23 green. **NOT live-validated yet:** STOP-1 supervised reconstitute,
> golden 0.153.0 bake (P3 registry-reachability probe from the vacation site is load-bearing), Viktor's
> two hub saves, and his C6 customer-restore UI run.
> **2026-07-20 — v0.152.0 + felhom-samba 1.1.0 (Megosztás on a Mac).** Closes **S-3**. **A capture
> on the box overturned the earlier guess:** macOS DOES send a correct NBNS query for `<20>` and
> nmbd DOES answer it correctly in 140 µs (flags `0x8580`, RCODE=0, right address) — macOS simply
> never acts on it. NetBIOS there feeds legacy browsing, not `smb://` URL resolution, so **the bare
> `smb://` can never work from a Mac** and nmbd was never the broken part (it is what serves
> Windows). felhom-samba 1.1.0 adds **avahi + dbus**, templating `avahi-daemon.conf` and the
> `_smb._tcp` service file from `FELHOM_SERVER_NAME` so a rename re-advertises; both daemons are
> non-fatal on failure. v0.151.0's card had offered `smb://` for Mac — the one dead form — now
> `smb://.local`; Windows keeps flat `\\`. Spiked live by hand and confirmed from the
> operator's Mac BEFORE publishing the image (the operator's call, and it chose the design too).
> **STILL OPEN: Finder-sidebar discovery is NOT shipped** — the record is published and answers
> browse queries, but was never observed working; likely a Finder Settings → Sidebar toggle, but
> unverified. **Windows was not retested.** Two test bugs fixed en route, neither a production
> defect: `TestRenderSambaCompose` pinned a literal image tag, and `TestFabUpload_GCAndIdleTimeout`
> asserted an async unlink synchronously (it passed alone, failed in the full package once the new
> render tests made `web` heavier). 23/23 green twice; 2 red-proofs.
> **2026-07-20 — v0.151.0 (Megosztás).** Closes **S-1/S-2/S-4-core/S-5** of
> `felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md`; **S-3 (no mDNS/Bonjour) stays OPEN**,
> awaiting Viktor's `smbutil lookup FELHOM` + `dns-sd -B _smb._tcp` from the Mac. **The `/sharing`
> page had been reload-looping at ~1.2 s for every customer with sharing enabled since v0.147.0** —
> `/sharing/status` coerced `idle`→`running` on the JOB phase channel, and the client answers a
> terminal `running` with a one-shot `location.reload()`, so the first poll of every steady-state
> page load re-armed it. The rule this leaves behind, now recorded against R-45 too: **a phase a
> client answers with a one-shot action is an EDGE — never synthesise it from a level, and serve it
> exactly once.** Both halves are server-side; `sharing.html`'s `