Files
felhom-controller/CHANGELOG.md
T

13488 lines
1.1 MiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
## v0.289.1 — the provider's rclone notice no longer reads as "0 snapshots" (found live on demo-felhom) (2026-10-03)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.127.0 (unchanged).
- **Found live in the first run through the pinned key:** the provider's rclone prints
`NOTICE: Config file … not found - using defaults` on every connection, restic forwards it into the combined output,
every `--json` parse failed, and the box recorded **0 snapshots as MEASURED** (`stats_known`) over a store holding 12 —
the shape the hub's R-431 detector reads as a mass deletion — and the clean-up step could not list snapshots.
`Manager.runner()` now strips exactly that notice line (`stripRcloneNotice`; an rclone ERROR line stays).
- **An unreadable snapshot count is no longer a measured zero (R-331 class):** `offboxRecordStats` returns
`(count, ok)`; on a failed or unparseable listing the box keeps the last measured count and reports `stats_known=false`.
- Tests: `TestStripRcloneNotice`, `TestRunOffbox_UnreadableCountIsNotZero` (red-proved). **Second release this session,
against the one-release rule, because v0.289.0 was live on a box and reporting a false zero.**
## v0.289.0 — the off-site key cannot delete: append-only transport, no password on the box, retention only inside a hub window (decisions 68–69, R-820, R-822) (2026-10-03)
**MinAgent: 0.131.0** (unchanged). **Needs hub v0.127.0** (the key registrar and the window endpoints; an older hub
answers 404 to `register-key` and the apply-bridge keeps retrying). No new household string.
- **The box never fetches the Storage Box password any more.** The apply-bridge (`offsiteapply`) sends its PUBLIC key
to the hub's registrar (`HubRegistrar`: `register-key`, `confirm-key`), which pins it in the sub-account to
`command="rclone serve restic --stdio --append-only <repo>",restrict`; the box then PROVES the key reaches the pinned
server (`PinnedProber`: exit 0 and rclone's output — an unpinned key gets the restricted shell, exit 8, measured)
before configuring anything. `HTTPConsumer`, `SSHCopyIDInstaller` and the `sshpass` path are gone. A box upgraded
from v0.288.0 re-applies once (`descriptorHash` gains `|pinned-v1`) and re-registers its EXISTING key.
- **Transport.** A hub-provisioned target (`Transport: "rclone-pinned"`, set by `ApplyOffsiteTarget`) uses restic's
`rclone:` backend with `-o rclone.program="ssh -p 23 … -i <key> … rclone"` — restic 0.14.0 suffices, rclone is NOT in
the image (it runs at the provider). An sftp-written repository reads, extends, restores and `check --read-data`s
through it (measured on the provider). The household's own SFTP NAS target is unchanged.
- **Retention leaves the box on the pinned tier (decision 68).** Both `forget --prune` sites (after a run; over quota)
go through `offsiteWindowRetention`: no window → nothing deleted; inside a hub-granted window the **fake-snapshot
guard (R-822)** refuses on any future-dated snapshot, any snapshot newer than the hub's bound, or a plan that would
remove a snapshot younger than 8 days; otherwise the OLDEST `max_remove` planned snapshots are forgotten by id and the
window is closed with counts. **Disagreement recorded:** the brief said abort when the plan exceeds a week's
removal; the first window after the interim legitimately does, so the box caps and takes the oldest instead.
- **Move-aside is the hub's** (`POST /offsite/move-aside`); **abandonment is deferred to the operator** on the pinned
tier — nothing deleted, `offbox_abandon_deferred` (operator-only) — because the key cannot delete (R-823).
- Transport failures of the `rclone:` backend (`error talking HTTP to rclone`) classify as transport.
- Bug found by the existing suite and fixed before release: the NAS move-aside path assigned a shadowed `newPath`.
- Tests: `offsiteapply` rewritten (fresh, upgraded, already-pinned, registered-but-not-pinned, wrong fingerprint,
host-key mismatch, idempotent, confirm failure); `offbox_window_test.go` (the lab's 13 future fakes refused; recent
removal refused; honest plan oldest-first capped; pinned run never forgets without a window; pinned move-aside asks the
hub; pinned abandonment defers). Red-proofs: `felhom.eu/documentation/audits/offsite-lock-build-2026-10-03/partC/`.
## v0.288.0 — remove tells the truth about the household's files (decision 67, R-800) (2026-10-02)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New string: `layout.userdata_marad` (both languages).
- **The household's userdata folders are named on remove.** `GET /api/stacks/{name}/hdd-data` and the results of
`POST /api/stacks/{name}/remove` and `DELETE /api/stacks/{name}` carry `userdata_kept` — the folders the app binds
through `${USERDATA_PATH}` that exist (always a list, never null). Both remove dialogs and both result messages show
„A fájljaid ezekben a mappákban megmaradnak — a fájlböngészőben látod őket:" / "Your files in these folders stay —
you see them in the file browser:" with the folders. Nothing new is deleted or kept: a remove never deleted userdata
(pinned since R-442 by `TestRemoveStack_R442_UserdataConventionFromPerAppPath`); it only said nothing about it.
- **The false note is gone:** for an app whose only drive folder is userdata (MeTube), "remove with data" no longer
answers „Az alkalmazás nem tárolt saját adatot külső meghajtón…" — the kept folder is named instead.
- Tests `TestRemoveStack_R800_UserdataKeptIsNamed`, `…_SSDAppUserdataKeptEmpty`; red-proofed (two mutants). Parity
fixtures regenerated: additions only (109 pages carry the shared layout).
## v0.287.0 — the family gate: family members with their own logins in front of chosen apps (decisions 63/64, R-780) (2026-10-02)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `family_gate.*`, `err.stacks.family_gate_failed`,
`err.stacks.needs_newer_controller` (both languages). **Catalog companion:** `family_gate:`, `family_gate_except:` and
`min_controller:` in `.felhom.yml` (a template with `family_gate` must say `min_controller: "0.287.0"`).
- **The family list** (`internal/family`): the household's dashboard admin adds, resets and removes family members on the
security page („Család" card). Each has their OWN name and a generated password (4×4 letters/digits, shown once, in the
answer to the press — never in a page, never logged). Sessions last 30 days and survive a restart (`family.json`,
0600, atomic). A reset (the member's generation moves on), a removal or a logout ends access at the next request.
- **The door** (`internal/stacks/family_gate.go`): an app whose template says `family_gate: true` gets a traefik
forwardAuth file BEFORE its first start (also when a removed app is restored); a life record (`family_gate:` in
app.yaml) keeps it while the app is installed — a catalog change never gates or un-gates an installed app. Priority
below the install hold, the setup gate and the sign-up block. `family_gate_except:` lists literal path prefixes
(e-reader and phone apps) routed WITHOUT the door — **anchored** `^/prefix(/|$)` (the spike's finding F1: an unanchored
`PathPrefix(/api/v1/opds)` let `/api/v1/opdsx` through); a matcher or regex in the template refuses the install.
- **The answerer** (`internal/web/family_gate.go`): `/__felhom_gate/family` (forwardAuth; the app cookie
`felhom_famgate` is host-only and names a store session); `/__family/start|login|logout` on the dashboard host,
outside the dashboard's auth (the family session cookie `felhom_family` is scoped to `Path=/__family`). The sign-in is
counted per VISITOR (`clientIP`, R-753: 5 per minute) and per NAME (10 per 10 minutes) — a stranger locks only himself;
a name under a spread attack waits minutes. The household's dashboard session vouches (decision 46's rule) as a
household session in the family store. **`RequireAuth` never reads a family cookie; no family page ever sets the
dashboard cookie.** While the controller is down a gated app answers an error (traefik's forwardAuth), never the app.
- **`min_controller:`** — a template that needs a newer box is refused before anything is written.
- Tests: `TestFamily_*` (store), `TestFamilyGate_*` (web: stranger, member-not-dashboard, reset/remove/logout, locks,
household, token, setup gate untouched, messages, card), `TestFamilyExceptRegexp_Anchored`, `TestFamilyGate_*`
(stacks: before the first start, bad exception refuses, restore + loop), `TestMinController`. Red-proofs RP-F1..RP-F7,
each seen failing (`felhom.eu/documentation/audits/family-gate-2026-10-02/A/`).
## v0.286.1 — R-772 found live: a stopped probe container is seen on the next tick (2026-10-01)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). No new strings. **v0.286.0 was never floored**
(scratch 9202 only); the fleet goes from 0.285.0 to 0.286.1 and takes everything in v0.286.0 below.
- **R-772, measured live on 9202 with v0.286.0:** with a HEALTHY record 2 minutes old, stopping paperless-webserver
left the record `healthy: true` until its 5-minute interval ran out — the not-checked branch sat behind the interval
check. Finding the probe container costs no network call, so it now runs first: a stack with nothing to probe is
recorded not-checked on the very next tick. Test `TestRunHealthProbes_AStoppedContainerIsSeenAtOnce`; red-proof
RP-D1b (the 0.286.0 order fails it). Evidence `felhom.eu/documentation/audits/visitors-2026-10-01/D/`.
- **Live on 9202 with v0.286.0 (unchanged in .1):** Part A — a stranger's 5 wrong dashboard logins through the simulated
tunnel, rotating a forged leftmost address, locked only him; the household signed in at once; a LAN or impostor
forgery counted as its own address (`…/A/L1-9202-live.txt`). R-773 — Karakeep removed (keeping backups) and restored:
record `opened_by: restore`, block file present, a stranger's `/signup` and `users.create` 403, the data back
(`…/D/r773-live.txt`).
## v0.286.0 — the box tells visitors apart (R-753); a health check that could not run is not "healthy" (R-772); a restore keeps the sign-up lock (R-773) (2026-10-01)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `login.msg.*` (3, both languages).
**Catalog companion:** 19 apps whose software reads the LEFTMOST `X-Forwarded-For` carry a router middleware that removes
the chain (app-catalog `04e9516`..`50e4fb4`), pushed BEFORE this release — without it, trusting the tunnel would let a
stranger write the address those apps believe. **Every installed box's traefik and cloudflared are recreated once** when
it takes this release (routing pauses a few seconds; the tunnel reconnects).
- **R-753 — `09` §3 decision 63 (operator ruling: extend the box's own gate; first the box tells visitors apart).**
Measured first (`felhom.eu/documentation/audits/visitors-2026-10-01/A/`, demo-hp's real tunnel): every tunnel visitor
reached traefik as cloudflared's one docker-assigned address; Cloudflare APPENDS the visitor to a client-written
`X-Forwarded-For`, passes a client's `X-Forwarded-Host`/`-Port`, strips a client's `X-Real-IP`, and refuses a
client-sent `CF-Connecting-IP` at the edge (403).
- `felhom-tunnel` network `172.16.253.0/29` (allocation confined to `.4/30` — measured: traefik joining first was
given `.2`): cloudflared alone at `.2`, traefik at `.3`. traefik's `websecure` trusts forwarded headers from
`172.16.253.2/32` only; the entrypoint middleware `felhom-forwarded@file` removes `X-Forwarded-Host/-Uri/-Method/
-Prefix/-Tls-Client-Cert(-Info)`, `Forwarded`, `True-Client-Ip`, `X-Client-Ip`, `X-Cluster-Client-Ip`, `Client-Ip`,
`X-Original-Forwarded-For` and fixes `X-Forwarded-Port: 443` (`internal/infra`).
- `EnsureBaseStack` RECONCILES a running traefik/cloudflared whose rendered files changed (recreate; refuses a rewrite
that would drop the running certificate resolver), writes the middleware file before `traefik.yml`, and moves
cloudflared only once traefik is on the tunnel network. No network → the old shape, nothing trusted.
- `clientIP` (`internal/web/clientaddr.go`): believed only when the TCP peer is traefik (resolved by name); the
rightmost `X-Forwarded-For` entry — the hop traefik saw; the tunnel hop → `CF-Connecting-IP`. `rateKey`: IPv6 per
/64. The dashboard login, claim code, share password and escrow re-auth counters key on it: **a stranger's five
wrong dashboard passwords lock only the stranger** (before: every tunnel visitor shared one key and the household
was locked out of its own dashboard for a minute). The setup gate logs the visitor.
- The dashboard login's messages are keys, informal voice, both languages (`login.msg.*`; the placeholder too).
- Tests: `TestClientIP_*`, `TestLogin_StrangerThroughTheTunnelLocksOnlyHimself`, `TestLoginMessagesFollowTheReader`,
`TestRateKey_IPv6Per64`, `TestPeerResolver_*`, `TestTunnelConstantsAgree`, `TestRenderTraefik_TrustsOnlyTheTunnel`,
`TestRenderCloudflared_AloneOnTheTunnel`, `TestRenderForwardedHeaders_*`, `TestEnsureTunnelNetwork_*`,
`TestEnsureTraefik_*`, `TestEnsureCloudflared_*`, `TestEnsureBaseStack_TunnelOrder`;
`TestLoginRateLimit_RotatingXFF_NotLimited` reversed on purpose (a direct peer's rotating XFF no longer evades).
Red-proofs RP-A1 (leftmost hop), RP-A2 (the tunnel hop as the key), RP-A3 (no reconcile) — each fails.
- **R-772:** a health probe that finds no container to probe records `healthy: false, not_checked: true` (was
`healthy: true`, then 5 minutes of silence), is looked at again on the 10-second cycle, and the app page says the
check did not run. The state stays the containers' (`probeSaysUnhealthy`), so R-630 holds. Red-proof RP-D1.
- **R-773:** a REMOVED app restored from its backup gets its sign-up lock back — the record (`opened_by: restore`) and
the block written before anything starts; the loop sets the app's own switch. An installed app the household never
closed keeps what it had (decision 49). Red-proof RP-D2.
## v0.285.0 — a box keeps two controller versions (decision 56); a crash in the update clean-up fixed (2026-10-01)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). No new strings.
- **Controller image retention — `09` §3 decision 56 (R-745).** Decision 53's sweep never touched the controller's
own repository, so every release a box ran stayed (demo-hp's guest: 83 controller images on 2026-10-01). Now, 3
minutes after every start (`stacks/controller_image_retention.go`), the box keeps the controller image it runs, the
one before it — by the swap's own record (`update-state.json` `previous_image` of a swap that ended on the running
version, `selfupdate.UpdateState.RecordedPrevious`), by version order only when no record names one present —
every version ABOVE the running one (a pulled target not swapped to yet) and every non-version tag (`latest`, `-rc`).
Older and untagged controller images are deleted by exact name/ID after the in-use check; never forced, never pruned;
nothing while the controller swaps itself; registry tags never. One INFO line per pass, one per deletion with size.
**Measured first:** the agent's swap rolls back to what `/etc/felhom-controller-image` named when the swap began — the
RUNNING image — and is never handed a previous image; so the self-update's own roll-back needs only the running one.
- **R-751 (found by the full suite):** the retention after an update runs in a goroutine and re-reads the stack after a
rescan; when the app was gone by then it dereferenced a nil stack — a panic in a goroutine ends the controller. Now
it returns. The two retention seams are no-ops in the package's tests (`TestMain`), where the goroutine outlived its
test and wrote into a removed temp dir.
- Tests `TestControllerRetention_*` (6), `TestRecordedPrevious`, `TestRetainImagesAfterUpdate_AppGoneDoesNotPanic`;
red-proofs: drop the previous from the keep switch, ignore the record, drop the swap check, drop the in-use check,
drop the success check, drop the nil-stack check — each fails (`felhom.eu/documentation/audits/rulings-2026-10-01/B/`).
## v0.284.2 — the image clean-up sees digest-pulled (untagged) images (2026-09-30)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). No new strings. v0.284.0 and v0.284.1 were never
floored (scratch 9202 only); the fleet goes from 0.283.1 to 0.284.2.
- **Found live on 9202:** the product pins `tag@digest` (v0.269), and an image pulled that way is stored UNTAGGED
(`repo:<none>`). `docker image ls` without `-a` did not list any of them, so the retention could not see most app
images: the one-time sweep deleted 3 images while dozens of old untagged app images stayed. Now `image ls -a`; an
untagged image keeps its repository name, so it is attributed like any other; a fully anonymous `<none>:<none>`
entry is never a candidate. The one-time marker is `image-retention-v2.done`, so the corrected sweep runs once on
every box (9202 ran the blind v1).
- Test `TestImageRetention_SeesUntaggedDigestPulledImages` (the fake Docker hides untagged images without `-a`, as
measured); red-proof: drop `-a` → it fails.
## v0.284.1 — the image clean-up also runs on the household's Remove button; one summary line per pass (2026-09-30)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). No new strings. **v0.284.0 was never floored** — it
ran only on scratch guest 9202, where this was found; the fleet goes from 0.283.1 to 0.284.1.
- **Found live on 9202:** v0.284.0 wired the remove half of decision 53 into `DeleteStack` only; the app page's Remove
(`POST /api/stacks/<n>/remove`) runs `RemoveStack`, so a removed app's images stayed (the "seam built but never wired"
class, a fifth instance). Now `RemoveStack` reads the app's image repositories before its `compose down` and runs the
retention after (seam `retainAfterRemoveFn`). Its old comment "keep images for potential redeploy" is superseded by
decision 53 (a redeploy pulls).
- **One INFO line per retention pass, whatever it did** (`pass over N image(s) of [repos] — C candidate(s), D deleted`):
v0.284.0 logged deletions only, so a pass that ran and kept everything was indistinguishable from one that never ran.
- Tests `TestImageRetention_TheRemoveButtonRunsIt`, `TestImageRetention_ADoneUpdateRunsItWithThePrevious`; red-proofs
(the call removed → each fails).
## v0.284.0 — a box deletes old app images (decision 53); an after_install app is held until its known login is replaced (R-741) (2026-09-30)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `app_info.install_hold_title`,
`app_info.install_hold_closed` (hu + en).
- **Image retention — `09` §3 decision 53 (R-736).** A remove ran `compose down --rmi local`, which never removes a
registry-pulled image; an update left the old version's image; nothing else deleted any (9202: 84 images, 53 GB
used by no container, an install refused). Now (`stacks/image_retention.go`): when a guarded Update ends `done`,
`previous_images` records what the app ran before; the app's images older than the running one and that previous one
are deleted. At `undone` the attempt's image goes. At remove the app's images go. **Never** an image any container
(running or stopped) uses, any installed app's live compose names, or any installed app's installed/previous record
names — a box-wide keep set read at delete time; deletion by exact image id, never forced, never `prune`; an id
with several repositories' names is left alone; a keep set that cannot be read deletes nothing; **no pass runs while
any update runs** (its undo's image is named by nothing then). One line per deletion (names, id, size). A one-time
sweep at the first start after this release (3 min after start, a marker file) applies the rule to every image the
catalog names, so the images of apps removed before this release go too; never the controller's or infrastructure's.
Tests `TestImageRetention_*` (6); red-proofs: the undo's image, a stopped app's compose, the update-in-flight pause,
the unreadable keep set, the shared engine image.
- **The install hold — R-741 (decision 45).** Measured 2026-09-30: calibre-web answered its public default login
through traefik for 1–18 s after a fresh install, before `after_install` replaced it. Now an `after_install:` app is
installed HELD (`stacks/install_hold.go`): before its first start a traefik file puts the setup gate's forwardAuth
door in front of every router it publishes, at a priority above the gate and the sign-up block. A stranger is refused;
the household (dashboard session) passes, so a failed `after_install` leaves it able to change the login by hand and
press "I changed it", which opens the hold. `after_install` succeeding opens it (record, then file). The gate loop
reconciles: stale files go, a record that says the login was replaced opens, an install whose hook a restart cut off
re-runs `after_install` once (only installs older than the process). The app page shows a card while held.
Tests `TestInstallHold_*` (6); red-proofs RP-IH1..4 (the file before the first start, the open on success, the
household's word, the re-run only after a restart). Found while testing: `DeployedAt` has whole seconds, so the
comparison uses the process start truncated to the second.
## v0.283.1 — a Stop holds during the nightly volume dump IN PRODUCTION, and at the crash recovery (R-721) (2026-09-30)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). No new strings.
- **Found live, on v0.283.0 itself (9202, 2026-09-30 08:42):** the household's Stop was recorded during the volume
dump, and the dump restarted the app 8 s later. v0.283.0 asked `WantsStopped` through an optional interface; the
unit test's fake answered it, but production hands the backup manager `cmd/controller`'s `stackAdapter`, which did
not — the assertion failed and the check silently skipped (the "seam built but never wired" class). The quiesce
path was fine (it is handed the stacks manager itself).
- `stackAdapter.WantsStopped` and `gatedAppStopStarter.WantsStopped` forward to the stacks manager; the startup
crash recovery (`AppStopGuard.Recover`) skips an app the household stopped, and still clears its marker.
- Tests: `TestR721_TheBackupsStackProviderAnswersWantsStopped` (the PRODUCTION types — adapter, gated starter, stacks
manager — each answer the question), `TestR721_CrashRecoveryKeepsTheHouseholdsStop`. Red-proofs RP43, RP44.
- A second release in one session, deliberately: the rule is one release per repo per session, and shipping a Stop
that is known to be undone every night serves no purpose of that rule.
## v0.283.0 — apps go off-site by themselves (decision 50), with a size warning; a Stop holds during a backup (R-721); page slips (R-724, R-725) (2026-09-30)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `backups_offsite.offer_text`,
`backups_offsite.offer_all`, `backups_offsite.offer_no`, `backups_offsite.fit_head`, `backups_offsite.fit_total`,
`backups_offsite.fit_quota`, `backups_offsite.fit_largest`, `backups_offsite.fit_choose`, `flash.offbox.enabled_all`
(hu + en). Evidence: `felhom.eu/documentation/audits/evidence-fixes-first-tester-2026-09-30/`.
- **Decision 50 (R-720):** a fresh install on a box whose customer has off-site switches the app's off-site copy ON
(`settings.DefaultOffboxOnForNewApp`, from the deploy-done hook). An app with an earlier backup choice keeps it (a
reinstall after a removal that kept the backups keeps the household's OFF). Apps already installed are never
switched by the release: both backup pages offer ONE press („Van alkalmazás, amelyről nem készül távoli mentés.
Bekapcsolod mindegyikre?" — `POST /backup/offbox/enable-all`, „Nem most" → `/backup/offbox/offer-dismiss`, shown
until answered).
- **The size warning:** the apps' off-site size (recovery unit + mandatory files) is estimated at the start of every
off-site run and, in the background, when the page is opened and the estimate is older than 6 h (the page never
runs `du`). Over the quota the Távoli mentés page names the three largest with their sizes and asks the household
to choose. **Measured first:** over the quota the box already refuses NEW pushes and runs only the ruled retention
(`forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6`), and an app whose files would cross the quota is pushed
settings + database only — it deletes no history to make room. Now pinned (`TestDecision50_OverQuotaDeletesNothingExtra`).
- **R-721 — a Stop holds:** the household's Stop pressed while a machine had the app down is honoured at that
machine's resume — the whole-guest backup's quiesce (`restartAll` asks `WantsStopped`), the nightly volume dump
(and its crash marker owes no restart), and the nightly update leg (`stopped_by_household` skip: an update would
bring the app up). Measured 2026-09-29: the backup's resume started a stopped app the same second.
- **R-724 / R-725 (box half):** Beállítások shows the schedule the box really runs (window + legs) and the update
check in local time; the dashboard's „Utolsó mentés" is in local time (R-500); the backups card reads „az előző
N órája készült" instead of a „0 órája" under „Következő mentés"; the restore-test card names the tier it read
from; the recovery-code wizard speaks „te" (formal-form ceiling 18 → 17).
- Tests: `TestDecision50_*`, `TestR721_*`. Red-proofs RP31–RP38, each seen failing on an assertion.
## v0.282.0 — the app's own sign-up switch after the setup (after_setup); "close sign-up now" (decision 49); probes read lists and a "done" status (R-715) (2026-09-29 evening)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `app_info.close_signup_text`,
`app_info.close_signup_btn`, `app_info.signup_native_failed`, `app_info.signup_window_restart`,
`err.setup_gate.close_signup_not_offered` (hu + en). Evidence: `felhom.eu/documentation/audits/signup-lock-2026-09-29/`.
- **`after_setup:`** (`.felhom.yml`): `env: {SIGNUP_CLOSED: "true"}` (or a command, with after_install's argv-safe
rules). Run once when the gate opens — after the address block is up — merged into the app's env and ONE
`compose up -d`. Measured: 9 of the 11 apps with a block have their own switch, all read from the environment; 6 of
them refuse even the household's first admin while it is on, so it goes on AFTER the setup. Two locks now: the app's
own switch and the address block. An installed version whose compose does not read the variable is reported, never
recorded as locked. The household's 15-minute window lifts both (one restart) and the loop closes both again (one
more). A failed switch retries at most every 30 minutes; the page says so; the block holds.
- **"Close sign-up now" (decision 49):** an installed app whose template has a lock and whose install has none (an
app installed before decision 47) shows one sentence and one press (`POST /apps/<slug>/close-signup`). It writes a
lock record (`opened_by: close-signup` — never a closed gate), the address block, then the app's own switch.
Offered once. The box never applies it by itself.
- **R-715:** a probe's `field` may index a list (`setup.0.status`, `0.done`); `done_status:` counts a fixed non-200
answer as done (gramps-web's 405); any other non-200 stays "cannot read" (fail closed).
- Tests: `TestAfterSetup_*`, `TestCloseSignup_*`, `TestCloseSignupPage_*`, `TestProbe_ListIndexesAndADoneStatus`.
Red-proofs RP25–RP30, each seen failing on an assertion.
## v0.281.0 — "Done" asks the app first; open sign-up closed once the first admin exists; a password is never read as code (2026-09-29)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `signup_closed.page_title`,
`signup_closed.page_body`, `app_info.setup_gate_confirm`, `app_info.signup_title`, `app_info.signup_closed`,
`app_info.signup_add_how`, `app_info.signup_open_until`, `app_info.signup_window_btn`, `err.setup_gate.probe_not_done`,
`err.setup_gate.no_signup_block` (hu + en). Evidence: `felhom.eu/documentation/audits/gate-rollout-2026-09-29/`.
- **„Kész, beállítottam" asks the app first.** For an app with a `setup_done_probe:` the press reads it and REFUSES
while it says "not done", and when it cannot be read (409, "The app says its first setup is not done yet. Create your
account, then try again."). Measured 2026-09-29: uptime-kuma pressed before its setup answered anyone. A probe app
now also shows the press. An app without a probe asks in the page first (`felhomConfirm`: "Press this only after
you have created your own account. After this, anyone can reach the app's address.").
- **Sign-up closed once the first admin exists (`09` §3 decision 47).** `.felhom.yml` `signup_block:` (a traefik
matcher for the app's own sign-up address). When the setup gate opens, the box writes
`<stacks>/traefik/dynamic/signup-block-<app>.yml` FIRST (a failed write keeps the gate closed), then removes the
gate: that address alone answers "sign-up is closed" (`/__felhom_gate/signup-closed`, a page or 403 JSON). The app
page's new sign-up card says how to add a family member (`app_info.add_people`, per app, hu + en) and has
„Regisztráció megnyitása 15 percre" (`POST /apps/<slug>/signup-window`); the loop closes it again. Only an app
whose gate this box opened gets a block — an installed app is never touched.
- **R-713:** `after_install` refuses a value that would land inside code (text with spaces, quotes or brackets around
the placeholder) when it holds a quote, a backslash, `$`, `{`, `}`, a backtick or a line break. Its own argument and
plain arguments (`--password=${X}`, `admin:${X}`) take any value. New `${NAME|base64}`.
- Tests: `TestSetupGateButton_*`, `TestSignupBlock_*`, `TestSignupClosed_*`, `TestSetupGatePage_ConfirmAndSignupCard`,
`TestR713_*`. Red-proofs RP17–RP24, each seen failing on an assertion.
## v0.280.0 — the setup gate: a new app is closed to strangers until its household set it up; "I changed it"; an installed app's password leaves the page HTML; a generator with a special character (2026-09-29)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `setup_gate.page_title`,
`setup_gate.page_body`, `setup_gate.sign_in`, `app_info.setup_gate_title`, `app_info.setup_gate_closed`,
`app_info.setup_gate_probe`, `app_info.setup_gate_button_hint`, `app_info.setup_gate_done_btn`,
`app_info.default_login_changed_btn`, `err.stacks.setup_gate_failed`, `err.setup_gate.not_closed`,
`err.setup_gate.open_failed` (hu + en). Evidence: `felhom.eu/documentation/audits/login-gate-2026-09-29/`.
- **The setup gate (`09` §3 decision 46, built after the spike PASSED — `audits/login-gate-2026-09-29/B/B-VERDICT.md`).**
A template with `setup_gate: true` is installed CLOSED: before the app's first start the controller writes a traefik
file-provider router per app router (same rule, priority 100000 + rule length, `forwardAuth` →
`http://felhom-controller:8080/__felhom_gate/auth`, then the app's own `<svc>@docker`). The answerer lets a request
through only with a host-only gate cookie; a browser without one goes to `https://felhom.<domain>/__gate/start`,
where a valid DASHBOARD session gets a 60-second one-use token bound to the app host, swapped on the app host for
the cookie. The dashboard cookie is never widened. A script or a phone app gets 401. The gate OPENS (record first,
then the file) when the app's own `setup_done_probe:` says so (read on the docker network every 20 s), or when the
household presses „Kész, beállítottam". The record (`setup_gate:` in app.yaml) is a life record: a restart rewrites a
missing file, a restore keeps it, a kept-data load never gates. A gate that cannot be written REFUSES the install.
The key is persisted (`<data>/setup-gate.key`, 0600): a restart does not re-gate a browser that passed.
- **R-710:** an app installed before its template gained an `after_install` was never warned about its live default
login (an absent record read as "not run yet" for ever — measured on demo-hp's bookstack). An absent record now means
"not run yet" only for 30 minutes after the install. And the default-login card has „Megváltoztattam" / "I changed it":
the household's word, recorded as `default_login: {changed_at, by}`; the card then goes.
- **R-709:** an installed app's `type: password` value is no longer written into its settings page; the eye fetches it
(`/stacks/<n>/auto-field/reveal`, now also for password fields of an installed app, never for a restore-generated one).
- **`generate: password:N:special`:** a lower, an upper, a digit and one of `-_.!@#%+=`, a letter or digit first — for
an app whose own policy demands it (calibre-web). The install page's „Generálás" button makes the same shape.
- Tests: `TestSetupGate_*` (stacks + web), `TestSetupGatePage_*`, `TestKnownLogin_*`, `TestR709_*`,
`TestGenerateValue_PasswordWithASpecialCharacter`; a `composeExecFn` seam so a deploy test never reaches Docker.
Red-proofs RP1–RP16, each seen failing on an assertion (`audits/login-gate-2026-09-29/C/redproofs/`).
## v0.279.0 — no app goes live with a login a stranger knows (after_install); an empty backup of a running app is an alarm; the night's chain on a button; R-706 (2026-09-28)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged; Part D rides the existing operator-only
`backup_run_failures` digest). New strings: `app_info.known_login`, `backups_apps.hollow_local`,
`backups_apps.hollow_offsite`, `debug.night_chain` (hu + en). Evidence: `felhom.eu/documentation/audits/logins-nvme-2026-09-28/`.
- **`after_install:` (`09` §3 decision 45).** A template may declare ONE command the box runs once in the app's own
container after a FRESH install (the deploy-done hook, ok=true) — to replace a known default login with a generated
`type: password` field, shown on the app page as the first password. `env:` names the deploy values filled into
`${NAME}` (an undeclared or empty one refuses — never an empty password); `success:` is a marker the output must
carry (claper's CLI exits 0 on an error). Retried while the app boots (6 × 20 s), recorded in `app.yaml`
(`after_install: {at, ok, detail}`), never logged expanded. Never on a restore or a kept-data load (R-694). The
deploy-done hook is now set on EVERY box (it was inside `if notifier != nil`, so a box with no hub would never run it).
- **The page says when a default login is in effect** (`defaultLoginInEffect`): the install dialog before the install
when nothing will replace it, the app page while it is in effect ("This app starts with a known, shared password: %s.
Change it right after the install."); the default-login card is HIDDEN once after_install replaced it.
- **Part D — a running app whose newest copy holds no data is an alarm.** At the end of the dump leg (the local unit)
and in the off-site loop (the pushed unit): an installed app that RUNS and has volumes, whose copy lists no dump and no
tar → the operator digest once per app per tier per day, and a sentence on the backup page until a later check finds
data. A held app that is stopped is not flagged. Measured before: yesterday's demo-hp nextcloud had a WARN line only.
- **R-705 (controller half):** debug action `POST /api/debug/backup/night-chain` — dump → Tier 2 → off-site → update leg,
one at a time, refused while a backup/restore op, a guarded update or another chain runs. The update leg gets its
normal length from its own start (`RunUpdateLegNow`; by day the night's W+5h has passed).
- **R-706:** a removal "with its backups" also deletes the app's off-site verification copy.
- Tests: `TestAfterInstall_*` (stacks + main.go wiring), `TestKnownLogin_*`, `TestPartD_*` (backup + page),
`TestR705_*` (web + stacks), `TestR706_*`. Red-proofs RP7–RP15, each seen failing on an assertion.
## v0.278.0 — a hold left by an earlier install no longer holds the new one (R-704) (2026-09-28)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: none. Evidence:
`felhom.eu/documentation/audits/kept-offsite-2026-09-28/` (E2b, redproofs RP4–RP6),
`…/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`.
- **Found live 2026-09-28.** demo-hp's fresh nextcloud (installed 10:13) carried an update hold from a nextcloud of
2026-09-13 — set before v0.242.0, whose removal did not clear it. The backup leg then skipped the new install's
volumes and unit ("the app is HELD stopped"), so its off-site snapshot held no data, and its page showed a two-week-old
failure. On 9202 a calcom crash-loop stop (decision 28) survived two removes and refused the new install's update.
- **Removal** (`settings.ClearUpdateHold`) now clears the crash-loop stop too, not only the update hold.
- **A new install** — a plain install or "use my kept data" — drops a leftover update or crash-loop hold of an app that
is not installed (`Router.dropLeftoverHold`), which also heals holds already left on boxes. A restore hold (R-379)
is never touched by either: it stays operator-cleared.
- The kept load's unit restore now goes through the same seam as the off-site load (`keptUnitRestore`).
- Tests: `TestR704_ClearUpdateHoldClearsTheInstallsHoldsOnly`, `TestR704_AFreshInstallDropsALeftoverHold` (the "use"
path end to end behind seams, the R-379 negative, and the plain install's call before `DeployStack`). Red-proofs
RP4–RP6, each seen failing.
- **Second controller release this session** (the rule is one): R-704 blocked the brief's off-site proof on demo-hp and
there is no product route to clear the leftover hold.
## v0.277.0 — „Use my kept data" and Load also bring the database back from the off-site copy (R-691 (2)) (2026-09-28)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: `kept.backup.offsite`,
`kept.choice.use.desc_from`, `err.kept.offsite_version_unknown`, `err.kept.offsite_not_usable` (hu + en). Evidence:
`felhom.eu/documentation/audits/kept-offsite-2026-09-28/`.
- **The choice looks off-site too.** `backup.KeptBestCopy` = the newest usable local copy (own unit, second drive) or
the off-site copy when it is NEWER or the only one (a tie goes local — no download). The off-site copy is asked with
ONE `snapshots --json` (bounded 20 s; the kept list asks once per page) and offered only when its newest snapshot
holds the app's recovery unit. Dated by its data time (`07` §6.6 A4).
- **The page names the copy and its date.** The install choice now says „Az adatbázist innen töltjük vissza: távoli
mentés, 2026-09-28 04:15. …"; the kept list says „távoli mentés, …" in its copy column (shared `backup.KeptCopyKey`).
- **The load (`LoadKeptOffsite`)** downloads the unit ALONE (the nightly proof's unit-only restore, into the proof
scratch), judges it — holds data, taken of THIS drive, and its data version is RECORDED (`07` §6.6: a unit with no
`data` block is refused, never loaded blind; mixed or mismatched is refused by `unitVersionCheck`) — and only then
moves a dated folder's files back and runs the one unit restore (`RestoreFromRecoveryUnitAt`). The downloaded copy is
removed on every path BEFORE the outcome is reported. A failed download or a refusal leaves the kept files as they were.
- Tests: `TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest`, `TestR691_OffsiteLoadDownloadsJudgesThenRestores`,
`TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone` (4 refusals), `TestR691_InstallChoiceNamesTheOffsiteCopy`
(hu + en). Red-proofs RP1 (0.276.0's local-only choice), RP2 (no version refusal), RP3 (the copy removed after the
outcome — seen in the first draft), each seen failing.
## v0.276.0 — a restore and a drive move keep the app's records (R-697, R-700) (2026-09-27)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: none. Evidence:
`felhom.eu/documentation/audits/records-carried-2026-09-27/`.
- **R-700 (new, found reading the code for R-697): a drive move unpinned the app.** `doFlipRedeploy` persisted
through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the
update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim
and the next start takes the newest version — past the ladder, and for a PostgreSQL app past its conversion step.
Now `persistDriveFlip` changes `HDD_PATH` and nothing else (load-then-save); the up-and-report tail is
`upFromAppConfig`, shared with `RedeployFromEnv`.
- **R-697: a restore dropped the kept pre-conversion copy's record, so the copy was never released.** The restore's
write (`PersistUnitRedeployConfig`) now carries the app's life records from the `app.yaml` it replaces
(`carryLifeRecords`): `conversion_copy`, `earlier_conversion_copies`, `desired_state` (dropped, a dead app after a
restore read as "unknown intent" and was never alarmed), `failed_update_step`, `last_update_undone`,
`last_auto_update`. Not carried: the pin (the restore pins to the unit), `installed_images` (an observation of what
ran before). And a second conversion after a restore to the old major no longer overwrites the first copy's record:
it moves to `earlier_conversion_copies`, released by the same rule.
- Tests: `internal/stacks/r700_records_carried_test.go` (4). Red-proofs RP1–RP4, each seen failing.
## v0.275.0 — a backup's data and its version travel together (R-696, `07` §6.6, D4 option A); R-695, R-691, R-694, R-699 (2026-09-26)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: yes (hu + en, 5 keys). Evidence:
`felhom.eu/documentation/audits/version-travel-2026-09-26/`.
- **A1 spike first (9202, 0.274.0, `A1/README.md`).** The five-minute status refresh re-captured the recovery unit's
DEFINITION (`compose/`, `image_pins`) two minutes after an update, over the previous version's DATA. Restored in
that window: an app step (docmost 0.95.0 → 0.96.0) came back only because docmost migrated the old data at its
first start; an engine step (PostgreSQL 16 → 18) poured the 16 datadir back, `postgres:18` refused it and the app
was **left down**. The off-site restore never wrote a definition at all.
- **The data records its versions (A2).** Every data file a leg writes is stamped at that moment (`stampDataFile`:
size, mtime, the definition's pins, `installed_images` as `ref@digest`, in the unit's `data-stamps.json`); the
capture folds them into the manifest's new `data` block and **keeps the definition the data belongs to** — a
refresh after a pin change no longer rewrites `compose/` until the next data run. `image_pins` stays the app's
current pins. Wired in the nightly DB and volume legs and the update's own "back up first" (new `dumpOne` seam).
- **A restore never mixes versions (A3).** The unit restores (own unit, second drive, kept-data Load) refuse, before
anything is touched, a unit whose `compose/` names other pins than its data or whose files were written by
different versions (`ErrUnitVersionMismatch`); the off-site restore writes the snapshot unit's definition (and pin)
right after the stop when its version differs from what runs. A unit without `data` (older) restores as before,
WARNed. The restore page's first sentence names the backup and the version („…visszaállt a(z) %s-i mentésből, a(z)
%s verzióra. Elérhető frissítés: …"), or, when no tested step leads on from that version, that the box will not
update it by itself (the automatic leg already refuses such an app — `LegSkipOlderThanLadder`).
- **The times tell the truth (A4).** Tier 1 = `data.at` (else the newest data file — never the manifest, never a
`pre-restore-*` undo copy); Tier 2 capped by the mirror's data time; Tier 3 capped by the data time recorded at
push (`settings.offsite_data_at`). `ReleaseConversionCopies` additionally requires a database dump written after
the conversion whose recorded engine is the NEW major (`ConversionCopy.Service` recorded from now on).
- **R-695:** the file-browser sync is single-flight (callers queued behind a running sync are covered by one sync
that reads the state after them); an EMPTY dated kept folder is never listed or bound.
- **R-691:** the read-only „Megőrzött adatok" view joins the owning GROUP of a group-readable kept folder another
user owns (`group_add`, never root's group, binds stay `:ro`, the household's files and modes untouched — decided
by CC unattended, `07` §6.5); a language switch re-syncs the file browser so the source's name follows it.
- **R-694:** a restore that GENERATED a `type: password` login (no guest app.yaml) records it (`restored_logins`); the
page no longer shows that value as "the first password set at install" and says to use the password valid at the
backup — measured per app: six of seven catalog apps keep the login in their data (code-server is the exception,
`loginAppliedEveryStart`).
- **R-699 (found live on 9202 during Part B):** a unit that is only a captured definition — a just-installed app's,
written by the status refresh before any backup — satisfied the update's precondition ("Tier 1 copy 2m0s old") and
tandoor's PostgreSQL was converted with no backup of its database. Such a unit stays listed (restorable as the
definition) but is never a precondition copy on Tier 1 or Tier 2 (`RestorePoint.DataProven`, `unitDataTime`); the
update then backs up first.
- Red-proofs (each seen failing, tree restored): `A5-redproofs/` RP1–RP6, `D2/`, `D3/`, `D4/`. Parity: one new
Hungarian fixture (`deploy_deployed_restored_login`); no existing fixture changed.
## v0.274.0 — kept data: a choice at reinstall, a list, a read-only view, a load (2026-09-25, `09` §3 decision 36); R-690, R-692
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: yes (hu + en). **A second release
in this session, allowed by the brief:** v0.273.0 (the PostgreSQL conversion) was released first because Part D needed
the floor before this part was ready.
- **Kept data — `09` §3 decision 36 (Part E).** A reinstall over an app's kept drive folder asks the household:
„A megőrzött adataimat használom" / "Use my kept data" (a load from the newest copy of THIS drive's install — own
unit or second-drive mirror — then the template's `after_load:`) or „Tiszta lappal kezdem" / "Start fresh" (the
folder is renamed into `<drive>/kept/<app>/<date>/`, with the removed app's unit; nothing is deleted). The install
API answers 409 `kept_data_choice` until one is chosen; `DeployStack` refuses too (`ErrKeptDataChoice`). New page
„Megőrzött adatok" / "Kept data" (`/kept-data`, linked from Tárhely): app, date, size, which copy can bring it back;
Load / Look / Delete (typed confirmation — the only deletion of kept data; the box deletes none by itself, D3 open).
FileBrowser gains a read-only „Megőrzött adatok" source (one `:ro` bind per item). The drive-full warning names the
kept folders and sizes. `<drive>/kept` is in `ProtectedHDDPaths` and outside every backup leg. `.felhom.yml`
`after_load:` (nextcloud: `occ files:scan --all`, catalog `9cc829c`).
- **R-690 fixed** — the removed-app restore (R-487) never found a unit on a DATA drive: `primaryUnitDirFor` and
`ListRestorePoints` asked `GetStackComposePath` (true for every catalog app), so nextcloud came back with no env,
no database and its files bound on the guest's root disk (measured on 0.272.0, `E/E1-README.md`). Now
`isStackDeployed`; pinned by a production-shaped provider (the R-487 fake answered it for deployed apps only).
- Red-proofs (each seen failing, tree restored): R-690 ×2, kept rules 1–4, the install guard, the install API,
the copy's drive check, the `:ro` bind — `E/redproofs/`. Parity: 7 Hungarian fixtures regenerated, the only
removed lines are the two edited deploy-page lines; 2 new (`kept_data_full`, `kept_data_empty`).
- **R-692 fixed before release** (found live on 9202, `0.274.0-rc1`): the list named romm's and paperless's leftovers
„Filebrowser" — the read-only view's own binds made the file browser look like their owner. An owner is now an app
that binds the folder through `${HDD_PATH}` (the folder or one inside it), never a protected stack.
`TestKept_OwnerIsNeverTheFileBrowser`, red-proofed.
- Proven live on 9202 (`0.274.0-rc2`, endpoint level, both languages): ask → start fresh → a unit → a file after it →
remove keeping data → „use" (the unit on the DATA drive, the rescan ran, the account and both files back) → start
fresh again → Load refused while a leftover occupies the folder → Delete refused for a wrong name and for an
unlisted path → Delete → Load (files moved back, database loaded, rescan) → a write into the read-only view refused
(`felhom.eu/documentation/audits/night-2026-09-26/E/E5-*`).
## v0.273.0 — the box converts a PostgreSQL major, docmost first (2026-09-25, `09` §3 decisions 35–38, §6.4 part 10); R-687
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged). New strings: yes (hu + en, three keys).
- **The conversion** (`internal/stacks/pgconvert.go`): a step whose ladder entry carries `engine_conversion
{service, engine: postgres, from, to}` gets a `converting` phase after the undo copy — the old engine alone, the
check (owners, encodings, roles, extensions, every table's row count), `pg_dumpall` validated by its completion
line, the DB volume emptied ONLY after the copy's marker is validated again, the new engine alone, the
entrypoint's empty databases dropped and existing roles' `CREATE ROLE` skipped, the load with `ON_ERROR_STOP`,
the check again + `PG_VERSION`. Any failure → the existing undo; a restart during `converting` → undone.
- **Refusals before anything moves:** a PostgreSQL major move with no mark (preflight AND job) — „Ez a lépés a(z)
%s adatbázisának fő verzióját váltaná, de nincs róla próba. Nem változott semmi."; too little room for the dump
beside the stack dir — „A(z) %s adatbázisának átalakításához %s szabad hely kell, de csak %s van. Nem változott
semmi."
- **The old datadir's copy is kept** after success (`app.yaml` `conversion_copy`) until a backup is proven after
the conversion; the hourly `conversion-copy-release` job removes it, logged by name.
- **R-687:** an empty update leg reports `"steps": []`, not `null`; a taken `files_may_change` step logs the whole
copy that allowed it.
- Proven live on 9202 with `0.273.0-rc1` (this release adds only the recovery line's wording and R-687): docmost
16 → 18 done in 42 s; load failure, unhealthy app and a SIGKILL during `converting` each undone to 16 with the
seed read back (`felhom.eu/documentation/audits/night-2026-09-26/D/`). Red-proofs: 9 (`…/B/`) + 2 (`…/F/`).
## v0.272.0 — the backup page says when a whole-box backup does not fit; three leftovers (2026-09-25, R-685, R-671, R-670, R-677)
**MinAgent: 0.131.0** (unchanged; the new sentence needs agent v0.134.0+ to have anything to say). Needs hub
v0.123.0 (unchanged). New strings: yes (hu + en).
- **R-685 (page half):** when the agent SKIPPED a whole-box backup because it cannot fit (agent v0.134.0,
`skipped: not enough space: …`), the tier row on the backup page says „A teljes rendszermentés nem fér el: %s
kell, %s szabad. Kevesebb régi mentés megtartása vagy nagyobb lemez segít." / "The full system backup does not
fit: it needs %s and %s is free. Keeping fewer old backups, or a bigger disk, fixes it." — the numbers read from
the agent's own sentence; a skip whose numbers cannot be read gets the sentence without them.
- **R-671:** a restore that lifts an update hold removes the undo copies that hold kept (never for a restore hold,
never while an update moves the app) — `backup.SetUndoCopyRemover` wired to `stacks.RemoveUndoCopies`.
- **R-670:** the undo and a step's `.felhom.yml` are read with `stacks.LoadProbeMetadata` (health check and
resources only), so no false `[ERROR] … backup block rejected … docker-compose.yml unreadable` on every undo.
- **R-677:** for a floating tag re-tested at a new digest, the badge's age counts from when that digest was tested
(`stacks.BehindSinceAge`, used by both badge producers).
- Also: the unprompted-work rule names DooPlex and ep0 only (Peti's box retired 2026-09-25).
- Red-proofs: five (`felhom.eu/documentation/audits/retire-peti-2026-09-25/B/`).
## v0.271.0 — automatic app updates: the update leg, the failed step remembered, fresh badges (2026-09-25 night, `09` §6.4 part 7, R-680, R-678, R-643)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged; the report's new `update_leg` is additive —
the hub stores the raw report). New strings: yes (hu + en).
- **The update leg** (`09` §3 decisions 11, 12, 14, 15, 20; §6.4.2 point 6 (a)–(d)): chained to the
`offbox-backup` job on every path (`chainUpdateLeg` — nil, error, panic; a box without an off-site target runs it
at W+105m). One app at a time, one tested step per app per night, through the public guarded Update only.
Skipped with a named reason: held, current (not counted), ahead, order unknown, unpinned, no test record,
older than the ladder (logged by name), not `proven`, `needs_person`, `files_may_change` without a fresh whole
copy, the step that failed before on the same ladder, W+5h reached, switched off. Transient refusals retried once.
- **Decision 20:** the full-system backup's gate defers while the leg runs, until W+5h; after it only for a step in
flight, never past W+5h30m. One shared constant (`backupwindow.UpdateLegStopOffsetMin`). The controller's own
self-update also waits for the whole leg.
- **Decision 12 — the switch:** `app_update.unattended` in `settings.json`, absent = ON; a card on the settings page
in both languages (`POST /settings/app-update`). `stacks.update_window` removed (never read).
- **R-680:** an undone or held step is recorded in `app.yaml` (`failed_update_step`, tied to the ladder's print); the
leg never presses it again until the catalog's ladder for that app changes. A person can.
- **R-678:** after `done` (and `undone`) the app's steps-left, badge inputs and pin are re-read before the update
says it finished.
- The household hears nothing new: no mail for a successful automatic step; the app page says „Automatikus
frissítés %s-kor — sikeres." / "Automatic update at %s — done.". The operator: one summary line per night in the
log and the hub report (`update_leg`).
- Decision 28: pinned that the crash-loop stop never fires during an automatic step (verify, undo). **Found:** a
DEPLOY's first start is NOT covered — `Deploying` clears when `compose up -d` returns (R-676 updated).
- Red-proofs: eleven (`felhom.eu/documentation/audits/night-2026-09-25/B/redproofs/`).
## v0.270.0 — no update for a current app; an interrupted install is reported; a restore brings back the right health check (2026-09-24, R-679, R-681, R-669, R-674)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (unchanged; R-681 rides the existing `app_deploy_failed`).
New strings: yes (hu + en).
- **R-679:** an Update pressed on an app already at the catalog head is refused BEFORE anything moves —
`409 already_current`, „Ez az alkalmazás már a legfrissebb elérhető változatot futtatja — nincs mit
frissíteni." / "This app is already running the newest version available — there is nothing to update."
"Current" is `CatalogOrder`'s verdict, so a re-tested digest for a floating tag still updates. Measured:
navidrome at the head was pressed four times, each a dump, a pull and a restart. (A test comment had called a
same-version Update "the repair path"; Restart is that path.)
- **R-681:** an install cut off by a controller restart is finished through the failure path and REPORTED: an
install marker is written before the compose-up and removed when the install ends; a marker found at start →
`compose down` (volumes kept), the stale pin records cleared, `app_deploy_failed` with the reason, and the
apps page says „A telepítés egy újraindítás miatt félbemaradt…" until the next install. An install that had
finished (its record says deployed) only loses the marker.
- **R-669:** the recovery unit captures the PINNED version's `.felhom.yml` (`applied-meta`), not the stack dir's
file the sync may already have replaced with a newer or failing step's; a restore makes the restored file the
applied record. Measured 2026-09-24: the undo judged the correct old version with the failed step's probe
and held the app.
- **R-674:** the ladder log says "AT THE HEAD" when the pin equals the newest step, not "older than the ladder".
- Red-proofs: five (REPORT).
## v0.269.1 — an installed app keeps the image it runs until an Update moves it (2026-09-24 night, Part B live finding)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 (as v0.269.0). New strings: none.
- **Found live on 9202 (Part B, `audits/night-2026-09-24/B/10-floating-tag.*`):** the catalog re-tested
`redis:7-alpine` at a new digest; v0.269.0's sync wrote that digest into the RUNNING app's compose before
anyone pressed Update. The next restart (a backup's stop/start, a power cut, a reboot) would have pulled the
new image with no backup and no undo — the guarded update bypassed.
- **Fix:** `stacks.CarryDigests` — for an INSTALLED app the sync keeps the digest the app's current file names
for each service whose reference did not change, and adds none. A fresh install still takes the ladder's
tested digest; only a guarded update (`advancePinTo`) moves an installed app's digest. The badge is
unchanged: it still reads „Frissítés elérhető" for a newer tested digest.
- Test: `TestDigest_SyncerKeepsTheRunningDigest` (an older digest kept; no digest stays none; the fix still
flows; the stored definition gets the same bytes). Red-proof: the pre-fix call site fails both cases
("the sync MOVED the running digest").
- One release per repo was the brief's rule; this is the second controller release tonight, because the first
would have shipped the bypass to the fleet with the floor. Decided by CC unattended — operator may reverse.
## v0.269.0 — the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; each step judged by its own file (2026-09-24 night, `09` §3 decisions 26–28, R-661, R-666, R-667, R-668, R-664, R-665, R-662, §6.4 part 6)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 for `app_stopped_unhealthy` (older hubs answer 400 and the
event is lost; the stop itself does not depend on it). New strings: yes (hu + en).
- **Decision 26 (R-661):** `backup.RestoreTier2Whole` — the second drive's „Teljes visszaállítás" for a file app
brings back the drive files by four rules (never delete; never overwrite a newer live file; bring back every
missing one; an older differing live file is replaced and KEPT beside as `<name>.felhom-<ts>`), then the unit
(settings + database) from the mirror. Refuses before anything moves without a proven, openable mirror with
file legs, or below the 2 GB floor. `WholeOnTier` counts the second drive as whole for such an app. R-538's
guard stays for every other caller.
- **R-668:** the same-disk check fails CLOSED — a storage path it cannot read is the SAME disk (a removed folder
became a file app's "second drive" on its own disk, measured on 9202).
- **Decision 27 (R-666):** while a held app's page says support is informed, the remove dialog offers only
"keep my data" and the API refuses data deletion (409). The no-whole-copy sentence is informal now.
- **Decision 28 + R-667:** crash loop = ≥ 6 restarts in 10 min (Docker's back-off caps a steady loop at ~1/min —
gokapi measured 7 in 7 min), counted from `RestartCount`; OOM storm = ≥ 20 kernel kills in 30 min. The box stops
the app, records an `unhealthy_stop` hold (same store as every hold), shows the sentence with a Start button,
and sends `app_stopped_unhealthy`. Start lifts it (one more try); a second stop within 24 h says support is
informed. Never inside a deploy, an update, a restore, a backup's stop or a quiesce. Default-on event, seeded
add-only.
- **R-664 / R-665:** the update judges the new version by ITS OWN `.felhom.yml` — the ladder step's
`steps/<key>.felhom.yml` or the catalog's, journaled — never the stack dir's, which a restore rewrites. The
step's memory request and applied record come from the same file.
- **R-662:** the dead „database and settings only" second step is removed (decision 25 ruled that restore out).
- **`09` §6.4 part 6, box half:** the compose the app runs pins `name:tag@sha256:…` from the ladder entry that
tested those refs (update and sync); pins and records stay digest-free; the badge reads „Frissítés elérhető"
for a floating tag only when the catalog's tested digest differs AND was tested after the install.
- Red-proofs: fourteen, each seen failing (REPORT).
## v0.268.0 — the undo finds the storage after a restore; a held app's page tells the truth; one press = one tested step (2026-09-24, R-658, R-659, R-660, R-651; `09` §6.4 part 5)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.122.0 for `app_hold_no_whole_copy` (older hubs answer 400 and
the event is lost; the household's mail is unaffected). New strings: yes (hu + en).
- **R-658 — the undo selects volumes from the app's definition, not from a label.** `DeclaredVolumeNames`
resolves the compose file's named volumes (`<project>_<key>` or `name:`), each checked to exist; the
`com.docker.compose.project` label is a logged cross-check. The unit restore now creates volumes WITH
compose's project, volume and version labels (never a guessed config-hash). The remove's report counts
unlabelled declared volumes (it said `volumes_removed: []` over volumes it did remove).
- **R-659 — operator ruling 2026-09-24, option A.** The hold names the newest WHOLE copy on any tier — for an
app with declared drive files only off-site, because both unit restores refuse it (measured from source:
R-538's guard sits in `RestoreFromRecoveryUnitAtWith`, which the second-drive unit restore calls too).
With none: `hold.update.no_whole_copy` (the operator's copy; English without „please", the house rule), no
Mentések button, and `app_hold_no_whole_copy` (critical) to the operator.
- **R-660 — a held app is not "down".** `classifyRunStates` gains a fourth suppression, the update-held set.
- **R-651 — remove deletes `applied-compose.yml` and `applied-meta/`.**
- **The ladder (`09` §3 decision 14, §6.4 part 5).** One press applies exactly one tested step with its own
definition (`steps/<StepKey>.yml` from the catalog clone; the newest step = the template). The app page
says how many steps remain. An installed version older than the ladder jumps as before, logged by name.
- Red-proofs: eleven, each seen failing (REPORT).
## v0.267.0 — tests never touch DooPlex's Docker, a cut-off copy is never loaded, two pages tell the truth (2026-09-23, R-650, R-640, R-499, R-518; R-626 measured)
**MinAgent: 0.131.0** (unchanged). No hub change. New strings: yes (hu + en).
- **R-650 — a unit test can no longer act on production Docker.** New `internal/dockerexec`: every docker
exec in the controller goes through it, and under `go test` a real docker is refused with an error
naming the command (opt-in `FELHOM_TEST_REAL_DOCKER=1`; a stub under the temp dir is allowed). The sweep
found **8 `api` tests** and the `stacks`/`web` fixtures running real `docker ps` / `docker compose
version` on DooPlex, and one `stacks` test reaching a real `docker-compose down`; those packages now run
under a silent stub (`RunWithStub`). `backup`, `appexport`, `system` tests read `docker ps`/`info` and
pass on the refusal. `TestR650_NoBareDockerExec` pins the invariant repo-wide.
- **R-640 — a cut-off database copy is refused before anything is touched.** `appbackup.CheckDumpComplete`
reads the end of the copy for the engine's completion marker. The unit restore and the off-site
restore refuse before the first mutation („…adatbázis-másolata csonka… nem indult el”); every replay
checks again right before the load, whatever the import seam is.
- **R-499 — the system-disk sentence says where THIS box's whole-system backup goes.** Four branches
(own drive / same disk / drive gone / cannot ask); „(PBS)” and „nincs külön teendő” only where true.
- **R-518 — the backup button states the measured stop** (about 8 minutes on a 12-app box). The brief's
„csak néhány másodpercre” was already gone since v0.243.0; the vague „általában néhány perc” is replaced.
- **R-626 measured, not reproduced:** remove → 390 s of `docker events` (the remove's destroy seen, no
create) + a controller restart + a guest reboot → no container, no volume.
- Red-proofs: eight, each seen failing (REPORT).
## v0.266.0 — a failed install removes what it started (2026-09-23, R-649)
**MinAgent: 0.131.0** (unchanged). No new strings. No hub change.
- **Operator ruling 2026-09-23 (R-649):** when the deploy's `compose up -d` fails for its own reasons (e.g. a
dependency whose healthcheck never passes after its container started), `runComposeDeploy` now runs
`compose down` before recording „not deployed" — so „not installed" never stands over running
containers (R-634's promise, completed). `down` WITHOUT `-v`: named volumes are kept. A failed `down` is
logged; the household's Remove clears what is left (v0.262.0 `halfStateEvidence`).
- Tests: `TestR649_AFailedInstallRemovesWhatItStarted` (stub compose, PATH = the stub only),
`TestR649_ASuccessfulInstallIsNotTakenDown`. Red-proof: the `down` removed → `calls=["up -d"]`.
## v0.265.0 — the cause of "runs but not installed", a held app that says so, a louder OOM storm (2026-09-23, R-634, R-625, R-636, R-647)
**MinAgent: 0.131.0** (unchanged). **Needs hub v0.121.0** (deployed first). New strings: yes (hu + en).
- **R-634 — the mechanism, diagnosed and fixed.** Reproduced on 9202: deploy `outline`, press the
whole-box backup 20 s later. The backup's list read the in-memory `Deployed` flag, which is true from
the moment a deploy is ACCEPTED, so the volume leg stopped the deploying app (`compose down`), dumped its
half-made volumes and ran a SECOND `compose up -d` beside the deploy's own; both failed and the deploy
recorded „not deployed" (on 2026-09-22 the backup's `up` won and the containers ran under it). Fix:
`ListDeployedStacks` leaves out a deploying app; the volume leg asks again right before the stop (SKIP,
never a failure); `StopStack` / `StartStack` refuse a deploying stack for every caller
(`ErrStackDeploying`). `sparkyfitness` did not reproduce alone (deployed in 47 s).
- **R-625 — a held app shows the truth.** The update badge of a held app reads „Megállítva —
visszaállítás szükséges" / "Stopped — restore needed" (`tag-error`, title = the hold's first sentence
in the reader's language). No Update button (the list already hid it; now pinned). Parity fixtures
`app_info_deployed` and `stacks_full` regenerated — measured diff: exactly the new badge, one line each.
- **R-636 — a repeating OOM problem gets louder.** The scan reads the kernel's `oom_kill` counter,
`memory.max` and `memory.peak` for flagged containers; 20+ kills in 30 min of the same container run →
ONE `app_oom_storm` (error, operator-only). `app_oom` unchanged. RomM's measured rate was ≈375 kills /
30 min; one hiccup is 1–3.
- **R-647 — the three leftovers.** A held update's error is the key `update.error.held`, rendered per
reader (page funcs now apply to Hungarian too; same bytes on a Hungarian box); the `app_update_held`
details carry the `copy_holds` KEY, not the Hungarian phrase; the undone event no longer logs
`hold recorded`, and the disabled-notifier line names the real severity of a health change.
- Red-proofs: eight in this repo, each seen failing (REPORT).
## v0.264.0 — the undo reaches the fleet, and the household is TOLD, in its own language (2026-09-23, `09` §6.4 parts 2–3)
**MinAgent: 0.131.0** (unchanged). **Needs hub v0.120.0** (deployed first). New strings: yes (hu + en).
- **Two new events (`09` §3 decision 15).** `app_update_undone` (warning): the update failed and the
box put the previous version and its data back — sent once, at the end of a successful undo.
`app_update_held` (error): the update, or its undo, failed and the app is held — sent once, also
when the hold could not be saved (the operator must hear exactly that). Details:
`{app, stack_name, from, to, at, copy_tier, copy_date, copy_holds}`. The held mail's line is the
hold sentence in the household's language (`RestoreHoldForLang`). Wired by
`Manager.SetUpdateEventSink` in main.go (pinned by `TestUpdateEventSinkIsWiredAtStartup`).
- **On by default, and on for existing boxes.** Both joined `DefaultEnabledEvents`; a box that already
has a list gets them once, add-only (`app_update_events_seeded`). Two new toggles on the
notifications page — without them, the next page save would push a list without the new types to
the hub and undo the hub's own migration. The parity fixture for that page was regenerated; the
measured diff is exactly the two toggles.
- **R-606 — every update sentence in the reader's language.** `UpdateError` is stored as a key + args
(`update.error.*`, `update.refusal.*`); phase labels (`update.phase.*`), the undo line and prefix,
the hold sentence (`hold.update.*`) and `UpdateCopyHolds` (`hold.copy_holds.*`) all render per
reader on both pages and in `GET /api/stacks/<name>`. The stored Hungarian is byte-identical
(parity gate green); a sentence stored by an older version renders as stored.
- **R-646 — the startup pass.** `BackfillAppliedMeta` records `applied-meta/.felhom.yml` for every
deployed, pinned app that is CURRENT with the catalog and has no record; a behind app is skipped
and named in the log (its pinned version's file is gone — guessing it is worse). Idempotent; never
overwrites.
- **R-620 — a disabled notifier says what it drops.** One WARN per event type per process
(`notifier disabled (no hub configured): DROPPED event <type> …`), then DEBUG.
- Red-proofs (REPORT): nine, each seen failing.
## v0.263.2 — the undo keeps the probe of the PINNED version, recorded when it was pinned (2026-09-23, R-637)
**MinAgent: 0.131.0** (unchanged). No new strings.
**The second thing the live proof on 9202 caught.** v0.263.1 still judged romm's serving old version
"did not start" — this time `last: health check failing` on port 8999. The undo's "old `.felhom.yml`"
was saved AT UPDATE TIME, from the stack dir — but `.felhom.yml` flows into the stack dir on every
catalog SYNC (`09` §5.4: it is never frozen), so by the time anyone presses Update the stack dir already
holds the NEW version's file. v0.263.0's unit tests wrote the new file at PULL time — the wrong moment
— and passed.
**Fix:** a new record, `<stack>/applied-meta/.felhom.yml` — the `.felhom.yml` of the PINNED version,
written when a version is pinned (deploy, pin adoption, the guarded update's pin advance) and put back
by the undo's pin-back, exactly as `applied-compose.yml` is for the definition. The update keeps THAT
file for the undo. **An app pinned before v0.263.2 has no record yet**: its undo probes with the
current file, and the job says so in the log; the record appears at its next pin (deploy or update).
The test fixture now puts the new `.felhom.yml` in place at SYNC time, before the press; red-proofs:
v0.263.1's read of the stack dir fails `TestUndo_UsesTheOldProbe`; a pin advance that records nothing
fails `TestUndo_PinningRecordsThatVersionsProbe`.
## v0.263.1 — the undo asks the OLD probe even when the new one has marked the app unhealthy (2026-09-23, R-637)
**MinAgent: 0.131.0** (unchanged). No new strings.
**Caught by the live proof on 9202, not by a test.** v0.263.0's undo put docmost's data and old
definition back, and then judged the old version "did not start" after the full 90 s — while it was
serving. The periodic health probe judges an app with its CURRENT `.felhom.yml`; the new version had
brought a probe the old one does not answer (the drill's wrong port, and a real case whenever a probe
moves with a version), so it flipped the app to `unhealthy` every 10 s — and the update's health wait
only probes an app whose state reads `running`. The old probe was never asked (`last: state unhealthy`).
**Fix:** with the undo's override probe, an `unhealthy` app IS probed, and the old check decides; it
is never settled on container state (the settle path still requires `running`). The normal update
path is unchanged. **Why the unit tests missed it:** `TestUndo_UsesTheOldProbe` injects the undo's
whole health wait, so it never reached the state gate. The new test drives the real wait loop with only
the network probe faked (new seam `Manager.probeRunFn`) and reproduces the live message verbatim when
the fix is switched off. The first live attempt ended in an honest HOLD (data put back, old version
not judged healthy) — kept as evidence.
## v0.263.0 — a failed update puts the app back by itself (2026-09-23, R-637)
**MinAgent: 0.131.0** (unchanged). New strings, all born as bundle keys (hu + en): `app_info.update_undone`,
`err.stacks.update_undo_copy_failed`, `err.stacks.update_undo_space`, `hold.update.undo_failed`,
`hold.update.undo_state.{untouched,half,not_started}`, `update.phase.{copying,undoing,undone}`.
**`09` §3 decision 15, built.** Until now a failed health check after an update stopped the app and
held it, and the household's only way back was a restore. Now the box UNDOES it: the previous version
and its data exactly as they were seconds before the update, checked with the previous version's own
probe. It holds only when that undo fails too — and the hold then says so, and what state the data is
in. This protects the manual Update button today and the automatic one later.
**How (decision 19, chosen by a bake-off — `felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/`).**
A new phase `copying` between the pull and `up`: the app stops (it stops there anyway to be
recreated), each named volume is copied `cp -a` into `<volume>.pre-update-<stamp>` by an `alpine`
helper that writes a finished-marker last, and the new version starts. On failure: `undoing` — every
copy validated before anything is put back, volumes refilled, definition and pin restored from the
job's OWN pre-update copies (R-639 — the recovery unit is never read: measured re-captured with the
failed definition 10 s after a hold was lifted, R-645), the OLD `.felhom.yml` probe, then `undone` or
the hold. Bind-mounted folders are never copied or touched. Measured cost on 9202: 1–5 s extra
downtime, ~420 MB/s, disk = the volumes (refused before anything moves near the 2 GB floor).
- **R-637** built (the eight-point list): pre-update copies kept incl. the old `.felhom.yml` (R-639);
the undo in `failAndHold` with `captureHoldLogs` first (R-621 unchanged); the copy is folder-level, so
the dump loader's two traps do not arise (R-638, R-640) and apps with no database server get their
last-second copy by construction (R-641); `app.yaml` `last_update_undone`; journal phases `copying`
and `undoing` with recovery (a cut while copying puts the old version back; a cut while undoing
resumes the undo — never "done").
- **R-642**: start/restart answer `requested — state now: <state>`, never "completed".
- `UpdateGuards.HoldAfterFailedUpdate` gains `undoState`; `settings.RestoreHold.UndoState`; the hold
sentence opens with the undo prefix + state clause in the box's language.
- A removal deletes the app's kept undo copies (label `felhom.undo-copy-of`).
- `docker_run_volume_path_gate`: the undo's four named-volume mounts allowlisted with their why.
Red-proofs (REPORT.md): nine mutations, each seen failing.
## v0.262.1 — the busy refusal is a CONFLICT, not a server error (2026-09-22, R-633)
**MinAgent: 0.131.0** (unchanged). No new strings.
**Caught by the live proof, not by a test.** v0.262.0's remove-busy guard fired exactly right on
9202 — `RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)`,
with the household's own sentence — and answered **HTTP 500**. The status mapping in
`router.go` greps the error TEXT for `not deployed` / `still running` / `not found` / `protected`,
and the busy sentence contains none of them, so it fell through to the default.
**A 500 tells the UI something broke. This is "wait a moment".** The refusal is now a typed
`*stacks.RemoveBusyError` the handler recognises with `errors.As`, answered **409**, carrying both
the Hungarian bytes and the bundle key. Its test asserts the sentence contains none of the words the
text mapping greps for — so the type is load-bearing rather than decorative.
## v0.262.0 — six defects two drill nights found in the update, remove and hold paths (2026-09-22, R-630/R-634/R-633/R-626/R-621/R-614)
**MinAgent: 0.131.0** (unchanged). **Two** new sentences, each BORN AS A KEY in both bundles and
registered in `i18n_go_keys.json`: `health.no_probe_container` and
`err.stacks.az_alkalmazason_mentes_vagy_visszaallitas_fut`. A third, `stacks.held_restore_needed`,
was written and then **removed** — it belongs to R-625, which is not in this release, and an unused
bundle key is a promise that was not kept. The go-parity gate caught both new keys as UNLISTED
before the push and was right to.
**The updates themselves were never the problem.** Two nights walked all 53 apps; what broke was the
machinery around them.
### R-630 — an app with NO health probe was STOPPED by a successful update (P1)
`waitUpdateHealthy` kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
`findProbeContainer` returned `""` its `else` set `last = "no probe container"` and **looped** — the
settle path that judges an app with no declared check sat in the outer `else`, unreachable. So
`verifying` could only ever time out, and `failAndHold` then stopped a working app. Measured on
paperless-ngx: all three containers `healthy`, `failed` at **+313.0 s**, front door 404 afterwards,
and the controller's own words — `not healthy within 5m0s (last: no probe container)`.
It now falls through to the same settle path with a WARN naming the candidates, and the journal says
`no probe container — settled on container state` so it is distinguishable from `no health check
declared` without reading the log. **A stack with no probe is not healthy and not failing — it is
settled on container state (`09` §3), and never a reason to stop a running app.**
**The target is decidable now.** `HealthCheckConfig.Container` (precedent
`InitialCredentials.Container`) and `findProbeContainerMeta` resolve by **exact stack name →
explicit `container` → a UNIQUE prefix → nothing, with the candidates returned**. The old rule took
the FIRST prefix match: for `immich` — four `immich-*` containers, no exact match — that is whichever
the container list yields, and it was seen live on `outline` probing `outline-postgres:3000` during
startup. A skipped stack now also records WHY instead of going silent.
### R-634 — an app running and serving while recorded as not deployed, and then unremovable (P1, half)
`RemoveStack` refused on `!stack.Deployed` — a **flag** — while the machine had containers, a compose
file and an `app.yaml`. Three apps were measured in that state, one of them answering its own
`/_health` with 200 on three containers, and both remove calls said `stack "x" is not deployed`. The
only exit was a shell, and a household has none. The refusal now asks whether anything **exists**.
**The mechanism that produces the bad record is still not diagnosed** and R-634 stays open for it.
### R-633 / R-626 — remove during a restore reported success and left a ghost
A remove sent while a restore was in flight tore down what existed; the restore's own `compose up`
re-created it seventeen seconds later. Both calls returned success, the record read `deployed:
false`, and a container restarted for hours with a live public route. `RemoveStack` now consults
`UpdateGuards.Busy` and `IsUpdating` and refuses with the app's own sentence — the product already
refused this clash for `update` and for `restore`, and `remove` was the one door without a lock. And
because `down` returning 0 is a request rather than a result, the project is watched for 25 s
afterwards, anything carrying its label is removed by name with its labels logged, and the answer
carries `verified`.
### R-621 — a hold destroyed the evidence of why
`failAndHold` now writes `compose logs --no-color --tail 400` into `<stackdir>/hold-logs/<ts>/`
**before** the `down`. Best-effort by design: a hold must never fail because its evidence could not
be written.
### R-614 — a removed app left its update phase behind
`RemoveStack` calls `ClearUpdateState`. The name is the only thing a new install shares with the old
one, so the record goes when the app does.
### Not in this release
**R-625** (a held app still renders an Update button that then refuses) is **not fixed here** — named
in the report rather than half-done.
## v0.261.0 — the controller no longer swaps itself out from under an app update (2026-09-21, R-608/R-609)
**MinAgent: 0.131.0** (unchanged). Hungarian pages byte-identical; the two new sentences are BORN AS
KEYS in both bundles and accounted for by the Go-parity gate.
**The collision was found by reading the clock, not by a failure.** The controller updates itself
daily at `self_update.auto_update_time` — **04:30 by default** — and again after any hub report once
a floor sits above the box, so at any hour. That swap restarts the controller container. The window
`09` §3b Q1 proposes for automatic app updates is **02:30–05:00**. It contains 04:30.
- **`stacks.Manager.AnyUpdating()`** — is a guarded update in flight for ANY app.
- **`Updater.SetAppUpdatingCheck`** — a sibling of `SetBackupRunningCheck`, deliberately the same
shape, consulted in the SAME three places: the dry run, `TriggerUpdate`, and `maybeAutoUpdate`.
One busy-gate pattern in that file, not two.
- **`Manager.SetSelfUpdatingCheck`** — the reverse direction. `UpdatePreflight` now refuses with
reason `self_updating`: „A vezérlő éppen frissül. Próbáld újra néhány perc múlva." Both halves are
wired in `main.go`, the only place holding both objects; `stacks` never imports `selfupdate`.
- **The gap was NARROWER than assumed, and measuring it is why the comment is right.** The update's
`backing-up` phase takes the backup single-flight (`RunAppBackupNow` → `acquireRunning`), so
`backupMgr.IsRunning()` ALREADY covered that one phase. It covered none of `checking`,
`safety-dump`, `pinning`, `pulling`, `starting` or `verifying` — and the last two are exactly where
the new version may already have touched the customer's data.
- **The lock must NOT latch, and that is the test that matters.** `Stack.Updating` is cleared on done,
failed AND held, so a held app does not block the controller's own updates — including the release
that might fix whatever held it. A latching gate would be a worse failure than the one prevented,
and a silent one.
**R-609 — the 409 carries a machine-readable `reason`, additively.** `UpdateRefusal.Reason` has
existed since v0.237.0 and never left the process. The body now carries `data: {"reason": "..."}`
beside the unchanged sentence, because the distinction is not decorative: `busy`, `updating`,
`deploying`, `migrating`, `self_updating` are **transient**; `held` and `downgrade` are **terminal**
until a person acts. A caller that cannot tell them apart either gives up on a passing backup window
or presses a refused button for ever.
- **Found while writing the test, not by reading: the router refuses a HELD app on its own line,
BEFORE `UpdatePreflight`** — so `held`, the reason that matters most, would have been the one
missing. That line now carries it too.
**Five red-proofs, each SEEN to fail.** `AnyUpdating` always false → the in-flight test fails;
`AnyUpdating` also true for a held app → the release-after-hold test fails; delete the preflight's
`self_updating` block → the refusal test fails with `ref = <nil>`; delete `TriggerUpdate`'s app-update
block → the swap proceeds over a live app update; drop `Data` from the refusal → `reason = ""`.
## v0.260.0 — a box AHEAD of the catalog reads „Naprakész", and the pin never moves backwards (2026-09-21, R-524)
**MinAgent: 0.131.0** (unchanged). Hungarian pages byte-identical; the two new sentences are BORN AS
KEYS in both bundles and the Go-parity gate accounts for them.
**The defect was measured, not imagined.** BIGNIGHT Phase 6, 2026-09-15: privatebin was updated
2.0.5 → 2.0.6 through the guarded Update, the catalog was then reverted to 2.0.5, and the box read
the tag „**Frissítés elérhető — ma**" with the title inviting the household to press Frissítés. The
comparison asked only "does the installed reference DIFFER from the catalog's?", so a catalog revert
— an operator act on our side — presented itself to a customer as an update, and the guarded Update
behind it would have advanced the pin 2.0.6 → 2.0.5, onto a datadir the newer version may already
have migrated, with §4's ruling saying that cannot be undone.
- **`stacks.CatalogOrder` — the comparison gains a FOURTH answer, and moves house.** Unknown /
Current / Behind / **Ahead**. It now lives in `internal/stacks/updateorder.go` rather than in
`web`, because TWO callers must reach the same verdict — the badge and the update's refusal — and
a comparison implemented twice is a comparison that drifts. `web.compareInstalledToTemplate` is a
thin wrapper; every property it carried is carried still (absent means UNKNOWN and never
„Naprakész"; it reads `CatalogImages` and never `TemplateImages`; it queries NO registry).
- **The badge.** An app ahead of the catalog reads „Naprakész" / "Up to date" with `tag-ok` — the
same word and the same class as level, because there is nothing for the household to do — and a
title that says why (`badge.update.ahead.title`, both bundles). No version number reaches the
customer, as before.
- **The refusal.** `UpdatePreflight` returns `downgrade` (409) with
„Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.", logged
with both image maps. **The API renders update refusals through `errText` now**, so a refusal
carrying a key reaches an English household in English — without that line the new key would have
been a seam built and never wired, which is a documented failure class in this repo.
- **Ahead is the NARROW arm, deliberately.** Every differing service must be orderable AND newer.
One service older, one tag unorderable, and the answer falls back to Behind — i.e. to exactly the
behaviour of v0.233.0..v0.259.0. This gate can BLOCK an update, so it errs towards letting one run.
- **Ordering is `util.Version.Compare` and nothing else** (the house rule: one comparator). The new
code is a tag NORMALISER in front of it: `X.Y` and `X.Y.Z`, optional leading `v`, a two-part tag
padded with `.0`, and a trailing suffix that must be IDENTICAL on both sides — so
`nextcloud:31.0.14-apache → 31.0.15-apache` orders, while `26.05.2-ls310 → 26.05.2-ls311`,
`postgres:16-alpine`, `kimai/kimai2:apache-2.57.0`, a date stamp and a digest pin do not.
**Measured against the real catalog**, not invented: 8 of the 66 pins float and `apache-2.57.0`
puts its version at the back.
- **Three red-proofs, each seen to fail.** (1) Make the Ahead arm return Behind →
`TestR524_PreflightRefusesDowngrade` fails with the update ALLOWED. (2) Treat an unorderable pair
as ahead → the floating-tag, different-image and digest-pin cases fail, which is the verdict that
would suppress a real „Frissítés elérhető" on the floating pins. (3) Delete the ahead arm from
`localeFuncs` → the English badge test fails.
**R-589 was NOT open, whatever the register said.** The row (P3-LOW, READY) claims the badge builds
from four raw Hungarian literals with no key. Those literals are real and DELIBERATE — they are the
parity guarantee — and the English form has been rebuilt from the bundle in `localeFuncs` since
**v0.258.0**, proven live on a fresh box the same day ("Up to date", English drill item 9). The row
is stale, not the code; it is closed in this session with that citation.
## v0.259.0 — the claim page and the backup warnings answer in the reader's language (2026-09-21, R-596/R-598)
**MinAgent: 0.131.0** (unchanged). The Hungarian pages are byte-identical; the Go-parity gate
measures every new key against the frozen base capture and was red-proofed on a single added full
stop.
The 2026-09-20 English drill ended one screen short. This release is that screen.
- **R-596 (P1) — the claim page's ANSWERS.** The page was English and its answers were Hungarian:
a mistyped code was met with „Hibás vagy lejárt kód", and an English-speaking household could not
tell a typo from a dead code on the one screen between them and their machine. **Fourteen call
sites carrying nine distinct messages** now go through `s.msg(r, "claim.msg.*")`, so they follow
the request's language exactly as the template does.
- **`data["Title"]` was DEAD and is deleted, not translated.** R-596 named it as one of the
Hungarian strings on that page. It is not: `claim.html` is a standalone page with its own
`<title>` (already bundle-backed), and `.Title` is read only by `layout.html`, which this page
never includes. Translating it would have left a permanent false signal that the page's title is
decided in the handler.
- The anonymous, cookie-less claim page resolves its language from `customer.language` in
`controller.yaml` — `langFor` → `settings.GetLanguage` → `configLanguage`. That chain was
verified at source before any edit and is now PINNED by
`TestAnonymousClaimPageFollowsTheCustomerLanguageWithNoCookie`; it was previously an unpinned
assumption.
- The lockout is unchanged and language-blind: the test asserts the COUNTER, not the text, in both
languages, so nobody gets a longer run by switching a cookie.
- **R-598 — the backup page's PROTECTION WARNINGS**, which are promises about whether the customer's
files are safe. `backupTargetDegradedText` / `OfferText` / `AbsentText` are now bundle KEYS;
`degradedMessageFor` returns the key (the decision stays language-free and in one place) and the
caller that knows the reader resolves the words. `buildTierViews`, `backupTargetLabel` and
`loadGuestBackup` take the reader's language the same way `buildDataPathCards` already did.
- The English says exactly what the Hungarian says, and a test pins the NEGATION —
"protects against corrupted files, **but not against a disk failure**". An English sentence that
promised disk-failure protection would be worse than no translation.
- **An apostrophe cost a render.** The first English absent-drive sentence read "The system
backup's drive…" and never matched on the page: `html/template` escapes `'` to `&#39;`. Caught
by the render test, not by review. All 23 new English values are now free of `' " < > &`.
- **The existing copy-contract tests were kept, not weakened.** Six assertions that compared
Hungarian WORDS would have silently become assertions about key spelling; each now resolves the key
through the real embedded bundle (`huText`), so they still convict on a reworded sentence and now
also convict if a key is ever dropped from `hu.json`.
- **This is the third and fourth instance of one defect class**, after R-573 and R-590: *a composed
sentence handed to a renderer as page data*. No template-parity fixture and no
`TestI18nEnglishPages` can see it, because the template renders the field faithfully — only the
value is wrong. Both new test files drive the real handlers and read the HTML, which is the only
instrument that can.
Red-proofs run this session: the wrong-code literal restored (the English test convicted on both
halves — the English absent AND the Hungarian present); one added full stop in `hu.json` (the
go-parity gate named both sides).
## v0.258.0 — the last four Hungarian things an English household met (2026-09-20, R-589/R-590/R-572/R-573)
**MinAgent: 0.131.0** (unchanged). The Hungarian pages are byte-identical — the parity fixtures and
the Go-parity gate both hold.
Slice 5's live proof found the residue this release clears: on an English app page the only
Felhom-authored Hungarian left was a badge and a sentence about backups. Both are built in **Go**,
which is why neither the template parity fixtures nor `TestI18nEnglishPages` could see them.
- **R-589 — the update badge.** `updateBadge` and `lifecycleBadge` gain bundle-backed forms in
`localeFuncs`, reusing `compareInstalledToTemplate` and `EffectiveLifecycle` so **the decision
cannot drift between the languages — only the words do**; the badge's `Class` is asserted equal
across languages for the same stack. The Hungarian forms in `templateFuncMap` are untouched, which
is the parity guarantee, and `TestLocaleFuncsHungarianBundleMatchesFuncMap` now pins that `hu.json`
says exactly what they say — because `i18n_go_keys.json` cites that test for these keys.
`lifecycleBadge` was not in the row; it is the same builder one function away, and leaving a known
identical defect there would have been a choice.
- **R-590 — the data-folder sentence, which is a PROMISE ABOUT THE CUSTOMER'S FILES.** It said, in
Hungarian and under an already-English folder label, that a drop-zone is temporary and unbacked. An
English household who cannot read it may leave originals in a folder the backup filter discards at
every tier. `consequenceFor` takes its lookup as a parameter, and the test asserts **the consequence
in both languages**: the excluded folder says temporary AND unbacked, a kept folder promises a
backup, an unclassified one promises nothing, and the two languages do not produce the same string
(which would mean the English fell back). The „%.1f GB szabad" free-space line went with it.
- **R-573 — the two channel banners.** The checker already CLASSIFIES the fault, so the banner is
keyed by that classification rather than by the sentence it happened to compose. The composed
sentence stays as the **fail-open fallback**: an unmapped reason renders it verbatim rather than a
blank banner or a raw key. These are the two most operationally important banners there are — the
2026-07-25 outage (R-77) ran 17.5 hours with this banner as the only signal.
- **R-572 — NOT what the row said, and the deletion is the finding.** The row said a template calling
`pruneLabel`/`nextPruneLabel` renders „vasárnap" on an English page. Measured: **no template and no
Go file called either.** They were dead func-map entries returning Hungarian. Translating dead code
would have added machinery with no reader and a test pinning a fiction, so they are **deleted**.
Deletion is the fail-loud direction, and that was proven rather than asserted: a template naming the
removed function makes `loadTemplates` panic at startup with `function "pruneLabel" not defined`.
`fmtDuration`, named in the same row, produces no Hungarian at all ("< 1s", "5s", "2m 3s") and is
left alone.
**Red-proofs, each seen to fail and then revert:** the `updateBadge` override removed from
`localeFuncs` (English read „Naprakész"); `MessageKey` dropped from the banner (English read the
Hungarian sentence); the folder sentence hard-coded back to Hungarian (the English arm convicted on
three assertions); one Hungarian byte changed in a badge key (the funcmap-equality test named both
sides); and a template naming a deleted helper (startup panic).
**The fourth red-proof says something about the gate, not just the code:** the Go-parity gate did
NOT convict that changed byte, because `localeFuncs` keys sit under `_preexisting` and are
deliberately not measured against the base capture. The citation to the test is what carries them —
so the test has to be real, and it is.
## v0.257.0 — the app catalog can speak English (2026-09-20, R-560 slice 5 Part A)
**MinAgent: 0.131.0** (unchanged).
The last Hungarian an English household meets on the dashboard is the text that comes from the app
catalog: each app's one-line description, its tagline, „mire jó" list, „első lépések", the labels and
help lines under every install setting, the prerequisites. 1 032 strings across 53 apps. This release
is the READ PATH for their English twins; the twins themselves arrive in catalog pushes.
**The format** (`10-localisation.md` §7, now measured rather than proposed). A sibling block inside
the SAME `.felhom.yml`, so a reviewer sees one app whole:
```yaml
description: "Titkosított jegyzet és szöveg megosztás"
i18n:
en:
description: "Encrypted note and text sharing"
```
- **`Metadata.For(lang)`** (`internal/stacks/metadata_i18n.go`) merges FIELD BY FIELD. An English
field that is absent — or blank — shows the Hungarian one, so a half-translated app is a legal,
shippable state. That is what makes the catalog batches safe to push one at a time.
- **`For("hu")` is the parsed struct with `I18n` cleared**, measured against all 53 real catalog
files in `internal/stacks/testdata/catalog/` (`TestMetaForHuIsIdentity`). The Hungarian product
cannot be changed by the presence of a translation.
- **Lists are replaced whole; everything else is matched by ITS OWN KEY** — `deploy_fields` by
`env_var`, options by `value`, `optional_config` groups by `match_group`, integrations by
`target`, data paths by `path`. Position matching would mistranslate silently the first time a
Hungarian field is inserted above another.
- **`For` never writes through.** Metadata lives in the stack manager's cache and is shared by
concurrent requests; an in-place merge would leak one household's language into another's page.
- **Pages reach copy only through `LocalizeStacks` / `LocalizeStackPtr` / `MetaFor`**, and
`TestNoDirectMetaCopyReadOnPages` keeps a named, reasoned allow-list of every direct `.Meta.<copy>`
read in `internal/web`, so a NEW page that reads catalog copy off the manager fails the suite
instead of quietly rendering Hungarian.
- Two copy producers sit in packages that have no request and therefore no language — the
integration rows and the initial-credentials note. Both are re-taken from the localised metadata
in the handler; identical bytes for Hungarian.
- **An older controller is unaffected**: `LoadMetadata` uses non-strict `yaml.Unmarshal` (this repo
constructs no `yaml.Decoder` at all, verified), so a pre-0.257.0 box drops the whole `i18n:` block.
That is what lets a catalog push reach the entire fleet ahead of the floor raise.
**Eight red-proofs**, each seen to fail and then revert: a mutated hu branch; deploy fields matched
by position; a blank English string winning over Hungarian; the defensive copy removed; optional
config, integrations and data paths matched by position; the untranslated half reaching the page.
**Two of them found a hollow test rather than a bug** — a struct copy shares its slices' backing
arrays, so the obvious `DeepEqual` mutation check passed a deliberately broken merge, and a one-entry
fixture cannot tell key matching from position matching. Both tests were rewritten.
## v0.256.1 — the push log says whether a household sentence was attached (2026-09-18, R-558)
**MinAgent: 0.131.0** (unchanged). One log line, no behaviour change.
Found while trying to PROVE v0.256.0 on a live box: there was nothing to look at. `Event pushed:` logs
the Hungarian message only, the hub logs the Hungarian message only, and the payload is inside TLS —
so "did the box attach a second sentence?" was unanswerable from either end. That is the fork in the
diagnosis when an English household reports a Hungarian mail, and it had no observable.
```
[INFO] Event pushed: app_deployed (info) [+household(en)] — Alkalmazás telepítve: PrivateBin
[INFO] Event pushed: app_deployed (info) [hu-only] — Alkalmazás telepítve: PrivateBin
```
The sentence itself is not logged — it is the same sentence twice and one copy is already on the
line. Pinned by `TestPushLogNamesWhetherAHouseholdSentenceWasAttached` in both languages.
**The general form: a feature whose only failure mode is "the wrong language arrived" needs an
observable that says which branch was taken. Without one, every diagnosis is a guess.**
## v0.256.0 — the box sends its own sentence in the household's language (2026-09-18, R-558 Part B)
**MinAgent: 0.131.0** (unchanged). **Needs hub v0.118.0+**, which shipped first and tolerates a box
that sends none of this — every box in the fleet is that box until this release reaches it.
Slice 3's other half. The hub now writes the household's e-mails in their language, but roughly a
third of those mails carry a sentence the BOX composed, naming a drive, an app or a number. The hub
cannot translate one. So the box sends it twice.
- **`message_customer`** on `POST /api/v1/event` — the same sentence in the household's language,
beside the unchanged Hungarian `message`. `omitempty`, and **a Hungarian household sends nothing
extra at all**: its payload stays byte-for-byte what it is today, so the hub's fallback path — the
one the whole fleet uses — keeps being the path production exercises, rather than becoming a branch
nobody takes.
- **19 producers converted** to bundle keys, rendered twice from ONE key: health ×3, controller
started/updated/update-failed, storage disconnected/reconnected, backup target absent/restored,
app deployed/deploy-started/deploy-failed/removed, crossdrive ×2, disaster recovery ×2, db-dump
completed. `message` stays Hungarian **always** — it is what the operator is mailed and what
`notification_log` records.
- **The Hungarian is byte-identical, and it is measured twice over**: `TestEventMessageWireTextIsFrozen`
(goldens captured at the slice-2 base commit) and the **Go parity gate**, which now checks all 19
keys against the base-commit literals and passes.
- **`customer.language` bootstraps a new box.** `GetLanguage` prefers the household's stored choice,
then `controller.yaml`'s `customer.language`, then Hungarian. The config value is **never written to
`settings.json`** — persisting it would record a choice the household never made, and a later
config pull could no longer be told apart from a deliberate switch.
- **The untranslatable tail** (a docker error, a validator's sentence) is appended to BOTH renderings,
so an English household gets an English sentence with a foreign tail rather than no sentence.
**Three guards had to learn the change, and one of them caught me:**
- The test seam `pushFn` now carries `messageCustomer`. A seam that cannot see a new field cannot test
it — the lesson this file already records one release back, when `PushEvent` had to become the seam.
- The R-329 severity contract walks call sites BY FUNCTION NAME. Converting producers moved two
dynamic-severity sites off `PushEvent`, and the register immediately reported them as no longer
existing — the "a stale register entry is also a failure" half doing exactly its job. The three new
push helpers are registered; the walk now checks **36** severity literals, up from 20.
- `TestConfigLanguageIsWiredInMain` reads `main.go` and asserts the setter is CALLED. `cmd/` is
gitignored, so ripgrep skips `main.go` and a reviewer's search finds nothing — which is precisely
how a seam gets built and never wired in this repo.
**A mistake worth recording:** the first pass of the conversion dropped the `displayName` argument
from three producers, which would have mailed customers `Alkalmazás telepítve: %!s(MISSING)`. It was
caught by reading the diff, and is now pinned by a test that refuses any sentence containing `%!`,
`MISSING`, `%s` or `%d` in either language.
## v0.255.0 — The globe on the sign-in-flow pages: styled, and inside the card (2026-09-18, fixes v0.254.0)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Nothing on the wire moved. One Hungarian
block changes, on the pages outside the dashboard chrome only — **zero dashboard fixtures were
re-captured**.
Two defects in v0.254.0's globe, both visible on a real browser and neither catchable by any test that
existed:
**1. The stylesheet was requested with no cache-buster, so the globe rendered UNSTYLED.** `layout.html`
has always used `/static/style.css?v={{.Version}}`; the shells did not. A browser that had the file
cached from before v0.254.0 kept serving a copy with no `.lang-globe` rules in it, and the globe came
out as a bare `<details>` — a stray triangle and two plain words, floating at the edge of the window.
**It was not three shells but FIVE**: `login`, `claim`, `recovery`, **and both guest share pages**. The
two share pages have no globe, but they have the same stale-stylesheet problem for any CSS change, and
their visitor is the likeliest of all to be holding an old copy. All five fixed.
**And `.Version` was not in the data of three of them** — which is exactly how the next shell would
have been missed. It is now set in `executeTemplateLang`, the single choke point every shell renders
through, so a new shell cannot forget it.
**2. The globe floated outside the card.** `.shell-lang` was pinned to the corner of the VIEWPORT, so
it read as part of the browser rather than part of the page. It now sits **inside the card, centred
under the footer**, with the menu opening upward — which is the shared `.lang-globe-menu` rule, so the
block adds no positioning of its own and the dashboard and the shells cannot drift apart.
**A third defect, found by an existing test while fixing the first two.** Putting `?v={{.Version}}`
on the guest share pages would have printed the controller version onto a page **a stranger holding a
capability URL can open** — `TestShareGuest_HeadersTilesNoAdminChrome` refuses exactly that, and
refused it. Those two pages now take an **opaque per-build tag** (`AssetTag`, four bytes of a hash of
the version): it busts a cache just as well and tells a visitor nothing about which build is running.
The fill lives in one function, `addShellAssetData`, called by `executeTemplateLang` **and** by the
parity harness — because a fixture rendered through a different data path is a picture of a page
nobody serves, which this release's predecessor got wrong twice.
`TestGlobeOnAnonymousShells` gained both halves — the stylesheet carries a non-empty `?v=`, and the
globe is inside the card and below the footer. Each was red-proofed by restoring the defect's exact
shape. The test server gained a version, because an empty `?v=` busts a cache exactly once and never
again, and baking that into a fixture would have hidden it.
## v0.254.0 — The saved notes follow the language, and the switch becomes a globe (2026-09-18, R-557 slice 2 release C — SLICE 2 CLOSED)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Nothing on the wire moved.
**Hungarian changes in exactly TWO places, both declared and both measured** — the sidebar footer of
every dashboard page, and the three pages a visitor meets. Everything else is byte-identical.
**1. The notes a background run SAVES now follow the language.** The line under last night's backup,
the last error, the proof result, the restore outcome — written in the BOX's language at the moment
they are written (`util.Text(lang, key, …)`, `m.note`/`s.note`). The consequence, stated rather than
hidden: **a household that switches sees the note from the run before in the old language until the
next run rewrites it.** That is the operator's §16 option 1; storing a code and rendering live would
need a dozen new persisted fields and a legacy path for each — the R-570 shape a dozen times over.
Converted: the off-site failure classes and quota notes, the tier-2 reasons and warnings, the undo
phrases, the reconstitute hold, the restore record, and the whole restore-outcome family.
**`EndRestoreOp` no longer receives a Hungarian literal from anywhere.**
**2. The language switch is a globe.** Two text links wrapped in the sidebar footer and asked the
reader to recognise „Magyar"/„English" as links; a globe is the one symbol every web user already
reads as "language", so nobody has to read Hungarian to find their way out of Hungarian. It is
`<details>`/`<summary>` — a menu with no script, which opens on click and on Enter and which a screen
reader announces. Language names inside are shown in their own language and are never translated.
**3. The sign-in, claim and recovery pages get the same globe — and a visitor's choice stays theirs.**
Those pages are met by someone who has not signed in. They have no setting to read, and they must not
write the household's: a box's sign-in page is reachable by anyone who can reach the box. So the
choice lives in the visitor's own browser (`felhom_lang`, display-only, one of two values), and
`langFor` reads it **only when there is no session** — a signed-in household can never inherit a
language a previous visitor picked in the same browser.
`POST /lang` is exempt from CSRF for a narrow, checkable reason, written down where it is made: the
only thing a forged request can achieve is to change the language of the page the victim's own browser
shows them. It writes one cookie, reads nothing, touches no setting, and `safeBackPath` keeps its
redirect on this box (a protocol-relative `//evil.example` is refused — "starts with /" alone is not
the test). **If that handler ever gains a second effect, it needs CSRF that day.**
**§16, the operator's default, taken: a claim CARRIES the visitor's language.** Someone who switched
the claim page to English and then claimed the box chose English deliberately. Only on success, and
only there — the one moment an anonymous visitor becomes the household.
**No globe on the guest share pages or the not-found page.** A share visitor is a stranger and the
household's setting is the wrong default for them; that is a promise about the share feature and an
operator decision (**R-577**). Pinned by `TestGuestSharePagesHaveNoGlobe` so it stays a decision.
**The two parity exceptions, measured rather than asserted.** 106 fixtures rendered and compared with
a real (LCS) diff — **exactly two change shapes**: the dashboard footer (89 fixtures) and the globe in
the shells (12), with 5 byte-identical, which are precisely the three pages that must not change. The
whole diff is in `felhom.eu/documentation/audits/i18n-slice2-2026-09-18/C/parity-exception-diff.txt`.
**Two defects this release found in its own measurement, recorded rather than tidied away:**
- The parity HARNESS rendered the three visitor shells through the DASHBOARD path, so the fixture
would have baked a globe posting to `/settings/language` with a CSRF field — a form the real page
never serves. Caught by reading the diff before re-capturing. The harness now branches on
`i18nDirectTemplates`.
- The first "is the change only the two blocks?" measurement compared LINE BY INDEX. An insertion
shifts every line below it, so it reported 60 520 changed lines and measured nothing. **A
line-index compare is not a diff.**
`R-575` folded in where it was one call (`tier2NoTargetReason` now returns a note in the box's
language); the soft memory-overcommit warning stays as its row says.
## v0.253.0 — Errors carry the key of the sentence they are (2026-09-18, R-557 slice 2 release B)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Every Hungarian sentence is byte-identical
to the literal it replaced, measured by `i18n_go_parity.py`, and nothing on the wire moved.
179 Hungarian sentences were built deep inside a package with `fmt.Errorf` and printed by whoever
caught them. They could not be translated where they are SHOWN (by then they are a finished string)
nor where they are MADE (that code has no request and no language). **Every one of them now carries
its key across that gap, and zero Hungarian error literals remain.**
`util.MsgError` does three things at once, and each was earned:
1. **`Error()` returns the Hungarian, byte for byte.** Every un-converted printer — a log line, a
`%v`, a third-party wrapper — keeps printing exactly what it printed before. That is what made
converting 179 producers safe without converting all their printers in the same commit.
2. **`errors.Is` still answers, for the kind AND for a wrapped cause.** `Unwrap() []error` returns
both. The predecessor `KindErrorf` returned the kind alone and dropped the cause; a converted
producer usually wraps one, and a caller testing for it would have silently stopped matching.
3. **An error ARGUMENT is rendered recursively**, so `„formázás sikertelen: %w"` translates the whole
chain. An inner error that is not ours (restic, docker, ssh, the stdlib) prints itself — that text
is not written here and is not translated here (R-553's rule).
- **Every display path goes through `errText`** (76 sites in `internal/web` and `internal/api`), and
`TestNoErrErrorInPageOutput` keeps it that way. A `strings.Contains(err.Error(), …)` is a
COMPARISON and is deliberately untouched — the four English ones are R-569.
- **The memory refusal is now an error where it is made.** `memoryVerdict` returns
`util.MsgErrorf(ErrNotEnoughMemory, …)` instead of a sentence a caller had to wrap, so the API's
409 and the household's language come from one value. `UpdateRefusal` gained a `Cause` so the same
error survives to the API through the update path.
- **Plurals, one rule, stated once:** a key that carries `.one`/`.other` forms in a language is a
plural key and its FIRST parameter is the count (`i18n.Bundle.form`). Hungarian never carries them
— it does not inflect after a numeral — so a Hungarian render is unchanged at every count. Not a
per-call-site flag: the one producer somebody forgot would read „3 app is not running" with nothing
to catch it. `TestPluralFirstArgIsNumeric` pins that every plural value really does take its count
first; `TestNoOrdinaryKeyEndsInAPluralSuffix` pins that `.one`/`.other` stay reserved — a real
collision (`alert.deadapp.one`) was caught by it on the day the rule landed.
**Two defects this release found in its own tooling, both recorded rather than quietly fixed:**
- The bulk converter **silently dropped the continuation of a multi-line concatenation**, damaging 7
producers. The parity gate did not catch it: every surviving fragment WAS a real base-commit
literal, so its question was answered yes while the call had lost text. Two behaviour tests
(`TestR356_ScenarioC`, `TestR379_ScenarioA`) caught it, because they assert the sentence a customer
reads. All 7 rebuilt as joined keys.
- The counting script was **case-sensitive**, so it reported "0 error literals left" while five
remained. The R-565 shape, in the instrument. Re-measured; the five are converted.
## v0.252.0 — The sentences the program writes follow the language (2026-09-18, R-557 slice 2 release A)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Nothing a Hungarian household reads
changes: every converted sentence is byte-identical to the literal it replaced, measured by a new
gate rather than read, and nothing on the wire moved.
Slice 1 translated the dashboard's MARKUP. The sentences the program BUILDS — flash lines, page data,
the JSON the page's own script reads, the alert banners, the country names — were still Hungarian
literals in Go, so an English household clicked an English button and was answered in Hungarian.
This release moves 226 of them into the bundle and gives them a way to reach the reader.
**The flash line was the hard part.** It travels to the page INSIDE THE REDIRECT URL, so it is
rendered by a different request from the one that wrote it — until now in the language of whoever
redirected. It now travels as a bundle key plus its parameters (`?flash=flash.share.enabled&fa=…`),
and the page renders it. A link minted by an older controller carries prose; that is shown verbatim,
never as a raw key and never dropped (`TestFlashKeyRoundTrip`, including the legacy and the
`<script>` case).
| what | how it reaches the reader now |
|---|---|
| flash lines (8 writers, 8 readers) | a key in the redirect URL, rendered by the page that reads it |
| page data, view-model labels | `s.msg(r, key, …)` / `s.msgLang(lang, …)`; the builders learned the language |
| API answers (`internal/api`) | `r.msg(req, key, …)` — the same helper, one package over |
| alert banners | `Alert.MessageKey` + args, rendered on the way out of `GetAlerts(lang)` |
| 237 country names | `country.<ISO2>`; the cloudflare TABLE is untouched and still validates codes |
| three app-named page titles + the tier-2 one (R-566) | a title key with a `%s`, plus `TitleArgs` |
**What deliberately stays Hungarian, and why.** Anything a HUB reads: the report's `health.warnings`
and `health.issues`, and every `notify` event message — the hub composes the customer's e-mail from
those and falls back to the controller's own sentence when it has no entry of its own
(`FormatCustomerEmail`). Those follow the household's language in slice 3 (R-558). Both are now
pinned by wire goldens (`internal/monitor`, `internal/notify`) so a later slice cannot translate them
by accident. Also untouched: the R-570 producer, and every `fmt.Errorf` (release B).
**The gate that makes "byte-identical" a measurement.** `controller/scripts/i18n_go_parity.py`
freezes every Go string literal at the base commit (7 467 of them, `scripts/i18n_go_base.json`) and
refuses a key whose Hungarian is not that text, byte for byte — reworded, re-punctuated, split or
joined. `scripts/i18n_go_keys.json` records what each of the 426 keys replaced. Registered in
`controller_gates.py`; three decoys in `test_gate_decoys.py`, each seen to convict. It caught a real
defect in its own first version: the capture filtered literals through an ASCII-Hungarian word list
and missed seven („Naponta", „5 percenkent", „Eletjel (Heartbeat)", …). The list is gone; every
literal is indexed.
- `Bundle.Msgf` fills a message's own printf verbs. English reorders with Go's explicit argument
indexes (`%[2]s`) — no second placeholder syntax, and hu.json stays byte-equal to the format string
the code always had, which is what the parity gate compares.
- `s.bundle()` falls back to the embedded `i18n.Shared()`, so a Server built without `loadTemplates`
renders sentences and never a raw key (red-proofed).
- The notifier's test seam moved into `PushEvent`: most producers called it directly, so a recorder
saw nothing from them and "pushed nothing" looked exactly like "pushed an empty message".
- `HU_FORMAL_CEILING` 16 → 18. Two „ön" forms the gate could not see before are now in the bundle,
word for word as the Go code said them. R-516 still owns the words.
## v0.251.0 — Nothing decides by reading a Hungarian word (2026-09-17, R-553 + R-563)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Nothing a household reads changes: every
Hungarian sentence is byte-identical, and the hub report is unchanged (pinned by a wire test).
Five places branched on their own customer-facing copy. Localisation slice 2 translates exactly those
words, and each one would then have taken the other branch, silently. Each now reads a signal set where
the message is made — `internal/util.KindErrorf` builds the same bytes `fmt.Errorf` did while carrying a
sentinel for `errors.Is`.
| # | decision | read before | reads now |
|---|---|---|---|
| 1 | deploy HTTP status (`api.deployStatusFor`) | „kötelező", „memória", "does not exist", "already deployed" | `stacks.ErrRequiredField` / `ErrPathMissing` / `ErrNotEnoughMemory` (400), `ErrAlreadyDeployed` (409) |
| 2 | off-site failure class (`ClassifyOffsiteFailure`) | „tárhelykeretet" | `backup.ErrOffsiteQuota` |
| 3 | where a disk warning is shown (`web/alerts.go`) | „meghajtó"/„adattároló" in the warning | `monitor.WarnKindStorageNotSeparate`, carried in `HealthReport.WarningKinds` |
| 4 | the stale „nothing selected" note (`offboxWarningDisplay`) | „nincs mentésre jelölt alkalmazás" in the PERSISTED text | `settings.OffboxTarget.LastWarningKind` = `backup.OffboxWarnNoAppsSelected` |
| 5 | the remote-backup poll (`backups_remote.html`) | the displayed word „Fut" | `data-status="running"` on the status element (R-563) |
- **The restic / ssh signatures in the classifier stay text matches on purpose** — that output is not
ours and is never translated. Listed as external in REPORT.md.
- **Site 4 legacy:** a box upgraded to 0.251.0 still carries the old persisted text with no kind until
its next off-site run. The text test survives ONLY for `kind == ""`, with a removal row (R-570);
**slice 2 must not translate that producer before R-570 closes.**
- **Site 5 is the one rendered byte that changed:** the `data-status` attribute (plus the poll line).
12 `backups_remote` fixtures re-captured; each differs from its predecessor by exactly those two
things, the other 94 are byte-identical. „Fut…" is now `{{T "backups_remote.fut"}}` (en „Running…"),
so the English page finally sees a running backup.
- **The wire is untouched:** `internal/report/builder.go` copies Status/Issues/Warnings only;
`TestR553_HubReportWarningsAreUnchangedOnTheWire` pins the field set and the bytes.
- New in REUSE.md §1: „Error kinds — never branch on a customer-facing sentence".
**Red-proofs** (`felhom.eu/documentation/audits/r553-2026-09-17/redproofs.txt`): every site's pre-fix
predicate restored and watched failing — deploy statuses fall to 500/400-for-anything, the translated
quota error reads „ismeretlen okból", the translated disk warning jumps to the top banner, the stale
note returns, and both languages lose the poll. The detector (`felhom.eu/scripts/i18n_inventory.py`)
was re-run: template compares 1 → 0, and planted decoys prove it still convicts.
## v0.250.0 — Every dashboard page in English; the language switch shown to every household (2026-09-17, R-556 release C)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** **A Hungarian household sees one new thing:
a „Magyar / English" switch in the sidebar footer.** Everything else in Hungarian is byte-identical.
- **Converted (608 markers):** `storage` (162), `storage_network` (70), `storage_init` (42),
`storage_attach` (27), `sharing` (74), `login` (9), `claim` (21), `launcher_shared` (4),
`launcher_share_password` (6), `catchall` (4), `debug` (189). All 31 slice-1 templates are done; the
English gap is 0.
- **Parity:** 24 fixture states captured from the UNCONVERTED templates (`0555091`). Those cases had
invented page titles; they now carry the handlers' real titles and the 12 affected fixtures were
re-captured from `0555091` (diff: the `<title>` line, and the nav highlight on the two drive wizards).
- **Sign-in, claim, guest share (both), catch-all follow the language** through `executeTemplateLang`
(no session CSRF, no escrow reminder). `renderLogin` takes the request. Tests:
`TestDirectRenderHandlersFollowLanguage` (the real routes, en and hu), `TestI18nDirectRenderPagesHaveNoAdminChrome`
(escrow reminder genuinely due; absent from the page data of all six pages and from the guest page in both languages).
- **The switch for everyone:** the hide condition in `addLanguageData` is deleted. 89 layout fixtures
re-captured; each differs from its predecessor by exactly one inserted switch form placed after the
version span (89 of 89 byte-for-byte after removing the form); the 17 standalone fixtures are unchanged.
`TestLanguageSwitch_EndToEnd` now asserts a Hungarian household with no `?lang=` sees the form.
- **Titles:** `TitleKey` for storage, network storage, both drive wizards, sharing.
`TestHandlerTitleKeysMatchHungarianTitle` pins every handler's Hungarian title literal to its hu.json value.
- By-eye review found six ASCII-only Hungarian JS fragments the extractor missed („, majd a(z)",
„FIGYELEM:", „jelenlegi:", „mp", „p", „ db"); converted by hand. No test would have caught them on the
English page (row filed).
- `HU_FORMAL_CEILING` 12 → 16 (R-516) — and the stem list sees only 4 of at least 22 formal forms on these pages (R-516 extended).
- `r400_debug_routes_test.go` and `debug_route_gate.py` still read `debug.html` raw: correct, they match
addresses and identifiers, and the bundles hold no `/api/` string (measured 0).
**Red-proofs** (`felhom.eu/documentation/audits/i18n-slice1-2026-09-17/C/redproofs.txt`): parity (one letter of
`sharing.megosztott_mappak` → 6 sharing cases), English page (`storage.eleresi_ut` removed → 2 cases),
marker coverage (a marker in an unrendered branch), title keys (reworded hu value; deleted handler line),
direct routes (all five back on `s.tmpl` → five English rows fail), no admin chrome (`addEscrowBanner`
added to `executeTemplateLang` → fails per page), the switch (hide condition restored → end-to-end test
and 89 parity cases fail).
## v0.249.0 — English for the backup pages, Hungarian byte-identical (2026-09-17, R-556 release B)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Nothing changes for a Hungarian household.
- **Converted:** `backups_apps`, `backups_remote`, `backups_restore`, `backups_restore_wizard`,
`backups_escrow`, `tier2_config`, `recovery`, and the restore-progress JS in `backups_shared`
(`restore_banner_js`, which v0.247.0 had left Hungarian).
- **Parity:** 41 fixture states captured from the UNCONVERTED templates (commit `f8ebc47`), plus
`backups_apps_stale_copy` captured from `f8ebc47` for a condition inside a confirm text.
- **`recovery` follows the language** through the new `executeTemplateLang` — the language-aware render
for pages outside the dashboard chrome, WITHOUT `executeTemplate`'s CSRF injection and escrow bar.
Tests: `TestI18nDirectRenderPagesFollowLanguage` (hu = fixture bytes, en = English) and
`TestRecoveryHandlerFollowsLanguage` through `renderRecovery` (red-proofed by restoring `s.tmpl`).
- **Deliberately not converted:** „Fut…" and the JS `indexOf('Fut')` on `backups_remote` — the page
decides its progress poll by reading that word (R-563). Code values (`snapshot_id=helyi`) stay.
- **Retrieval-promise gate, English:** 13 English claims registered with reasons; six mirror Hungarian
registrations. **Seven exposed a Hungarian blind spot**: split verbs („állíthatók vissza", „hozod
vissza") escape the Hungarian stems, so those sentences were never scanned in Hungarian (row filed).
- Extractor review found ASCII-only Hungarian it missed: Konfig, Megtartva, helyi, pl. (×3), Befejezve,
automatikus, (jelenleg:), kedd/szerda/szombat, „1. szint"/„2. szint", and a confirm text split by
`{{if .Tier2StaleCopy}}`; converted by hand, words added to its list.
- Title keys for the backup sub-pages and the escrow page. `HU_FORMAL_CEILING` 10 → 12 (R-516).
**Red-proofs** (`felhom.eu/documentation/audits/i18n-slice1-2026-09-17/B/redproofs.txt`): parity (one byte of
`backups_escrow.kod_letrehozasa` → both escrow cases name the line); recovery routing (render back on
`s.tmpl` → English lost, test fails).
## v0.248.0 — English for apps and settings: ten more pages, Hungarian byte-identical (2026-09-17, R-556 release A)
**MinAgent: 0.131.0** (unchanged). **No hub release needed.** Nothing changes for a Hungarian household:
the switch is still shown only on English pages (decision 6 stands until release C).
- **Converted:** `dashboard`, `stacks`, `logs`, `monitoring`, `app_import`, `app_export`, `deploy`,
`settings_system`, `settings_security`, `settings_notifications` (572 markers). `app_row`,
`meta_badge`, `icons` carry no copy (verified). 16 repeated strings now share `common.*` keys
where the meaning is the same everywhere; „Mentés" (Save/Backup), „Frissítés" (Update/Refresh),
„Indítás" (Start/Boot) deliberately stay per page.
- **Parity:** 25 new fixture states captured from the UNCONVERTED templates (commit `6be55a0`) before
any of these pages carried a marker. Five v0.247.0 fixtures were re-captured from v0.246.0
(`89dd3e94`) because their one-word Hungarian data had to become multi-word (below).
- **New test — `TestI18nParityCoversEveryMarker`:** every marker occurrence must be rendered by some
parity case. It found two gaps in v0.247.0's own pages (the app page's Install button; one „Due"
tag), closed with two new cases captured from v0.246.0.
- **English page test widened:** only fixture strings of 2+ words or 12+ characters are masked; titles
go through `addLanguageData`. Measured: with one English nav key deleted, the old test flagged 8
cases and missed the 6 launcher cases; the new one flags all 14.
- **Retrieval-promise gate reads English** (`EN_PATTERNS`, `ALLOWLIST_EN`): HEAD passed a planted
„Your old backups can be restored at any time." in `en.json`; the gate now convicts it and accepts the
same place carrying a sentence that promises nothing. Decoys in `test_gate_decoys.py`.
- **Go-side copy on these pages:** titles via `TitleKey` (dashboard, stacks, monitoring, the three
settings pages); `infraMeta` descriptions (and samba's name) from the bundle, Hungarian pinned equal
to `inframeta.go`.
- **Found, not converted:** `backups_remote.html` decides with `indexOf('Fut')` against its own
Hungarian status text (row filed); converting that word would silently break the English page's
progress poll. Composite page titles (logs, deploy, export) and `lifecycleBadge`/`updateBadge`
remain Hungarian — slice 2.
- **Extractor:** ASCII-only Hungarian it missed on these pages — Mind, Folyamat, Titkos, Processzor,
Hamarosan, aldomain, „Az adatai itt voltak", Figyelem, pelda@email.hu — converted by hand; words added
to its list (not „Mind": an English homograph). It also learnt that a JS string can start inside an
attribute value.
- `HU_FORMAL_CEILING` 6 → 10: the converted copy carries four more „ön" forms (R-516), unchanged by rule.
- Old wording tests (`TestDiskMeter_*`, `TestNetworkBadgeTemplates_DistinctStrings`,
`TestSettingsSectionInventory`) read the Hungarian expansion (`huTemplateSource`), not the raw file.
- `loadTemplates` logs its parse duration per language (DEBUG).
**Red-proofs** (`felhom.eu/documentation/audits/i18n-slice1-2026-09-17/A/redproofs.txt`): parity (one byte
of `dashboard.lemezek_allapota` → 4 dashboard cases name line and bytes); coverage (a marker in an
unreachable launcher branch → named, while parity stayed green); mask widening (8 vs 14 cases);
retrieval gate (HEAD rc 0 on the planted promise → fixed gate convicts, negative control accepted).
## v0.247.0 — the dashboard can speak English: mechanism proven on three pages, Hungarian byte-identical (2026-09-17, i18n spike)
**MinAgent: 0.131.0** (unchanged — nothing here talks to the agent). **No hub release needed:** the
report gains an additive `language` field the hub stores raw and does not read yet.
**Nothing changes for a household.** Hungarian stays the default and the only language shown; the
switch appears only on an English page or a request carrying `?lang=`.
- **Mechanism — a flat message bundle, expanded BEFORE the template is parsed.**
`internal/i18n/locales/{hu,en}.json` (embedded; hu authoritative). Converted templates carry
`{{T "key"}}` markers; `loadTemplates` parses one template set per language after `i18n.Expand` has
replaced them, so the Hungarian set is parsed from the same bytes, in the same escaping contexts, as
before. Chosen over `golang.org/x/text/message` because the inventory found no gender and only
count-plurals (English `key.one`/`key.other`; Hungarian has one form), and over a runtime `T` func
because the contextual escaper would change Hungarian bytes (`+`, `'`, `"`) and need per-context
handling inside `<script>`. A key missing in English shows Hungarian and is counted; an undefined key
fails the template load.
- **Three pages + the layout converted:** the launcher, `/backups`, `/apps/<slug>`, and `layout.html`
(menu, banners, the layout's delete/remove modals) — 252 markers, 235 Hungarian keys, all translated.
Go-side copy on those pages: page titles (`TitleKey`), and English overrides for `stateLabel`,
`timeAgo`, `timeAgoStr`, `nextRunLabel`, `statusText`. `<html lang>` follows the language.
- **The setting:** `settings.json` `language` (`hu`|`en`, empty = hu); `POST /settings/language`;
`?lang=` per-request override (never persisted); reported to the hub as `"language"` at all four
report build sites.
- **Gates.** New `i18n_missing_gate.py` (keys exist, no orphans, English gap and formal „ön" count are
ratchets, no "please"/"kindly"). **Found while converting:** the retrieval-promise gate went STALE the
moment copy moved to the bundle, and the emoji / native-confirm / secret-in-markup gates went BLIND
to it — they now read templates expanded (`scripts/i18n_bundle.py`); mojibake scans the bundles.
Decoys for all of it in `test_gate_decoys.py`; the HEAD versions of the emoji and retrieval gates
were run against the bundle decoys and passed them.
- **Tools:** `scripts/i18n_extract.py` converts a template. Its first pass missed eight ASCII-only
Hungarian strings; they were found by eye and the word list extended.
**Tests, each red-proofed (seen failing, then passing):** `TestI18nParity` — 14 fixture states
captured from a clean worktree of `origin/main` 89dd3e94b1de; one changed byte in hu.json fails it
naming the page and line. `TestI18nEnglishPages` — no Hungarian template copy, raw key, marker or new
empty element on the English render; removing one English key fails it. `TestLanguageSwitch_EndToEnd`
(real router: hidden for Hungarian, `?lang=` door, save, 400 on `de`, back to hidden) — fails with
`addLanguageData` removed. `TestI18nJSContextValuesAreSafe` — fails on a quote in a JS-context English
value. `TestBundleParametersMatchAcrossLanguages` — fails when English drops `{{.RecoveryAbandonDate}}`.
`TestLocaleFuncsHungarianBundleMatchesFuncMap` — fails when hu.json's state word differs from the funcmap.
## v0.246.0 — an interrupted restore is told, and the recovery-code reminder waits until the box can take it (2026-09-17, R-550 / R-546)
**MinAgent: 0.131.0** (unchanged — the readiness check reads the escrow preflight, served since agent
0.88.0; nothing here needs agent 0.132.0). **Requires hub v0.117.0** for `restore_interrupted` to be
accepted; an older hub 400s that one event and nothing else changes.
- **R-550 (operator ruling "fix", 2026-09-17) — the restore record survives a restart.** A DESIGN
REVERSED: `opstatus.go` was in memory by choice ("same precedent as notification cooldowns"). Chaos
night round 10 measured the cost — a restore accepted, the box hard-reset four seconds later, and
`/api/backup/restore-status` answering the Go zero value; nobody could learn whether it finished.
Now `restore-status.json` in `DataDir`, written atomically at both ends of an op
(`internal/backup/restore_record.go`). At startup a record still marked running becomes a failed,
`Interrupted` result — „A visszaállítás megszakadt (a doboz újraindult) — indítsd el újra." — kept
per app until that app's next restore, shown as a „Megszakadt visszaállítás" card on
`/backups/restore` and in the off-site wizard's outcome card, and raised ONCE as
`restore_interrupted` (warning, for the household). Cooldowns stay in memory.
- **R-546 — the recovery-code reminder waits until the box can run the ceremony.** The R-543 bar now
consults the agent's OWN preflight `ok` (every blocking item — not a copy of `pbs_storage_id`), cached
60 s and probed only while paused; held back while not ready. `/backup/escrow` shows a waiting card
that polls and reloads itself instead of red crosses and the agent's English diagnostics.
`POST /api/escrow/start` refuses 409 with the same sentence BEFORE staging or starting — the direct
path chaos night used, which ran the ceremony and returned the raw `-storage` stderr. Unknown
readiness keeps the bar (fail loud). The transition is logged at INFO.
**Red-proofs, each seen failing then passing:** the restore record across a restart (blank status,
`StartedAt:0001-01-01` — the chaos-night symptom exactly); a finished restore stays finished; the next
restore of that app clears its notice; `main()` calls both startup functions (AST); the startup helper
with loading skipped pushes nothing; the restore page card with the handler line removed; the bar held
back (readiness check removed → bar shown); the waiting card (never set → checklist); the start refusal
(gate removed → 200 and `stage,start` ran). Controls: a ready box shows the bar; unknown readiness shows
the bar; a box with no interruption shows no card. The three existing start-order tests now expect
`preflight` first — their load-bearing assertion (stage BEFORE start) is unchanged.
## v0.245.0 — the household is asked for the recovery code, and the page says „szünetel" until then (2026-09-16, R-543)
**MinAgent: 0.131.0** (unchanged — nothing here needs a newer agent)
- **R-543 — off-site backup is ON by default and does NOT RUN until the household creates its
recovery code, and nothing anywhere asked them to.** The pause is the DESIGN, not a defect: the
escrow is zero-knowledge (07-backup-architecture §6.3), the household's code is the only key, and a
run started without one would produce a copy nobody could ever open. Nothing in this release
touches that mechanism. What was missing is the ASKING. Measured on a fresh box 2026-09-16: the
off-site tier sat at „Kulcsletétre vár" from the first minute, no snapshot was ever written, and
the app-backup page told the household their files were protected by that very copy.
- **A reminder bar on every dashboard page** while the off-site tier is configured and its escrow
is not complete: „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." with
the route to `/backup/escrow`. It is the **R-241 bar, second instance** — same session-cookie
dismissal („Most nem"), back at the next visit, gone for good when the state is `escrowed`. **No
second banner system was built.** It hangs off `executeTemplate`, the single render choke point,
so it reaches every authenticated page rather than the three pages a per-handler helper would
have covered; the login and claim pages render through `ExecuteTemplate` directly and never pass
through it, and a session check keeps it off the public guest share page.
- **The tier-1 sentence now renders by STATE, not by the app's shape.** v0.244.0 fixed the label
(R-537) and put a new promise in its place: „a távoli másolat … védi", printed whenever an app
keeps files on the drive. On a paused box that named a protection which had never run once. It
now reads „…**védené** — a távoli mentés a helyreállítási kód létrehozásáig szünetel" with the
route out; with no off-site and no second drive it says there is no copy and names both ways to
get one. It takes `tier3State`'s own vocabulary rather than re-deriving the state, so the row and
the sentence cannot disagree.
- Red-proofs, each seen failing: delete the one `addEscrowBanner` line from `executeTemplate` →
`TestR543_A_PausedBoxAsksOnEveryPage` fails on BOTH pages („does not tell the household the
off-site copy is PAUSED"); compute the sentence from the app's shape alone as v0.244.0 did →
`TestR543_Tier1Sentence_PausedStateDoesNotPromise` fails quoting the exact promise a fresh box
read all day. Negative controls: an escrowed box is not nagged, an unconfigured box is not
nagged, and an app with no drive-side files gets no sentence in any state.
## v0.244.0 — the backup page stops promising what it does not hold, and a restore refuses to lie (2026-09-16, R-537 / R-538 / R-536)
**MinAgent: 0.131.0** (unchanged — nothing here needs a newer agent)
- **R-537 — the app-backup page labelled a Tier-1 unit „DB + Konfig + Adatok" and printed the app's
data-drive size beside it, over a unit that holds no copy of those files.** The label was computed
from the APP's shape (`HasHDDData || HasVolumeData`) and rendered on all three tier rows, so one
string stood for three tiers that capture different things. It is now computed per tier from what
the tier actually carries: Tier 1 says „Adatok" only when the app's data really is inside the
volumes the unit captured, and a class-A app gets one sentence saying where its files ARE
protected. **The design is unchanged** (07-backup-architecture §6.1/§6.2: a Tier-1 unit has no
file-copy step; the file legs of calibre-web, immich, nextcloud and paperless-ngx are carried by
Tier 2 and Tier 3) — what changed is that the page now says so. Measured on a fresh box
2026-09-16: five photos in Nextcloud, in no backup, with the page saying they were.
Red-proof: restore the old app-shaped label → `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`
fails at „the Tier-1 label claims it holds the app's data".
- **R-538 — a unit restore replayed a database over files it did not have, reported success, and
destroyed the app's own wastebasket on the way.** `RestoreFromRecoveryUnitAt` now REFUSES before
anything is touched when the app's files live on the data drive, names the route that can return
them (the off-site wizard's „Teljes visszaállítás (fájlok + adatbázis)", or the second drive's
„Fájlok visszaállítása"), and says plainly when there is no copy at all. The explicit
database-and-settings-only path is a separately-worded second step
(`UnitRestoreOptions{AcceptMissingFiles}`), never a sibling control. The refusal runs before the
stack is stopped, because the measured harm included the trash going unreachable.
Red-proof: disable the guard → `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles` fails at
„a restore that cannot return the files must refuse".
- **R-536 — the hub was told „Alkalmazás telepítve" when the deploy was merely ACCEPTED.** The API
now emits `app_deploy_started` beside its 202, and `app_deployed` is emitted from the async path's
own end (`stacks.SetDeployDoneHook`), with `app_deploy_failed` (warning) when it ends badly —
which used to be silence. Measured 2026-09-16: mealie was recorded as installed in the same second
its deploy was accepted, then killed 5 s in, and ended `not_deployed` with nothing correcting the
event. The accept-time `app.yaml` is deliberately NOT deleted on failure: it is the crash-safe
record written with `Deployed:false` and it carries the settings the customer typed.
Red-proofs: put the old call back → `TestDeployAcceptance_DoesNotClaimTheAppIsInstalled` fails;
remove the success-path hook → `TestDeployDoneHook_FiresAtTheEndAndSaysWhichEndItWas` fails at
„the deploy ended and nothing was told about it".
- Requires hub **v0.116.0**, which registers `app_deploy_started` / `app_deploy_failed` in both
`allowedEventTypes` and `customerMessages`; against an older hub those two POSTs 400 and the
events are simply absent (`app_deployed` keeps working).
## Unreleased — after v0.243.0 (2026-09-15)
- **R-517 follow-up — the whole-system tile printed „0 B" for a backup whose size is unknown.**
Measured live on demo-hp guest 9201 right after the agent's self-update to 0.131.0: the local tier's
success was read back from storage (the agent's in-memory record was empty), and the top card showed
„✓ … 0 B · Helyi tároló (local)". It now shows „–" (`guestBackupView.SizeUnknown`). The per-tier rows
were already right. Red-proof: without the flag, `TestBuildTierViews_SuccessFromStorageAfterRestart`
fails at "tile would print a size it does not know". Ships with the next release; v0.243.0 is live.
## v0.243.0 — the file manager gets a real password, the backup page tells the truth per tier, an absent tier stops no app, an OOM-killed worker is seen (2026-09-15, R-513 / R-517 / R-518 / R-514)
**MinAgent: 0.131.0** (the per-tier backup status and tier storage presence are agent v0.131.0; the controller supervisor rides the same release)
- **R-513 (P1, SECURITY) — FileBrowser no longer accepts `admin` / `admin`.** Measured first
(2026-09-15, `gtstef/filebrowser:1.3.3-stable`, `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/B1`):
the config key `auth.adminPassword` / env `FILEBROWSER_ADMIN_PASSWORD` sets the password on a fresh
and an existing database — but RE-APPLIES IT ON EVERY START, overwriting a password set by hand; the
API (`PUT /api/users?id=<id>` with `X-Password: <current>`) changes it once and it sticks. So ONE
mechanism for fresh and existing boxes, the API: every base-stack tick until decided,
`Manager.EnsureFileBrowserAdminPassword` logs in as admin/admin against `http://filebrowser:80`;
**200** → generates `password:16`, sets it, verifies new=200 AND admin=401, stores it AES-encrypted in
settings (`filebrowser_admin_state: generated`); **401** → records `operator` and never touches it
(the HP and the N100 were changed by hand on 2026-09-15). Unreachable → nothing recorded, retried.
The FileBrowser app page shows „Kezdeti belépési adatok": user `admin` and the password behind
„Megjelenítés" (the R-254 reveal: POST, no-store, logged, never in the page), or „az üzemeltető
állította be". Cost, stated: on a brand-new box admin/admin works from FileBrowser's first start until
the next tick. Red-proofs: without the PUT, admin/admin still logs in; without the operator branch, the
page does not say who set it.
- **R-517 (P1) — „Rendszermentés" speaks per tier.** The tile read the agent's single LATEST record, so
a failed 0-byte PBS attempt on absent storage became „✗ · 0 B · PBS · Naprakész" and ticked
„Távoli rendszermentés — külön hardveren". With agent ≥ 0.131.0 the page shows, per tier, the newest
SUCCESSFUL backup (date, size — unknown when read back from storage after a restart), a failed attempt
✗ „sikertelen" UNDER it, and a tier whose storage does not exist as „nincs beállítva". „Naprakész" is
computed from successes (a set-up tier whose newest success is older than 1.5 × its cadence is
„Esedékes"); the remote tick needs a current PBS SUCCESS. An older agent renders exactly as before.
Red-proof: reading a tier's success from its last attempt loses the local backup's size.
- **R-518 (cheap half) — a tier whose storage does not exist is not attempted, and no app is stopped for
it.** `quiesce.skipAbsentTiers` drops tiers the agent reports `storage: absent` from the manual and the
scheduled run (unknown/legacy never skipped), logs every skip and pushes `backup_tier_skipped`
(warning, operator-only, once per absence). Button copy now tells the truth: „A mentés alatt az
alkalmazások leállnak — általában néhány perc, nagyobb adatnál több." Per-tier quiesce (the other half)
stays open. Red-proof: without the skip, the manual run starts felhom-pbs and the scheduled run stops an
app for it.
- **R-514 — an OOM-killed worker inside a running container is visible.** BIGNIGHT: Paperless's worker
was killed by the memory limit, the container ran on, the app read „Fut". The 30 s dead-app check reads
`State.OOMKilled` for running app containers in one `docker inspect`; the dashboard row shows „Memória
elfogyott" (with a what-to-do title) and the hub gets `app_oom` (warning, operator-only), once per
container run. **The task named the tag „memória elfogyott — újraindítva"; the controller does NOT
restart the app** (an automatic restart is an unmeasured mechanism that could cut a household's
upload), so the tag does not claim it. Red-proof: skipping the `true` lines hides the killed worker.
- Hub ≥ v0.114.0 registers `backup_tier_skipped` and `app_oom`; deploy the hub first.
## v0.242.0 — a removed app is listed with its kept backup, and five small ones (2026-09-13/14, R-487 / R-491 / R-490 / R-489 / R-476 / R-456)
**MinAgent: 0.129.0** (unchanged)
- **R-487 (P2) — a removed app whose backups were kept is listed, with the restore that brings it
back.** After „Töröld az adataimat is" with the backups kept, the unit sat on the drive and was
restorable through `POST /backup/restore` — and listed on NEITHER backup page (measured on demo-hp
with adventurelog). The local lists are now keyed on the DRIVES the way R-237 keyed the off-site
list on the store: `Manager.ListRemovedAppUnits` walks `backups/primary/` on the system path and
every connected registered drive and lists the units whose app is not deployed. The Mentések page
renders them after the deployed rows („Eltávolítva — visszaállítható", one action: *Visszaállítás
a mentésből*), the Visszaállítás picker lists them in their own group, `GET /api/backup/snapshots`
answers for them instead of 404, and the restore opens the unit WHERE IT SITS
(`primaryUnitDirFor`) — a unit kept on a data drive used to be unreachable, because the drive of a
removed app is unknown and the fallback named the system path.
- **R-476 — a Tier-2 copy is dated by its data, not by its manifest.** The manifest moves only when
the app's definition changes, so a copy holding a dump from 00:30 today was dated the day before.
`Tier2Coverage.UnitDataDate` (newest dump in the mirrored unit) is what a refreshed leg names; a
PRESERVED package keeps the manifest date (R-403).
- **R-491 (P2) — a removal clears the app's update hold.** An app held after a failed update and then
removed kept its hold in the store, so a reinstall under the same name would begin refused with a
sentence about a copy that no longer existed (measured on demo-hp with nextcloud). `removeStack` now
clears an UPDATE hold (`Settings.ClearUpdateHold`); an R-379 restore hold stays operator-cleared.
- **R-490 — the monitoring page's memory-distribution card can render.** `/api/system/info` was eaten
by the web layer's `/api/system/` prefix (404 „ismeretlen végpont"); an exact-path mount on the API
router wins in ServeMux. `systemInfo` also gained the default-storage-path fallback every other
reader of the empty `cfg.Paths.HDDPath` already had (R-465).
- **R-489 — `volumes_removed` reports the volumes a removal removed**, `[]` when none — never `null`:
the project's volumes are listed before and after `down --volumes` and the difference is reported
(compose's progress lines are printed to a TTY it does not have here). **Measured limit, live on
9202 the same night (row kept open):** the listing filters on the compose project LABEL, and a
volume recreated by a unit restore (`docker volume create <name>`, `restore.go`) carries no labels
— compose still removes it, and the response says `[]`. A fresh compose-created volume is
reported. Next release lists by the `<project>_` name prefix as well.
- **R-456 — the boot-orphan rule is pinned:** an absent member does not make a stack degraded, so a
partly-dead stack is not a boot orphan; a present-but-dead member is (`internal/bootrecon/r456_partly_dead_test.go`).
**Tests / red-proofs:** `TestR487_` (lister over two drives, picker, unit dir, row builder, two
render tests with negative controls), `TestR476_` (date rule + coverage reader), `TestR491_` (wiring),
`TestR490_` (mount + behaviour), `TestR489_` (difference, never null, the docker listing), `TestR456_`;
red-proofs in `felhom.eu` `audits/v0242-2026-09-14/`.
## v0.241.0 — a bind-data app leans on off-site before its own unit, and the hold says what the copy holds (2026-09-13, R-479)
**MinAgent: 0.129.0** (unchanged)
**Operator ruling 2026-09-13 (R-479).** For an app whose data lives in bind-mounted files, the
recovery unit holds the definition and the database dumps but not the files, so a route back that
names it restores settings and not data (measured on demo-hp with gokapi: „a beállítások
visszaálltak … adatot nem"). Two changes:
- **The tier order depends on the data layout.** An app with classified binds (`DataOutsideUnit`)
walks second drive → **off-site** → own unit (`updateTierOrderBindData`); an app whose data is in
named volumes keeps second drive → own unit → off-site. `UpdateTierOrderFor` is the one place that
decides; `UpdateRestorePoints` reads it.
- **The hold sentence names what the copy holds.** `RestoreHold.CopyHolds`, computed at hold time by
`UpdateCopyHolds` from the app's layout and the chosen tier, e.g. *„… ebből a biztonsági mentésből:
saját meghajtó, 2026-09-13 17:34 — ez a másolat csak a beállításokat és az adatbázist tartalmazza,
a fájlokat nem."* A hold written by v0.239.0–v0.240.0 (tier, no phrase) keeps its sentence
(`UpdateHoldTierFmt`).
**TESTS.** `internal/backup/r479_tier_order_test.go` — the order per layout, the CONSEQUENCE (Tier 2
absent, unit and off-site both fresh: a bind app leans on off-site, a volume app on its unit), the
phrases, the sentence verbatim, the tier-only sentence for older holds; the adapter wiring test
demands the phrase travel through `HoldAfterFailedUpdateHolding`. **Red-proof:** making the order
layout-blind fails the consequence case (felhom.eu `audits/v0241-2026-09-13/`).
## v0.240.0 — seven defects the any-tier proof and the first nightly rotation found (2026-09-13, R-486 / R-484 / R-485 / R-480 / R-477 / R-478 / R-474)
**MinAgent: 0.129.0** (unchanged)
Four were found live proving v0.239.0 in the afternoon; three more the same evening by the first
"be a customer for the night" walk on demo-hp (`adventurelog`, `felhom.eu` `audits/nightly-2026-09-13-adventurelog/`).
Each is fixed, tested and red-proofed.
- **R-486 (P1) — removing an app with its backups KEPT no longer forgets its second-drive copy.**
`removeStack` cleared the cross-drive record on every removal, so the „Teljes visszaállítás" the
customer kept the backups for was refused with „nincs másodlagos fájlmásolat" over an intact 236 MB
mirror. The record — with the rest of the app's backup preferences — now goes only with
`remove_backups`.
- **R-484 (P2) — PostGIS, pgvector and TimescaleDB images are Postgres.** `dbTypeForImage` matched the
substring `postgres` only, so `postgis/postgis:16-3.5-alpine` was "not a database": no nightly
dump, no pre-update safety dump, no DB-only replay window for the app.
- **R-485 — the backup card (`GET /api/stacks/{name}/backup-data`) sizes the recovery unit and the
Tier-2 mirror(s)**, not a `db-dumps` directory and a pre-v2 `secondary/<app>/rsync` path. It
answered `has_backups:false` over 484 MB.
- **R-480 — the card no longer tells a customer a running app is stopped.** After a held update, the
hold's sentence stayed as `update_error` after a successful restore cleared the hold, and after
removal. The stack remembers that its last update ended held; `fillHoldReason` hides that outcome
once the hold is gone or the app is not deployed. A pull failure keeps its (still true) sentence.
- **R-477 — the update's off-site lookup is one `snapshots` call.** It went through
`OffsiteInventoryList`, which also runs one `stats` per app; on demo-hp that spent the whole 15 s
bound and killed a size call for an unrelated app. New `OffsiteSnapshotTimes`; both share
`offsiteNewestPerTag`.
- **R-478 — a copy older than this install's deploy does not count.** A reinstalled app leaned on a
unit left by its removed predecessor. `usableRestorePoint` = the age rule plus "not older than
`deployed_at`"; the update after a restore backs up first.
- **R-474 — "delete backups" deletes the backups.** The whole recovery unit, the app's Tier-2
mirror on any registered drive (`Tier2MirrorDirsForApp` / `RemoveTier2Mirrors`, exact-path
checked, never another app's or the shares mirror) and the app's backup preferences. Off-site
snapshots are not touched. `volumes_removed` still reads `null` over removed volumes — that half of
R-474 stays open.
**TESTS.** `internal/stacks/r480_r478_test.go`, `internal/stacks/r485_backup_card_test.go`,
`internal/backup/r477_r474_test.go`, `internal/api/r474_remove_wiring_test.go` (R-474 + R-486),
`internal/appbackup/dbservices_test.go` (R-484 cases). **Red-proofs** in the felhom.eu audit dir
`audits/v0240-2026-09-13/`: each fix removed in turn fails its test.
## v0.239.0 — any backup tier lets an app update (2026-09-13, R-475)
**MinAgent: 0.129.0** (unchanged)
**Operator ruling 2026-09-13: every backup counts.** Until now the update's precondition was the
Tier-2 unit predicate alone, so an app with no second-drive copy could never be updated. On demo-hp
`gokapi` and `nextcloud` were in exactly that state, each with a fresh recovery unit on its own drive.
**WHAT CHANGED.**
- **`backup.Manager.UpdateRestorePoints`**, beside `Tier2UnitRestorePoint`. It walks the tiers in the
ruling's order: Tier 2 (second drive), Tier 1 (the app's own unit, `ListRestorePoints`, „helyi"),
Tier 3 (`OffsiteInventoryList`, bounded by 15 s). It returns the first copy the caller accepts and
stops there, so a fresh Tier-2 copy never reaches the network. An unreachable off-site repository
counts as ABSENT, with a WARN. A box with no off-site target is plainly absent, with no WARN.
- **The age rule is one rule.** stacks passes `freshRestorePoint` (`update.backup_max_age`) as the
acceptance test, so the limit applies to whichever tier is chosen. The first FRESH copy wins, so a
stale Tier-2 mirror never forces a backup while the app's own unit is minutes old.
- **No copy anywhere → back up first**, then re-read every tier. The preflight refuses `no_backup`
only when there is no copy on any tier AND `CanBackUpApp` says no backup can be taken now. The
Hungarian sentence changed to say that; it no longer tells the customer to switch on the 2nd backup.
- **`RunAppBackupNow` tolerates no Tier-2 target.** A Tier-2 failure after the unit capture is a WARN.
The captured unit's manifest is marked proven current, because the capture's checksum skip leaves it
untouched on a quiet app, and Tier 1 is aged by the newest artifact's mtime. Without that mark an
app with no database and no volume would be refused forever — the ProvenCopyTime trap, one tier down.
- **The hold names the tier.** `settings.RestoreHold.CopyTier`; the sentence now ends *„Visszaállítható
a Mentések oldalon ebből a biztonsági mentésből: <második meghajtó | saját meghajtó | távoli mentés>,
<dátum>."* A hold written by v0.237.0–v0.238.1 has no tier and keeps its original sentence
(`UpdateHoldLegacyFmt`). The update journal records `proven_tier`, so a resumed update names the
right copy.
- **A successful off-site restore lifts an update hold.** That path never went through
`RestoreFromRecoveryUnitAt`, so a hold naming „távoli mentés" could otherwise never be cleared.
- **Unchanged:** the backups page still calls `Tier2UnitRestorePoint` for „Teljes visszaállítás".
**TESTS.** `internal/stacks/update_tiers_test.go` — G (Tier 2 chosen when present), H (own unit alone),
I (off-site alone), K (nothing anywhere: backed up first and the update completes; a failing backup
moves nothing), L (an existing copy still carries an app that cannot be backed up; control: without it
the app is refused), M (a stale copy on each tier is backed up first; a stale Tier 2 does not block a
fresh Tier 1; the limit is inclusive on one clock). `internal/backup/update_tiers_test.go` — the tier
order and early stop, H/I at the source, J (unreachable → absent + WARN; no target → silent; a hanging
repository ends at the bound), the hold text per tier and the legacy text, the tolerant pre-backup
tail, `RunAppBackupNow` really calling it, `CanBackUpApp`, and the off-site restore clearing the hold.
`cmd/controller/r475_wiring_test.go` — the tier numbers agree across packages; the adapter reads every
tier and never `Tier2UnitRestorePoint`; the hold is told the tier. Slice 4's tests were moved onto the
new seam (`TestSlice4_C` is now the L refusal).
**Red-proofs** (felhom.eu `documentation/audits/rulings-r472-r475-2026-09-13/`): **M** — age checked
only for Tier 2: three M cases fail (a 30-hour-old own unit and off-site copy carry the update with no
backup). **L** — refuse even with a copy: the L test fails. **Tail** — the proven-current mark misses:
the tail test fails on the mtime. **Off-site clear** removed: the hold stays and the test fails.
## v0.238.1 — the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (2026-09-13, slice 4 follow-up)
**MinAgent: 0.129.0** (unchanged)
**Found live, not in review.** v0.238.0 Scenario F on demo-hp: an update to `alpine:3.20` (an image
that exits at once) waited its 5-minute health timeout and was held at 10:18:02 — correctly. But the
periodic recovery-unit capture ran at **10:17:09**, inside that wait, when the app was updating and
not yet held, and wrote the never-started definition into the app's **primary** unit:
```
primary-unit: image: alpine:3.20 manifest "created_at": "2026-09-13T10:17:09Z"
tier2-mirror: image: louislam/uptime-kuma:2.4.0 manifest "created_at": "2026-09-13T10:09:51Z"
```
The Tier-2 mirror the hold text names survived only because the Tier-2 run is daily. A nightly Tier-2
falling inside a verify window would have mirrored the broken definition over the very copy the
customer is told to restore from.
**Fix.** `backup.Manager.isHeld` — the predicate the capture sweep, the Tier-2 run and the volume dump
already consult since v0.237.0 — is also true while a guarded update is moving the app, through a new
`SetUpdatingCheck` seam wired in `main.go` to `stacks.Manager.IsUpdating`.
**Tests.** `TestSlice4_NightlyLegsLeaveAnAppMidUpdateAlone` (all three legs skip the updating app; the
other app is still mirrored — positive control), red-proofed by removing the clause (the app is then
dumped, captured and mirrored); `TestSlice4_UpdatingCheckIsWiredAtStartup`.
## v0.238.0 — the page follows the update, and a held app offers no way to start it (2026-09-13, update arc slice 4 Part 4)
**MinAgent: 0.129.0** (unchanged)
**No behaviour change on the box — the surface only.** v0.237.0 made the update a job; this release
makes the page show it.
- **Pressing `Frissítés` follows the job.** The button text becomes the phase label —
`Ellenőrzés…`, `Biztonsági mentés készül a frissítés előtt…`, `Adatbázis pillanatkép…`,
`Új verzió letöltése…`, `Indítás az új verzióval…`, `Működés ellenőrzése…` — polling
`GET /api/stacks/{name}` every 3 s, and the page reloads when `updating` goes false. It no longer
reloads the instant the request returns.
- **An updating card offers no lifecycle button** and shows the phase as a progress tag.
- **A held card** (failed update OR failed restore) shows the hold sentence and a `Mentések` link, and
offers nothing that would start or update the app — `Eltávolítás` stays.
- **A failed update that held nothing** (backup failed, pull failed, precondition vanished) shows its
sentence above the buttons.
- **The updating and held checks come BEFORE `isOperational`.** That predicate counts `restarting` as
operational, which is how the 2026-09-01 spike saw a green `Frissítés` beside a crash loop. Pinned
with fixtures in `StateRestarting`; red-proofed by moving the checks after it (both tests fail).
- The app info page shows the same three notices under its header. **No new CSS**, no version number.
## v0.237.0 — the Update button takes a backup first, and tells the truth (2026-09-13, update arc slice 4 — R-448, R-443, R-439)
**MinAgent: 0.129.0** (unchanged)
**What it replaced.** `UpdateStack` advanced the pin, pulled, ran `up -d` and returned. No copy first,
no check of memory, disk, a running backup or a held app, and success the moment `up` returned —
measured in `SPIKE-app-update-2026-09-01` §4 as **HTTP 200 over an app that was already
crash-looping** (R-443). A held app could be updated at all (R-439). **`UpdateStack` is deleted**; its
only caller was the API.
**The guarded update** (`internal/stacks/update.go`), a job answering **202** at once, with phases
`checking → backing-up (only if stale) → safety-dump → pinning → pulling → starting → verifying →
done | failed` on `GET /api/stacks/{name}` (`updating`, `update_phase`, `update_phase_label`,
`update_error`, `hold_reason`):
1. **Cheap refusals first, each a 409 with a Hungarian sentence, before the intent is recorded:** held
(R-439 — `update` joins `start`/`restart` in the router's hold check), a backup/restore/app-data op
or quiesce holding it, a migration, already updating, deploying, memory (the deploy's own check,
extracted as `memoryVerdict`, releasing the app's current request), disk (**fixed 2 GB floor** on the
Docker data root — the image size is not known without a registry query).
2. **The precondition is the existing verified backup** (operator ruling 2026-09-02): an openable
Tier-2 recovery unit with a PROVEN copy date — `backup.Tier2UnitRestorePoint`, **extracted from the
backups page handler, not duplicated** (the page renders identically, pinned). No such copy →
refused. Copy older than `update.backup_max_age` (default `24h`) → `RunAppBackupNow` runs this app's
own legs (DB dump → volume dump → unit capture → Tier-2) first; it must succeed AND yield a fresh
restorable copy, or nothing moves.
3. **Safety dump before the pin moves** (`WriteUpdateSafetyDump` = R-361's `writeSafetyDump`).
4. **Pin before pull**; a **pull failure puts the pin and definition BACK** (nothing ran).
5. **Health, not the exit code:** the app's `.felhom.yml` check through the existing probe, or — with
none declared — every container running, none restarting, for 60 s; bounded by
`update.health_timeout` (default `5m`). Not healthy → **the app is stopped and HELD**
(`settings.RestoreHold`, new `reason: update_failed` + `copy_date`, same store and same gate as
R-379), the pin stays on the new version, and the page carries the hold sentence naming the backup
to restore from. A successful unit restore lifts an update hold (only that kind).
**Measured before it was designed, and the design changed because of it.** The copy's age is the last
SUCCESSFUL Tier-2 copy, not the unit manifest's `created_at`: on demo-hp bookstack's mirror held a
database dump from 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z, because a capture
rewrites the manifest only when the app's DEFINITION changes. Aged by the manifest, a quiet app would
be "stale" forever and "back up first" would not fix it.
**Crash safety is a journal** (`<data>/update-journal.json`, fsynced, written before each phase).
`RecoverUpdates` runs before the boot sweep: interrupted before the pin → dropped; while pinning or
pulling → pin put back; after `up` → the app is marked Updating (the boot sweep and the dead-app alarm
leave it alone) and `ResumeInterruptedUpdates` re-runs `up` + the health wait once the backup side is
wired, ending healthy or held.
**"A hold that only one path honours is not a hold" — three unattended paths did not honour one, and
now do:** the drive-return gate (`intermediary.go` restart + boot recreate), and the nightly volume
dump (`DumpAppVolumesSafe` ends in `StartStack`). The nightly capture and Tier-2 run also skip a held
app, so they cannot overwrite the restore point the hold text names.
**Deliberately NOT built:** putting the old version back automatically. Measured per-app
(`SPIKE-upgrade-test-2026-09-06`) and unpredictable; the route back is the restore. The multi-major
jump is not gated here (Slice 6); it fails health and is held honestly.
**Tests.** stacks: A–G (success only after health, stale → backup first, backup failure, no unit,
seven cheap refusals, pull failure pin-back, health failure hold, unsaved hold, three recovery
shapes), hold text on GetStacks, the exact phase labels. backup: proven copy time (incl. the measured
bookstack case), hold text, restore clears only update holds, busy, nightly legs skip held apps. api:
R-439, R-443 (202, never "completed"), no-backup 409 records no intent. web: backup row unchanged by the
extraction, drive-return skips a held app. cmd: guards wired, recovery before the boot sweep, resume
after the guards, boot gate + dead-app alarm know about updates. **Six companion red-proofs, each seen
to fail** — three first ran INERT or did not compile and were fixed before they counted (REPORT.md).
## v0.236.0 — "delete my data too" deletes the data, or says that it could not (2026-09-13, R-442)
**MinAgent: 0.129.0** (unchanged) — backfilled 2026-09-13 (R-470): read from v0.232.0's header; `internal/agentapi/` has no commit since 2026-09-01
**The defect.** A customer removes an app and ticks *also delete my data*. The box says it worked;
**the data is still on the disk** — measured 2026-09-01 on demo-hp: `HTTP 200` with
`hdd_paths_removed: null, hdd_paths_preserved: null` and 128 MB of Nextcloud left at
`/mnt/felhom-drives/hdd_1/appdata/nextcloud`. `DeleteStack`, `RemoveStack` and `GetStackHDDData`
resolved the drive from the GLOBAL `cfg.Paths.HDDPath`, which has no default and is set on no box, so
`ParseComposeHDDMounts` returned nil on its first line and the removal reported, truthfully, that it
removed none — as a success.
**Fix 1 — the source of truth.** The three lookups now read the app's OWN recorded `HDD_PATH` from
`app.yaml` (`Manager.appHDDPath`, via `LoadAppConfigByName`) — the rule `07-backup-architecture.md`
states (~L437: *"the drive if the app declares one (`HDD_PATH`), the system data path otherwise"*) and
that deploy, the start gate and the backup destination already implement. **Deliberately no fallback
to the global** when the per-app value is empty: that silent fallback is the exact path this closes.
**Fix 2 — success is never reported over inaction** (the R-443 class). When the data was asked for and
the box cannot resolve where it is, the removal is REFUSED with a typed `stacks.RemoveRefusedError`
that the handler maps to **HTTP 409** and the exact sentence
„Az alkalmazás adatainak helye nem állapítható meg, ezért semmit nem töröltünk. Az alkalmazás nem
lett eltávolítva.” The refusal happens BEFORE `compose down`, so a refused removal has touched nothing
— **the app is kept too**: an app gone with its data left behind cannot even be re-run from the UI.
A recorded drive that is not mounted right now refuses the same way („A(z) %s tárhely jelenleg nem
elérhető — az alkalmazás nem távolítható el, amíg a meghajtó vissza nem csatlakozik.”), via the same
`DriveLive` signal the belt and the boot reconciler use.
**"Declares no drive" ≠ "cannot resolve the drive".** An SSD-resident app (no `HDD_PATH`, no
`${HDD_PATH}`/`${USERDATA_PATH}` bind in its compose) is NOT refused: `hdd_paths_removed: []` — an
empty list, never `null` — plus `hdd_note` saying the app kept no data on a drive. A recorded folder
that is already gone is listed under `hdd_paths_missing` and stated, not refused.
**The backup half gets the same honesty.** The "outside expected directory" refusal now reaches the
response as `backup_paths_refused` (path + reason), not only a WARN line. Its base is resolved by the
same rule (the app's namespace root — its own drive, or `<system data>/felhom-data` otherwise), which
is what the router's `AppNamespaceRoot` produced the paths from; under the global lookup every backup
path of every app on every box was being refused, silently.
**Tests (15 new, two red-proofs run):** `internal/stacks/delete_r442_test.go` asserts the folder is
gone / present on a real tree, the marker-leaving stub compose never ran on a refusal, `[]`
serialisation, the userdata root untouched; `internal/api/remove_r442_test.go` drives the real
handler → real `Manager` (constructor + `ScanStacks`) → 409 + exact sentence + `app.yaml` still on
disk. Red-proof 1 (pre-fix silent fallback): C fails with `err=nil … app.yaml present=false` and the
handler test with `HTTP 200 ok=true`. Red-proof 2 ("no drive" refuses): the SSD test fails with the
refusal sentence. UI: both success modals render `hdd_note` and „Nem törölt mentések”.
**Observation:** `cfg.Paths.HDDPath` still has readers outside removal (`report/builder.go`,
`monitor/healthcheck.go`, `api/router.go` system-info, `web/server.go`, `main.go` auto-discovery +
metrics) — left alone; not deleted here.
## v0.235.0 — freeze the version, keep the fixes flowing (2026-09-06, update arc slice 3)
**MinAgent: 0.129.0** (unchanged) — backfilled 2026-09-13 (R-470): read from v0.232.0's header; `internal/agentapi/` has no commit since 2026-09-01
**OPERATOR RULING, 2026-09-06 — Option 1.** R-447 was `BLOCKED` because R-438 established that
`RestartStack`'s use of `up -d` to pick up template changes was **chosen** and written down in its own
comment; reversing a chosen behaviour is a decision, not a bug fix. The decision is made and this
release implements it.
**The rule, in one sentence: while the catalog is offering the same version you are running, its fixes
flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose
to update.**
**Nothing was added to any of the thirteen `compose up -d` call sites.** They are made safe by
removing the reason, not by gating them — the most important of them are REPAIRS (the boot reconciler,
the drive-return gate, the app-stop guard), and a repair path that refuses to repair leaves a
customer's app down, which is worse than the problem.
### The pin (`app.yaml.pinned_images`)
`AppConfig.PinnedImages`, service → image ref. **It is not `InstalledImages`.** That field is an
OBSERVATION ("what is running"); this one is a DECISION ("what should run"), written only by an act
entitled to move a version. Letting an observation feed a decision would make a bad reading become a
bad deployment — the category error `desired_state` exists to avoid (R-166). **Absent means UNPINNED,
and unpinned means the app behaves exactly as it did before this release.**
**Four writers** (`internal/stacks/pin.go`): the deploy path; `UpdateStack`; the restore
(`stackAdapter.RecreateStackDefinitionFromUnit`); and the one-time adoption pass. Each also stores the
exact definition the pin came from as **`applied-compose.yml`** in the stack dir — a name
`Syncer.copyTemplates` does not copy, written atomically.
**`UpdateStack` advances the pin and re-renders BEFORE the pull, and the ordering is load-bearing:**
`pull` and `up -d` act on the file on disk, so the catalog's definition has to BE that file first.
A pin set afterwards would pull the frozen version and report success. **A failed pin write REFUSES
the update** — the opposite of `recordInstalledImages`, because this field is intent.
### The render (`internal/sync/sync.go`)
One new nil-safe seam, `SetRenderPlanFn`, in the same shape as `rescanFn`/`postSyncHook`. The syncer
never reads `app.yaml`. `copyTemplates` now applies a table rather than copying:
| app state | result |
|---|---|
| not deployed / protected / no seam | catalog verbatim — today's behaviour |
| deployed, **unpinned** | catalog verbatim + one DEBUG |
| deployed, pinned, catalog images **equal** | catalog verbatim — **fixes flow, self-healing works** |
| deployed, pinned, catalog images **differ** | the **stored definition** — frozen WHOLE |
| pinned, differ, nothing stored | catalog verbatim + one WARN |
| mid-deploy | the compose file is left alone this cycle |
**`.felhom.yml` is always copied verbatim** — it holds no image, and it carries `catalog_since`, which
the badge needs. That asymmetry is a known limitation, filed as **R-458**.
**The frozen branch writes a WHOLE file, never a substitution of refs into a newer template:**
`wger 2.6` needs a full DB configuration the older template cannot supply, so a new template around an
old image is a third state nobody chose.
**And this is NOT "skip deployed apps"** — that was option B, rejected, because it also stops
health-check fixes, memory limits and new deploy fields, and destroys the self-healing measured live
in `SPIKE-app-update-2026-09-01` §3.
### Adoption, and the startup ordering
`Manager.AdoptPins` runs once at boot, immediately after `BackfillInstalledImages`, and pins every
deployed app to what it is already running. It reads and writes **files only** — no container is
started, stopped or touched. It skips, loudly, when the observation is incomplete (reusing
`observationCoversTemplate`, not a second rule) or when the app is running something the current
template no longer offers and no stored definition exists. Those apps keep pre-v0.235.0 behaviour.
**`syncer.Start()` moved to after adoption.** It fires an immediate sync; at its old position that
first sync ran while every app was still unpinned and would have copied the catalog over a deployed
app once per boot — precisely the behaviour this release removes.
### The badge had to change or slice 2 would have inverted silently
`Stack.TemplateImages` is read from the app's LIVE compose file, which is now the **rendered** one. On
a frozen app that file names the OLD version, so `compareInstalledToTemplate` would have found
installed == template and answered **„Naprakész" on exactly the apps that are behind** — with every
test still green, because the two fields have the same type. The comparison now reads a new
`Stack.CatalogImages`, taken from the syncer's git clone. **No readable catalog entry renders
nothing.** No Hungarian string, badge state or partial changed.
### A fix must also refresh the stored definition — found by the LIVE validation, not by review
The stored definition is written when the pin is written. A fix delivered afterwards landed in the
live compose file but **not** in the stored one — so the first time the catalog moved a version, the
freeze would have reverted every fix delivered since, **silently undoing the half of the ruling that
says fixes keep flowing**. The equal-images branch now refreshes the store as it delivers. The images
cannot move there by construction, so no version moves and no intent is rewritten.
`TestFixRefreshesTheStoredDefinition` asserts both halves: the fix reaches the store, and it survives
the subsequent freeze.
### Tests
+17 (1729 → 1746). 28 packages green, 0 FAIL. New: `internal/sync/render_test.go`,
`internal/stacks/pin_test.go`, Group G in `internal/web/updatebadge_test.go`.
**Three companion red-proofs, each run, observed failing, and reverted (2026-09-06):**
| # | mutation | observed failure |
|---|---|---|
| 1 | the frozen branch returns the catalog template | `TestGroupB` — *"a pinned app must NOT receive the catalog's new version"* |
| 2 | adoption's completeness guard removed | `TestGroupE/incomplete_observation` — *"pinned 1, want 0 — only 1 of 2 services was observed"* |
| 3 | the badge reads `TemplateImages` again | `TestGroupG` — *"THE FEATURE IS INVERTED"*, plus „Naprakész" with no catalog entry |
**Wiring:** `TestGroupH` walks the AST of `cmd/controller/main.go` for `SetRenderPlanFn` and
`AdoptPins` **and asserts their order** against the backfill and `syncer.Start()` — a
`strings.Contains` would match a commented-out call, and the render is inert without the seam.
**One hole was found by a test rather than by review:** the syncer trusted the applied path handed to
it and would have written an empty compose file over a live app. It now re-reads and falls back to the
catalog. `TestRenderTable_EmptyStoredDefinitionIsTreatedAsAbsent`.
## v0.234.0 — the label now appears on an app nobody has touched (2026-09-03, update arc slice 1b)
**MinAgent: 0.129.0** (unchanged) — backfilled 2026-09-13 (R-470): read from v0.232.0's header; `internal/agentapi/` has no commit since 2026-09-01
**Found by the operator on demo-felhom the morning after v0.233.0, and it is a real gap, not a
misunderstanding:** OpenGist had been up 15 hours, was running exactly what the catalog pins, and
showed **no badge at all**. v0.233.0 wrote the record only from the four bring-up paths, so an app
nobody restarts carried no record — and therefore no label — **indefinitely**. On a quiet box that is
every app, which is the box we most want to be able to see. v0.233.0's own architecture note called
this a known limitation; a day of it showed the limitation was the feature not working.
### `Manager.BackfillInstalledImages` — seed the absences, once, at startup
`controller/internal/stacks/installed.go`, called from `cmd/controller/main.go` beside
`BackfillDesiredState` and before the boot reconciler.
**It READS. It starts nothing, restarts nothing, upgrades nothing and writes no compose file.** That
is what makes a backfill safe here, and it is the same shape R-166's desired-state backfill already
uses — with one deliberate difference:
- **It never overwrites an existing record.** The bring-up paths own updates; this only fills gaps.
An app that already has a record is not even observed.
- **It REFUSES to seed a partial observation, and that is the whole reason this is not a three-line
loop.** `compareInstalledToTemplate` reads a service-count mismatch as BEHIND, so seeding what can
be seen on a degraded or crash-looping app would render „Frissítés elérhető" over an app that is
perfectly current — **a confident wrong answer, which is worse than the silence it replaces.** The
bring-up paths do not have this problem: they run right after a SUCCESSFUL `compose up -d`, where a
missing container is real news and already logs a WARN. A backfill meets whatever state a box is in
at boot, so it is stricter. An app it cannot observe COMPLETELY keeps no record — unknown, which
renders nothing, which is the honest answer.
- Protected and undeployed stacks are skipped, and the summary line is a **positive observable**:
`N recorded, M already had a record, K left unrecorded`.
**Nothing else changed.** No behaviour, no new endpoint, no auto-update, buttons byte-identical.
### A test of mine was a calendar bomb, and it went off overnight
`TestGroupD_BadgeRendersOnBothSurfaces` hardcoded `catalog_since: "2026-07-18"` **and** asserted
`"Frissítés elérhető — 46 napja"`. The pure tests inject a clock; **the RENDER path calls the funcmap
entry, which uses `time.Now()`.** So the test was green on the day it was written (2026-09-02) and
**red the next morning** — 47 days, not 46. It is now derived: the fixture's `catalog_since` is
computed as *today minus 46 days*, so it asserts the real number through the real clock and cannot
rot. **Filed as R-457** — six other test files mix a hardcoded date with `time.Now()` and are named
there as unchecked candidates, not accused.
### Tests
+5 (1724 → 1729). 28 packages green, 0 FAIL. **Companion red-proof (run 2026-09-03):** delete the
`observationCoversTemplate` guard and `TestGroupG_BackfillRefusesAPartialObservation` fails with
*"backfilled 1, want 0 — a partial observation must NOT be seeded"*. Reverted.
**Wiring:** `TestGroupG_BackfillIsWiredAtStartup` walks the **AST** of `cmd/controller/main.go` for
the call and asserts its ORDER — after the desired-state backfill, before the boot reconciler — because
a backfill nothing invokes seeds nothing, and a `strings.Contains` would match a commented-out call.
## v0.233.0 — the box writes down what it installed, and one label says whether it is current (2026-09-02, update arc slices 1 & 2)
**MinAgent: 0.129.0** (unchanged) — backfilled 2026-09-13 (R-470): read from v0.232.0's header; `internal/agentapi/` has no commit since 2026-09-01
**Neither slice changes any behaviour.** The Frissítés button, the restart path, the sync and the boot
reconciler are byte-identical. This adds a RECORD and a LABEL, because the behaviour work (slice 3)
is easier to judge once the fleet's real state is visible. Opened by
`felhom.eu/documentation/audits/SPIKE-app-update-2026-09-01.md`; the reasoning now lives in
`felhom.eu/documentation/architecture/09-update-architecture.md` (**R-438**, an absence that was itself
a finding).
### Slice 1 — `app.yaml` gains `installed_images`
`AppConfig.InstalledImages map[string]InstalledImage`, keyed by **compose SERVICE name**, each entry
carrying `ref`, `digest` and `at`. Written by `Manager.recordInstalledImages`
(new `controller/internal/stacks/installed.go`) after a successful compose up from `StartStack`,
`RestartStack`, `UpdateStack` and `runComposeDeploy`.
- **Read from the CONTAINER, never from `docker-compose.yml`.** That file is the value that has
already moved — the syncer overwrites a deployed app's compose on a 15-minute cycle with no
deployed check, and the two disagreed for 25 minutes in the spike's own measurement.
`checkLocalImages` (a line scan of that file) is deliberately not reused.
- **A failed write NEVER refuses the action**, and that is the deliberate opposite of
`SetDesiredState`. Intent refused, observation logged: refusing to start a customer's app because a
note could not be written trades a real outage for a bookkeeping gap. One ERROR line, app stays up.
- **Not called from `StartStackServices`** — the R-47 DB-only window would overwrite a complete
record with a partial one.
- **An unchanged observation does not rewrite `app.yaml`** (the `SetDesiredState` rule — that file
holds encrypted secrets), and `at` is carried forward so it answers "running since", not
"last looked at".
- Its own docker seam, `Manager.installedExecFn`, with a **context and a 30 s timeout** —
`composeExecCustomEnv`/`execCommand` have neither, and REUSE.md's trap table says so.
- `ParseComposeImages` is a real yaml.v3 `services:` parse, never a line scan (immich's top-level
`immich_ml_cache:` has exactly the shape a scan misreads as a service).
### Slice 2 — one Hungarian badge, and no version number anywhere
`.felhom.yml` gains optional `catalog_since: "YYYY-MM-DD"` (`Metadata.CatalogSince`, backfilled across
all 53 catalog apps in `app-catalog-felhom.eu@69761cf`). `web.updateBadge`
(new `controller/internal/web/updatebadge.go`) compares the RECORDED reference per service against
what the CURRENT template pins, and returns a `*MetaBadge` rendered by the existing `meta_badge`
partial on `app_info.html` and `stacks.html` — **no new markup and no new CSS**, which is what
`metabadge.go`'s own comment asked of its second user.
| state | badge |
|---|---|
| every service matches | „Naprakész" (`tag-ok`) |
| any service differs, age known | „Frissítés elérhető — N napja" (`tag-warn`) |
| any service differs, age unknown | „Frissítés elérhető" |
| **no record, or template unreadable** | **nothing rendered** |
- **NO RECORD MEANS UNKNOWN AND NEVER MEANS CURRENT** — the R-166 lesson applied to an observation.
Every `app.yaml` written before this version has no entry, so a fall-through to „Naprakész" would
have told the whole fleet their months-old apps were current. Red-proved.
- **No version number reaches the customer** (operator ruling: a household cannot act on `26.05.2`).
- **No registry is queried** — a box must not need eight upstream registries to render a page.
- `catalog_since` is tolerant in the `lifecycle` style: absent, empty, malformed **or future-dated**
all degrade to a badge with no age plus one WARN. The future case is not pedantry — a box whose
clock lags the catalog would otherwise print „-3 napja".
### Known limitation, stated rather than hidden
For the **23 floating pins** (`postgres:16-alpine`, `mariadb:11.6`, …) the reference can be identical
while the image behind it has moved upstream — measured live in the spike §5, where `mariadb:11.4` and
`mariadb:12.3` had both already moved. **Those apps read „Naprakész" when they may not be.** Closing it
needs a registry query and a digest comparison; filed as a register row, not left implicit.
### Tests
+17 test functions (1707 → 1724). New: `controller/internal/stacks/installed_test.go`,
`controller/internal/web/updatebadge_test.go`. Full suite green, 28 packages, 0 FAIL.
- **Wiring, through the REAL caller:** `TestGroupE_RestartStackReachesTheRecorder` drives
`RestartStack` end to end (only the compose and `docker ps` process boundaries are stubbed), and
`TestGroupE_EveryBringUpPathCallsTheRecorder` walks the **AST** of `manager.go`/`deploy.go` for the
four call sites — a `strings.Contains` would match a commented-out call, which is the exact shape of
the seam-built-but-never-wired class.
- **Companion red-proofs, all three run and reverted (2026-09-02):** (1) making a recording failure
refuse the action → `TestGroupC` fails with "restart must SUCCEED"; (2) the no-record case falling
through to `updateCurrent` → three sub-tests fail with a „Naprakész" badge on an app nobody measured;
(3) deleting the `{{template "meta_badge" (updateBadge …)}}` line from each template → the matching
render sub-test fails.
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
absent). **All four found by accident.** The gates enforce everything else here and were the one part
nothing had checked.
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
proves the cause was the listing rather than the decoy.
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
**In this repo:** six gates (`emoji`, `native-confirm`, `app-row-dedup`, `template-id`,
`secret-markup`, `retrieval-promise`) now walk instead of listing one directory; `debug-routes` and
`app-row-dedup` strip comments before matching — a dispatcher case left in a commented-out block
counted as a live handler, which is R-400 reached through the one door its own gate could not see.
New: `controller/scripts/test_gate_decoys.py` (10 decoys). R-425 (`offbox-rename`'s fixed FILES list)
is left OPEN with its decoy recorded rather than quietly fixed.
## scripts — the golden NOTICE, where the debt is created (2026-09-01, R-404) — NOT A RELEASE
**No version heading on purpose.** This changes no Go code, builds no image and bumps nothing. Giving
it a version would create the exact golden debt the change is about.
`controller/scripts/golden_notice.py`, registered in `controller_gates.py` as the first
**non-blocking** gate. Until now this repo — where a release actually happens — had no
golden-currency check at all, while `felhom.eu` ran one on every push including documents-only ones
that can neither create the debt nor clear it. So the person who could act heard nothing and the
person who could not act was blocked, thirteen `--no-verify` uses' worth (R-404, R-417).
- **It is ADVISORY IN EVERY CASE and that is the only correct behaviour**, not timidity: at the
moment a release is committed the golden legitimately does not exist yet, so blocking there would
refuse the commit that starts the process — and blocking later is the mistake being undone.
- **No second implementation.** It IMPORTS `felhom.eu/scripts/golden_currency_gate.py` and calls that
gate's own `released_versions()` / `newest_baked()`, so it is the same comparison read in the other
direction. Cross-repo shape copied from `instructions_gate.py`; the script is never duplicated here.
- **Absent sibling clone → INCONCLUSIVE and silent about currency.** It never guesses.
- **`controller_gates.py` gained a fifth `blocking` field** — it could not express a reporting-only
gate before, so the capability was added rather than the notice compromised (R-420). False for
exactly one gate; a test asserts that.
Tests: `controller/scripts/test_golden_notice.py` (N1–N4, with a positive control that every other
gate is still blocking). **Red-proof run:** making the debt branch return 1 fails N1 — in production
that would refuse the commit that starts a release.
## v0.232.0 — one writer at a time, and a check that can actually run (2026-09-01, R-411/R-408/R-407, R-414, R-412a)
**MinAgent: 0.129.0** (unchanged)
**THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID.** R-411 named one function missing
`acquireRunning`. Fixing it and then pinning the invariant with an AST walk surfaced **four** in total:
| function | why it mattered |
|---|---|
| `RestoreOffboxScratch` | the reported one |
| `OffboxRestorePrepareFull` | **the second request in the customer's own two-step full restore, and the one that actually shells `restic stats`.** The UI reaches it FIRST, so flagging only the restore would have left the collision reachable by the ordinary path |
| `RestoreSharesScratch` | R-411's exact shape on the **shares** tier: `unlockStale` + `resticStep`, a live web caller, and its sibling `PlaceSharesRestore` has always taken the flag |
| `RestoreOffbox` | no production caller today, but the same pattern — flagged so a future caller inherits the guard, not the defect |
`OffsiteInventoryList` is **registered exempt with its reason**: it issues only `restic snapshots
--json`, measured not to take a lock, and flagging it would make browsing a page refuse during a
backup for no safety gain.
- **THE REAL DELIVERABLE IS THE WALK, not the acquire.** `offbox_integrity.go:28` asserted *"Every
off-site operation takes `acquireRunning`"* since v0.227.0, nothing checked it, and it was **false
for months** — the **ninth** instance of this project's most-repeated class. `resticStep`'s licence
to run `unlock --remove-all` rests entirely on that sentence, so a false sentence there is a licence
to delete a live operation's lock. It is an **AST walk, not `strings.Contains`** — a commented-out
call still contains the string. **Red-proofed twice:** removing the acquire fails it naming
`RestoreOffboxScratch`; an unregistered fake entry point fails it naming the fake.
- **R-407 — two sentences corrected in place, not deleted** (R-360's rule). `check` *does* write a lock
file; and **`restic stats` takes one too**, which is the fact nobody had and the one that made R-411
possible at all.
- **R-414 — the proof could not run at all on a box with no registered drive.** The determination came
out as **neither** "missed" nor "deliberate": R-356's own test comments say the scratch resolver
*"still resolves … only the DESTINATION moves"*, so it was **out of scope**, and it was never ruled
out on state-only grounds — the one comment about a `systemDataPath` fallback belonged to
`PlaceOffsiteRestore`, concerned bulk **userdata**, and R-356 overruled even that. So §6.3's rule
applies and **now has a fourth consumer**.
- **The fallback is SCOPED**, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a **unit-only** restore may fall back to the system data
path (§7 records as `[FACT]` that a driveless app's unit already lives there indefinitely, and that
the same-device placement is *"intended, not a defect"*); a **full** restore keeps today's refusal,
because it pulls bulk userdata onto a state-only tier.
- **And the silence ends either way:** a proof that cannot start records `cannot_run` rather than an
`Err`, so `last_proof_result` is **never absent** — absent already means *"controller too old"*, and
a second meaning on one field is the `StatsKnown` trap one level up. Recorded **without** advancing
per-snapshot due-ness, so the app stays retryable once a drive is registered.
- **A defect I introduced and live validation caught:** the fallback resolved a scratch the cleanup
then refused to delete (*"not inside a proof root"* — its accepted-roots list is built from
registered drives, which a driveless box has none of). Every nightly proof would have left a copy
behind on exactly the boxes the fallback exists for. Fixed, and pinned by a pair of tests — one that
the copy IS deleted, one that a path outside every proof root is still **refused**.
- **R-412 leg 1** — a per-app push whose unit carried no dump and no tar now says so, at `WARN`.
Wording only; no guard, and the capture is untouched (08 §8.2). **Leg 2 stays OPEN.**
- 18 new tests, 1689 → 1707. Red-proofs run and reverted byte-identical for the walk (×2), the
driveless verdict, the push wording and the scratch leak.
---
## v0.231.0 — the box proves its own off-site copy still holds something (2026-08-31, R-87)
**MinAgent: 0.129.0** (unchanged)
**The weekly check proves the stored bytes are the bytes we stored. It cannot tell us we stored the
WRONG thing.** A hollow recovery unit backs up cleanly, checks cleanly at 100 % depth, restores
cleanly and gives the customer nothing back — R-403 measured that on demo-hp on 2026-08-31, 120 082
104 B to 7 036 B in one nightly run, recorded as a success. **Nothing in the product asked that
question, on any tier, at any cadence.** Now one job does, every night, on one app.
**Scope is the SPIKE's verdict, not the row as filed.** `audits/SPIKE-restic-restore-test-2026-08-31.md`
measured that an unattended scratch restore would have caught **one** of the five restore-path defects
drills found in six days. As a test of our restore code it is not worth an evening; as the only thing
asking *"is there anything in there?"* it is. **This does NOT prove a restore puts data back into a
running app** — that stays drill work, and `07` §8 matrix row 4 does not move.
- **THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP.** "Check the unit against its own
packing list" passes on an empty package — a hollow unit declares nothing, so everything it declares
is present. So: (1) everything declared is present, **and** (2) the manifest declares what the app is
supposed to have. Part 2 is the whole value. `TestR87_HollowUnitForAnAppWithADatabaseFAILS` is the
fence, red-proofed against the naive rule.
- **THE EXPECTATION COMES FROM INSIDE THE UNIT, never from the live box** — the snapshot may predate
the app's current shape, and `GetDockerVolumes` enumerates from live Docker, which answers a
different question. The unit's own `compose/docker-compose.yml` answers both halves: `DBServiceNames`
(the same discriminator `RestoreFromRecoveryUnit` uses) and `ParseComposeNamedVolumes`.
- **THE VOLUME HALF IS AN EXISTENCE CHECK, NOT A NAME MATCH, deliberately.** Volume tars are
`<project>_<volume>.tar` and `ResolveDockerVolumeNames` derives the project from the compose file's
parent directory — which inside a unit is the literal string `compose`. Measured on all eight real
units on demo-hp the counts match exactly and `<stack>_<volume>.tar` held every time, but "held on
eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented.
- **THREE OUTCOMES, NOT TWO:** pass, fail (readable and empty), and **cannot judge**. Collapsing the
third into a pass hides a real gap; into a failure, it alarms on our own blind spot. An app that
legitimately has neither a database nor volumes **PASSES** — `TestR87_AppWithNoDatabaseAndNoVolumesPasses`,
red-proofed by alarming on any empty unit.
- **IT NEVER WRITES TO THE REPOSITORY, and that is asserted as a NON-EFFECT.** `--no-lock`, no
`unlockStale`, and `m.runner()` rather than `resticStep` so the `unlock --remove-all` escalation is
unreachable rather than unlikely. The spike measured that `restic restore` takes no lock and `restic
check` does (R-407). `TestR87_NeverWritesToTheRepository` asserts the argv.
- **IT TAKES THE SINGLE-WRITER FLAG ITSELF and skips rather than waits.** `RestoreOffboxScratch` does
not take it (R-408) while `offbox_integrity.go` states that invariant as universal; this job does not
wait for that to be fixed.
- **DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock.** `ProvedSnapshots[stack]` holds the
snapshot ID proved. A timestamp re-proves the same snapshot forever — red-proofed, and it also breaks
the rotation.
- **ITS SCRATCH IS A SEPARATE ROOT** (`backups/offsite-proof`), and that is not tidiness: the job
deletes its copy on every path, and sharing `backups/offsite-restore/<app>` would mean a nightly
background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to
placement, so a proof copy can never be pushed into a live app.
- **NEW EVENT TYPE `offsite_proof_empty`, severity `error`, operator-only** — deliberately NOT
`backup_integrity_failed`, whose Hungarian template says the store is damaged. Here the store is
sound and the content is absent: a different fact, a different action. **The hub half ships in the
same commit** (allowlist + `operatorOnlyEvents`), because an unallowlisted type is 400'd and vanishes.
- Verdict published on `OffboxReportStatus` with the `StatsKnown` absence rule: `last_proof_result`
absent means NOT RECORDED, never "failed".
- Slot **05:30**, chosen from the live schedule read off demo-hp (offbox-backup 04:15 / 2m52s,
abandon-sweep 05:10, offsite-integrity 06:00 / 40.3 s). Measured cost: **one app ~2.3–4.0 s**, all
eight 25 s — cheaper than the weekly check beside it.
- 33 new tests. `unitOnlyHeadroom` extracted so the proof and the customer restore share one gate and
one Hungarian refusal.
---
## v0.230.0 — a poorer copy must never delete a richer one (2026-08-31, R-403)
**MinAgent: 0.129.0** (unchanged)
### The loss was MEASURED first, then fixed
R-403 was filed yesterday from a code reading with the dangerous half explicitly recorded as
**unverified**. It was run before anything was built, on the shipped **v0.229.0**, on `demo-hp`:
```
BEFORE secondary db-dumps: 4 volume-dumps: 3 size: 120082104 bytes
(a hollow primary unit, produced through the R-102 restore path exactly as on 2026-08-31)
RUN [backup] Tier 2 copied docmost → …/backups/secondary/docmost (14.9 KB, 0 leg(s), 0s)
AFTER secondary db-dumps: 0 volume-dumps: 0 size: 7036 bytes
```
**Four database dumps and three volume tars — the customer's last surviving package — deleted in one
nightly run, which recorded itself a SUCCESS.** Evidence:
`felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/`.
The mechanism was three individually-correct lines: `RunTier2` guarded the unit leg with `os.Stat`
alone (*does the folder exist*), `rsyncMirror` is `rsync -a --delete`, and nothing between them
compared source to destination. **An empty recovery unit is a folder that exists.**
### The guard
- **One predicate, `unitCarriesData` / `unitIsHollow`** (`internal/backup/r403_hollow.go`). It **asks
the MANIFEST and never the byte size**: a unit with a big compose tree and no dumps is dangerous, a
tiny unit for a tiny app is fine. Fail-closed on an absent or unparseable manifest.
`TestR403_SizeIsNeverConsulted` is the guard that keeps `dirSizeBytes` out.
- **`RunTier2` skips the unit leg** when the source is hollow **and** the destination is not. The
other legs still run, the run is **not** failed, and the skip is recorded **for the surface**
(`CrossDriveBackup.UnitLegSkipped` + `UnitPackageDate`) as well as logged loudly.
- **`--delete` STAYS and shrinking stays legal.** `07-backup-architecture.md` §8 row 5's derived-copy
rule ("Migration = rebuild, not preserve") is unchanged, and `tier2.go`'s own header records that a
classified app's copy legitimately shrinks as `export` drops out. The fence is exactly one shape.
`TestR403_DataLegShrinkIsUnaffected` is the guard on the guard.
### The honesty — a preserved copy must not read as a fresh one
A preserved package is older than the run that preserved it. The per-app card carries a notice, and
the unit-restore confirm names the **package's** date — read from the mirrored manifest's own
`created_at`, not from the status record — plus a clause saying why it is older. **The outcome
sentence names the same date**, because §2.3's rule is "not a plain green success anywhere".
> **The live run caught a defect the unit tests did not.** The first draft also flagged "older" by
> comparing the package's date to the run's — but a unit is ALWAYS captured shortly before the run
> that mirrors it, so that was true for **every healthy app on the box**. Measured: bookstack, kimai,
> opengist and privatebin all had manifests at `12:03:49Z` against a run at `12:14:24Z`, and all four
> would have been told their package was stale. **A warning that fires on everything costs the same
> as the comforting lie it replaces.** The flag is now `UnitLegPreserved` and nothing else;
> `TestR403_AHealthyAppIsNeverCalledStale` pins it.
### The cause — the primary is filled back in
`RestoreTier2Unit` now refills an **absent or hollow** primary unit from the mirror it just restored
from, **inside the call, before it returns**. The hollow manifest was written **two seconds** after a
restore by the 5-minute capture job; any follow-up job or goroutine races it.
`TestR403_RehydrateHappensBeforeTheCallReturns` asserts the ordering, never a timer.
Never over a **complete** primary (it may be newer — that is R-403 pointed the other way) and never
after a **failed** restore. **The capture job is deliberately NOT guarded:** a capture that describes
an empty drive as empty is correct, and with the primary refilled there is no hollow state left to
describe. Guarding it would make the manifest lie.
### The rider — a credential reader that ends a three-time mistake
`felhom.eu/scripts/read_credential.py`. Values in `~/.config/credentials` are single-quoted; stripping
only `"` sends two extra characters and the failure looks exactly like a stale password. **2026-07-20**
it was diagnosed as drift and written into memory; **2026-08-31** it recurred and was caught;
**2026-08-31, hours later, it recurred again and rewrote a live box's `password_hash`.** Between them
the project already had a memory file, a worked recipe and a session report — none of it stopped
occurrence three. The rule now lives in the code path: one matching quote pair is unwrapped, the
result is **refused** if it still carries a quote, `--expect-length` gives the caller a second
opinion, and the value goes file→file at 0600 with only its length on stdout.
### Live validation on `demo-hp`
| Scenario | Result |
|---|---|
| **A** — the loss on v0.229.0 | **CONFIRMED** — 120 082 104 B → 7 036 B |
| **B** — the same state on v0.230.0 | **PRESERVED** — all 7 files, all sha256 identical, WARN quoted |
| **C1** — both sides complete | mirrors as before (`114.5 MB, 0 leg(s)`, no skip) |
| **D** — the surfaces | only the skipped app carries the notice; the other **seven are the control** |
| **E** — the rehydrate | primary real the instant the call returned, and still real after **3** capture cycles |
### Tests
**24 new Go tests (1632 → 1656)** plus 2 Python tests. Red-proofs run and reverted: **A6** (predicate →
size threshold), **B1** (guard removed → the copy's 3 files are DELETED and the seam is called), **B6**
(a general never-shrink rule → the shrink case fails), **C2** (only-when-hollow dropped → the complete
primary is overwritten), **E1** (quote assertion removed → three cases fail by name), plus the
over-eager-stale red-proof. `recordTier2Success` and `tier2UnitConfirmMsg` keep their old signatures as
thin callers, so **no existing test needed editing**.
## v0.229.0 — the second drive's copy becomes a way back (2026-08-31, R-102 + R-103)
**MinAgent: 0.129.0** (unchanged)
### The Tier-2 unit mirror is restorable (R-102)
Tier-2 has written a full `recovery-unit/` mirror — the app's definition, its portable secrets, its
database dump and its named-volume tars — to `<dest>/backups/secondary/<app>/recovery-unit/` on every
run for months. **No code path read it.** Every reader of a recovery unit could only name a path under
`backups/primary/`, because `appbackup/paths.go` joined that segment literally.
**The failure that made it worth doing:** Tier-2 exists for the loss of the primary drive, and in
exactly that loss the primary unit is gone while the mirror survives — unreachable by any customer
action (`07-backup-architecture.md` §6.3, §7.2). For the **45** class-B apps in the catalogue (Part 3
below) that is the whole of their data.
- **`appbackup`** gains four unit-directory-relative primitives — `UnitComposeDir`,
`UnitManifestFile`, `UnitDBDumpDir`, `UnitVolumeDumpDir`. The four `(nsRoot, stackName)` helpers
become thin wrappers over them and return byte-identical strings; every existing caller compiles
untouched. Pinned by `TestR102_PathWrappersAreByteIdenticalToToday` against hand-written literals,
not against the helpers under test.
- **`Manager.RestoreFromRecoveryUnitAt(stack, unitDir)`** holds the whole body;
`RestoreFromRecoveryUnit(stack)` is the thin caller naming the primary unit. ONE implementation, two
callers — the rule `restoreDockerVolumesFrom` already states beside itself, and for the same reason.
- **THE SOURCE MOVES; THE DESTINATION DOES NOT.** `unitDir` changes only where the manifest, the
compose capture, the `.sql` and the tars are READ from. Data still lands in the live Docker volumes
and the live database container, and the definition still in the guest. A restore that also
relocated the app's data would be a migration.
- **Unchanged and pinned:** the R-47 mutation order (stop → volumes → recreate → DB-only start →
replay → start), the secret reconciliation with unit-over-guest precedence, the fail-closed data-key
gate, and the no-unit fallback to `RestoreApp` with its `CountsUnknown` handling.
- **`Manager.RestoreTier2Unit(stack)`** resolves the recorded copy, refuses **fail-closed** unless the
mirror carries a parseable `manifest.json` — *a directory that exists is not a package* — and
delegates. The single-writer flag is taken inside `RestoreFromRecoveryUnitAt`, not beside it.
- The 35-minute DB-replay bound is now named once (`dbReimportTimeout`), so the two bounded entry
points cannot drift in how long a wedged import may hang a restore.
- The unit DIRECTORY is now logged on every restore. Which copy a restore read from is a real question
with two answers, and an absent log line is not evidence.
### The refusal becomes an action (R-103)
An app with no file legs but a full mirror was told to press a button on a **different page**.
- **`POST /backup/tier2/unit-restore`** + `backupTier2UnitRestoreHandler`: same guards, same
`restoreOpBlocked()` refusal (R-351b), same async shape as the file restore beside it, plus the
fail-closed pre-flight so the app is **never stopped** for a mirror that could not be opened.
- **`Tier2Coverage` gains `UnitRestorable`, and `CanRestore()` is NOT widened.** It still answers only
*"can the additive file restore run?"*. One predicate answering two questions is **R-356**, which
refused 40 running apps for months. `HasUnit` also keeps its old meaning — a half-copied mirror is
still unread data the file restore must disclose, even though the unit restore refuses it.
- **The row offers the action where the refusal was**, in `btn-danger-outline`, as a SEPARATE button.
The two are not merged: one adds what is missing, the other overwrites. **The confirm carries that
difference in words** and names the copy's date — and says so differently when that date is only an
ATTEMPT (R-101). It is assembled from named Go constants rather than inside an HTML attribute, so a
test asserts it verbatim; `fmtTimeStr` now delegates to a package-level `fmtRFC3339Local` so the
confirm and the outcome cannot render one date two ways.
- **`tier2NoCoverageMsg` is NARROWED** to the case that remains — no legs *and* no openable unit — and
still names the route that works. **`tier2UnitNotCoveredMsg` is NOT deleted:** it is appended where
the FILE restore ran and is still exactly true of it.
- The outcome reuses `unitRestoreOutcomeMsg` unchanged and appends which copy overwrote the live data.
### The count, settled (Part 3)
`07-backup-architecture.md` §6.2 recorded **two** Tier-2 coverage counts that disagreed — 9/43/1 and
7/45/1 — both unresolved. Counted at catalogue `459766cb16395fd1d1a66282f5cc6da59ead5924` by running
the PRODUCTION rule (`LoadMetadata` → `ParseComposeClassifiableBinds` → `ClassifyBinds` →
`ComputeCaptureSet` at `TierSecondary`) over all 53 templates:
**A = 7 · B = 45 · C = 1.** The INV Part B.1 enumeration was right. The two apps the C9-F1 Phase-0
count put in A are **radarr and sonarr**: both bind `${USERDATA_PATH}` paths **writably**, so the
`:ro`-default rule Phase 0 says it applied to plex/jellyfin/emby/navidrome does not catch them — they
are excluded by an **explicit** `class: excluded` entry instead. Class C is **bentopdf**, which
declares no volumes and no namespace binds at all. No catalogue file was changed.
### Live drill — demo-hp, endpoint level
`documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (in `felhom.eu`). docmost, class B, its
Tier-2 run reporting **0 leg(s)**. With the **primary unit moved aside** the restore returned 3 volumes
of 3 and 1 database of 1 from the secondary mirror in **28.65 s**; an accented Hungarian filename came
back byte-for-byte (verified as hex, not as rendered text — R-364) and the app read its own row **over
TCP with its own credential**. The post-backup discriminator was **gone**, so the replay was real.
**Scenario D** repeated it with the guest's `app.yaml` moved aside: `secrets recovered=2/2` from the
mirrored unit — this closes `00-capability-map.md`'s open *"not exercised live"* clause for Tier-2's
own cross-drive copy of a secret-bearing unit.
**Filed, not fixed — R-403.** Two seconds after a restore that ran with the primary unit absent, the
5-minute status refresh (`captureAllRecoveryUnits`) rewrote the primary unit from a drive with no
dumps, producing a manifest carrying `"db_dumps": []` and `"volume_dumps": null`. The ordinary restore
then read it and honestly reported that the backup held only settings. The dangerous half — that the
next Tier-2 run would mirror that hollow unit over the good secondary copy, since `rsyncMirror` carries
`--delete` — **was not tested and is recorded as unverified.**
### Tests
26 new test functions (1606 → 1632). Red-proofs run and reverted: **A1** (change one wrapper's join →
fails on all three fixtures), **A5** (swap the volume replay and the recreate → fails on the sequence),
**B2** (point the Tier-2 reader back at the primary → fails, and fails again with `permission denied`
once the primary tree is unreadable), **C1** (widen `CanRestore` to include `HasUnit` → the unit-only
cases fail), **D6** (drop `EndRestoreOp` from the handler's goroutine → *"the restore never published a
result"*). The Tier-2 fixtures build their mirror with the production `RunTier2`, so the claim under
test is *the copy Tier-2 writes is the copy this restore reads*.
## v0.228.0 — the check reads the data, and the debug page stops lying (2026-08-31, R-399 + R-400)
**MinAgent: 0.129.0** (unchanged)
### The off-site check now re-reads the stored data (R-399)
`monitoring.integrity.read_data_subset` **defaults to `100%`**. Every box whose `controller.yaml` has
no `integrity:` block — which is every box in the fleet — now downloads and re-hashes the whole
off-site store on its weekly check instead of reading only the catalogue.
**The fact that forced it:** on `demo-hp`, 2026-08-30, a pack was damaged **without changing its
size**. Plain `restic check` — the structure-and-index check every box ran — reported
`no errors were found` and exited clean. Every `--read-data*` form caught it. A structurally-verified
store is one whose rot is found at restore time, with a customer waiting.
**The cost**, same store and day (140 829 678 B / 2 651 blobs / 67 snapshots): structure 35.0 s,
10% 35.9 s, 50% 37.3 s, **100% 39.2 s**.
- **`off` (any case) is the new off token.** Absent or empty means *not configured*, therefore the
default; a setting with no off switch is not a setting, and emptiness cannot mean both things.
- **A malformed value falls back to the DEFAULT, never to structure.** Downgrading on a typo would
silently remove the protection this row adds — R-357's shape, a guard that opens quietly.
- **A completed check over 5 minutes logs a WARN** naming the duration, the depth and R-401. Operator
log only: no hub event, no customer alarm, and it never changes the depth by itself. A skip or an
unreachable store never warns — neither has a duration to judge. The threshold is deliberately
imprecise (≈7.6× the only full-depth number that exists) because a notice changes no behaviour,
while a precise number invented from one measurement on one 134 MB store would not.
- **The depth is now recorded with the verdict** — `settings.OffboxTarget.LastIntegrityDepth` and
`OffboxReportStatus.LastIntegrityDepth`. Empty means NOT RECORDED (a pre-0.228.0 controller), never
"structure": absence means the box cannot answer, following the `StatsKnown` precedent beside it.
- **R-401 filed with a TRIGGER, not a date:** revisit the depth when the timing WARN fires on any box.
No rotation schedule, size threshold or bandwidth budget is built here — every one of those would be
a number invented from a single data point.
**R-87 (the restic tier is never restore-tested) stays OPEN.** Reading the bytes back out of the store
is not a restore.
### The debug page stops lying (R-400)
The shipped debug page referenced **24** `/api/debug/...` addresses and the dispatcher answered **17**.
Seven controls did nothing — and three of those seven were not buttons at all: `dr/infra-status` and
both `storage/watchdog-status` calls fetch on page **load**, so whole panels had been permanently
blank and nobody had to click anything to be misled. This is the page an operator opens when something
is already wrong.
| control | disposition | why |
|---|---|---|
| `backup/crossdrive` | **IMPLEMENTED** | `Manager.RunTier2` is live; only the route was missing |
| `backup/infra` | **DELETED** | the disk-tier infra backup moved to the host agent in slice 8C |
| `hub/infra-push` | **DELETED** | `Pusher.PushInfraBackup` was removed 2026-06-16 (it pushed plaintext secrets) |
| `dr/infra-status` | **DELETED** | it rendered the two retired mechanisms above; fetched on page load |
| `storage/watchdog-status` | **DELETED** | the slice-8C watchdog is retired; the drive-gate reconcile replaced it and publishes no such status. Fetched on page load, twice |
| `storage/simulate-disconnect` | **DELETED** | no backing capability, and it WRITES storage state — a button that fakes a drive disconnect on a customer's machine is a foot-gun |
| `storage/simulate-reconnect` | **DELETED** | same |
Each deletion took its panel and its JavaScript with it; the "Tárhely teszt" section went entirely.
A panel left behind renders nothing forever, which is how this class hides.
**`controller/scripts/debug_route_gate.py` makes the class impossible.** Two lists and a difference: it
fails when the template references an address the dispatcher lacks, **and** when the dispatcher answers
one nothing references. Registered in `controller_gates.py` **after** the seven were resolved — a
registered-but-failing gate refuses every push. Ten lines on purpose. Red-proofed in both directions.
Tree after the change: **18 referenced addresses, 18 dispatched, none orphaned.**
### Three corrections
- `internal/report/types.go` — the dead-field warning block said *"the controller runs no integrity
check, and `NotifyIntegrityOK`/`NotifyIntegrityFailed` … are called from nowhere"*. **Both halves
became false in v0.227.0.** The fields stay dead and unrendered (`TestBackupReport_DeadFieldsStayZero`
still passes unmodified); only the *reason* changed.
- `configs/controller.yaml.example` had **no `integrity:` block at all** — an operator could not
discover the setting exists. Added, with both keys, the default, the off token and the measurement.
- `internal/backup/offbox_integrity.go` — `integrityCheckTimeout`'s comment said read-data "ships OFF
… whoever turns it on must revisit this number". Rewritten: it is now the number a large store meets
first, and the slow notice exists to say so long before it does.
### Superseded tests
Two R-359 tests asserted the ruling this release reverses, and are replaced rather than weakened.
`TestR359_StructureCheckPassesNoReadDataFlag` → `TestR399_AbsentConfigRunsFullDepth`.
`TestR359_MalformedReadDataSubsetIsTreatedAsOff` → `TestR399_MalformedFallsBackToTheDefault` (its NAME
was the defect: treating a typo as "off" is the quiet downgrade). Both are recorded in place, so a
later reader does not re-derive the old ruling from an absence.
## v0.227.1 — the damage classifier matched restic's ordinary progress output (2026-08-30, R-359 follow-on)
**MinAgent: 0.129.0** (unchanged)
**A patch rather than a rebuilt 0.227.0, deliberately** — 0.227.0 was already deployed to `demo-hp`
when this was found, and re-pushing changed bytes under a tag that is already running somewhere is the
`:latest` hazard with extra steps.
`looksLikeRepositoryDamage` matched bare `"pack "`, `"tree "`, `"snapshot "` and `"blob "`. **A healthy
`restic check` prints `check all packs` and `check snapshots, trees and blobs`**, so any check that
failed for a NON-damage reason — a connection dropped mid-run, say — would have been classified as a
corrupted repository and told the customer their backups may be damaged. That is the false alarm that
teaches an operator to ignore the true one.
The signatures are now phrases from restic's own error wording (`does not match`, `not found in index`,
`ciphertext verification failed`, `repository contains errors`, …), and the negative control that
caught it — `TestR359_HealthyRealOutputIsNotDamage`, built from the real bytes of a real passing
check — is what keeps it caught.
## v0.227.0 — the off-site store gets checked, and the check that was advertised becomes real (2026-08-30, R-359 + R-397)
**MinAgent: 0.129.0** (unchanged)
### R-359 — nothing ever verified that the off-site copies are still readable
The whole-guest tier has verify jobs. The tier holding the customer's documents and photos had none:
the complete set of restic verbs this controller used was `restore, snapshots, backup, unlock, stats,
init, forget, prune, cat, config` — **no `check`**. We would have found out at restore time, with a
customer waiting.
### ⚠ AND THE MEASUREMENT CHANGED WHAT THE FEATURE IS WORTH — read this before the rest
Part 5's positive control corrupted one pack of a throwaway repository **without changing its size**
(64 zero bytes at offset 1024). Both depths were run against it:
| depth | exit | verdict |
|---|---|---|
| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** |
| any `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` |
**The structure check reported a corrupted store as healthy.** It verifies the index, the pack
inventory and the snapshot graph — real failure modes, and it catches missing packs, broken indexes and
unreadable snapshots. It does **not** re-hash pack contents, so it cannot see rot inside a pack that is
still the right size. **R-399 is therefore not only a bandwidth question: at the shipped default a
class of damage is not checked at all.**
**The cost curve, measured against the live store (134.3 MB, 67 snapshots) rather than reasoned:**
| depth | wall | over structure-only |
|---|---|---|
| structure only | 35.0 s | — |
| `10%` | 35.9 s | +0.9 s (+3%) |
| `50%` | 37.3 s | +2.2 s (+6%) |
| `100%` | 39.2 s | **+4.2 s (+12%)** |
At this size, re-reading **all** the data costs four seconds more than reading none — the wall clock is
dominated by SFTP round-trips, not transfer. **These figures do not extrapolate**: the structure
check's cost tracks the index, a read-data run's tracks the data. The default is still not chosen here;
R-399 now has numbers instead of guesses.
### The hazard that shapes the whole design
`resticStep` self-heals a crash lock by running **`unlock --remove-all`** and retrying, and its own
comment records why that is safe: *every caller holds the in-process single-flight mutex, so any lock
it meets is stale.* **A check that did not take that flag could meet a LIVE `forget --prune`'s lock
from this same box, remove it, and retry over the top of it** — on the tier holding customer data.
So the check **takes the flag and SKIPS rather than waits**. Waiting would pin the nightly backup
behind it; a skip costs nothing because due-ness makes tomorrow try again.
`TestR359_SkipsWhenRunningFlagHeld` asserts the **non-effects** — restic never invoked, `unlock` never
in any argv — and its red-proof prints the real thing: restic running `check` with the flag held.
**Proven live too.** The intended demonstration could not be run (`POST /api/backup/offbox/run` is a
404 — that is **R-279 and stays open**), so the same flag was exercised by its other holder: two checks
6 s apart. The second returned `skipped: true`, `"a backup or restore is already running"`,
**`duration_ms: 0`** — it never ran restic at all.
### Due-ness, not a weekday
A daily job asking *"is the last successful check older than 7 days?"*, not *"is it Sunday?"*. A box
switched off on its check day is checked the next day it is on. **R-341 is exactly the other shape** —
a dated check quietly missed and never caught up. No `Weekly` primitive was added; due-ness is smaller
and is what R-86 already chose for restore-tests. Registered at **06:00**, chosen from the live
schedule read off `demo-hp` (db-dump 02:30, tier2 + fill-watch 03:30, metrics-prune 04:00, offbox-backup
04:15, abandon-sweep 05:10, whole-guest gate 04:30–08:30).
### Three outcomes, not two
`Skipped`, `Unreachable` and failed are different facts. **"I could not look" is not "I looked and it
is broken"** — R-339 already owns reachability, and a second alarm for the same fact trains the
operator to discount the one alarm that means the backups are damaged. A timeout is unreachable, never
damage. A failure advances due-ness (a broken store must not be re-checked nightly); a skip and an
unreachable store do not.
### R-397 — the notifiers get their caller
`NotifyIntegrityOK` / `NotifyIntegrityFailed` existed with **no caller**; the hub allowlists both event
types and carries the Hungarian customer text for both; the settings checkbox exists; the debug button
posts to `/api/debug/backup/integrity` and **the dispatch had no such case**. Everything was built
except the part that runs. **Sixth instance of that shape in this project.**
Success is severity `info`, which `severityNotifies` drops before either leg — it **mails nobody, by
design**. A weekly success e-mail is how people stop reading their alerts. Observed live:
`Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)`.
The customer gets a **sentence**; restic's words go to the log, truncated (R-379: 615 bytes of raw
database text reached a customer once). Published on `OffboxReportStatus`, **not** on
`report.BackupReport`'s `IntegrityOK` — those were retired by R-331 the day before, and
`TestBackupReport_DeadFieldsStayZero` still passes unmodified.
`backup_integrity_failed` was checked against `perAppCooldownEvents`, `operatorOnlyEvents` and
`DefaultEnabledEvents` and **deliberately left out of all three** — it is already in the right shape.
### Two defects this work introduced and then caught
**The damage classifier matched restic's ordinary progress output.** The first draft looked for bare
`"pack "`, `"tree "`, `"snapshot "` — and a healthy run prints `check all packs` and
`check snapshots, trees and blobs`, so a check that failed for a non-damage reason would have alarmed
the customer that their backups were corrupt. **Caught by the negative control**
(`TestR359_HealthyRealOutputIsNotDamage`) using the real bytes of a real passing check. The signatures
are now phrases from restic's own error wording.
**An exit code I misread.** An early run showed `exit=0` on the subset forms while they printed
`Fatal: repository contains errors`. That was not restic — the commands were piped through `tail`, so
`$?` was tail's. Re-measured without pipes, every read-data form exits 1. The project's own
"exit codes that lie" trap, caught by re-measuring rather than reasoning.
### Part 0 was NOT built, and R-398 was my own mistake
R-398 (filed by me yesterday) said `resticStep` is not a seam so no test can drive a restic-backed
path. **The first half is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` /
`m.runner()` has been injectable since the off-site tier shipped, and other tests already drive restic
paths through it. A `resticStepFn` seam would have been **worse** here — it would replace the
`unlock --remove-all` escalation and hide it from the assertions that must observe it. R-358's AST
ordering test is converted to a real execution test instead, which immediately surfaced something the
AST walk could not: `unlockStale` legitimately runs before the restore. R-398 is **corrected, not
closed**.
**Green gate:** 28 packages, rc 0. Four red-proofs (A2, B1, C2, D2), each printing the pre-fix
behaviour. Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/`.
## v0.226.1 — the unknown that the v0.226.0 fix drew as a zero (2026-08-30, R-353 follow-on)
**MinAgent: 0.129.0** (unchanged)
**A patch release rather than a rebuild of 0.226.0, deliberately.** 0.226.0 had already been deployed
to `demo-hp` by hand when this was found, and re-pushing a changed image under a tag that is already
running somewhere is the `:latest` hazard with extra steps — two different images, one name, and no
way for a box to tell which it has.
The defect and its fix are described in full under v0.226.0's *"One defect the fix itself introduced"*
heading: `RestoreFromRecoveryUnit`'s no-unit fallback returned a ZERO `UnitRestoreResult`, which is
Scenario B's shape, so it would have told a customer whose volumes had just been restored that their
backup held only settings. `CountsUnknown` carries the unknown instead, and a fourth sentence states
it. Pinned by `TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty`; the A5 seam test corrected
alongside it.
## v0.226.0 — four ways the restore screen could mislead a customer (2026-08-30, R-353/R-357/R-358/R-360)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
> **Version note.** The task specifying this work targeted **v0.224.0** against baseline `f8c9390`.
> Both were consumed earlier the same day by R-330 (v0.224.0) and R-331 (v0.225.0). The drift was
> re-confirmed against live Gitea before the first edit, the operator authorised proceeding, and every
> symbol the spec named was re-verified present at the real baseline `e5eee50`. This is that work at
> **v0.226.0**.
All four defects were proven on `demo-hp` during the 2026-08-21 backup-truth drill and all four were
still in shipped code. They share one acceptance idea: **a restore surface must state what it actually
did, and must refuse what it cannot do.**
### R-353 — a restore that gave back nothing still said it worked
`RestoreFromRecoveryUnit` returned only `error`, so the surface reported `<app> visszaállítva
(<snapshot>).` — a sentence equally true of a run that returned an app's entire dataset and of one that
returned nothing. On 2026-08-21 an `opengist` restore printed it over a unit holding `manifest.json`
and `compose/` and nothing else.
**The count already existed and was thrown away one line deep.** `restoreDockerVolumesFrom` has always
returned it; the wrapper `restoreDockerVolumes` discarded the int. It now returns `(int, error)`, and
`RestoreFromRecoveryUnit` returns `(UnitRestoreResult, error)` carrying volumes replayed, DBs replayed,
**and what the manifest LISTED** — because zero-replayed has two causes and they are opposite news:
| state | sentence |
|---|---|
| something came back | „A(z) X: 2 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult." |
| nothing came back, unit listed nothing | „…FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem." |
| nothing came back, unit listed dumps | „…FIGYELEM — a mentés N adatkötetet és M adatbázis-mentést sorol fel, de egyik sem állt vissza. Az adataid változatlanok maradtak." |
**Every sentence is a claim about THE BACKUP, never about the app.** „ennek az alkalmazásnak nincs
adata" is forbidden here and the reason is recorded: the off-site twin has `SafetyDump` as an honest
discriminator, this path has none, and 07-backup-architecture §6.3 records that an absent dump has
causes that say nothing about the app (R-361 destroyed apps' canonical `.sql` files for four months).
That is R-355's rule, now extended to the Tier-1 path.
The snapshot id is dropped from the sentence deliberately: it named WHICH backup ran and said nothing
about what came out of it, which is the question the sentence exists to answer.
### R-357 — the destructive restore had no free-space gate
`offbox_reconstitute.go` contained **zero** references to `offboxFree`. All three existing headroom
gates guard NON-destructive paths. The one path that stops the customer's app and overwrites their live
data had none — and on 2026-08-21 it stopped an app, ran out of disk half-way, left 2 of 5 planted
items in place and restarted the app.
The gate now sits **before `mapOffsiteRestorePaths`, before `writeSafetyDump` and well before
`StopStack`**, so a refusal costs the customer nothing. **Position is the whole fix**, which is why the
test asserts `StopStack` was never called rather than asserting the error string.
**No headroom multiplier**, matching `PlaceOffsiteRestore`: this is a local copy whose size is known
exactly, unlike `OffboxRestorePrepareFull`'s ×1.1, which is predicting a download. Stated in a comment
so it is not "fixed" later. **Fail-closed on either probe returning ≤ 0** — without that, `free < need`
with `need == 0` is FALSE and an unmeasurable scratch sailed straight through: a gate present and
inert, which is worse than no gate.
### R-358 — a failed download was offered as a good one
`OffboxFullScratchReady` answered "the directory exists and is non-empty". A restic run that dies
part-way leaves exactly that, so „Teljes visszaállítás indítása" was offered over a part-copy and
reported success. Its doc comment — *"PlaceOffsiteRestore re-validates per-path completeness"* — is
what made the weak gate look adequate; that call stats top-level placements, not the files inside them.
`RestoreOffboxScratch` now clears any stale marker **before** restic runs and writes
`.felhom-restore-complete.json` (0600, tmp+fsync+rename) **only after** restic returns nil. The gate
reads it. Absent, unreadable, wrong schema or `full:false` → **not ready**, with a WARN naming which.
**Both handlers refuse server-side.** The wizard's `PlaceEnabled`/`RestoreEnabled` flags control a
button, and a hidden button is not a guard — a direct POST over a part-copy used to be accepted.
**Scenario F's open question is answered, and the answer is worse than the question assumed.** The
task asked whether a unit-only scratch is reachable through the real UI flow. **It is, by the most
ordinary route available:** „Ellenőrző visszaállítás" (`mode=unit`, advertised as non-destructive)
writes the SAME directory — `offboxRestoreScratchDir` ignores `full`, and `--include` limits what
restic extracts, never where — so a customer who ran the SAFE verification restore was then offered the
destructive one over a unit-only copy. Filed as **R-396**; the marker closes it.
### R-360 — the delete refused only while a BACKUP ran
`offboxVerifyCopyDeleteHandler` guarded on `s.backupMgr.IsRunning()`, which is FALSE for the whole of a
verification restore. Its five siblings on that surface all use `restoreOpBlocked()`, which consults
both flags; this one was missed. **Its doc comment claimed it refused during a restore, and that
sentence is why nobody looked** — it is corrected in place rather than deleted.
Observed live 2026-08-21 22:35: `RestoreStatus().Running == true` while `IsRunning() == false`, and the
delete of the copy the restore was writing into went through. No app-name comparison was added:
refusing during ANY restore is strictly stronger and matches the other five handlers.
### One new test seam, and why it was necessary
`SetOffboxLatestSnapshotFn` overrides the restic snapshot lookup. Without it R-357's gate could not be
tested at the level that matters — reaching it requires getting past `offboxLatestSnapshot`, which
shells out to restic, and a gate proven only by reading the code is the assurance class this project
has been burned by. Nil in production.
### Red-proofs — each printed the pre-fix behaviour
| mutation | observed failure |
|---|---|
| revert the outcome to `stackName+" visszaállítva ("+snapshotID+")."` | `THE PRE-FIX SENTENCE REACHED THE CUSTOMER: "opengist visszaállítva (snap-123)."` |
| delete the R-357 headroom gate | `THE APP WAS STOPPED (1 call(s)) … (err=<nil>)` on all three gate tests |
| restore the old non-empty scratch check | `a part-copy was reported READY` + the unit-only and unreadable-marker cases |
| revert the delete guard to `IsRunning()` | `THE VERIFICATION COPY WAS DELETED while a restore was writing into it` |
| remove both server-side scratch refusals | both handlers redirected with „elindult" over a part-copy |
**The first R-357 red-proof exposed a hollow test of my own and is recorded rather than quietly
fixed:** the fixture refused earlier, at the placement stat pre-pass, so `stops == 0` passed against
the pre-fix code. The scratch is now populated the way a completed download leaves it, and the
assertions are ordered so a removed gate reports the outage rather than "no error returned".
### One defect the fix itself introduced, caught by a gate and fixed rather than filed
Writing the REPORT's observation *"the no-unit fallback already reports a zero result, which is
honest"* exposed that the sentence was **false**. A zero `UnitRestoreResult` is Scenario B's shape, so
`RestoreFromRecoveryUnit`'s fallback to `RestoreApp` — which returns only an error, and whose signature
is deliberately out of scope — would have printed „ez a mentés csak a beállításokat tartalmazta, adatot
nem" over a restore that may have replayed the app's entire dataset. **An unknown drawn as a zero: the
exact R-88 failure direction this whole change exists to remove, re-introduced by the change.**
`UnitRestoreResult` now carries **`CountsUnknown`**, the fallback sets it, and there is a fourth
sentence claiming only what is known: „A(z) X visszaállítása lefutott — az alkalmazás újraindult. Ehhez
a mentéshez nem tartozik mentési egység, ezért nem tudjuk megmondani, mi állt vissza belőle." Pinned by
`TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty`; the A5 seam test was corrected too, because
its fixture has no unit and so exercises exactly this path while asserting the wrong sentence.
**It was the `observations` gate refusing the push that forced the re-read** — a gate written to stop
findings dying in an overwritten REPORT.md caught a live defect instead.
**Green gate:** `go build ./... && go vet ./... && go test ./...` — 28 packages, rc 0.
## v0.225.0 — the hub could not tell an empty off-site store from an unmeasured one (2026-08-30, R-331)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
### R-331 (controller half) — forward `stats_known`
One field, and the reason it is a release. The hub's operator Backup card read
`Snapshots 0 · Repo Size 0 MB · Integrity Unknown` for **every customer** because it rendered the
report's `backup` object, whose snapshot/size/integrity fields have had **no producer** since
disk-tier restic moved to the host agent (slice 8C) — `buildBackupReport` says so in a comment and
leaves them zero. Measured on `demo-hp` 2026-08-30, while that night's log said
`[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`.
The live numbers were always in the report's `offsite` object (`OffboxReportStatus`), which the hub
already reads for its Offsite page and its fill/staleness alarms. **The hub fix is to render that
object — and that made exactly one field mandatory that was not being forwarded.**
`snapshot_count: 0` means two opposite things: *this repository holds nothing* and *nobody has ever
measured this repository*. **R-225 measured that confusion inside this repo** — a rebuilt box rendered
„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67` — and
`settings.OffboxTarget.StatsKnown` is what fixed it for the controller's own UI. It was never put on
the wire, so the hub was free to make the identical mistake one layer up, and did.
`OffboxReportStatus.StatsKnown` now carries it. It is `omitempty`, so a controller below this version
sends no key at all and a hub parsing it sees `false` — **which must mean "cannot answer", never "the
answer is zero"**. Absence is ignorance, not emptiness; that direction is pinned by
`TestOffboxReportStatus_AbsentStatsKnownParsesAsUnknown`.
### The four dead `BackupReport` fields are now labelled and pinned
`SnapshotCount`, `RepoSizeMB`, `LastIntegrityCheck` and `IntegrityOK` stay on the wire so historical
reports in the hub's store keep parsing, but they now carry a warning naming R-331 and pointing at
`Offsite` instead. `TestBackupReport_DeadFieldsStayZero` fails the moment a producer appears for any
of them — the prompt to update the hub card in the **same** change rather than shipping a field
nothing renders. That is the "seam built but never wired" class, hit five times in this project.
`IntegrityOK` has no source at all and cannot get one by accident: the controller runs no integrity
check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are called from nowhere. The hub
card drops the row rather than re-sourcing it.
**Tests:** `offbox_statsknown_report_test.go` asserts the JSON the hub actually sees, not the Go
struct — the measured-empty and never-measured cases must differ on the wire, which is the entire
point of the field. **RED-PROOF:** dropping `StatsKnown: t.StatsKnown` from `OffboxReportStatus()`
fails with `a MEASURED empty repository reported stats_known=<nil>`.
## v0.224.0 — the backup alarmed about the apps it was holding down (2026-08-30, R-330)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
### R-330 — 61 e-mails about apps that were never broken
**Measured live on `demo-hp` 2026-08-30, controller 0.223.0.** Every night both demo boxes e-mailed
the customer `app_start_failed — "Telepitett alkalmazas nem fut: <App>"` about healthy apps. Both
bursts were the box's own backup:
| leg (UTC) | window | events |
|---|---|---|
| `db-dump` 00:30 | W (02:30 CEST) | Docmost, Paperless-ngx, RomM |
| `offbox-backup` 02:15 | W+105m (04:15 CEST) | Docmost, Paperless-ngx |
`DumpAppVolumesSafe` stops a stack (`docker compose down`), tars its volumes and starts it again —
**~13 s per stack, measured** — while `deadapp-check` scans every **30 s**. The scan caught whichever
stack was mid-cycle. Every other scan that day logged `0 currently down`, and the same night's
`[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s)` proves the backup itself was healthy.
**The defect is not a missing mechanism — it is a mechanism that was never consulted.**
`quiesce/suppress.go` solved exactly this in v0.179.0 (R-97b) and works. But `classifyRunStates` read
only `Loop.SuppressedStacks()`, and the quiesce loop covers the **whole-guest** (vzdump/PBS) backup.
The **per-app** legs stop stacks through `Manager.DumpAppVolumesSafe`, which registered with nothing.
Two mechanisms in this product stop a customer's app on purpose; only one told the alarm. That is the
**"seam built but never wired"** class, and this is its fifth instance — the first where the unwired
half was a *consumer* rather than a producer.
**The fix sits on `AppStopGuard`, not in a fourth registry.** The guard already brackets every
deliberate stop in the product — `Begin` before the stop, `End` after a successful restart — at all
three call sites (volume dump, off-site reconstitute, `.fab` export), and `main.go` hands the SAME
guard object to the backup manager and the exporter. The fact the alarm needs already existed there
with exactly one writer. `scanDeployedAppRunStates` now takes the union of both suppression sets.
**All three per-app stop paths are covered by the one change**, not just the nightly one that was
reported. A `.fab` export and an off-site restore stop an app the same way and would alarm the same
way; fixing only the observed leg would have left two loaded guns.
### It must never latch — the harder half
Permanent suppression trades a loud false alarm for a silent real one, which is F-CRIT-1 and R-88
Scenario D over again. `AppStopGuard.End()` runs **only on a restart that succeeded**, so an
open-ended hold is a real hazard here in a way it is not for the quiesce loop, which always releases.
Three independent things stop the window latching:
1. **`ReleaseFailed`** — a restart that was ATTEMPTED AND BROKE drops the entry **immediately**, so
the app alarms on the very next scan with no delay at all. Wired at every failure path: the volume
dump, the off-site reconstitution's `restartStack`, and the exporter's restart defer (the seam
interface grew the method rather than the exporter keeping its own bookkeeping).
2. **`Begin` REPLACES the set.** The marker file holds one operation, so a new `Begin` proves the
previous one is over; a set stranded by an operation that died mid-window cannot survive into a
later one.
3. **`appStopMaxHold` (6 h)** — a backstop for a hold nothing ever released, logged at WARN when it
fires. Longer than any real hold (13 s dump, minutes for an export, hours at the outside for a
multi-gigabyte reconstitution) and far shorter than "forever". Exceeding it means something is
wrong, and the right answer when something is wrong is to let the alarm through.
The grace after a successful restart is **180 s, deliberately the same constant as
`quiesce.quiesceAlarmGrace`** and by the same derivation (120 s deploy health timeout; Mealie's 60 s
`start_period` plus check intervals). Two suppression windows over one alarm that disagreed on how
long a restart takes would be a bug waiting to be found on whichever path used the shorter one.
**The suppression is deliberately NOT persisted.** After a crash the guard holds nothing: `Recover()`
either brings the apps back or leaves them genuinely down, and a down app must alarm. Reviving a
suppression across a restart would silence the exact case the alarm exists for. The durable crash
marker is untouched by all of this and stays the recovery record — `ReleaseFailed` drops the
suppression and **keeps** the marker, and a test pins that.
### Tests, and what each red-proof actually printed
`internal/backup/appstop_suppress_test.go` drives the **real** `DumpAppVolumesSafe` and asserts the
suppression set the dead-app scanner actually reads — the consequence, not a log line. Three
red-proofs were run and are recorded in `REPORT.md`:
- deleting `markStopped` from `Begin` → `suppressed at stop = map[]`, the exact pre-fix shape;
- deleting `ReleaseFailed` from the volume dump → `suppressed = map[bookstack:true] after a restart
that FAILED`;
- passing `nil` instead of `appStopGuard` in `main.go` → the AST wiring test fails.
That third one is the point: the component was never the broken part, so a test that only injects it
directly would have passed against the shipped defect. `cmd/controller/r330_backup_suppression_test.go`
walks main.go's AST rather than matching a string, because a commented-out call satisfies
`strings.Contains` — a sibling test in that package records paying for exactly that.
### Proven live on `demo-hp`
`POST /api/backup/run` (the endpoint the "Mentés indítása" button invokes), 8 stacks stopped and
restarted over 87 s, **3 dead-app scans ran inside that window** — 16:09:26 on `docmost`, 16:09:56 on
`paperless-ngx`, 16:10:26 on `romm`, the same three apps that alarmed the night before — and **zero**
`app_start_failed` events were pushed. The scan count is the positive control: an absent alarm is
equally consistent with "suppressed correctly" and "the scanner stopped".
A first attempt was **discarded and said so**: it fired 52 s after a controller restart, inside
`deadAppBootGrace` (90 s), where the scan returns early and could not have alarmed whatever the code
did. `demo-felhom` is deployed but not independently proven — its single app cycles in ~1 s, too fast
for any 30 s scan to land inside.
## v0.223.0 — the alarm we had just built reached nobody, and the stop nobody heard (2026-08-23, R-329 + R-386)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
### R-329 — one word, and every `app_start_failed` was delivered to no one
`NotifyAppStartFailures` emitted severity **`"warn"`**. The hub's vocabulary is exactly
`{info, warning, error, critical}` and it **silently coerces** anything else to `info` at ingest;
`severityNotifies` then drops `info`, returning **before both** the operator and the customer leg. So
the banner appeared, the hub stored the event, the POST returned 200, and no mail left the building.
**This is the SECOND time.** `DiskAlertKind.Severity` emitted `"warn"` until v0.215.0, and its own doc
comment records that every Figyelmeztetés-level disk alert the product ever produced went to nobody.
The lesson was written down. Nothing enforced it. It happened again — and it hid for months only
because R-384's ordering defect meant the event could not fire at all, so a broken severity had
nothing to break.
**The guard is now an AST walk over the whole controller**
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`), not a comment and not a `strings.Contains`:
the word "warn" appears legitimately nine times in `internal/monitor` and `internal/selftest` as a
*healthcheck status*, and only one of those hits was the defect.
**The sweep found exactly one bad severity and is reported in full**, including its limits. The walk
cannot follow a variable, so the six call sites that pass one are **registered by name with the values
each can take** — a new dynamic site fails the test rather than slipping past it. Two of those six
were found by the guard itself, not by the hand sweep that preceded it.
**A latent one, found while sweeping and pinned rather than left:** `fillwatch.Band.Severity()`
returns `""` for `BandOK`, which would coerce to `info` and vanish. It is unreachable — `Check()`
notifies only on an *escalation* — but that safety lives in a different function from the one that
looks unsafe. `TestR329_FillwatchNeverEmitsTheEmptySeverity` pins the **consequence**, not the mapping.
### R-329, part two — the app-down alarm gets a customer switch, DEFAULT OFF
`app_start_failed` is now a customer-facing toggle („Alkalmazás nem fut"), and is **deliberately NOT
in `DefaultEnabledEvents`** (operator ruling, 2026-08-23).
**The operator is emailed either way.** `processOperator` consults only `operatorOn`, the address and
a one-hour cooldown — never customer preferences. The toggle governs the customer leg alone. It is
also deliberately **not** added to `operatorOnlyEvents`: that would make the switch visible,
flickable and structurally incapable of delivering, which is worse than not offering it.
### R-386 — ask the field that knows, instead of guessing from the state
`classifyRunStates` decided "the customer stopped this" from `st.State == StateStopped`. **Every**
stopped stack was therefore assumed deliberate. Measured live on `demo-hp` 2026-08-23: `privatebin`
stopped out of band, nine dead-app scans over four minutes, **zero events and zero banner lines** —
while the comment beside the code claimed an out-of-band stop *"still alerts"*.
The product already records the answer. `DesiredState` is what the customer asked for, it has exactly
**one writer** (their own action), and it is tri-state. The ruling, operator-approved:
| Intent | Verdict |
|---|---|
| `Stopped` | the customer asked → **no alarm** (unchanged) |
| `Running` | nobody asked → **ALARM** (the fix) |
| absent | **UNKNOWN, never "running"** → no alarm, **and say so** |
**The unknown case keeps today's behaviour deliberately.** Reading absent as "nobody asked" would, on
the first cycle after upgrade, email about every app any owner has ever deliberately stopped —
fleet-wide, from a field that predates the intent being asked of it. The backfill cannot help: it
seeds `Running` only from an observed-UP reading, so anything stopped at upgrade time stays unknown —
which is exactly the ambiguous population.
**The gap is BOUNDED, not silent.** Every suppression sets `AppRunState.IntentUnknown`, and the
scheduler logs the names at INFO on the heartbeat cadence, so an operator can answer *"how many apps
am I blind to, and which?"*. **A rule without a mechanism is a wish** — the log line is the mechanism.
It closes itself as apps are started and stopped through the interface.
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: removing it
re-opens F-CRIT-1's indefinitely-silent dead app. Every existing suppression — quiesce grace, boot
grace, crash-loop threshold, the `Deploying` skip — is untouched. **No new `DesiredState` writer was
added**; the field keeps its single owner.
### Two settings toggles secretly governed two alarms each
„Lemez figyelmeztetés (90%+)" also wrote `disk_critical` — the drive-is-**failing** alarm. A customer
silencing a disk-nearly-full notice silenced "this drive is dying", and the label claimed only the
first. „Elvárt mentés elmaradt" had the same shape across the file-backup and database-dump misses.
Each is now two honestly-labelled toggles; **12 toggles became 15.**
**The risk was never the split, it was the migration.** Every stored list was written by the OLD form
names. A no-op save now travels a different path, and a settings page that rewrites a setting while
merely rendering it would be worse than the defect. So a save whose event **set** is unchanged stores
the **existing slice verbatim** — byte-identity by construction, not by argument. The legacy compound
form names are still read, so a stale browser tab cannot drop a key.
**This was not theoretical:** with the guard removed, the `defaults` case reorders. The red-proof
caught it.
### Tests
`internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
Test count **1504 → 1522**.
**Red-proofs: five planted, five seen failing — and ONE PASSED FIRST TIME AND IS REPORTED.** The
fillwatch mutation `next <= prev` → `next < prev` is **inert**: an earlier `if next == prev { continue }`
had already removed the equal case, so the code's behaviour did not change and the test was right to
pass. Removing the de-escalation guard outright convicts it. **A red-proof that passes needs the
mutation checked before either verdict is believed.**
## v0.222.0 — an app whose database dies raised no alarm, because the wrong question answered first (2026-08-23, R-384 + R-383)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**`IsDownState` did NOT change.** That is the first thing to say, because the obvious fix here is the
wrong one. `unhealthy` is still excluded from the down set, byte-identical, for the reason recorded at
`manager.go:41-53`: an unhealthy container is *running*, and folding it in reintroduces the flapping
that exclusion exists to stop. No new container state was minted either — `StateDegraded` already
means exactly this and every consumer already handles it. **The defect was the ORDER of two questions,
and only the order moved.**
### R-384 — a dead database hid behind its own unhealthy front end
"Is a SUPERVISED member of this app dead?" and "is a RUNNING member failing its healthcheck?" are two
different questions. `aggregateState` returned `StateUnhealthy` the moment `unhealthy > 0`, and the
R-51 mixed-case block that asks the first question sat **below** it — so the second question was
answering the first, and always won.
A two-container app whose database exits goes unhealthy seconds later *because* it cannot reach that
database. So the very symptom the dead database causes was what suppressed the alarm for it.
`unhealthy` is not a down state, so `classifyRunStates` never marked the app down, no banner appeared
and `app_start_failed` never fired.
**Measured live on `demo-hp` 2026-08-22:** `bookstack-db` stopped at 21:27:01 and the F-OBS heartbeat
printed **"0 currently down"** throughout. This is R-51's own 18-hour immich failure returning through
a different door — the dead member hiding behind an *unhealthy survivor* instead of behind three live
helpers.
**Two things had to move, and either one alone leaves the defect standing:**
1. **The order.** The supervised-down test is hoisted above the `unhealthy`/`starting`/`restarting`
returns.
2. **The guard.** "Some members are up" now means **any member not in the down bucket** — running,
unhealthy, starting or restarting. The old guard was `running > 0`, counting `StateRunning` alone,
which made the R-51 block **unreachable in exactly the case it was written for**: an unhealthy
survivor beside a dead database counted as nothing up.
**The benign case is untouched.** A one-shot init/migrate container that has finished has policy
`no`/`on-failure`, `supervisedPolicy` returns false, and the stack reads exactly as before. Without
that filter every app with a migration step would alarm on every start.
**The priority comment was rewritten**, because it asserted an ordering the code no longer has, and a
comment asserting an invariant the code does not provide is what cost this project four months one
version ago.
**Three existing subtests were AMENDED, and this is reported rather than buried.**
`TestAggregateState_UnchangedBranches` asserted that an unhealthy / starting / restarting member beat
an `exited` peer that was on `unless-stopped` — that is, it pinned the defect as if it were settled
behaviour. They keep their original intent ("the live member's state wins over a down member") with
the down member given a BENIGN policy, which is the only situation in which that sentence was ever
true. The supervised versions now assert `degraded`.
**Every `IsDownState` consumer was walked and is named in `REPORT.md`.** Two change deliberately
(`classifyRunStates`, the intended fix; and `bootrecon`, which will now repair a half-started stack at
boot instead of calling it recovered). `isObservedUp` is an allow-list of `{running, starting}` and is
unaffected — verified, not assumed. The quiesce suppression is cycle-keyed and state-blind, so
R-97b's guarantee is untouched.
### R-383 — the double-failure message promised an undo copy it never looked for
When BOTH a database replay and its rollback fail, the message ended „a korábbi állapot mentése
megvan: <file>" — *the previous state's backup exists*. It was built from the path `writeSafetyDump`
returned, **without ever asking the filesystem**. One of the two ways the rollback fails is that the
file is gone, so the sentence was most likely to be false in precisely the case it was printed.
Measured twice live, on v0.220.2 and v0.221.1.
`undoCopyPhrase` now describes the undo copy **from disk**: present (named), partially present (both
halves named), missing (says so, and still names where it should have been), or never written. A
zero-length dump counts as missing — a 0-byte file restores nothing. The filename is **not** simply
dropped: R-351's lesson is that a refusal naming nothing forces a person to remember what the product
already knows.
### Tests
`internal/stacks/degraded_test.go` (R-384 unit + production-path wiring),
`cmd/controller/r384_dead_db_alarm_test.go` (the classifier consequence + the quiesce suppression),
`internal/backup/r383_undo_phrase_test.go` (the phrase + an AST seam test that the message is still
wired to the builder). Test count **1494 → 1504**.
**Red-proofs: four planted, four SEEN FAILING.** The two halves of R-384 convict independently —
reverting the hoist reads `"unhealthy"`, the exact state the live box reported, and narrowing `up`
back to `running` reads `"unhealthy"`/`"starting"`/`"restarting"`. The classifier mutation returns an
empty banner. The R-383 mutation prints the false claim verbatim, and for an empty set printed
`megvan: .` — naming a file that never existed. Guards sit at the layer each defect lives in: the
ordering at `aggregateState`, the consequence at `classifyRunStates`, the claim at the phrase builder.
## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**This version was BUILT, BAKED and VOUCHED on 2026-08-23 and had no CHANGELOG entry of its own until
now (R-385).** Its fix was written inside the v0.221.0 entry instead, so the record named a version
that was not the one running. Nothing is wrong with the image; the record was. The reasoning below is
moved here verbatim from that entry — it is not new text, and v0.221.0 no longer claims it.
**One change made another unreachable, and only the live box showed it.** Excluding the undo copies
from `db_dumps` made that list STABLE across restores — which is correct — and that made
`CaptureRecoveryUnit`'s already-current early return start firing where it never had. The undo-copy
prune sat *after* that return, so the cap silently stopped applying: **four copies on disk against a
cap of three, counted on `demo-hp` minutes after the change.** The prune now runs above the check,
where it belongs — it is housekeeping on the dump directory and has nothing to do with whether the
manifest needs rewriting.
**Why it is safe above the check:** pruning cannot disturb `dbDumps`, which after v0.221.0 no longer
contains those names.
**Test:** `TestR361_UndoCapHoldsWhenTheUnitIsAlreadyCurrent`
(`internal/backup/r361_canonical_dump_test.go:245`). **Red-proof:** move the prune call back below the
already-current early return — the cap fails at 5 copies against a cap of 3.
**Shipped as commit `810b18a`.**
## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-361)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**A comment asserted an invariant the code did not have, and the comment was believed for four
months.** `writeSafetyDump` called `DumpOne` into the app's OWN unit directory and renamed the result
to `pre-restore-*` afterwards. `DumpOne` writes `<stack>-<dbtype>.sql` — **the app's canonical dump,
the name the replay loop matches exactly** — so every safety dump **overwrote the app's real backup
and then moved it away**. The comment beside it said the rename meant it *"can never overwrite the
app's real dump"*. It was false as written.
**The consequence, not the mechanism:** between a restore and the next nightly run the app had **no
database backup of its own**. A local restore-from-unit in that window finds no `.sql` and tells the
customer the app never had a database.
**Measured live before the fix, on `demo-hp` 2026-08-22:** `docmost` and `bookstack` each held only
`pre-restore-*` files in their unit and **no canonical dump at all**.
**The fix is a destination.** `DumpOneTo` takes the final path and derives its own `.tmp` from it;
`DumpOne` keeps its signature and calls it with the canonical name, so its other callers do not move.
`writeSafetyDump` now asks for `pre-restore-<stamp>-…` **directly** and the rename is gone. The
comment states the invariant and how it is enforced.
**The scratch file matters too:** `.tmp` is derived from the FINAL path, so a nightly dump and a
safety dump running into the same directory cannot share it.
**The manifest no longer lists the undo copies** (`db_dumps`). They are local material for a restore
that went wrong, not part of the app's recovery set. **Every consumer of `Manifest.DBDumps` was
grepped and named: there are three, all inside `recovery_unit.go`** — the declaration, this
enumeration, and the change-detection compare. Nothing reads it for recovery; no hub or agent
consumer exists. Excluding them also makes that compare stable, since the copies come and go with
every restore and prune. **The files are neither deleted nor hidden** — their visibility is a
recorded design decision and it stands.
**Tests:** `internal/backup/r361_canonical_dump_test.go`, additions to
`internal/appbackup/r381_undo_naming_test.go`, `cmd/controller/r361_classifier_control_test.go`.
Test count 1485 → 1493.
**Red-proofs: five planted, and TWO PASSED FIRST TIME — both reported.** The behavioural tests inject
the dump seam, so a mutation *inside* `DumpOneTo` was invisible to them; and Part 1.3 initially had no
test at all. Guards were added at the layer each defect lives in and both mutations then convicted.
**The assertion that convicts R-361 is the canonical dump's bytes, unchanged, across a restore** — a
test asserting merely that the undo copy exists passes just as well when the app's backup was
destroyed.
## v0.220.2 — the operator route to clear a hold now says the restart is required (2026-08-22, R-379)
**MinAgent: 0.129.0** (unchanged)
`--clear-restore-hold` runs as a SECOND process. It clears the hold in `settings.json`, but the
RUNNING controller holds its own in-memory `Settings` and keeps refusing until it reloads. Measured on
`demo-hp`: the clear succeeded, the file was correct, and the customer's start button still refused —
until the controller was restarted, after which the app started normally.
The command now prints the restart it needs. **A route the operator believes worked, and did not, is
worse than no route**, and this was found by using it rather than by reading it.
Recorded with it: there is a lost-update window while both processes hold the file. Restarting
promptly closes it. Clearing through the running controller would remove both problems and is the
right shape later; it needs an operator tier the controller's HTTP surface does not have today — it
authenticates as the customer, and a customer clearing their own hold is what the hold exists to
prevent.
## v0.220.1 — the rollback poured the undo into a container that no longer existed (2026-08-22, R-379)
**MinAgent: 0.129.0** (unchanged)
**Found by v0.220.0's own live walk, on its first real run, an hour after it shipped.** The undo FILE
is stable; the container it must be poured into is not. `writeSafetyDump` captures its
`DiscoveredDB` **before** the stop, and by the time the rollback runs the stack has been stopped and
the DB service re-created with a new id.
Measured on `demo-hp`: `docmost-postgres` was captured as `9adbc14f9af6` at 16:05:44, re-created as
`309795897b82` at 16:05:47 by the DB-only start, and the rollback's `docker exec` against the dead id
sat in `waitDBReady` until it timed out 30 s later. **So the app was HELD for an infrastructure reason
while its data was perfectly recoverable** — the hold worked exactly as designed, on a case that
should never have reached it.
The rollback now re-discovers the containers and matches each undo file to a live one by
`{stack, engine}`, which is what `reimportDBDumpsFrom` already did for the replay. A database whose
container cannot be found fails **closed** rather than pouring an undo into something unidentified.
**Why no unit test caught it:** every rollback test injects the import seam and never looks at
container identity. The new one asserts the identity handed to the import, and its red-proof — using
the captured id again — convicts.
## v0.220.0 — when a database restore fails, the customer's own copy goes back (2026-08-22, R-379/R-380/R-381/R-382)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**R-379 and R-380 were ONE failure and they get ONE fix.** Both ended with a half-restored database.
The only difference was whether it looked broken: on Postgres the database was emptied and the app
crash-looped; on MariaDB part of the dump applied, the rest did not, and the app reported
`health=healthy, running=true, restarts=0` with its schema-version table holding zero rows. Measured
live on `demo-hp` on 2026-08-22 (`audits/DRILL-r356b-driveless-db-restore-2026-08-22/`).
**The undo copy was already being taken, and was already good.** It was proven good by hand that day
on both engines — `docmost` recovered, `bookstack`'s `migrations` went 0 → 102 rows. What no product
action could do was apply it: `pre-restore-` files are skipped at three call sites so they are never
mistaken for a replay source, and the filename appeared only inside an error string. **The fix is the
product making the same `ImportDump` call a person made by hand.**
**When the replay fails, the undo set is now re-applied automatically**, before any restart and with
the DB service still up, so the app never observes the half state. The app then starts and the
message says **both** things: the restore failed, *and* the data is back as it was. A message that
reported only the failure would leave the customer believing their data was gone when it is not — the
omission of a gain misleads exactly as much as the omission of a loss.
**THE WHOLE undo set, not the first file.** `writeSafetyDump` returned one path for an app with two
databases; a rollback built on that would have restored one and left the other half-written — this
defect, one database over. It now returns the set, matched on **this run's stamp**: four `pre-restore-`
files accumulated on one app in one afternoon, so a prefix match would replay an arbitrary older state.
**When the rollback ALSO fails, the app is HELD STOPPED — operator ruling, 2026-08-22.** A running app
on a half-written database lets the customer type into it and turns a recoverable state into a
permanent one. The hold is persisted, every start path refuses it (the customer's button, the app-stop
`Recover()` starter, and the boot sweep — through the shared `driveStartGate`, checked **above** its
driveless early return because these apps have no drive), the app-stop marker is ended so nothing
auto-restarts it at the next boot, the row goes **red** rather than green, and the operator is told.
Clearing it is `--clear-restore-hold <app>`, an operator CLI route reached through `docker exec`.
**`--single-transaction` is a BELT, not the fix.** Added to the Postgres import so a replay is
all-or-nothing at the engine. **It does NOT make MariaDB atomic** — MariaDB's DDL is not
transactional, so a partial apply there is unavoidable at the engine level. That is precisely why the
rollback exists, and this flag does not make it optional.
**R-381 — the failure message stops pasting the engine's output at the customer.** It was 407 bytes on
Postgres (a caret diagram and `exit status 3`) and **615 on MariaDB, whose middle was an `INSERT INTO
migrations VALUES (…)` listing — rows out of the customer's own database, HTML-escaped, on their
dashboard.** The full engine text now goes to the operator log, which never had it before: the
diagnostic is **added**, not removed.
**R-382 — the summary log line prints the volume count it already held.** It said "0 file(s) placed,
1 DB dump(s) replayed" on a run that returned a 52 MB Postgres data directory.
**The undo copies are named for what they are, and bounded.** `pre-restore-20260822T140924Z-docmost-postgres.sql`
used to derive the phantom stack `pre-restore-20260822T140924Z-docmost`; it now resolves to `docmost`
and carries `IsUndo`. **The reported symptom — that they render as apps on the customer's backup page
— did NOT reproduce**: the live page was read first and contained zero `pre-restore` strings, because
`buildAppBackupRows` iterates deployed apps and only reads that map by key. The phantom name was real
as a map KEY; the row was not. Their VISIBILITY is unchanged and deliberate. A cap of **3 per app**
now applies, pruned from the capture side and never from the restore path — a delete on the failure
path is how an undo goes missing at the moment it is needed.
**Tests:** `internal/backup/r379_rollback_test.go`, `cmd/controller/r379_hold_gate_test.go`,
`internal/appbackup/r381_undo_naming_test.go`. Test count 1468 → 1483. **Eight red-proofs; one PASSED
and is reported rather than omitted** — the R-381 behavioural test injected below `ImportDump` and so
could not see a leak reintroduced inside it. A guard at that layer was added and the mutation then
convicted.
## v0.219.0 — the off-site restore refused every app that has no data drive (2026-08-22, R-356)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**What broke.** Forty of the fifty-three catalogue apps could not be restored from the off-site copy
at all. The customer opened the restore page, pressed the button, and was told the app **"nincs
telepítve"** — while it was running in front of them — and instructed to reinstall it "ugyanerre a
helyre", a place those apps never offer. The instruction could not be followed, so the restore never
started.
**Why.** `ReconstituteFromOffsite` and `PlaceOffsiteRestore` both resolved the restore destination
with the RAW `HDD_PATH` (`stackProvider.GetStackHDDPath`) and read an empty answer as "the app is not
installed". One predicate was answering two questions. Measured in the catalogue at `459766cb1639`:
**53 templates, 13 declare `needs_hdd: true`, 40 declare `false`** — and for those 40 the answer is
*correctly* empty, permanently. The capture side never had this defect: `CaptureRecoveryUnit` resolves
via `GetAppDrivePath`, which falls back to the system data path — which is why the 2026-08-21 refusal
was able to print `/mnt/sys_drive`, a destination the backup had recorded and the restore refused to
use.
**The fix separates the two questions.** *Installed?* is now asked directly, of
`ListDeployedStacks()`, through a new `Manager.isStackDeployed` that fails CLOSED on a nil provider —
"cannot tell" must not become "go ahead" when the caller's next act is a write. *Where?* is then
answered by `GetAppDrivePath`, the same resolver the capture side wrote the snapshot with, so the
restore aims at the place the backup came from.
**The 13-class behaviour is unchanged.** A drive app still resolves to its own drive, the placement
mismatch check still fires when the recorded drive differs, and `ackPlacementChange` is still required
to pass it. That claim is not an absence claim: the red-proof plants the system-data fallback into the
drive path, and three tests convict it.
**A third refusal now exists, with its own sentence.** Installed, driveless, and the box cannot name
its own data root ⇒ refuse and name the storage page. Widening `nincs telepítve` to cover this would
send a customer to reinstall a running app and hide the real fault.
**Why now.** v0.218.0 (R-354) taught the reconstitution to replay an app's named volumes. Those forty
apps keep **all** of their data in exactly those volumes, so that fix could not reach the apps that
need it most until this one shipped.
**Changed:** `internal/backup/backup.go` (`isStackDeployed`), `internal/backup/offbox_reconstitute.go`,
`internal/backup/offbox_restore.go`. **Not changed, deliberately:** `offboxCaptureSet`'s raw
`GetStackHDDPath` call — capture resolves an app's declared `userdata`/`import` file legs against that
value, and a system-data fallback there would write a snapshot claiming to hold files it does not.
**Tests:** `internal/backup/r356_hot_only_restore_test.go` (Scenarios A–E, both entry points, and the
positive/negative control on the deployment predicate) and `cmd/controller/r356_deployed_seam_test.go`
(AST walk over the production adapter wiring). Five companion red-proofs, each SEEN failing.
**Fixture correction in the same commit:** several existing fixtures marked an app "installed" by
giving it an HDD path. That is the conflation this change removes, so they now state deployment as its
own fact. No assertion was weakened.
## v0.218.0 — the database nobody backed up, and the restore that returned most apps nothing (2026-08-22, R-354/R-355)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
Both fixes were found by watching a real machine on the night of 2026-08-21, and both are confirmed
the same way. Neither is a refactor: each removes a case where the product told a customer something
that was not true about their own data.
**The database that was dumped and then abandoned (R-355).** `paperless-ngx` runs its PostgreSQL in a
container called `paperless-postgres`. `deriveStackName` (`internal/appbackup/dbdump.go:770`) strips the
role suffix to `paperless`, finds that is not a deployed stack, finds no known stack is a prefix of the
container name — **and then returns the unresolved candidate anyway**. The nightly dump therefore landed
in `backups/primary/paperless/db-dumps/` — a directory for an app that does not exist, on the SYSTEM
drive — while the app's own recovery unit, on the data drive, recorded `db_dumps: null`. Nothing
collected it, nothing off-sited it, nothing restored it. Measured live: **284 617 bytes, 72 tables,
valid, unreachable**, and reproduced unattended by the box's own 02:30 cycle.
**And it compounded, which is why it went first.** `writeSafetyDump` filters the discovered databases
with the same value, so a destructive restore of that app found no database, **took no undo copy**, and
the fail-closed refusal that protects every other app could not fire — it was never reached. A customer
could press restore, lose the live database and have no copy anywhere. Verified live on 2026-08-21:
`find /mnt -name "pre-restore-*"` was empty both before and after a full restore over a live 72-table
PostgreSQL.
**The fix is to stop guessing.** Every container the controller starts carries
`com.docker.compose.project`, and that label IS the stack name by construction: compose is run with
`cmd.Dir` set to `/opt/docker/stacks/<stack>` and no `-p` (`internal/stacks/manager.go:1218`). The label
is now read and preferred whenever it names a deployed stack; the old derivation stays as the fallback
for containers not started by compose, and an attribution that resolves to no known stack is now **loud**
instead of silent. **A catalogue-wide sweep — proven able to convict by planting a second mismatch,
watching it caught, removing it and watching it clear — reports exactly one affected app of 53.** The
fix is in the controller, not the catalogue: renaming the container would have fixed this one app and
left the guessing in place for the next.
**The restore that returned most apps nothing (R-354).** `ReconstituteFromOffsite` skipped every
placement flagged as the unit, and the named-volume archives live INSIDE the unit — so the off-site
restore had a files leg and a database leg and **no volume leg at all**. Proven live with planted,
hash-recorded files: calibre-web's `calibre_web_config.tar` was in the unit, in the off-site snapshot
and in the verification folder, and the restore returned the five declared files, reported success, and
did not replay it. **For the 40 of 53 catalogue apps that declare no data drive, that archive is
everything the customer owns.**
**The comment beside the skip was half false and is corrected rather than left.** It justified the skip
by saying the snapshot's dump is replayed from the scratch unit so nothing is lost — true of the
database, false of the volumes, and the reason a reader would not look. **The half that still holds is
named:** the live unit is the LOCAL restore path's own source and must never be clobbered. Scenario D
now fingerprints the whole live unit across the operation and compares.
`restoreDockerVolumesFrom` is the local path's own replay with an explicit directory — the same shape
`reimportDBDumpsFrom` already had, and deliberately ONE implementation with two callers, because a
second copy of that loop is what produced the divergence. Volumes replay **before** the database, so a
logical dump still wins over a volume-tar copy of the same database, and inside the stopped window,
because Docker will not replace a volume a container holds.
**And it reaches the sentence.** `VolumesReplayed` is on the result and in the message: „5 fájl **és 1
adatkötet** visszaállítva". A restore that replayed an app's entire dataset and mentioned only its file
count is how a silent loss reads as a success. A snapshot with no volumes produces the byte-identical
sentence it produced before.
**„Ennek az alkalmazásnak nincs adatbázisa" is no longer inferred from a counter.** `DBsReplayed == 0`
has two causes — the app has none, or it has one and the snapshot carried no dump — and both printed the
same confident sentence over a live 72-table database. The undo copy is the honest discriminator, and
the second case now says so and names the undo.
**Seven red-proofs, each mutation asserted applied and reverted.** The name fix reverted returned the
empty database record with both divergent paths printed; the message predicate reverted returned the
false „nincs adatbázisa" sentence verbatim; the refusal removed was seen letting a restore proceed with
no undo; the volume leg removed returned the silent loss; the volume count dropped returned the
true-but-incomplete sentence verbatim. **Two of them found defects in the tests rather than the code:**
scenario D PASSED with the unit guard removed, because the fingerprint had been narrowed to the volume
directory and was blind to a placement writing into the unit root — the R-181 class, in my own test, and
the reason the red-proof is mandatory.
## v0.217.0 — the restore knows where the data lived (2026-08-21, R-351/R-352/R-353)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
**Found by walking the screens on a rebuilt `demo-hp`, not by reading them.** A person restored an
app onto a rebuilt machine and had to remember two things the backup already held: the web address
and the data folder. Neither can be changed after installation without deleting the app and its data.
**The blindness (R-351a).** Every recovery unit's `manifest.json` has carried `drive` and
`namespace_root` since schema 1 (`internal/backup/recovery_unit.go:48-49`), written at capture from
the app's own live placement. `grep -rE '\.Drive\b|\.NamespaceRoot\b' --include=*.go` found **no
non-test reader anywhere in the repository**. The reconstitution opened that very manifest
(`offbox_reconstitute.go:235`) purely for the coherence stamp, then resolved its destination from the
LIVE app instead. **A restore into a destination different from the one the backup recorded therefore
succeeded silently, under a green message.** One side wrote the fact; the other never received it.
Now: `internal/backup/offbox_placement.go` — `CheckPlacement` (pure, total), `RecordedUnitForStack`,
`PlacementMismatchMessage`. The comparison happens **before the safety dump and before the first
byte**. A mismatch is **named** — both values, never "a destination differs" — and refused; the
customer may proceed deliberately via `ack_placement`, a **separate** field from `confirm=1`, because
one click must not carry two decisions. An **unknown** recording is never a mismatch: refusing on an
absence would strand every pre-field unit. The not-installed refusal (R-253) now names the recorded
drive. The deploy page prefills the address and folder **from the app's own backup**, labelled as
such, and still editable — a memory, not a lock.
**The second press (R-351b).** All seven restore handlers gated on `backupMgr.IsRunning()` — the
CONCURRENCY flag, which the restore goroutine acquires *inside* itself (`offbox_reconstitute.go:180`)
**after** the handler has returned. Established with a test before any change: the reconstitute and
place handlers both answered „…elindult" and **overwrote the first restore's op and stack**. The
wizard had read the correct flag since v0.154.0 and explained why in a comment; the handlers were
never moved over. `Server.restoreOpBlocked()` now reads **both** flags — the display flag covers the
whole off-box restore, the concurrency flag is the only one the nightly backup holds.
**The invisible result (R-351c).** The page *does* refresh; the defect was the RESULT. The banner
gated its terminal state on a page-local `sawRunning`, so a restore that finished before the page was
opened — or inside one 3 s poll — was shown to **nobody**. The 2026-08-21 OpenGist restore took
**8.666 s** and no screen ever said it completed. `RestoreOpStatus.LastRecent` now carries the
server's verdict, and `RestoreResultWindow` moved to `internal/backup` with `internal/web`'s constant
as an alias: **one expression, two surfaces.** The wizard's self-contradiction — that the state
refreshes automatically *and* that you must refresh the page — is gone.
**Where the data goes, said out loud (R-352).** Measured: **40 of 53** catalogue templates declare no
data path, and for those the deploy page offered no storage field and the configured default was
never consulted — `GetDefaultStoragePath()` has three non-test callers and **none of them places
data**. The deploy page now **states where the app's data will live before the button is pressed**.
**Visibility only: no placement changed and nothing was migrated.** The specification for the rest is
filed at `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
**The 16-second page (R-351d).** Measured on the live off-site target before changing anything:
`snapshots --json` 2605 ms once, then `stats` **2697 ms per app, sequentially** — 2605 + 5×2697 ≈
**16.1 s**. The per-app size calls now run concurrently, **bounded to 4**: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns 0, which silently *under-reports* the
customer's data rather than failing visibly. The bound is asserted by test, not just the speed.
**Red-proofs, each mutation asserted applied and reverted.** B with **both** guards removed **was
seen starting a restore with no drive attached** — no error, full 3.00 s run, writing into
`/tmp/mutant-destination`. C returned the silent divergent restore; E returned the fabricated empty
prefill; A returned the blank form; D forced on broke **8** ordinary reconstitute tests, proving the
guard is reachable in both directions; Part 4 reverted to sequential failed the concurrency assertion.
**Known and NOT fixed here, filed as R-353 and named as the next session's first item:** a restore
whose unit carries no `db_dumps` and no `volume_dumps` reports a bare completion. The OpenGist unit
contained configuration and nothing else, and the outcome said only that it had finished.
## v0.216.0 — one physical disk, one verdict (2026-08-14, R-335)
**MinAgent: 0.129.0**
**Found on live hardware within two hours of the v0.215.0 deploy, by reading the artefact the release
itself introduced.** The first two hourly checks on demo-hp logged *"3 disk(s) evaluated"* while the
persisted `disk-health-state.json` held only **two** records. The discrepancy was the bug: demo-hp's
`c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same
`diskKey`, so one disk was walked twice in a single run.
**Why that was not cosmetic.** The loop writes a disk's new record before the next entry reads it, so
the SECOND copy of an aliased disk consumed the FIRST copy's write as its prior. The disk therefore
**sustained against itself** and reached **Hiba on a FIRST sighting** — defeating truth-table row 6,
the rule the whole v0.215.0 ladder is built on and the one thing standing between a one-hour benign
excursion and a false critical alert. It would also have emitted **two identical events** for one
drive. Nothing fired on demo-hp because all three entries are healthy with zero counters, so the fault
was latent, not active — but any aliased disk developing a single pending sector would have gone
straight to Hiba.
**Fix:** `RunDiskHealthCheck` evaluates each `diskKey` **once per run** (`internal/web/disk_health.go`).
Both entries are still marked `seen`, so neither is mistaken for a disappeared disk, and the card still
renders **both** storage rows — the dedup is about state and alerts, not display.
**Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`**, which asserts the first sighting stays
Figyelmeztetés and silent, that the stored prior is not this run's own write, that the second run emits
exactly ONE event, and that the card still shows two rows. Companion red-proof run and reverted:
deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors — the exact
false critical described above.
This is the shape §3 of the workspace rules warns about: the release's own positive observable
("N disks evaluated") disagreed with its own persisted artefact, and only reading BOTH exposed it.
## v0.215.0 — the disk alert that never sent (2026-08-14, R-328..R-333)
**MinAgent: 0.129.0** (unchanged — every field this reads has been on the wire since agent v0.94.0/0.95.0)
**Lead with the one-word defect, because it is the one that decides whether anything arrives at all.**
`NotifyDiskHealthDegraded` set `severity := "warn"`. The hub accepts an exact-match lowercase
vocabulary — `{info, warning, error, critical}` — and **coerces anything else to `info`**
(`felhom.eu/hub/internal/api/handler.go:2121-2126`); `severityNotifies`
(`felhom.eu/hub/internal/notify/dispatcher.go:89-96`) then routes only warning/error/critical. `"warn"`
is not in that set. **Every Figyelmeztetés-level disk alert this product has ever produced was filed as
an informational notice and emailed to nobody**, on the customer leg and the operator leg alike. The
function's own doc comment said *"The hub applies its own per-event-type cooldown"* — an invariant
asserted in a comment, with nothing pinning it. Now `"warning"`, and `DiskAlertKind.Severity()` is
exported so the contract is checkable from any package rather than duplicated as a literal.
**PROVEN LIVE, side by side, through the real hub endpoint** (2026-08-14, demo-hp, synthetic events):
pushed at `"warning"` → stored `warning`, `notification_log` id 689, channel `operator`, status
**`sent`**. The same push at `"warn"` → stored **`info`**, and **no `notification_log` row exists at
all**. That pair is the whole task in two rows.
**A drive can never fail its own verdict on bad sectors, so the ladder stopped trusting it.** Attributes
187/197/198 all carry `thresh: 0` and a normalized SMART value floors at 1, so `smart_status.passed`
is **structurally incapable** of failing on unreadable sectors — the real drive stayed `PASSED` at 352
pending sectors with 1001 reported-uncorrectable reads. `DiskVerdictFor` now takes an
`agentapi.DiskPrior` and implements a 14-row top-down truth table (`internal/agentapi/diskverdict.go`).
Hiba is reached by: sustained unreadable sectors (present again at the next check), unreadable +
remapping together, a count >= 64, overheating (>= 60 °C), NVMe's own critical flag, or spent rated
endurance. **No fourth label** — predicted failure is **"Hiba"**, the word a broken drive already gets;
a fourth word sharing a root with "Figyelmeztetés" would make the more severe state read as the milder
one. `DegradedAttributes` now names the counters behind a Hiba reached from counters (it returns nil
only for a drive-reported FAILING, which has no single triggering counter).
**Provenance of the numbers, recorded in the code because a number without a reason becomes permanent.**
64 — the observed benign excursion peaked at 16 and cleared completely inside an hour; the terminal run
passed 64 at 13 Aug 11:28 and never came back. Sustain sits ABOVE the count in the table because on the
real drive it fires a full day earlier; the count is the backstop for a box that was powered off across
the sustain window. 55/60 °C — the operator's existing Prometheus bands on DooPlex, adopted unchanged.
**It spoke once, and forgot on restart.** The baseline was in-memory, so a controller that restarted
while a disk was failing re-baselined it silently and never alerted again; and between 8 and 352
sectors the verdict never changed level, so nothing further was emitted. State is now persisted to
`disk-health-state.json` in `cfg.Paths.DataDir` (atomic tmp+rename, the `selfupdate.SaveState` shape;
a missing file is normal, a corrupt one is LOGGED and treated as no-prior — never fatal). The decision
compares against the **last ALERTED** verdict, not the last observed one, which collapses a
Warn→OK→Warn flap to a single alert while leaving a genuine escalation free to fire immediately; and a
disk already at Hiba re-alerts once it has BOTH doubled its unreadable-sector count AND waited out a
24h cooldown (an AND — clearing one bar emits nothing).
**The card replays the prior that produced the verdict, not the one the check went on to write.**
`diskRecord.PriorSawUncorrectable`. Without it the dashboard chip read one level MORE severe than the
alert for the same disk — caught by the Scenario B test, not by review. The shared-verdict-function
guarantee is now pinned rather than asserted.
**Hourly, and that was measured rather than preferred.** On demo-hp the real `/disks` fetch costs
min 0.805s / median 0.821s / max 0.841s over 10 calls, all HTTP 200, 3 physical disk rows — 6x under
the 5s bar that would have kept it at 6h. The benign excursion lasted about ONE HOUR, and a 6-hourly
sampler can land either side of one and then catch the terminal run half a day late.
`cmd/controller/main.go`. The check also logs a POSITIVE observable every cycle
(`disk-health check complete: N disk(s) evaluated, M alert(s)`) — zero events from a check that never
ran is not evidence of health.
**Message shapes.** A customer told "the drive overheated" must not be told to arrange a replacement.
Five shapes: Figyelmeztetés (wording unchanged, now actually delivered), drive-reported FAILING,
Hiba-from-sector-count, Hiba-from-heat, and the still-worsening re-alert.
**Tests.** 12 scenario groups A-L, 1391 -> 1413 test functions, suite green. **11 of 12 companion
red-proofs failed as required; red-proof A PASSED and is reported as a finding** — removing truth-table
row 6 leaves the real drive tripping row 8 at 352 sectors, so that mutation cannot fail a test built on
the real drive's values. Row 6 IS pinned, by `TestLadder_SustainIsWhatFires` and
`TestDiskLadder_SustainDrivesTheEscalation`, which hold the counters at 8 and vary only the prior; both
fail under that mutation. Group L builds the Server through `web.NewServer` — the same call
`main.go` makes — over a real file, and runs TWO checks after the restart, because one cannot
distinguish a loaded state from a silent re-baseline.
**Not live-validated:** the Fail-from-counters path has never fired on real hardware (R-332).
Evidence: `felhom.eu/documentation/audits/DIAG-smart-passed-trap-2026-08-14.md` + two fixtures.
## v0.214.0 — the recovery screen stops hedging about a code it can now check (2026-08-12, R-311)
**MinAgent: 0.129.0**
**What was already right, and is worth saying first.** The screen did NOT bluntly accuse a customer
holding an older code: R-222/R-226 already hedged, naming both possible causes and the kept package.
That sentence was honest — *"innen nem tudjuk megkülönböztetni őket"*, we cannot tell them apart from
here. **It could not tell them apart because nothing ever looked.** Agent v0.129.0 looks, so the hedge
can become an answer.
**New failure class `RecoveryCodeOpensRetained`** on HTTP 422, gated by `FeatureRetainedRecoveryClass`
(MinAgent 0.129.0). The gate is the R-224 twin and is SEPARATE from `trustRefusal` on purpose: the two
name different agent versions (0.126.0 and 0.129.0) and a box can sit between them, where a 422 is a
shape we did not design and must not be read as a verdict. `ClassifyRecoveryFailure` therefore takes
both flags; the compiler found every call site.
**The message, and what it deliberately does not say.** It states the code is correct, names the
supersession date, says the earlier package is kept, and — the half a customer will otherwise assume
wrong — says the CURRENT backups are unaffected. It does **not** promise the older history can be
reopened from this screen: there is no in-product route to a set-aside store (the restore machinery
resolves its repository from settings and its password from one file), and the retained package may
itself predate the repository-password field. A conditional promise that turns out false on this
screen is worse than saying less — the R-202 lesson, on the highest-stakes copy in the product. It
routes to support, which CAN do it: the 2026-08-12 drill did exactly that by hand.
**An older agent keeps the hedged sentence.** Unknown → claims less → heals itself on update.
**The claim guard grew a surface (and immediately convicted something).** `retrieval_promise_gate.py`
scanned `internal/web/templates` only — while every recovery message is a Go string in a handler, i.e.
the highest-stakes copy in the product had never been scanned. It now scans `recovery_handlers.go`
too, with Go comments stripped for the same reason template comments are. On its first run it found a
PRE-EXISTING unregistered claim (`RecoverRefused`'s "reopening would overwrite it") — now registered
as an explanation rather than a promise. `visszanyit` joins the stems: the new message uses a fourth
verb for the same claim, and the gate's own history is what happens when it chases words not claims.
Six handler tests asserting which SENTENCE the customer sees, with red-proofs asserted applied —
including: make 422 unconditional and an agent that never looked is read as having looked; route 400
to the new class and a mistype is congratulated.
---
## v0.213.0 — the banner promises only what the box can still see is true (2026-08-12, R-302) — MinAgent 0.127.0
**The abandon countdown told every customer who had given up their off-site history: *„Addig még
visszaszerezheted őket a helyreállítási kóddal."* Unconditionally, on every page. It is false on a
reachable state — and it rendered on the same screen as the orphan card correctly saying we cannot tell.
**Why the obvious condition was rejected, recorded so nobody re-proposes it.** The natural proxy —
*does the hub hold a key different from the one this box uses?* — asks about the WRONG key. The
set-aside copies were written under an OLDER key the box no longer has, which is why they were set
aside. On a twice-rebuilt box the proxy answers “yes, promise it” about copies no key on file can open:
right in the ordinary case, wrong in the very case that started the investigation. **Demonstrated, not
argued** — under the proxy both Scenario B (package replaced) and Scenario D (legacy countdown) flip
back to promising.
**Instead the fact is recorded at the one moment it is a fact.** `startAbandonCountdown` pins the hub’s
escrow key fingerprint as cached AT THE DECISION (`AbandonPinnedEscrowKeySHA256`). From then on the box
asks one exact question — *is the hub still holding that same package?* — rather than guessing which key
is which. Written once, never refreshed: a field re-read at render answers a different question. Same
shape as R-300’s ownership record two sessions ago.
**⚠ IT IS A RECORDED ASSUMPTION, AND IT SAYS SO.** Nothing on the box records which key wrote the
set-aside copies. The pin presumes the package held at the decision is that one — true in the ordinary
rebuilt-box story, not provable, and wrong on a twice-rebuilt box. Written into the field comment and
into R-302 so it can be narrowed later rather than hardening into a fact.
- **The certain half always renders**: the deletion and its date. Only the retrieval clause is conditional.
- **Empty is not a match**, on either side — the hub sends “” for a package sealing no repository password.
- **A countdown started before this release carries no pin and takes the cautious branch.** Not
backfilled: that would assert as recorded-at-the-decision something read long afterwards.
- **A FOURTH and FIFTH instance of the same promise were found by sweeping every template.** The backups
page block (`„a mentéseid visszaszerezhetők, és a törlés elmarad"`) got the same condition — fixing the
strip and not the page would leave one contradicting the other. The abandon CONFIRMATION screen
(`recovery.html`) was deliberately left: it renders at the moment of the decision, where the promise is
true by construction, because that is the package about to be pinned.
**New gate — `retrieval_promise_gate.py`, and it pins the CLAIM rather than the word.** A string ban was
tried twice and failed twice (singular vs plural; then one verb vs another). It cannot simply be
broadened either: **the honest replacement copy contains the stem**, inside a question about whether the
thing is knowable. So every retrieval-claim occurrence across all 36 templates is now REGISTERED with a
reason, and unregistered ones fail. Proven by planting all three historical wordings in turn — each
convicted, each cleared on removal.
## v0.212.0 — the second promise (2026-08-12, R-299) — MinAgent 0.127.0
**R-299 — the orphan card’s OTHER sentence made the same unevaluable promise, and the spec said it was
fine.** v0.211.0 fixed the confirm block; the EXPLANATION paragraph above it
(`internal/web/templates/backups_remote.html`) still ended *„a hozzájuk tartozó helyreállítási kóddal
később **visszaállíthatók lehetnek**”* — the identical claim in the plural.
**It survived for two independent reasons, and both are the interesting part:**
1. `SPEC-orphan-card-copy-2026-08-10.md` §1 listed that line as *“Accurate; keep”*. The spec has been
corrected.
2. **The regression guard matched one INFLECTION.** It asserted `visszaállítható lehet` (singular); the
card carried `visszaállíthatók lehetnek` (plural), which does not contain that substring at all. **A
guard matching one inflection of a Hungarian verb guards one sentence, not the claim.** It now
matches the stem `visszaállíthat`, so any conjugation fails. Proven by planting the exact shipped
plural: the stem guard convicts and quotes it, while the old singular guard does not match it.
**And it was the ALWAYS-VISIBLE half.** The paragraph fixed in v0.211.0 renders only after the customer
clicks „Új távoli mentés indítása…”. On first view the explanation is the only text they read — so
until now, the sentence a customer actually saw was the one still promising.
**The two accurate halves are kept**, because declining a promise must not turn into telling the
customer less than we know: the store IS orphaned (and why), and new backups genuinely cannot be
written. New ending: *„A meglévő mentések nem sérültek. Azt viszont ez a gép nem tudja megállapítani,
hogy később megnyithatók-e — ez attól függ, megvan-e még a hozzájuk tartozó kulcs. Ha szükséged van
rájuk, írj nekünk.”*
Also fixed in the guard itself: its failure message sliced the rendered HTML at a BYTE offset, which
cuts Hungarian mid-character and printed a replacement char — a garbled failure message reads like an
encoding bug in the product. It now slices on rune boundaries.
## v0.211.0 — the wall a rebuilt box could not get past (2026-08-10, R-280 / R-294 / R-295) — MinAgent 0.127.0
**R-294 / R-202 — the orphan card stops promising what it cannot know.** The card told a customer,
at the moment they had just lost their off-site history, that the old copies *„a hozzá tartozó
helyreállítási kóddal később visszaállítható lehet"*. The discriminator is
`host_escrow_superseded.identity_blob` and it lives on the **hub**; the box caches only
`HubEscrowIdentityPresent` (the CURRENT escrow) and no report or ACK field carries superseded-blob
retention. **The renderer could not evaluate the condition it was stating**, and for everything set
aside before hub v0.93.0 (2026-08-04 ~11:11Z) it is false and unfixable. Copy replaced verbatim from
`documentation/design/SPEC-orphan-card-copy-2026-08-10.md` §4: it states what happens, declines the
claim it cannot evaluate and says why, and names a route (write to us). Four render tests, per branch
of the gate.
**SPEC DEFECT FOUND AND NOT ACTED ON — `backups_remote.html:98` makes the same promise.** The spec
lists that line as *"Accurate; keep"*, but it ends *„a hozzájuk tartozó helyreállítási kóddal később
visszaállíthatók lehetnek"* — the identical claim in a different conjugation, which the spec's own
regression guard (`visszaállító` + `lehet`, singular) does not match. Left as-is deliberately: the
instruction is not to improvise Hungarian at the customer. **Needs a wording decision → R-296.**
**R-295 — one name per secret (controller half).** The claim page called the SAME three-word
dashboard code „Beállító kód" on the first-time branch and „Visszaállító kód" on the reset branch,
while the TEN-word escrow code is „Helyreállítási kód". Two near-homographs for two different
secrets; the collision cost a real code. „Visszaállító kód" is **retired**: the dashboard code is
„Beállító kód" on both branches (`claim.html`) and in both operator-facing strings (`claim.go` —
the `print-reset-code` output and the lockout message), and where the one secret serves two
situations the **name is constant and the sentence changes**. **Naming only — no acceptance logic
moved**, pinned by `TestResetCode_StillAcceptedOnTheSetupPage`.
**Gate fix (instrument, not product).** `secret_in_markup_gate.py` treated a Go template comment
`{{/* ... */}}` as a rendered expression and convicted the prose explaining a fix for containing the
word "secret". Template comments are stripped by `html/template` and cannot reach the response body,
so they are now skipped — `<!-- -->` comments deliberately are NOT, because those do ship. Proven in
both directions: the gate passes the comment and still convicts a planted `{{.RecoveryPassword}}`.
**R-280 — after a reinstall the data drive can be re-attached, and the page stops promising a click
that does not exist.** Measured on the rebuilt demo-hp 2026-08-09: the restore page diagnosed the
situation perfectly, said *„Ez két kattintás"*, and pointed at a picker holding nothing. It was zero
clicks; getting past it needed an internal path no customer could produce.
**Why the obvious fix would not have worked, kept here because it cost the session an hour.** The
agent builds BOTH `initialize` and `attach` from its unclaimed-DISK scan, and widening that scan is
the fix the finding proposed. But the filesystem a rebuilt box must re-register is an **in-guest**
one — on demo-hp `/mnt/sys_drive`, the guest's own 70 GB data volume, which is what the escape hatch
actually registered. The agent enumerates HOST block devices and would have offered the 1 TB NVMe
(the `felhom-backup` target): the wrong drive, non-destructively attached, customer data still
unreachable. **So this ships in the controller and the agent is unchanged** — MinAgent stays 0.127.0.
- **`attach` now also carries the controller's own mounted-but-unregistered filesystems**
(`internal/web/attach_sources.go`), read from its own mount table — the controller runs in-guest
with `/mnt` bind-mounted in, so what it can see is what it can register. `initialize` is passed
through **untouched**: the format wizard's system/backup protection lives in the agent's scan and
is not widened by a single line here (pinned by `TestMergeAttachCandidates_InitializeIsUntouched`,
whose red-proof put the customer's data volume in the FORMAT list).
- **A union, not a replacement.** The agent's entries serve the case this wizard was built for — a
fresh external drive carrying a filesystem, not yet mounted — which a mount table cannot report
precisely because it is not mounted. Dropping them would fix the reinstall and break the USB.
- **These candidates are REGISTERED in place, never mounted** (`POST /api/storage/register-mounted`).
The posted path is re-derived server-side and refused if it is not currently offered, so the route
cannot register an arbitrary directory.
- **The „két kattintás" sentence is now conditional on the picker being non-empty**, and the false
branch says what is true and names a route. Both branches are render-tested.
- **Exclusions, each with a reason in the code:** the guest rootfs at `/mnt`, the intermediary-model
parent `/mnt/felhom-drives`, tmpfs/overlay, anything outside `/mnt/`, non-ext4/xfs, already-
registered paths — and **any bind alias of the rootfs, excluded by DEVICE**, because a bind
republishes a filesystem under a second path and registering that one would put app data on the
box's own root. That last guard was found by trying to build the live reproduction, not by review.
- **Fail-safe:** an unreadable mount table yields an EMPTY list, never a permissive one — and because
the page gates its promise on that list, "we could not look" renders as "we cannot offer this".
## v0.210.0 — two pictures that were not true (2026-08-08, R-259 / R-258) — MinAgent 0.127.0
Both are the same shape: something the box already knows, drawn as its opposite.
**R-259 — a disk we failed to read was drawn as a healthy empty disk.** `readDiskUsage`
(`internal/system/info_linux.go`) logged a `statfs` failure at DEBUG and returned, leaving the
caller's `TotalGB/UsedGB/AvailGB/Percent` at zero — and `usageColor(0)` is `"nominal"`. The
dashboard's most-looked-at meter therefore rendered „0.0 GB / 0.0 GB (0%)" with a 0%-wide bar in the
healthy colour. **"We could not look" and "there is plenty of room" were the same picture.**
`readDiskUsage` now returns whether the measurement succeeded; `SystemInfo` gains `DiskKnown` and
`HDDKnown`; and the template draws **no figure, no percentage and no meter fill** when unknown,
saying „A tárhely mérete most nem olvasható ki." instead. A healthy box is byte-identical to before,
colour band included.
**This session rules the convention** (`felhom.eu/CONTEXT.md` S-39): an explicit `…Known bool`
companion beside the figures, checked in the template — the shape `Offbox.StatsKnown` already uses,
whose own comment says *"a 0%-wide bar over an unread store is a picture of emptiness, and a picture
is a claim"*. Pointers and separate error fields are both legitimate Go, but a codebase with three
dialects cannot be gated. Existing call sites were **not** converted.
**R-258 — the per-app backup tick was green on presence, and red only on a global condition.**
`buildAppBackupRows` set `Tier1LastStatus` from `status.LastDBDump.Success`, which is the box's
single most recent dump **run, whichever app it belonged to**. An app whose own dump failed showed a
tick as long as some other app dumped successfully afterwards, and an app with no database took the
`nil` branch and went green on the mere existence of a restore point.
Now `appDumpVerdict` reads THIS app's own entries in `DBDumpStatus.Results` (matched on
`DumpResult.DB.StackName`, failure = non-nil `Error`). **Three states:** any failing database →
`error`; all clean → `ok`; **no result recorded → no verdict and no icon**, with the title
„Erről a mentésről nincs eredményünk." The recovery unit carries no per-run outcome of its own, so
green cannot honestly be derived from presence. The global `tier1DBStatus` label is untouched.
**RECENCY IS DELIBERATELY NOT ADDED.** A tick over a three-week-old restore point is a real
weakness, but an age threshold means inventing a number and the time is already printed beside the
icon. Recorded as an observation.
**An existing test was asserting the defect and was corrected, not deleted:**
`TestBuildAppBackupRows_Tier1FromRestorePoints` expected `"ok"` for a status with no `LastDBDump` at
all — green from nothing but a file's existence. It now expects no verdict; its real subject, the
`Tier1LastRun` time, is unchanged.
Four red-proofs, each with the mutation asserted applied: reverting to the global field returns app
X's false green; mapping "no result" to `ok` returns green-on-presence; ignoring the known-flag
returns „0.0 GB / 0.0 GB (0%)"; forcing the flag false shows a healthy box losing its numbers.
**No new tag on any declared wire** — `report/builder.go` maps into its own types and is untouched;
`wire_contract_gate.py` confirmed green.
## v0.209.0 — the box stops saying a false thing about its own recovery package (2026-08-08, R-247 / R-260) — MinAgent 0.127.0
**R-247, and it is the third instance of one shape: the answer was on the wire and was discarded at
the boundary.** The hub has sent `escrow_stale` in the report ACK since v0.57.0
(`json:"escrow_stale,omitempty"`). `report.EscrowStatus` had no field for it, so `encoding/json`
dropped it, and an empty `restic_pw_sha256` had exactly one possible reading here — *hash-less
supersession*.
On `demo-hp` that reading was **false in every clause** for four days, and the box said so in its own
words: the hub HAD the hash and was withholding it because the escrow row carries a stale flag
(R-246); there had been no supersession; and the bundle DID cover the password — the hashes matched
exactly.
**Fixed by receiving the field.** `EscrowStatus.Stale` now decodes, and `reconcileEscrowed` tells the
two conditions apart. A withheld hash now reports that the hub has flagged the row and is withholding,
that **this box therefore cannot verify its bundle either way**, and that it is *not established* that
the bundle fails to cover the password. The genuinely hash-less case keeps its original wording.
**Deliberately NOT changed:** the stale verdict itself (the hub's flag is still the hub's verdict, and
runs still continue), and the customer-facing card copy. Clearing the wrong flag is an operator act
hub-side and is R-246; re-wording the Hungarian card is UI work with its own review path. This change
is the wire and the diagnosis.
`felhom.eu/scripts/wire_contract_gate.py` (G-1) now refuses any new field of this shape on the three
declared wires.
## v0.208.0 — the last two secrets leave the page source, and a gate so there is no fourth (2026-08-08, R-254) — MinAgent 0.127.0
v0.207.0 removed a password from one page. The census that fix required found two more sites; this
closes both, and adds a check so the next one is caught rather than searched for.
### 1. An app's first-login password (R-254 site one)
`app_info.html` rendered `{{.InitialCreds.Password}}` into a `hidden` span — a **real per-install
credential**, read live out of the running container, in the response body of every render. `hidden`
stops a browser DRAWING it and nothing else.
The page now carries the non-secret half (username, note) plus a boolean; the value comes from
**`POST /apps/<slug>/initial-credentials/reveal`**, which **re-reads the container** rather than
serving a cached copy — caching it in the handler would put it straight back in the body one layer in.
`no-store`, CSRF-covered, and **logged as an act**. Both buttons (Megjelenítés *and* Másolás) go
through it; neither keeps the value between presses.
A reveal can now legitimately fail (container stopped, file deleted after first login) and **says so**
— an empty string would have rendered as a blank password.
### 2. The deploy form — established before changing (R-254 site two)
The task named the hidden input. **It is not the defect, and it was left alone:** it fires only on the
PRE-DEPLOY form, and `README §318` documents why the value must round-trip — the customer is shown the
generated secrets so they can note them down, and submitting them back is what makes the saved value
the same one they saw ("no silent re-generation on submit"). A form must carry what it submits.
**The defect was the neighbouring readonly display input.** On an ALREADY-DEPLOYED app the hidden
input is correctly omitted — nothing is being submitted — yet `<input type="password" value="{{$val}}"
readonly>` still rendered the secret into a page the customer merely opens. That is fixed by
**`POST /stacks/<name>/auto-field/reveal`**, authorised by requiring the field to be a `type: secret`
auto-generated field of *that stack's* catalog metadata. Both directions are pinned by tests: the
deployed page must not carry the value, and the pre-deploy form must still submit it.
**The premise that this contradicted a repo rule does not hold.** The rule is `CONTEXT.md:2070`,
*"Password fields require explicit input — prevents accidental empty-password deployments"*: it is
about EMPTINESS, not auto-fill. No line anywhere in the repo says "no silent auto-fill".
### 3. A gate, because three instances in two days is a pattern
`scripts/secret_in_markup_gate.py` (registered in `controller_gates.py`) reads all 36 templates and
convicts any `{{ … }}` whose expression names a secret, unless allowlisted with a stated reason.
**Its limits are measured, not estimated, and are in its own docstring.** It catches a launder through
a local variable (the assignment names the secret). It is **blind to a secret arriving under a neutral
page-data key** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly,
verified both ways. That is the shape of site two, which this gate would NOT have caught.
The complementary net is the runtime body assertion, which catches all of them — but needs each page's
data to be constructible, and **only 4 of 27 page templates have that today**. The other 23 have no
runtime coverage: **R-255**, filed rather than glossed. Two nets, different holes, both named.
### A correction to v0.207.0's report
It stated that HTML comments ship in the response body. **They do not, here** — this package renders
with `html/template`, which strips comments (measured: `text/template` keeps them, `html/template`
does not). A red-proof that plants a secret in a comment therefore correctly does **not** fail.
## v0.207.0 — a password stops living in the page source, and two refusals learn to say what to do (2026-08-08, R-249/R-252/R-253) — MinAgent 0.127.0
Three items the fifth walk exposed by passing. None of them touches the recovery path it proved; all
three are about what the product *says*.
### 1. The retrieval passphrase leaves the page body (R-249)
`settings_security.html` rendered the passphrase into a `display:none` span behind a „Megjelenít"
button. **That toggle stops a browser DRAWING the value and nothing else** — the plaintext was in the
response body of every render, so a `curl` of the page returned it. It was found by doing exactly
that: it landed in a session transcript on 2026-08-07 while driving the documented rebuild path.
**The product already had this rule and this page did not follow it.** `escrow_handlers.go` states it
for the recovery code — *"reveal (claim XHR only — R is NEVER templated server-side into HTML)"*. The
passphrase now follows the same shape: the page carries only `HasRetrievalPassword`, and the value
comes from **`POST /settings/retrieval-password/reveal`**, behind the same RequireAuth + CsrfProtect
every other POST sits behind, `Cache-Control: no-store`, **and logged as an act** — reading it off the
markup left no trace anywhere, where the hub's equivalent break-glass reveal has always emitted an
event.
**POST for a read, deliberately:** a GET would be re-fetchable from history, pre-fetchable, cacheable,
and — since CsrfProtect only covers unsafe methods — uncovered by CSRF.
**The test asserts the raw response body, not a rendered view**, because that is precisely why this
survived: every test that asked what the customer *sees* passed while the bytes carried the secret.
**Census (§7.1), reported not fixed:** the render-then-hide pattern appears **twice more** —
`deploy.html` (an auto-generated app secret in a `type="password"` input's `value=`; unavoidable on
the pre-deploy form, which must post it, but not on an already-deployed app's page) and
**`app_info.html`, which puts a per-install generated app password inside a `hidden` span** — the same
shape with a real secret. Filed as **R-254**.
### 2. „nincs elérhető adatmeghajtó" now names the reason and the route (R-252)
A rebuilt box's drives survive; their **registration** does not. Every restore then refused with a
sentence that named no next step and read like data loss. The restore page now states the precondition
**before** the customer presses anything, says the backups and the drives are both still there, and
links to Tárhely → Meghajtók. The refusal string says the same.
The page asks the question through the backup manager's own `HasRestoreDestination()`, which reads the
**same** `GetSchedulableStoragePaths()` the resolver reads — a second copy of that predicate is exactly
how a page ends up promising what the handler refuses, which is the next item.
### 3. The page no longer promises a reinstall the restore cannot do (R-253)
The restore list said **„Nincs telepítve — a visszaállítás előbb újratelepíti."** Three lines later the
restore refused *because* the app was not installed. Two shipped sentences, in the customer's own
language, contradicting each other at the last step of a recovery.
**The promise was the wrong half, and this is why:** reconstitution writes to the app's own data path
(`GetStackHDDPath`), which exists only once the customer has chosen a drive at deploy time. An
automatic reinstall would mean the product picking that drive for them — the one decision this whole
recovery path exists to leave with the customer. So the copy now says to install it first and routes
to `/stacks/<app>/deploy`; the refusal was reworded to match.
**A healthy box renders exactly as before** — both notices are conditional, and a test fails if either
becomes unconditional.
## v0.206.0 — the box does not mint a key over a sealed package, and abandoning ends the question (2026-08-07, R-241) — MinAgent 0.127.0
**R-241 was ruled a MINTING defect, not a screen-predicate defect** (`SPIKE-r241-recovery-offer-2026-08-07.md`),
and that reversed the fix. The recovery screen was telling the truth: there genuinely was nothing
recoverable under the key the box held, **because the box minted that key itself, over the top of a
sealed package it already knew the hub was holding.** Mending the screen would have papered over a
machine quietly making its own backups unopenable.
### 1. It stops minting
`WriteOffboxSecrets` auto-generated on **one** input — does the file exist. Its two neighbours in the
same file, `OffsiteRecoveryOffer` and `needsOffsiteCredential`, both consult
`GetHubEscrowIdentityPresent()`. **The same fact was available on three paths and used on two.**
Measured on the final walk: the credential self-heal reached it at 03:18:06Z and minted `9b4a9a9d…`
over a package sealing `30ef574f…`. The flag was not merely available at that moment — it was the
**precondition of the chain that reached the function**, logged at 02:48:03Z, six ticks earlier.
The guard is a **conjunction** (a package held AND no key present), so a first-time box mints exactly
as before. The refusal is a **holding state, not a failure**: the transport is still written, so the
recovery screen can bring the tier up the instant the key arrives (R-219). Returning an error instead
would have left the hub re-staging a consumed credential for ever. New declared state
`offsite.state=awaiting_recovery_key`, shown inert to every existing hub reader from their code.
### 2. The comparison it already made now drives the offer
`EscrowAutoConfirmer.Reconcile` has compared the hub's `restic_pw_sha256` against the local key on
every ACK since SLICE 3. On the venue it logged the mismatch at **03:28:03Z — thirty-five minutes
before the customer looked** — and threw it away. It is now persisted, and `OffsiteRecoveryOffer`
gains **shape (c)**: the hub holds a package for a key other than the one we are using.
**§7.2, decided deliberately:** a **known difference offers however old the reading** (age is not
gated on — gating would make a box offline from the hub silently stop offering); a **hash never
learned falls back to (a)/(b)**, because an empty hash is the hub positively saying its package seals
no key, not an unknown.
### 3. Abandoning is now a finishable thing
Setting the old history aside used to touch neither the escrow nor the key, so the hub went on holding
a package for a key nobody used and the question returned at every login. It now starts a **14-day
countdown**, visible and reversible, at the end of which the set-aside store **and the sealed package
that protects it are removed together** — after which shape (c) has nothing to compare and the offer
falls silent **because the state is right, not because something remembers it once was not**.
The grace is real: the recovery offer stays reachable throughout. The two halves cannot be atomic
across two machines, so it is a two-phase commit whose confirmation rides the **same ACK** that
carries the request. Needs hub **v0.98.0**.
### 4. The surface, and the trap that does not survive this session
The full page appears **once per entry into the offered state, not once ever** (an epoch, so a box
rebuilt months later is a new situation). Three dismissal levers with three scopes — a per-visit
session cookie, a durable epoch-scoped reminder opt-out, and the existing "most nem" — and **none of
them removes the entry point on the backups page.**
**§7.3 / Q7:** while a recovery is outstanding, „Helyreállítási kód létrehozása" is now **unavailable**
rather than merely captioned. Creating a new code seals the current key, demotes the package that
opens the earlier history to retained custody no shipped path can read (R-199), and re-enables the
screen while invalidating the code it accepts. A warning beside a button is a warning people click
past.
**The abandon confirmation changed with the behaviour (§2.4):** it used to promise *„félretesszük —
nem töröljük"*, and after this the history **is** deleted, on a date it now states.
### 5. Reminders and operator levers
Escalating emphasis at 1/3/7/14 days for an undecided box, 5/3/1 days remaining for an abandoning one.
`--abandon-status` / `--abandon-extend=N` / `--abandon-stop` on the controller CLI, because the path
that actually happens is the customer telephoning. Both levers **refuse rather than no-op** when
nothing is running or the store is already gone.
**The automatic 30-day abandonment is recorded and NOT built** → R-245.
### Caught by tests rather than review
Two real bugs in this change: `OffboxAwaitingRecoveryKey` omitted `t.Enabled`, so a customer who had
switched off-site off would have declared a holding state (caught by the existing
`TestOffsiteDeclare_DisabledTargetIsNotStranded`); and `recoveryInterrupts` returned early when the
offer was false, so the **falling** edge was never recorded and the page never came back — the exact
defect the epoch exists to fix, reintroduced inside the fix.
## v0.205.0 — a backup that skipped an app the customer chose is not „Rendben" (2026-08-06, R-234) — MinAgent 0.127.0
**Two defects, and the one that actually produced the measured sequence was NOT the one filed.**
### The verdict now counts a skipped selection
The R-203 verdict block already carries the sentence *"a warning beside a success is read as a
success"* — and applied it to **one of the two shapes it describes**. An app missing a declared
mandatory FOLDER made the run `incomplete`; an app skipped **entirely**, with nothing of it in the
snapshot at all, still reported `ok` with a warning beside it. The smaller gap moved the verdict and
the bigger one did not. It does now.
**Which skips count (§7.2), decided by measurement rather than assumption:**
| skip | counts? | why |
|---|---|---|
| selected + **deployed**, no recovery unit | **yes** | the app the customer chose is not protected |
| selected but **not deployed** | **no**, but NAMED with what to do | a box left amber forever by an app somebody removed is a status nobody reads |
| drive disconnected / decommissioned | **no** | it has its own card and its own signal |
| nothing selected at all | **no** | unchanged: the existing zero-selection notice |
The operator signal is **reused, not mirrored** — a skipped app is reported through the existing
mandatory-gap notification as a whole-unit gap, so one vocabulary covers both.
`LastSuccess` and `SnapshotCount` still record what WAS captured: half a backup is not no backup.
### The measured cause: a manual run silently dropped by the single-flight
Reproduced on demo-hp: **the pre-dump phase (`captureAllRecoveryUnits`) writes a unit for every
deployed stack before the push**, so "selected but no bundle yet" does not normally survive a run —
moving a unit aside and running recreated it and reported `ok`. So the filed mechanism could not have
produced the 2026-08-06 sequence.
What did: the customer pressed „Távoli mentés most”, the handler answered „A távoli mentés elindult”, `acquireRunning` refused because a run was already going, and the run
returned **nil** — no error, no signal. The card then showed the **previous** run's „Rendben", which
reads as covering the app just selected. It did not, and the restore refused minutes later.
The single-flight decision is now taken **synchronously in the handler**, before the goroutine, and a
dropped request says so. The nightly path is deliberately unchanged: returning nil is right for it —
nobody asked, and the next scheduled run retries.
### §7.3 — the wait, measured before deciding
`CaptureRecoveryUnit` writes compose config + a manifest (**a few KB**, per `admission.go`'s own
note), **enumerates** dumps already present rather than creating them, is idempotent, and **does not
stop the app**. It already runs for every deployed stack inside the off-site run's own pre-dump phase,
through `admitApp`. **So the inline capture this task contemplated already exists — nothing was built**,
and for a deployed app there is no wait to remove.
### Hungarian
- „Ezek az alkalmazások NEM kerültek be a távoli mentésbe, mert még nincs helyi mentési egységük: %s. A következő mentés általában már elkészíti — ha a második futás után is itt szerepelnek, szólj az üzemeltetőnek.”
- „Ezek az alkalmazások ki vannak jelölve távoli mentésre, de nincsenek telepítve, ezért nem menthetők: %s. Ha már nincs rájuk szükséged, vedd ki a kijelölésüket a Távoli mentés oldalon.”
- „Már fut egy távoli mentés — ez a kérés nem indított újat. A most látható eredmény még a korábbi futásé; várd meg, míg ez befejeződik.”
- sibling (mandatory folders), extended so both read alike: „… Ellenőrizd, hogy a mappák megvannak-e a meghajtón; ha igen és ez a következő mentés után is látszik, szólj az üzemeltetőnek.”
## v0.204.0 — what you can restore is decided by the store, not by what happens to be installed (2026-08-06, R-237 / R-238) — MinAgent 0.127.0
**A household that had just lost its box was shown nothing to restore.** Measured live on the R-201
re-walk (`felhom.eu/documentation/tests/part4-rewalk-2026-08-06/journal.md`): after a rebuild, with
the key recovered, the tier configured and the escrow re-sealed, `/backups/restore` said „Nincs
telepített alkalmazás" and the wizard refused every app with „Ez az alkalmazás nincs távoli mentésre
kijelölve" — while the repository held their snapshots the whole time.
**The list was keyed on the wrong thing.** It was `buildOffboxApps()` filtered on `.Enabled`: apps
**currently deployed** AND **currently toggled on for FUTURE off-site backups**. A rebuilt box has
neither. That is a circular dead end at the worst possible moment — to restore an app you must select
it, to select it you must have installed it, and to know what to install you must see the backup you
cannot see. **The toggle is a statement about future backups; requiring it to look at a past one
conflates two different questions, and that conflation was the defect.**
**The store is now the source of the list** (`internal/web/offsite_restore_list.go`), built on the
existing R-193 inventory (`OffsiteInventoryList`) which already reads the repository and groups by
app tag. Installed-ness became a property OF a row, never a filter on it: it changes what restoring
implies, not whether the row exists. Every case is answered rather than hidden —
| case | what the customer sees |
|---|---|
| snapshot present, app NOT installed | listed and restorable, plus „Nincs telepítve — a visszaállítás előbb újratelepíti." |
| app installed, no snapshot | listed, „Nincs mentése a távoli tárolóban — nincs mit visszaállítani." |
| store unreadable | „Nem tudjuk elolvasni a távoli tárolót, ezért **nem tudjuk, mi van benne**. Ez nem azt jelenti, hogy üres…" — **and the action is still offered**, because "we could not look" is not "there is nothing" |
| no target yet (the pristine rebuilt shape) | „A távoli tároló kapcsolódási adatai még nem érkeztek meg ehhez a géphez… Ez magától rendeződik." |
| store genuinely empty | „A távoli tároló üres — nincs mit visszaállítani." |
That unknown-is-not-empty rule is **R-225's, one screen over**, and it now points both ways: a read
failure must not be rendered as an empty list, and it must not silently withhold the action either.
**Two marker tags are excluded from the app list**: `felhom-offbox` rides on every snapshot, and
`_shares` has its own restore entry with no per-app wizard — listing either would have offered a
restore of something that does not exist. (The R-193 unlock listing still shows them; that is
recorded, not fixed here.)
### R-238 — the size gate no longer refuses in silence
**Classified as a harness artifact, and the residue fixed anyway.** `POST /backup/offbox/restore`
with `mode=full` and no `confirm=1` is step 1 of a deliberate two-step: it computes size + headroom,
**starts no job**, and redirects carrying `&full_prep=<app>` so `deriveWizardStep` reveals the
commit. A driver that does not carry that parameter forward lands back on the intent step — which is
`deriveWizardStep` working exactly as its precedence comments describe, and is why the endpoint-level
run read as "the button does nothing". **The operator's browser run completed the same restore.**
**What was genuinely wrong: neither branch of that step wrote anything to the log.** `restore-status`
is empty by design (no job), and `offboxRedirectTo` only flashes to the page — so a customer refused
a disaster restore, **including a refusal by the headroom gate**, left no trace on the box at all.
Both branches now log, and so does the concurrent-op refusal. Nothing about the wizard's precedence
rules was re-keyed: a stale parameter must still never resurrect a commit button mid-restore.
## v0.203.0 — the box collects what the hub staged for it (2026-08-06, R-218 consume half / R-220 message) — MinAgent 0.127.0
**R-218's declaration half shipped in v0.201.0 and works. Its consume half never existed.**
A stranded box says `offsite.state=needs_credential`; the hub's `offsiteheal` re-stages the one-time
secret and logs *"the box re-consumes on its next cycle"*. **There was no next cycle.** `Reconcile`
ran exactly twice in a process's life — once at start-up (whose own comment said *"retries on next
config refresh/restart"*) and once when the recovery screen drives it (R-219) — and **both fire before
the hub has anything staged, because the hub stages in RESPONSE to the declaration those runs
precede.** So the hub held a credential the box would never fetch.
**Measured on the R-201 re-walk, 2026-08-06:** unlock reconcile **11:43:07** · hub staged **11:44:57**
saying "next cycle" · a full report cycle ran **11:55:46** · **still unconsumed at 12:06**. A guest
command line applied it in **18 seconds** — proving the credential, the target and the key were all
correct and only the trigger was missing. That was the first of the two dead ends that kept the
recovery journey failing.
**The fix: re-run the SAME reconcile, on a tick, for exactly as long as the box says it needs a
credential.** `Bridge.RetryIfDeclared` is driven from the box's own published declaration
(`OffboxReportStatus().State`) — **the very statement the hub acts on**, so the two can never disagree
about whether a retry is wanted.
**Why a poll and not an ACK flag.** The deciding criterion was the promise the customer is given: the
no-target message says *"amint megvannak"* (no deadline) and the backups card says *"ha egy napon
belül nem áll be"* — **within a day**. A 5-minute tick is inside both by a wide margin and needs **no
hub change**. If either promise ever tightens to minutes, revisit.
**It stops by construction.** The instant a target exists the declaration goes false: a healthy box
does no work and **logs nothing** (asserted). **The settle gate is deliberately kept** — the retry goes
through `ReconcileWhenSettled`, so the day-0 floor race it guards is unchanged.
**The marker was investigated and left alone.** `applied_marker` lives at `<DataDir>/offbox/` — inside
the guest's data dir, which a rebuild destroys — so it cannot suppress a legitimate post-rebuild
re-run. It is not part of this defect.
### R-220's customer-facing half — a refusal that named an impossible action
The deploy refusal said *"Válasszon a listából csatlakoztatott meghajtót"* — choose an attached drive
from the list — **while the list was empty**, on a rebuilt box, for a reason the customer had no part
in. That is the I3 breach the campaign recorded. It now says what is true (the drive is not registered
**on this machine**, which is what a rebuild causes), points at the page where re-attaching happens
rather than at a possibly-empty list, and **promises no outcome**, because whether the drive can be
re-attached is not knowable from there. The NAS refusal is a different situation and is untouched.
Tests: `internal/offsiteapply/retry_test.go` (a credential staged after start-up is collected; a
healthy box does nothing and logs nothing; a nil bridge is silent; the settle gate holds) and
`internal/settings/refuse_message_test.go`. **Red-proofs:** removing the retry leaves the credential
uncollected — the re-walk's dead end reproduced; dropping the stop condition makes a healthy box
hammer the hub; calling `Reconcile` instead of `ReconcileWhenSettled` bypasses the settle gate;
restoring the old sentence brings the impossible action back.
## docs — the "CI is still owed" claim was stale; corrected (2026-08-06, R-229 part 2) — no version bump
**One sentence, no code.** This file asserted that continuous integration was still owed
(`felhom.eu` `OPEN-ITEMS.md` R-168). **R-168 was CLOSED on 2026-08-02** — a Gitea Actions runner
re-runs each repo's gate entry point on every push and emails the operator on failure. Found while
confirming this session's own push by run ID, which is the check that caught it.
The same stale sentence was in four instruction files across all four repos and is corrected in all
four. In `felhom-agent/CLAUDE.md` it **contradicted the same file's release section**, which already
said R-168 mails the failure — a contradiction inside one instruction file, which is the exact class
the R-229 work exists to find.
## docs — CLAUDE.md split into a core plus path-scoped rules (2026-08-06, R-229) — no version bump
**Documentation and gate registration only. No Go changed, no image built, no deploy.**
`CLAUDE.md` went from 215 lines to 110 (92 *effective* — block-level HTML comments are stripped
before injection and never reach the model, verified empirically on Claude Code 2.1.222 with a
control and a treatment run). What moved, rather than what was cut, is the point:
- Four new `.claude/rules/*.md`, each carrying a `paths:` glob list so it loads only when a matching
file is read: `gates.md`, `ui-hungarian.md`, `backup-paths.md`, `agent-coupling.md`.
- The `## Layout` tree was deleted as derivable (`ls controller/internal/`); `REUSE.md` already owns
the per-package seams and traps its annotations stood in for.
- The host/access table was deleted in favour of a pointer to `documentation/operations/nodes.md`.
**It carried three defects at once:** `demo-felhom` given as plain `root@192.168.0.162` (the LAN
*fallback*, not the route), a pinned `agent 0.93.0` against the project's own no-versions-in-docs
rule, and the claim that no drill VM was provisioned on `demo-hp`. **Measured live 2026-08-06:**
`qm list` shows VM `300 drill-r50`. `felhom-agent/CLAUDE.md` was right; this file was wrong.
- The seven session-critical invariants, the F9 live-validation fence and the end-of-session
checklist were kept verbatim — they are the file's highest-value content.
`controller_gates.py` now registers **`instructions`**, the shared
`felhom.eu/scripts/instructions_gate.py`. Never copied here; an absent sibling clone FAILS.
Full per-block accounting: `felhom.eu/documentation/audits/LEDGER-instruction-trim-2026-08-06.md`.
## v0.202.0 — the customer is blamed only after a real attempt refused their code (2026-08-06, R-224/R-226/R-225/R-227/R-228) — MinAgent 0.126.0
**CAMPAIGN-11's headline defect had moved, not gone.** v0.201.0 stopped an agent that is too OLD from
being reported as a wrong recovery code. An agent that is **stopped**, and a hub that cannot be
**reached**, still fell through to a message about the code — measured live on 2026-08-05 with a
**correct, current** code at **0.0299 s** and **0.0556 s**, against ~1.0 s for a genuine unseal. The
machine had not tried, and told the customer their code was wrong.
**And the inverse was true at the same time (R-226).** The one message that says *"check your ten
words"* was tested AFTER the retained-earlier-package message, so on any box that has re-escrowed —
precisely the box whose customer has just been handed a new code — a genuine mistype could never
reach it. Three unrelated failures got the accusation; the one that deserved it got something else.
**One defect from both ends: nothing on that path asked WHY it failed.** `rerr` was never inspected.
### The rule, and it is the whole change
> **The customer is blamed only after a real attempt refused their code. Every other outcome —
> including one we cannot classify — says something else.**
Five classes, **from the value and never the text**:
| class | from | what the customer is told |
|---|---|---|
| hub-unreachable | 502/503 | the connection failed; **the code was NOT used** |
| agent-unreachable | no agent verdict at all | the machine's own service is not answering; **NOT used** |
| no-bundle | 404 | nothing is held for this machine; not about the code |
| bundle-too-old | 409 | the code WORKED; the package predates the field |
| **asked-and-refused** | **400** | **the only class that may mention typing** |
| unknown | anything else | **neutral — claims neither that the code was wrong nor that it went unused** |
`agentapi.RecoveryRefusal` carries the status as a value; `refusalError` flattened it into a sentence,
and a sentence is not something a caller can branch on.
**THE SAFE DEFAULT IS THE POINT.** `RecoveryUnknown` is the zero value, and an unrecognised status
lands there rather than in an accusation. **That is the rule whose absence let this survive being
fixed once.**
**Coupling — `MinAgent 0.126.0`.** An agent below it answers **400 for both** a fetch failure and a
wrong code, so a 400 from one cannot be read as a refusal. `FeatureRecoveryFailureClass` withholds
that reading and the 400 degrades to **neutral**. The gate **blocks nothing** — the unlock is
attempted either way — it only decides whether the customer may be told to check their typing, and
"not sure" means they may not. It heals itself when the agent updates.
**Elapsed time is logged** (it is what diagnosed this, and it is the cheapest tell for the operator)
**and is never the classifier.** Time is a symptom; the status is the fact.
### R-225 — unknown is not zero
An unread store rendered `Tároló méret · 0 pillanatkép` and `Tárhelykeret: 0 / 50 GB (0%)` **directly
above a card saying the store held backups under another key**. An SFTP listing found snapshot
`f3d9cd67` and **12 535 KB** really there; `snapshot_count` and `repo_size_bytes` were simply ABSENT
from `settings.json` and the zero value spoke for them. `StatsKnown` is **named**, for the same reason
`OffsiteInventory.Empty` is — zero is what an unread store and an empty one both look like, and both
fields are `omitempty` ints, so on disk "absent" and "0" are the same bytes. The fill bar renders only
when the fill is known: **a 0 %-wide bar is a picture of emptiness, and a picture is a claim.** A
*measured* zero still says zero.
### R-227 — the gateway error
**Which layer answers: traefik**, whose config this repo generates. A branded proxy page is therefore
possible here, but traefik v3 serves no static files, so it would need a **new always-up container**
for every 502 on the box — out of proportion to this finding, and **scoped in the report rather than
built**. Shipped instead: the unlock posts via `fetch`, so a gateway failure is answered in Hungarian
without leaving the page. **Progressive enhancement** — with no JS the plain POST is unchanged and
still shows the proxy's error, and this entry says so rather than implying otherwise.
### R-228 — the set-aside history is visible, and honest
A customer who chose "I do not want the old data" was told the backups would be **kept**, and the
move-aside did exactly that — 12 535 KB, byte-exact. The box recorded the path in
`orphaned_renamed_to` and **a census found zero references to it in any template or handler.** It is
surfaced now as two facts and stops.
**It does NOT promise the history can be reopened, and the confirmation copy was corrected for the
same reason.** *"a helyreállítási kód nélkül többé nem lesznek megnyithatók"* implied that **with**
the code they could be; serving a superseded package is an unbuilt link (R-199's inventory), so it
cannot be opened by the customer, the operator, or anyone. The field's own comment called it
*"recovery-code-recoverable"* — the same over-promise, in the code.
### Tests
Scenarios A-H at the **handler** and as **render tests per branch of each gate**. Red-proofs, each
demonstrated failing and restored: delete the 502 case (A) - remove the mistype clause (C) - default
to the accusation (D) - route an instant transport failure to the typing message (E) - remove the
`StatsKnown` guards (F) - delete the set-aside block (H).
**Two existing tests encoded the defect and were corrected rather than deleted.** The web fake
returned a **bare** error for "wrong code" — which is the shape of a failure we cannot classify, and
now correctly renders the neutral message; saying "wrong code" in a test requires saying it the way
the agent says it. And R-222's test forbade **any** mention of typing on a superseded box, **half of
which R-226 deliberately reverses**: what stays forbidden is the bare accusation, not the hint.
## v0.201.0 — a correct recovery code is never called wrong again (2026-08-05, CAMPAIGN-11) — MinAgent 0.125.0
CAMPAIGN-11 walked the whole recovery journey end to end for the first time. **The data came back
byte-identical; the journey did not exist** — four operator interventions stood between a customer and
their files, and the first one was the machine telling them their perfectly correct recovery code was
wrong. This closes six of the nine findings.
**R-216 — the headline. A 404 was reported to the customer as a bad code.** The unseal lives in the
agent behind `POST /escrow/recover-offsite-password`, which ships in agent **v0.125.0**. On an older
agent the route 404s, and the unlock was attempted anyway: *„A megadott helyreállítási kódot nem
fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan…"* — measured live in 0.134 s, far too fast for
`age`'s scrypt, against a code that was perfect. **It was the default state, not a misconfiguration**:
the day-0 manifest vouches agent 0.120.0, and a reinstall actively *downgrades* a hand-fixed box back
to it. Three parts:
- `FeatureOffsiteKeyRecovery` now exists in `featureProbes` + `featureMinAgent` (0.125.0) with a
`Supports` gate at the unlock entry point — **the convention `features.go:72-76` and
`publish-train-rules.md` rule 4 already required, which v0.195.0/v0.196.0 declared in their headers
and v0.200.0 dropped when it put the same coupled feature behind a customer-facing button.**
- **This gate FAILS CLOSED**, alone in that table: anything but `SupportYes` says *the machine* cannot
ask yet, and no unlock is attempted. An attempt that cannot succeed must never be made, because its
failure is what got attributed to the code. The asymmetry is documented beside the row.
- The hub's own guard could not catch it (see the hub changelog, v0.97.0).
**R-218 — succeeding at recovery stopped the box asking for what it still needed.**
`needsOffsiteCredential` short-circuited on a repository password existing — and installing one is the
recovery screen's whole job. 32 seconds after the hub re-staged the credential, the customer's success
switched off the mechanism that would have delivered the coordinates for the key they had just
recovered. **The short-circuit is deleted**; the declaration now stops when the tier works, not when a
key exists. Scenario E (a deliberately disabled target stays silent) is pinned by its own test.
**R-219 — the promised listing could never render on the shape the screen exists for.** Listing needs
a target; a target cannot exist without a repository password; shape (a) is defined by having none. The
unlock now **finishes the job** — place the key, bring the tier up, then list — and a tier that is not
up yet says *waiting for the connection details* rather than reporting a failure.
**R-217 — an unreadable store claimed to have opened.** The failure path passed
`backup.OffsiteInventory{}`, whose `Empty=false` the template read as *„A tároló megnyílt, és van benne
tartalom, de nem tudtuk alkalmazásokhoz rendelni."* — three assertions, none of them known. The field
built to prevent exactly this says so in its own doc comment. Opened / empty / unreadable are now three
distinguishable states.
**R-222 — reaching for a retained earlier package read as a wrong code.** The hub's ACK carries
`superseded_present` / `superseded_at` (hub v0.97.0); the screen now names that situation. It states
the two facts the hub knows and **does not promise the earlier package can be opened** — that read path
does not exist.
**R-215 — `GET /recovery` rendered the recovery story on a box that never had backups.** The predicate
was right and the page never asked it. Gated on the same predicate as the interception.
**Four messages where there was one**, each saying what to do next and only one mentioning typing.
Left open deliberately: **R-214** (the console never stops showing a stale pairing code), **R-220**
(drives cannot be re-enrolled after a rebuild), **R-221** (a rebuilt box cannot run the escrow ceremony
— a real blocker for that flow).
Tests: `recovery_gate_test.go` (Scenarios A/F/G/H/I, handler-level), `offbox_declare_test.go`
(Scenarios D/E). **Five red-proofs, each demonstrated failing and restored** — removing the capability
gate brings the accusation back verbatim.
## v0.200.0 — the recovery screen: unlocking, and only unlocking (2026-08-05, R-193)
R-193's remaining half. The drill proved a customer's file comes back after a machine is destroyed;
the sessions since removed every step that needed the operator. What was left is that the customer had
no way to *begin* without a command line.
### It unlocks, and only unlocks (operator ruling, 2026-08-05)
The page explains the situation, takes the recovery code, opens the repository, and shows what is in
there — which apps, from when, how big. **It restores nothing.** Restore is already per-app and already
lives in the backups area; putting files back is a separate item and is not started here. A screen
that unlocks and then offers to overwrite is two decisions wearing one button.
### One core, two callers
`RecoverInstallCore` is split out of `RecoverAndInstall`. The CLI wrapper keeps its exit codes and
printed lines **byte-identical** — every pre-existing CLI test passes unchanged — and the web handler
drives the same function. Two implementations of the one operation that can permanently lose a
customer's data would drift, and only one of them would ever be tested. **Both directions are asserted
from source by AST** (comments dropped, so a commented-out call cannot satisfy them), plus a third test
that the routes and the landing-page interception exist at all.
### When it appears — two facts, and a third shape the task did not name
`backup.OffsiteRecoveryOffer` requires the hub to hold a sealed package **and** this box to be unable
to open what it protects. The second half has **two** shapes:
- **(a) no repository password at all** — the pristine rebuilt box. This is the literal reading of
"the data area is fresh".
- **(b) a password exists but the inherited history will not open under it** (`RepoState == orphaned`).
**Shape (b) had to be added, and the reason is load-bearing.** `WriteOffboxSecrets` AUTO-GENERATES a
repository password when none is present — that is precisely R-193's orphaning mechanism — and since
hub v0.96.0's credential self-heal the re-apply now happens by itself within ~15–30 minutes. Shape (a)
alone would have made this screen appear only inside a half-hour window that closes on its own, so the
customer who logs in the next morning — the actual customer — would never have seen it.
Shape (b) is also the state the shipped move-aside requires, which is what lets "I do not want the old
data" reach the existing handler rather than needing a new one.
**Claimed and behind the household password.** Added after a test caught the omission: a legacy-open
box (no password anywhere) reaches `ServeHTTP` through `RequireAuth`'s pass-through, so without an
explicit `authEnabled()` check the interception fired for an unauthenticated visitor.
### Three ways out, and none of them is a dismiss button
- **Recover** — the main path.
- **„Most nem"** — the full page stops interrupting. **The backups-area entry point stays,
permanently**: it is bound to `recoveryOffer`, never to the postpone flag, because the data is still
there whether or not anyone clicked and a one-shot notice a flustered person clicks past is a notice
that never happened.
- **„Nem kérem vissza a korábbi adatokat"** — the exceptional path, not an equal third button. Two
confirmations, the second naming exactly what happens, then the **shipped** `/backup/offbox/reset`,
which sets the store aside and never deletes. Offered only when that handler can actually run.
### The recovery code is handled no more loosely than on the command line
POST body only (`PostFormValue` — never a query string), never logged at any level, never persisted,
never echoed into any message, cleared on every path, the page served `no-store`, and the field
`autocomplete="off"`. **No lockout** (§8.4): the code is a ten-word EFF phrase, guessing is not the
risk, and locking a customer out of their own data for a typo is a worse failure than anything a
lockout prevents. Failed attempts ARE logged locally — without the code — so a box being probed is
visible in the debug ring.
### Two defects the tests caught before they shipped
1. **An unclaimed box would have been shown the page** (above).
2. **The inventory nil-dereferenced when no off-site target was configured** — which is exactly the
pristine rebuilt shape. It now returns a named error and the page says the key is back and the
listing will appear once the box has re-connected, rather than showing a failure the customer
cannot act on.
## v0.199.0 — a rebuilt box asks for its credential back (2026-08-05, R-204 item 4 / R-193)
The last of the four manual interventions the 2026-08-04 drill needed. A rebuilt box has no off-site
credential of its own — its predecessor spent the one-time provider password — and everything after
that point is already self-service. **This is the box's half: it now DECLARES what it needs.**
### Why a declaration and not an inference (the operator ruling, and the whole design)
From the hub, an **absent** off-site object has FOUR meanings — never configured, mid-restart, a
transient config read failure, and rebuilt-and-stranded — and the hub cannot tell them apart. **The
box can**, from two local facts it holds with certainty. So it says so, in its ordinary report, and
the hub acts on a stated request instead of on a silence.
### The two halves, and why neither is sufficient
`backup.needsOffsiteCredential` requires **both**:
- **a fresh data area** — no repository password on disk. Alone this is simply a box that never had
off-site backups, and declaring on it would make *every un-configured box in the fleet* ask for a
credential. That is the plausible wrong fix, and `TestOffsiteDeclare_NeverHadOffsiteSaysNothing` is
the guard that catches it.
- **a hub-held recovery package** — the report ACK's `escrow.identity_blob_present`. Alone this is a
healthy box that has run its ceremony.
A target that exists but is merely **disabled** is the customer's own choice and never declares.
### The ACK field stopped being discarded
`EscrowAutoConfirmer.Reconcile` now records `identity_blob_present` **first, before every gate**. Those
gates return immediately when the box is neither pending nor escrowed — which is exactly a rebuilt
box — so the one fact that distinguishes it from a box that never had off-site backups was thrown
away on every cycle. It is recorded through the confirmer because that is already the ONE place the
ACK's escrow object arrives and is already wired; a second consumer would be a second wiring point,
and this project's count of features built but never wired is six. `TestMainWiresRecordPresence`
asserts the wiring from main.go's AST.
The recorder is **last-write-wins, not set-only**: a customer RESET that removes the hub's escrow row
must be able to turn the declaration back off. A nil ACK escrow object records nothing — absence of a
statement is not a statement of absence.
### Inert to every existing reader
The declaration carries `enabled:false` and zero sizes. Established from the hub's code rather than
assumed: `OffsiteChecker.isStale` returns early on `!Enabled`, and `fillBand` returns OK on a zero
quota/size — so it raises no staleness and no fill alarm, on a new hub **or an old one**, and an
unknown `state` string is ignored by `encoding/json`. **A configured box's report JSON is
byte-identical to v0.198.0's** — there is no `state` key at all.
The one reader that would have misread it is the hub's `reportHasOffsite`, whose comment asserted
*"presence == applied-on-the-box"*; felhom.eu v0.96.0 tightens it to require `enabled:true`.
### Rider — the pre-push hook refuses a clone outside the workspace
`.githooks/pre-push` gains one assertion, identical in all four repos. The workspace root was already
written down and was drifted from anyway; a rule that has failed once as a reminder is not fixed by
writing it down again. A push is the right trigger — throwaway `/tmp` clones for probes never push.
The only bypass is the documented `--no-verify`.
## v0.198.0 — the four steps a customer would have hit alone: two of them closed (2026-08-05, R-204 items 1 & 3)
The 2026-08-04 recovery drill (R-201) passed — and it only passed because a person was there. Four
manual interventions stood between "the key is recoverable" and "the file is back". None of them is in
any design document. Two of the three defects are in this repo.
### Item 1 — a freshly minted reset code now works on the first attempt
`--print-reset-code` runs as a **separate process** (`docker exec`): it loads settings itself, mints a
code, persists it and exits. The running server's cache was never told, so it kept validating against
the previous hash. **The code the customer was told to type was refused until the controller
restarted, and nothing said so.** During the drill that cost two failed attempts with an operator
present; a customer alone stops there.
`effectiveClaimCode` now READS THROUGH to the persisted state (`settings.ReloadClaimCode`) before
applying the settings-vs-config precedence. **The precedence rule is unchanged and deliberate** — the
defect was the freshness of the settings value, not which source wins.
- **Read-through, not a watcher, a signal handler or a TTL.** A TTL is worse than the bug being fixed:
it opens a window in which a SUPERSEDED code still works. That is the mutation
`TestClaimCode_SupersededByASecondMint_RefusedImmediately` exists to kill, and its red-proof is
exactly that TTL — demonstrated failing with "the SUPERSEDED code was accepted".
- **Fail closed.** An unreadable persisted state keeps the gate up, refuses the claim and logs why. An
ABSENT file is not an error (a box before its first save falls back to the controller.yaml bake).
- Cost: one small file read per request **only while the box carries no password** — `claimGateActive`
returns on `authEnabled()` before touching it, so a claimed box never reads.
### Item 3 — a restore now says what it did NOT restore
The default restore (`mode=unit`) recovers the recovery unit: the app's definition, its configuration
and its database dumps. It does **not** recover the customer's own files — `RestoreOffboxScratch`
passes `--include <unit path>`, and the userdata that is in the same snapshot is excluded by it. The
old outcome was one sentence for both modes and named neither scope, so on the last step of a disaster
recovery the customer was told „visszaállítva" after the thing they were looking for had not been.
- `restoreScratchOutcomeMsg` (pure, unit-testable) now states, for a unit restore: what came back, that
the customer's own files did NOT, and the next step that gets them. The full case says the files came
with it — otherwise the absence of the warning would be the only difference, and an absence is not a
statement.
- The wizard's intent card 1 states its scope **before** the choice, not only in the outcome.
- **The full-restore size gate is untouched** — still compute, reveal, confirm, re-check at execution.
Pinned by `TestOffboxRestore_FullPathUnchanged`, which asserts no restore runs before the confirm.
- **The default stays `unit`.** All three wizard forms set `mode` explicitly, so the `mode==""` fallback
is reachable only by a hand-crafted POST: changing it would alter nothing the customer sees while
silently changing that POST's behaviour. The defect was silence, and silence is what was fixed.
`TestOffboxRestore_DefaultModeGetsTheScopedOutcome` pins the mode-less POST to the scoped wording.
Item 2 of R-204 (a re-issue marking a healthy escrow stale) is the hub's half — felhom.eu v0.95.0.
Item 4 (a rebuilt box cannot obtain an off-site credential unaided) is R-193 and remains open.
## v0.197.0 — the app and its backup look in the same place, and "ok" means it (2026-08-04, R-203)
Found when the R-201 drill halted at its pre-wipe backup rather than wiping a machine: the run
reported `ok` with three snapshots and the sentinel file was in none of them.
### Part 1 — one resolver, five callers
`appbackup`'s path helpers take a **namespace root**; five call sites passed a bare **drive** path.
On an enrolled drive the two coincide — which is why this survived. On the system-data fallback (a
**supported, named** arrangement: *"the SSD-only system-data fallback"*, `paths.go:26`) they differ by
exactly the `felhom-data` segment, so the app bound `/mnt/sys_drive/userdata/media/books` while the
off-site capture set looked for `/mnt/sys_drive/felhom-data/userdata/media/books`.
**The rule now has ONE expression** — `appbackup.NamespaceRootFor` / `IsEnrolledDrive`. There were
already **two** copies and they differed: `backup.Manager.namespaceRoot` compared without
`filepath.Clean`, `stacks.Manager.inGuest` with it, so a trailing slash from config would have flipped
the mode in one package and not the other. Both now delegate.
Routed through it: `stacks/deploy.go` `withPathVars` → `${USERDATA_PATH}` (the live defect);
`appexport/fabplan.go` + `export.go` (via a new `GetStackNamespaceRoot` provider method);
`web/handlers.go` FileBrowser mounts (**latent** — the system drive is deliberately never a
registered `StoragePath`, so it is the identity today); and `stacks/delete.go`'s `ExportDataMounts`.
`ComputeFabBuckets` now receives the namespace root, which is what `ComputeCaptureSet` has always
received — so the export's classified paths and the backup's capture set describe the same
directories by construction rather than by coincidence.
**The census found FIVE sites, not the four the spec named.** The fifth is the FileBrowser mount
builder — the customer's own file browser would have shown the wrong directory on a non-enrolled path.
**`ExportDataMounts` lives in `delete.go` and is NOT a delete path.** Its only production caller is
the `.fab` export adapter; nothing deletes on its result. The delete path's own guard,
`ProtectedHDDPaths`, is layout-agnostic by construction — it protects **both** `<hdd>/…` and
`<hdd>/felhom-data/…` — so deletion was never affected. That note is now in the function's doc comment,
and the change shipped as its own commit anyway.
### Part 2 — a run that missed a mandatory directory is not a successful run
The gap was already **detected**, and warned about, in Hungarian, naming the app and the folders —
that warning is what stopped the drill. The defect was that the run still reported **`ok`** beside it,
and a warning standing beside a success is read as a success.
`last_status` gains **`incomplete`**: minted, because the existing vocabulary (`ok` | `error` |
`running`) had nothing meaning *"it ran, and this app is not fully protected"*. **Not `error`** — the
rest of the run worked and the data captured is real, so `SnapshotCount` and the `LastSuccess` anchor
still record it. Half a backup is not no backup, and reporting it as none would be its own lie.
It reaches the **operator** via the existing per-run digest (`backup_run_failures`), not only the
page: a new event type would be a two-repo change and the hub drops anything outside
`allowedEventTypes`. The Hungarian customer warning is unchanged.
The stat-filter gains the `ClassMandatory` check Tier 2 already had (*"optional-missing is silent"*).
**It is a no-op today** — `TierOffsite`'s `tierKeeps()` admits mandatory only — so **no customer-visible
warning disappears**. Demonstrated: widening the tier filter alone keeps the tests green *because of
this check*; widening it and removing the check makes an optional gap start reporting.
### Anticipated live effect
**`calibre-web` on demo-hp has exactly this gap.** Its off-site status becomes `incomplete` and the
operator digest fires. That is correct — it genuinely is not fully backed up — and it is the fix
telling the truth for the first time, not a regression.
## v0.196.0 — the recovered key installs itself (2026-08-04, R-200 plumbing half) — MinAgent 0.125.0
`--recover-offsite-install` is the sibling of `--recover-offsite-check`: same fetch → unseal → extract
through the agent, same STDIN discipline for R, but it **places** the recovered repository password via
`InjectOffboxPassword` so a rebuilt box reopens the off-site history it inherited.
**Why this is code and not a manual step.** The alternative is recovering the password, reading it off
a terminal and pasting it into the injection endpoint by hand — which puts the offsite DATA key through
a human's screen, clipboard and shell history. In-process, the value goes agent → this process → the
0600 file and is rendered nowhere.
**The confirmation is a second invocation, on purpose.** Without `--confirm-install` it prints both
hashes and writes nothing, so the operator sees the comparison before any write is possible. A single
interactive prompt would have had to share stdin with R.
**Three outcomes, named distinctly**, because "it did nothing" and "it refused" are different facts:
**installed** (no local password — the rebuilt-box shape), **unchanged** (identical key already present,
nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current
repository is encrypted under, and which history to keep is not this command's decision; no force
option is offered). Exit `2` for the refusal, distinct from `1` for a step that failed.
The install re-reads the file afterwards rather than trusting the write — the observable is the file's
state, not the call's return.
**Red-proof observed:** removing the confirmation gate makes the dry run write the password and fails
`TestRecoverAndInstall_InstallsOnABareBox`. The R-persistence test carries a **positive control** — a
planted copy of the code is found by the sweep, then removed and not found — because an absence check
is worth only what its sensitivity is.
**Nothing customer-facing:** no page, no card, no form. `ensureOffboxRepo` and the orphan classifier
are untouched.
## v0.195.0 — prove the offsite key comes back (2026-08-04, R-200 plumbing half) — MinAgent 0.125.0
**The question, answered for the first time: is the offsite repository password actually recoverable
from the hub's sealed bundle?** Yesterday's hub v0.93.0 made the key SURVIVE a re-ceremony. Nothing
handed it back. The chain's last stretch had no client at link 6, a `--selftest`-only caller at
link 7, and nothing at all at link 8.
**`--recover-offsite-check`** — a `docker exec` diagnostic in the shape of `--print-reset-code`:
it reads the customer's recovery code from **STDIN**, asks the agent to fetch this host's sealed
bundle and open it (agent ≥ v0.125.0, `POST /escrow/recover-offsite-password`), and reports whether
the recovered key matches the one on disk — **by sha256**. It prints two hashes and a verdict. Never
a password, never the recovery code, never a blob.
docker exec -i felhom-controller /app/felhom-controller --recover-offsite-check < /root/r.txt
**R comes from stdin and not a flag** because a flag value is visible in `ps`, in shell history, in a
container's command line and in any transcript of the session that ran it — and R is the one secret
in this system that cannot be rotated, re-issued or recovered.
**IT COMPARES; IT DOES NOT INSTALL.** The recovered password is never written to `offbox/repo_password`.
Comparing proves recoverability; installing changes a live box on a path nobody has walked, and
"the existing repository opens under a recovered key" is a separate link with a drill around it.
A test asserts the on-disk password and the whole data dir are byte-unchanged after a check, and its
red-proof — adding the install call — fails it.
**Exit codes are load-bearing:** `0` match, `2` a clean MISMATCH, `1` a step failed. "It failed" and
"it worked and disagreed" must never share a status, because only one of them is a finding about the
system rather than about the run. A box with no local password is reported distinctly too — that is
the rebuilt-box shape, where the next step is to install rather than to compare, and reading it as a
mismatch would be wrong.
**Deliberately NOT in this release: anything a customer can reach.** No card, no form, no preview, no
wizard. Building an interface on top of a chain nobody has walked is how the preceding three weeks
went wrong; the interface comes next, on proven ground, and so does the wipe-and-restore drill.
## Changelog
### v0.194.0 — one operator e-mail per backup run, and nothing dropped without a trace (2026-08-03, R-182) — MinAgent: none
**The defect, measured rather than supposed.** On 2026-08-03 nine per-app
`recovery_unit_capture_failed` events reached the hub and **two operator e-mails went out**. The hub's
operator cooldown key is `customerID + ":" + eventType + tier-suffix`, and that event carries `app`
but **no `tier`** — so the key held no app identifier. The first refused app took the hour's slot and
**every other app's failure was discarded before anything was written down**, leaving no row on any
channel. A machine deciding not to tell you and nothing happening at all looked identical.
**The obvious fix was ruled against, and the reason is worth keeping.** Putting `app` into the key
fixes the swallowing by producing **one e-mail per failing app**, which on a full disk is a dozen —
the volume problem wearing the correctness problem's clothes.
**What ships instead: ONE digest per run, and every failure recorded when it happens.**
- **`internal/backup/runsummary.go`** — a per-run collector with exactly `admissionSet`'s lifetime
(created where the run begins, cleared when it ends), fed by all three write legs. It emits
`backup_run_failures` once at the end, **only when something failed**. A clean run emits nothing —
not an empty digest.
- **The per-app event stays and becomes the RECORD.** The hub now routes it *record-only*: stored and
written to the notification log every time, never competing for an e-mail slot. The record and the
notification are now different things, which is the durable half of this change.
- **Deliberate skips are not failures.** A disconnected or decommissioned drive has its own alert and
is excluded, because a nightly e-mail about an unplugged drive is one the operator learns to ignore.
- **A manual run always reports**, even if the nightly one already wrote that hour: the digest carries
a unique `run_id` that the hub's cooldown cannot collapse. Someone pressing the button is actively
trying to get a backup.
**THE PERIODIC SWEEP GETS A DIGEST TOO, and that is not symmetry for its own sake.** `GetFullStatus`
captures units outside any run. With the per-app event now record-only, a capture failure found
between runs would have been recorded and **never notified** — a new silence introduced while closing
one. So that path emits a digest as well, deliberately with **no `run_id`**, so the ordinary 1-hour
cooldown caps it exactly as before while the mail now lists *every* failing app instead of whichever
one happened to be first.
**A refusal is recorded ONCE, where the verdict is taken**, not at each of the three legs that consult
it — R-181's contract is that one verdict covers all three. Noting it per leg listed a single refused
app three times and produced counts like *"2 of 1 apps failed"*. **Found by the digest's own test, not
in review.**
**Why silence is safe, checked rather than assumed.** A digest is only safe if the absence of a mail
cannot mean "the run never finished". It cannot: the hub's daily deadline check raises
`expected_backup_missed` / `expected_dbdump_missed` (`hub/internal/monitor/deadline.go:396,417`) from
the box's **report freshness and stored events**, independently of any mail this box chooses to send.
**Tests: 7 new, plus 4 red-proofs demonstrated failing then restored.** One of them —
the `main.go` seam walk — **did not fail on its first attempt**, because the AST test walked the
backup package and not `main.go`; the test was fixed and the mutation re-run rather than the pass
being recorded.
### v0.193.1 — the refusal's size estimate is rendered in bytes, not `0.00 GiB` (2026-08-03, R-181 follow-on)
**Found by the live proof run for v0.193.0, not by review.** The refusal message printed the estimate
fixed to two decimal GiB, so **every app under ~10 MB rendered as `estimated 0.00 GiB write`** — which
reads as *"no estimate was available"* and is the exact opposite of what happened. Observed live on
demo-hp at 08:59:46: opengist's real **178 KB** estimate printed as `0.00 GiB`.
Shipped in the same session it was found because it is the same defect class the whole of R-181 is
about — a message that an operator cannot rely on is worse than no message. **The arithmetic is
unchanged and still in GiB** (the reserve's own unit, so the comparison against `FloorFreeGiB` reads
directly); only the rendering moved to `humanizeBytes`. `estimatedWriteGiB` → `estimatedWriteBytes`,
with the GiB conversion done once at the point of comparison.
Verified live after redeploy: the same refusal now reads `estimated 178.0 KB write`.
### v0.193.0 — the reserve guards the write that fills the disk, and its promise is true (2026-08-03, R-181) — MinAgent: none
**The defect, found on live hardware and not by review.** v0.192.0's capture floor (B2) shipped as the
deliberate replacement for the bulkhead the `mp1` partition used to give, and it was consulted in
**exactly one place** — `captureAllRecoveryUnits`, which writes a manifest and three compose files: a
few KB. The two legs that write the **bulk** into the same `backups/primary/<app>` tree — the database
dump and the volume dump — ran **first** and **unguarded**. Measured on demo-hp 2026-08-03 06:40:03:
opengist's volume dump wrote **2.0 GB with no check**, free fell to 1.0 GB, and the floor then refused
the cheap write it had already lost the argument to.
**Second limb: the refusal message asserted something the code did not provide.** It printed *"the
previous unit is untouched and NOTHING was deleted"*. *Nothing was deleted* held. *Untouched* was
**measured false** — that app's tar had gone 182,272 B → 2,147,666,432 B under a `manifest.json` whose
`created_at` and `checksums` had not moved. **This is the sixth entry in `CLAUDE.md`'s table of
shipped guarantees the code did not provide**, and the fourth of those found on live hardware.
**The fix: ONE admission verdict per app per run, taken before that app's FIRST write, covering all
three legs** (`internal/backup/admission.go`). The three write under one per-app root, which is
exactly why one verdict can honestly cover them — and why the message may now claim what it claims.
- **Decided lazily, at the app's first write — NOT once at the start of the run.** Space changes
during a run: app A's 2 GB dump can put app B under the reserve, and a run-start verdict would wave
B through on a reading that was true before the disk filled. That is the same class of mistake,
moved one level up.
- **Remembered for the run, never re-decided between an app's own legs.** Re-deciding reintroduces
the split this closes (DB admitted → volume admitted → capture refused, with the bulk written).
Reset per run: a set carried between runs answers tonight's question with last night's disk.
- **Placed ahead of `DumpAppVolumesSafe`, which stops the stack as its first act** — a refusal decided
inside it would already have bounced the app it is refusing to back up. It sits *after* the
volume-less check, because an app with no named volumes has no first write in that leg to gate.
- **Exactly one operator alert per refused app per run.** Three legs must not mean three emails.
- **The leg order is unchanged** — volume dumps still precede the capture so the manifests enumerate
the fresh tars (`backup.go`'s load-bearing comment).
**The floor is now SIZE-AWARE, not merely headroom-aware.** It asks *would this app's write leave the
filesystem below the reserve?*, not only *is it below the reserve now* — which is how an app was
admitted at 96% used and then allowed to write 2 GB. The estimate is the app's **previous** `.sql` and
`.tar` already on disk: free to read, and the next write is usually close. **No history →
headroom-only**, deliberately, or the first backup would be the one that can never happen; the alert
says so when that applies. Both post-write terms are evaluated, because a large write crosses the
percentage bound on a small volume and the free-byte bound on a large one.
**A container-based `du` per volume was measured and REJECTED, not assumed.** 66 timed runs on
demo-hp's guest 9201: **median ~355 ms per volume** (341–404 ms) on volumes holding tens of KB — the
cost is container start-up, not the walk, so it does not shrink for small apps and only grows for real
ones. Decisive on top of that: `docker run` needs the writable layer, so the measurement mechanism can
fail under exactly the disk pressure the reserve exists to handle. The previous-dump estimate also
measures the **artifact** that will be written rather than the live volume, which is the truer
predictor. The figure and the decision are recorded rather than left as a "we could do better".
**THE MESSAGE WAS NOT WEAKENED — the behaviour was moved so the wording became true.** It still says
the previous unit is untouched and nothing was deleted, and now adds *which* term bound (headroom or
size) and the estimate that produced a size refusal.
`TestAdmission_EveryClaimInTheRefusalMessageHoldsAgainstTheTree` checks **every** claim against a
sha256 fingerprint of the tree it describes — not against the log line, because a log line is exactly
what lied here.
**IT STILL REFUSES AND NEVER DELETES.** Unchanged and load-bearing: nothing on this filesystem is
generational, so "prune the oldest" could only mean destroying a **different** app's only local
recovery unit. `pruneStalePrimaryDirs` is not a retention policy and must never be repurposed for
headroom.
**Tests (11 new, all through the production functions; 4 red-proofs demonstrated failing then
restored).** The refusal assertions are **tree fingerprints before and after**, never log lines. The
DB leg cannot run without Docker, so its gate is pinned by an **AST walk** of `backup.go` asserting
`admitApp` precedes `DumpOne` — `strings.Contains` is insufficient, a commented-out call still
contains the string. Red-proofs: both dump-leg gates removed (= v0.192.0) → Scenario A red, tree shown
changing; the size term removed → Scenario D red; a prune injected into the refusal path → Scenario F
red; the floor moved above the warning band → Scenario G red.
### v0.192.0 — the capture floor replaces the bulkhead (2026-08-03, R-165 · decision B2) — MinAgent: none
**Ships BEFORE the disk-layout merge it exists for, and is harmless on a box that never gets it.**
The `mp1`→`mp0` merge (R-165 / decision D-a) removes a wall that was quietly doing a second job: the
20 G backup partition kept a runaway recovery-unit capture from filling the space the container
runtime itself needs, because `/var/lib/docker` was a **different filesystem**. After the merge it is
the same one, and a full Docker data-root is a stopped box, not a slow one. Decision **B2** is that
bulkhead, done deliberately instead of by accident.
**The floor, in `captureAllRecoveryUnits`, checked BEFORE anything is written.** If the target
filesystem is below the reserve, that **one app's** capture is refused, its previous unit is left
**byte-identical**, the operator alert wired in v0.191.0 fires with the used/free figures, and the
loop continues to the next app.
**Two terms, whichever binds first — 97% used or 1 GiB free** — the same shape as `internal/fillwatch`,
which proved live on 2026-08-02 that a percentage alone is not enough (its critical alert fired on the
free-byte term at 91% used, where a percent-only rule stayed silent).
**They sit deliberately BEYOND fillwatch's critical band (95% / 2 GiB), so the customer is ALWAYS
warned before a refusal can happen.** A floor that fires before its own warning is a silent failure
wearing a threshold; `TestFloorSitsBelowTheCriticalWarningBand` pins the whole ordering
(warn → critical → refuse) on both terms, and a red-proof setting the floor equal to the critical band
fails it.
**IT IS ABOUT THE FILESYSTEM'S HEADROOM, NEVER THE UNIT'S SIZE.** A per-unit cap would be R-163 rebuilt
inside one volume — the wall moved rather than removed — so a 120 GB app on a filesystem with 180 GB
free is captured. A red-proof substituting `UsedGB > 20` for the headroom predicate fails two tests.
**IT REFUSES; IT NEVER DELETES, and the reason is recorded because the question will be asked again.**
Nothing on this filesystem is generational: a unit is ONE fixed path per app
(`backups/primary/<app>`) refreshed in place, and a DB dump is `<stack>-<dbtype>.sql`, also fixed. So
"prune the oldest" could only mean deleting a **different** app's only local recovery unit to make room
for this one. `pruneStalePrimaryDirs` is **not** a retention policy — it removes ORPHANED directories
left when an app moves drives and has no notion of age — and must never be repurposed here.
**A nil usage read neither refuses nor warns** (§8.4): an unreadable filesystem is the drive gate's
business and already has its own alert, and refusing on it would block every capture on a box whose
drive merely blipped.
**Tests:** 1184 → **1191** (+7). New `unitSpaceFn` seam so a filesystem's occupancy is a test input
rather than something a test must manufacture on a real disk. One fixture was **strengthened** during
the red-proofs: `TestFloor_TheOld20GCeilingIsGone` originally sat at exactly 20 GB and therefore
survived a literal `UsedGB > 20` cap — a hollow test that passed the very shape it forbids. Its
figure is now 120 GB and the mutation fails it.
### v0.191.2 — a quiet fill check now says so (2026-08-02, R-167) — MinAgent: none
**Earned during v0.191.1's own live validation, which is the strongest evidence it was needed.** After
the customer had been warned about `/mnt/sys_drive`, the controller was restarted and produced **zero
`fillwatch` log lines** — and that was read, correctly, as unusable: an absent line is equally
consistent with *"the check ran and chose silence"* and *"the check never ran"*. Proving the checker
was alive required deliberately crossing into the critical band.
That ambiguity is **permanent, not rare, for this check specifically**: it is edge-triggered, so the
HEALTHY STEADY STATE IS A QUIET RUN. Standing rule 3 aimed at the one place it costs most.
`Check` now logs a per-RUN summary — `checked N filesystem(s), M unreadable/skipped, K
notification(s); bands: …` — on every run, healthy or not. **Unreadable is counted separately from
healthy**, so a drive that has quietly gone unreadable for weeks cannot read as "all fine". Two tests
pin it, including that a second, equally quiet run logs again (the observable is per run, not per
change).
### v0.191.1 — the fill check also runs at startup (2026-08-02, R-167) — MinAgent: none
**Found while live-validating v0.191.0 on guest 9201: the fill check was reachable only on its daily
schedule, and neither `sched.Daily` nor `sched.Every` fires on registration — both wait for their
first tick.** So a box that BOOTS with a filesystem already over the line would have stayed silent for
up to 24 hours. That is the R-100 shape — a real fault visible only after a deadline elapses — and the
hub's own checkers handle exactly this case deliberately, leaving already-breached keys unseeded at
init so their first `Check` emits (the F2 lesson, `monitor/storage_fill.go`). A daily-only schedule
here would have been the same gap one component over.
The watcher now also runs **once, 90 s after startup**. The delay lets mounts settle and the drive
gate tick first, so a drive still coming back reads as unreadable and is skipped (§8.4) rather than
warned about. It is safe to add because the check is edge-triggered against PERSISTED state: a
filesystem the customer has already been warned about stays silent, so this adds a warning only where
one is genuinely owed. Pinned by an AST assertion in `TestMainWiresTheFillWatcher` — the schedule
registration alone no longer satisfies it.
**Disclosure:** this also made the flow live-validatable at all. There is still no operator-triggerable
"run the fill check now" path; that is recorded as an observation, not fixed here.
### v0.191.0 — warn before the wall comes down (2026-08-02, R-167 · R-158 · R-174) — MinAgent: none
**Storage monitoring and backup alerts, decision D-c, landing BEFORE the `mp1`→`mp0` merge (D-a /
R-165) rather than with it.** D-a's own condition (2) says the monitoring ships in the same step and
never after, because the merge removes a wall that currently fails safely. Landing it first is
strictly better and costs nothing: the warnings go in and get proven on hardware while the wall is
still standing. **No disk layout is touched in this release.**
**R-167 — the customer is warned BEFORE a filesystem fills, and the warning is the pair that already
existed.** New `internal/fillwatch`. Nothing warned before this: the first sign of a full filesystem
was a backup that did not happen, and the only related signal — the healthcheck's generic
`health_degraded` at 90% — looked at REGISTERED STORAGE PATHS ONLY, so the docker area (`mp0`) and
the system-data area (`mp1`, which holds every driveless app's retained recovery unit) were invisible,
and it never reported free bytes or named a drive.
**`disk_warning` / `disk_critical` WERE ALREADY A COMPLETE PIPELINE WITH NO PRODUCER** — allowlisted
in the hub, carrying Hungarian copy, sitting in `settings.DefaultEnabledEvents`, with a UI checkbox
(`event_disk_alerts`) — and `grep` across all four repos found **zero emitters**. The sixth "built but
never wired" instance in this project. This release is their producer; minting a new near-duplicate
type would have left the pair inert forever.
- **Two threshold terms, whichever trips first** — used ≥ **85%** OR free < **5 GiB** (critical: 95% /
2 GiB). A percentage alone lies at both ends of this fleet's size range: 85% of a 20 G backup area
leaves 3 G, less than one DB-backed app's recovery unit (up to ~2× its data — measured 21.1 GB →
40.2 GB, `07-backup-architecture.md` §7.5), while 85% of a 4 TB media drive leaves 600 G.
- **Edge-triggered on ESCALATION ONLY**, state persisted across restarts. De-escalation is silent and
re-arms. **Hysteresis dead zone** between clear (75% / 7 GiB) and warn holds the previous band, so a
filesystem on the line does not flap; the gap is pinned by a test, because a warn and clear
threshold that can be edited into equality is a flapping bug waiting to be introduced.
- **The hub owns cooldown — no controller-side timer** (the `offboxEnlargeBlockedNotify` precedent).
- **A nil usage read is NEVER a warning.** An absent, unmounted or unreadable filesystem is the drive
gate's business and already has its own alert; calling it "full" would be a false alarm with a
misleading cause. It also does not CLEAR an existing warning — a blipping drive must not silently
retract a true alarm.
- **Per FILESYSTEM, never per app** — one full disk holding ten apps would fire ten times, nine of them
noise. Watched: the app-data volume, the system-data volume, and every registered non-decommissioned
drive, de-duplicated by path and resolved at check time (a drive added between checks needs no
restart). Daily at **03:30**, deliberately before the nightly app-data legs so a customer about to
lose a backup to lack of space hears about it with a night's margin.
**R-158 — the operator hears about a failed per-app backup.** `captureAllRecoveryUnits` logged
`[WARN] Recovery unit capture failed for %s` and stopped there; the manager carried three notify seams
and none for the unit capture. `/backups/apps` is the page a person opens to ask whether ONE app is
backed up, and it was the one page that never said. New `unitNotify` seam + `SetUnitNotify`, fired
**per app with the loop continuing** (one app's failure neither aborts nor silences its siblings), and
carrying the target filesystem's used/free bytes at the moment of failure — the overwhelmingly likely
cause is a full filesystem, and those numbers answer "why" without an operator logging in. `UnitSpace`
is **nil when the filesystem is unreadable** and renders as *"unavailable"*, never as zeros: "0 GB
free" and "we could not look" are opposite diagnoses.
**It is OPERATOR-TIER (`recovery_unit_capture_failed`), deliberately NOT `backup_failed`.** That type
carries a `customerMessages` entry AND sits in `DefaultEnabledEvents`, so reusing it — which is what
R-158's own proposal said — would email the customer, in Hungarian, that their backup failed, about
something they cannot act on. D-c routes it to the operator and overrides the proposal. Operator-only
is enforced by the hub's `notify.operatorOnlyEvents` register, **not** by the absence of a
`customerMessages` entry; v0.78.0 claimed the latter and was wrong, and there is a red-proof here
demonstrating the customer receiving it when the register entry is removed.
**R-174 — the app-stop guard stopped starting apps onto missing drives. A regression in v0.189.0 code,
found by review on 2026-08-02 and closed the same session.** `appStopGuard.SetStarter(stackMgr)`
handed `Recover` the RAW stack manager, whose `StartStack` has no drive gate. The guard runs **at
startup** — exactly when an external drive may not have come back — so a backup that stopped an app,
followed by a power cut and a drive that did not remount, ended with the app started onto a missing
drive. **This is R-171 one path over**, and the rule is not new: the API's own
`startGatedByMissingDrive` already refused this to the customer.
- The starter is now wrapped in `gatedAppStopStarter`, using the **same** drive predicate the boot
sweep uses. `bootDriveGate` could NOT be reused whole and the reason is recorded in the code: its
holder #2 reads `bootAppStopGuard.HeldStacks()`, which during `Recover` is **the guard's own marker**
— it would refuse every recovery it was meant to perform — and holders #1/#2 read package-level vars
assigned *after* `Recover()` runs, so a whole-gate reuse would be correct only by accident of
nil-safety. Holder #3 (the drive) is extracted into `driveStartGate`, which now has two callers and
one implementation; a test pins that `bootDriveGate` keeps delegating to it.
- **A refusal is not a failure.** New `ErrStartRefused` + an `AppStopRecovery.Refused` bucket. Both
keep the marker — the operation is genuinely unfinished — but only `Failed` alarms. Collapsing them
would push a deliberately-held app into `NotifyBackupFailed`, a customer-enabled type, producing
exactly the R-171 false alarm this fix exists to prevent. `main.go` now guards the notify with
`Alarming()` rather than `!= nil`, and the pre-existing seam test was **tightened** to require it.
- Fail-safe per R-171's contract: cannot determine ⇒ do not start. The check stays per-caller and was
deliberately NOT pushed into `Manager.StartStack` (fourteen callers, most legitimate).
**Tests:** 1157 → **1184** (+27). Every red-proof demonstrated failing and restored — the gate removed
from the guard's starter; the `unitNotify` call removed; the edge trigger removed (fires twice); the
clear thresholds edited into equality; and `recovery_unit_capture_failed` removed from
`operatorOnlyEvents`, which showed the customer receiving an operator event. Seam wirings are pinned by
walking `main.go`'s **AST**, not `strings.Contains` — a red-proof that comments out `SetNotify` fails
the test while the string is still in the file.
### v0.190.0 — the boot-recovery story finished, and a regression v0.189.0 opened (2026-08-02, R-157 A · R-170 · R-171)
**R-171 — a regression introduced by v0.189.0, found by reading the diff and CONFIRMED on hardware
before anything was written.** v0.189.0 correctly replaced `isBootOrphan`'s `len(Containers) > 0`
term with the customer's recorded intent. But the **drive-absent gate** stops apps with `compose
down` (zero containers) and never touches `desired_state`, because it is not the customer — so a
gate-stopped app began reading as a boot orphan. Observed live on guest 9201 with the drive held
unmounted:
```
[gate] drive ABSENT /mnt/felhom-drives/hdd_1 — stopped+blocked 2 app(s): [calibre-web immich]
[bootrecon] Boot reconciliation: 1 boot-orphaned app(s) found: [calibre-web] — up to 2 attempt(s)
[bootrecon] attempt 1/2: start "calibre-web" failed … attempt 2/2: … gave up
[bootrecon] recovered=[] still down=[calibre-web] (the dead-app alarm now owns these)
```
The **write** hazard did not materialise: compose failed with `mkdir /mnt/felhom-drives/hdd_1/
userdata: permission denied`, because the unbound mountpoint is host-root-owned and the guest is
unprivileged. **That protection is accidental** — no code chose it, no test pinned it, and it is one
`chown` (or one privileged guest) away from gone. The harm that DID occur is real on every box: two
wasted attempts and a **false dead-app alarm for an app the drive gate is deliberately holding**.
The fix is not a new rule. The **API's own start path already refuses this** —
`startGatedByMissingDrive` returns a Hungarian refusal to the customer — and the sweep bypassed it by
calling `Manager.StartStack` directly. New consumer-side seam `bootrecon.StartGate`, wired in
`main.go`, gives the sweep the same question to ask. Fail-safe by contract: **cannot determine ⇒ do
not start.** New `Manager.DriveLive` reuses the userdata belt's own `isMountPoint` seam so the two
cannot drift. Held apps are reported as `HeldByDrive`, deliberately **not** as `StillDown` — that is
the alarm's bucket and putting them there is the false alarm being removed.
Evidence: `felhom.eu/documentation/audits/DIAG-bootrecon-drive-absent-2026-08-02.md`.
**R-157 mechanism A — the sweep that looked once.** `runBootReconcile` waited 5 s and swept exactly
once, deriving its candidate set from a fleet docker was still restoring; measured failing on **three
of six hard resets**. It is now a **settle-then-sweep window**: sample the fleet (name, state,
container count) every **5 s**, call it settled after **3 identical samples**, and sweep **once**, at
the end, on a settled fleet. The window terminates on whichever comes first — settled, or a **50 s
budget** — and the log says which, because "settled and found nothing" and "ran out of time still
churning" are different facts about the box.
**The budget is 50 s and not 60 s because a test said so.** `bootReconcileSettle` (5 s) + budget +
one `DefaultRetryDelay` (30 s) must stay under `deadAppBootGrace` (90 s) so a successful recovery is
SILENT. 60 s was the first choice; `TestBootWindow_CommonCaseFitsInsideTheDeadAppGrace` rejected it
at 95 s. Extending the grace to fit was rejected outright (§8.3) — that hides a late recovery instead
of reporting it. A window that genuinely overruns now emits a **`LATE RECOVERY` WARN naming the
apps**, so a stale alarm never stands without counter-evidence.
**The sample REFRESHES first, and that was found by live validation rather than review.**
`GetStacks()` returns the Manager's in-memory map, which the scheduler refreshes on its own **10 s**
cadence — so sampling it every 5 s without refreshing means two consecutive samples can be identical
because *the cache did not update*, not because the fleet settled. Observed on 9201: a container
removed ~5 s before the window closed was still in the sampled fleet, and the sweep logged
`no boot-orphaned apps` for an app that had none. `sampleBootFleet` now calls `RefreshStatus()` first
(≈10 extra cheap `docker ps` calls per boot); a refresh error degrades rather than aborting.
**Sampling is otherwise read-only and there is still exactly ONE sweep.** Sweeping per sample was rejected: the
sweep's own `StartStack` changes the fleet, so it would never observe a settled one. The per-app
attempt bound is untouched — this widens a bounded window, it does not remove the bound.
**Widening the window made two more holders reachable (§8.2), so the gate covers all three.** The old
T+5 s sweep never overlapped a **quiesce** (starting an app mid-backup defeats the point of
quiescing) or a **running app-data operation** (restarting an app under its own tar). Both are now
refused through the same seam, reusing `quiesce.SuppressedStacks()` and a new read-only
`AppStopGuard.HeldStacks()` rather than second implementations.
**R-170 — the second boot gate stops guessing.** `shouldRecreateOnBoot` still ended in
`&& hasContainers`, so the two boot gates disagreed about the same question. It now reads
`desired_state` with the identical three-way table: `stopped` → never; `running` → recreate whatever
the container count; **absent → exactly the pre-v0.190.0 `hasContainers` behaviour**. Its comment
argued at length *for* the container count and has been rewritten — a correct implementation under a
comment arguing the opposite is worse than either alone. **`presentStable` is untouched and still
load-bearing**: an app whose drive is absent is never recreated here, which is the very term the boot
sweep was missing. The agreement between the gates is pinned from **both sides** against an identical
fixture table (`TestBothBootGatesAgreeOnIntent` / `TestShouldRecreateOnBoot_AgreesWithBootrecon`),
because the two cannot be called from one package without an import cycle.
**Tests: +25 across 3 packages (27/27 packages green).** Timing is tested by shrinking the window
constants, never by sleeping. Red-proofs, each observed FAIL then restored: A (restore the
single-sweep shape), B (remove the budget → the test **hangs**, the unbounded shape), C (drop the
`desired_state: stopped` branch → the customer-stopped app is started **twice** by the widened
window), D (restore `&& hasContainers`), G (remove the start-gate check → the drive-absent app is
started), H (comment out `SetDriveGate` → fails while the string is **still present**, which is what
the AST walk is for).
No hub change, no agent coupling, no user-visible string, no backup/restore/catalogue change.
### v0.189.0 — the box stops guessing what the customer wanted (2026-08-02, R-166 / decision D-b)
**The defect.** When an app was not running, the controller had to work out *why*, and it worked it
out by **counting containers**: zero containers meant "the customer stopped it" (leave alone), some
containers meant "something broke" (recover). That inference is wrong in two ways, and both were
silent:
- a **power cut mid-compose** or an **interrupted deploy** also leaves an app with zero containers —
read as a deliberate stop, so the app simply stayed gone until a human noticed (**R-157 mechanism
B**);
- a **backup that stops an app** to copy it safely, then dies, leaves it stopped with **nothing on
disk** recording that a backup stopped it or that it was owed a restart.
Neither is a guess the controller should be making, because the one fact that settles it — what the
customer actually asked for — **was written down nowhere**. `app.yaml` recorded that an app was
*installed*; it never recorded whether it was meant to be *running*.
**Part 1 — desired state, owned by the customer's action.** `AppConfig` gains `desired_state`, a
**tri-state** `""` / `running` / `stopped` (`yaml:"desired_state,omitempty"`), with named constants.
`Manager.SetDesiredState` is the only writer, and its callers are the only places a human's decision
enters the system: the `/api/stacks/{name}/{action}` switch (`start`/`restart`/`update` → running,
`stop` → stopped), `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, and the `.fab` import
adapter. Intent is written **BEFORE** the act, and an action whose intent cannot be recorded is
**REFUSED** — proceeding would recreate the ambiguity being removed.
**`StartStack`/`StopStack` are deliberately NOT writers.** A census on 2026-08-02 found **14 call
sites, of which exactly 2 are the customer**; the other twelve are machines (quiesce, the backup
volume dump, offbox reconstitution, app export/restore, the storage drive-absent gate, the migration
engine, the boot reconciler itself). Recording intent in the primitive would make a nightly backup
indistinguishable from the customer pressing Stop — the exact confusion this release ends.
**ABSENT MEANS UNKNOWN, NEVER "running" — the single most important line in the change.** Every
`app.yaml` on every existing box predates the field, so absent is what the whole fleet reads on
upgrade. Treating it as running would start, on the first boot after the upgrade, every app its owner
had deliberately stopped. Where intent is unknown the boot reconciler falls back to the **old
container-count rule, byte-for-byte**, rather than inventing an answer.
**The boot-orphan decision (`bootrecon.isBootOrphan`), replacing the container-count term:**
| `desired_state` | containers | result |
|---|---|---|
| `stopped` | any | never an orphan |
| `running` | 0 | **ORPHAN** — the R-157 case, invisible before this release |
| `running` | >0 + down | ORPHAN (unchanged) |
| `running` | >0 + up | not an orphan |
| absent | 0 | not an orphan — **exactly** the pre-v0.189.0 behaviour |
| absent | >0 + down | ORPHAN — **exactly** the pre-v0.189.0 behaviour |
`Protected` and `Deploying` guards unchanged. A **running-only** startup backfill converges apps that
are deployed AND observed up; `stopped` is **never** backfilled, from any signal — inferring it from
zero containers is the defect itself, so an ambiguous app stays ambiguous and keeps legacy behaviour
until the customer next presses a button.
**Part 2 — the app-stop crash marker (`backup.AppStopGuard`).** `<data_dir>/appstop-state.json`,
atomic (tmp + **fsync** + rename, 0600), modelled on the quiesce marker and deliberately **its own
file** — same shape, different owner, different lifetime; sharing would give one file two writers.
Written **before** the stop, cleared only after a restart that **succeeded**; a FAILED restart keeps
it so the next startup retries. `Recover()` runs at startup and **completes before** the
boot-reconcile goroutine is launched, so an app the marker explains is not also reported as an
unexplained boot orphan. A corrupt marker is quarantined loudly, never silently skipped.
**A `defer` is not the mechanism, and the code says so.** Campaign 8 fault 10 established on live
hardware that a SIGKILL runs no deferred function; the marker is what covers the hard crash. Its
test simulates a real abort (an unwind that skips the restart statement) rather than a graceful
return — an earlier version of that test called `Begin` itself and **survived the red-proof that
deleted the production call**, which is exactly the hollowness §10 exists to catch.
**All three stop-and-restart sites are covered**, with no uncovered sibling to imply the class is
handled: `DumpAppVolumesSafe`, `offbox_reconstitute.go` (all four bring-up paths, via one
`restartStack` closure so the success path cannot silently skip the clear), and `appexport`'s export
— the last through a two-method consumer-side seam so the exporter shares the ONE marker file instead
of opening a second. The reason string lives only in `backup`; the adapter in `main.go` supplies it.
**Also fixed, and it would have silently eaten this feature: `SaveAppConfig` rebuilt `AppConfig`
field-by-field.** That is the R-100 shape (v0.181.0 shipped with two live instances of it). The
literal named five fields, so the sixth — `desired_state` — would have been **dropped on every save**,
and nine call sites share that path: a customer's Stop would have been erased by the next unrelated
`app.yaml` write. Replaced with copy-and-overlay (`saveCfg := *cfg`), safe by construction. Measured
and documented: `app.yaml` does **not** round-trip keys the struct does not model (the trip goes
through the struct), pinned by `TestSaveAppConfig_UnknownYAMLKeysAreDropped`.
**Operator visibility (§2.4).** An interrupted operation rides the **existing** `backup_failed` event
type. A new type would need the hub's `allowedEventTypes` + `customerMessages` pair changed — a wire
change, and this release ships **no hub change and no hub version bump**. `Recover()` **returns** its
outcome rather than pushing it through a notifier seam, because it must complete before the boot
reconciler (`main.go:~236`) while the notifier is not constructed until `~307`; a seam wired after the
fact is a seam that never fires.
**No user-visible string changed** — N/A for UI work. No template, funcmap, notifier-type, event-type,
backup-content, retention, tier or restore change. **No agent coupling; MinAgent unchanged.**
**Tests: +37 across 5 packages (27/27 packages green).** Red-proofs, each observed FAIL then restored:
B (restore `len(Containers) > 0`), C (treat absent as running — the fleet-wide upgrade regression),
D (drop the up-state guard from the backfill), E (delete the production `Begin` call), H (restore the
field-by-field `SaveAppConfig` literal), §8.2 (move the intent write below the action switch), and
the seam test I — which fails while the commented-out call **is still present as a substring**, the
distinction that made the controller's first version of that test pass its own red-proof in
2026-07-21.
### CI — the gate entry point runs on every push (2026-08-02, R-168) — NO VERSION BUMP
**No version bump, no build, no deploy** — this adds a workflow file only. Stated explicitly so the
omission reads as a decision rather than a miss.
**`.gitea/workflows/gates.yml` (new).** Triggers on `push`, `runs-on: felhom-gates`, obtains the
source with a shallow `git fetch` of the **exact pushed SHA** from the in-cluster Gitea Service, and
runs this repo's entry point with `--fast` — nothing else. **No `uses:` step anywhere**: JavaScript
actions need a node runtime the host-mode runner does not have, and probe P3 measured a plain
`git fetch` as sufficient. No `|| true`; the entry point's exit code IS the job's result.
**It REPORTS, it cannot REFUSE**, and the workflow header says so: this repo pushes straight to
`main` with no pull request, so there is no merge for a status check to stand at. The refusing half
is `.githooks/pre-push`, which is per-clone and `--no-verify`-able; this half notices when that was
skipped. Making CI blocking needs branch protection plus a PR workflow → felhom.eu `OPEN-ITEMS.md`
R-169, an operator decision.
**A failed run emails the operator** via Resend and prints the provider's accepted id, because probe
P5 measured that Gitea itself sends nothing at all on a failed run. Demonstrated end to end on a real
red run (`RESEND-ACCEPTED id=…`), not assumed. Full detail:
`felhom.eu/documentation/audits/SPIKE-ci-runner-2026-08-02.md`.
**CI reproduces the workspace's SIBLING LAYOUT on purpose.** This repo's entry point invokes the
shared `reuse_refs_check.py` that lives in the `felhom.eu` clone next door and is deliberately never
copied here, and this repo's `REUSE.md` cites `wgsync/reconciler.go`, which lives in the hub. The workflow
therefore clones `felhom.eu` as a sibling; without it the gate fails **closed** with
`gate is MISSING` — correctly, but for the wrong reason. Verified that CI and the local hook then
agree exactly: 133 cited paths, 126 exact / 6 suffix / 1 cross-repo, 0 failures.
### Gate enforcement — one entry point + pre-push hook (2026-08-02) — NO VERSION BUMP
**Deliberately no version bump, and no build or deploy.** Nothing compiled changed: this touches
`controller/scripts/` and `.githooks/` only, so no behaviour on any box moves. Stated explicitly so
the omission reads as a decision rather than a miss.
**`controller/scripts/docker_run_volume_path_gate.py` — one allowlist entry, in its own commit
(`c432f70`).** The gate was RED, flagging `internal/appexport/estimate.go:179`. The finding is
benign: `realVolumeSize` mounts a **named Docker volume** read-only into a throwaway alpine to `du`
it from a container view — no host path is involved, the daemon resolves the volume name daemon-side,
and it is structurally identical to the already-allowlisted `internal/backup/backup.go` entry. The
gate was right to demand review; that diff **is** the review, on its own, because burying an
allowlist widening inside a feature commit is how an allowlist stops meaning anything.
`realVolumeSize` was not touched — the code is correct; the allowlist was incomplete.
**`controller/scripts/controller_gates.py` (new) — THE entry point.** A census of all thirteen gate
scripts across the four felhom repos found that every check a `CLAUDE.md` names was passing and two
of the four nobody is told to run were failing. This repo had seven gates and `CLAUDE.md` named
two; four more were reachable only through a line in `REUSE.md`, and the docker-`-v` gate through one
line in `REUSE.md` and nothing else — while RED. The runner invokes all seven plus `reuse_refs_check`
on the repo root, streams each gate's own output, and exits worst-wins non-zero. `--fast` selects the
gates that touch no network and no container runtime; today that is all eight.
**The shared checker is never copied here.** `reuse_refs_check.py` lives in `felhom.eu/scripts/` and
is invoked across the workspace at `<repo-root>/../felhom.eu/scripts/`. A copy would recreate exactly
the drift it exists to detect. If the sibling clone is absent the gate **FAILS** and prints the path
tried — fail-closed. On this repo it now resolves 133 cited paths: 126 exact, 6 by suffix, and
`wgsync/reconciler.go` cross-repo into the hub.
**`.githooks/pre-push` (new)** — runs `controller_gates.py --fast` and refuses the push. Per-clone
(`git config core.hooksPath .githooks`; a manual run WARNS when the clone is unarmed) and
`--no-verify`-able on purpose; both limits are written into the hook. CI is the unbypassable half and
is owed — `felhom.eu` `OPEN-ITEMS.md` R-168.
**`controller/scripts/test_controller_gates.py` (new, 4 tests)** — a SEAM test asserting each member
gate's own distinctive stdout, never the runner's summary line, which an inert runner prints while
calling nothing. Red-proofed: replacing `run_gate`'s body with `return 0` still prints "all controller
gates OK" and exits 0, and turns the seam test red.
### v0.188.0 — D5: an app restore works from the drive alone (2026-07-30) — MinAgent 0.113.0 (unchanged)
**Tier-1/Tier-2 no longer depend on the whole-guest tier.** Until now the recovery unit on the
customer's drive was secret-free, which made the two-lane split *look* independent while it was not: the
app's files sat safely on the drive and could not be brought back, because the secrets that make them
readable went down with the guest. After this, restoring an app needs **the drive and nothing else** —
not the server, not the operator, not the offsite copy.
**Part 0 first: the brief's own recommendation did not survive the test it asked for.** It proposed that
only `data_key`-flagged secrets travel. Two findings overturned it, both evidenced before any code:
1. **The `data_key` flag is not a trustworthy classification.** Only 5 fields across 4 apps carry it,
yet the catalog's own Hungarian labels contradict the flag elsewhere: `n8n/N8N_ENCRYPTION_KEY`
(„Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"),
`calcom/CALENDSO_ENCRYPTION_KEY`, `bookstack/APP_KEY` — same label as `adventurelog/SECRET_KEY`,
opposite flag. Travelling "only data keys" would omit real data keys, and the fail-closed gate would
not fire for them → a restore that succeeds onto unreadable data. Filed **R-127**.
2. **A DB password is not resettable in practice — proven, not argued.** `DumpAppVolumes` dumps every
compose named volume with no DB exclusion, so `immich_postgres_data.tar` is captured and restored.
Probe on `postgres:16-alpine` (seed → drop container, keep volume → redeploy with a regenerated
password): the **replay succeeded** (`docker exec psql`, no password — verbatim what `ImportDump`
does, and the image's local socket is `trust`), the **app path failed** over the compose network
(`FATAL: password authentication failed`), and the **old** password still worked — `POSTGRES_PASSWORD`
is ignored once PGDATA is non-empty. So the restore reports success, the dump replays, the rows are
there, and the application cannot reach them. 18 DB/root-password fields affected. MariaDB fails
louder: `getMariaDBPassword` reads the regenerated value from container env against a datadir holding
the old hash, so the replay itself gets Access denied (`nextcloud`, `romm`).
**The rulings (operator, 2026-07-30).** `type: secret` travels; `type: password` never does; minus a
code register. Plaintext, as the data already is.
- **TRAVELS (45 fields):** the 5 declared data keys, 18 DB/root passwords, 22 internal signing/encryption
secrets. Every one decrypts data on the SAME drive or authenticates to a container on an internal
compose network with no external listener — possessing it adds nothing to possessing the drive, which
is exactly D2's argument for plaintext DATA.
- **WITHHELD (8):** the 7 `type: password` admin/UI logins + `vaultwarden/ADMIN_TOKEN` via the
`nonPortableSecrets` register. These authenticate against published services, so their reach is NOT
bounded by the drive. **Excluding this class is what licenses the plaintext ruling; the two are coupled
and must not be relaxed independently.** The register is code, not a catalog flag — a boundary a
catalog push can silently move is not a boundary (cf. R-97a).
**What a customer must possess to complete a Tier-1/2 restore after this change: the drive.**
**Implementation** — one place per side, no parallel path. `stacks.PortableSecretEnvVars` is the whole
boundary; `GetStackRecoveryInfo` decrypts the portable class through the SAME `LoadAppConfigDecrypted`
the restore side uses; `buildUnitAppYaml` (was `buildStrippedAppYaml`) writes it at **0600** and names
the withheld class in the header so an operator can see WHY a credential is absent rather than suspect a
capture bug; `readUnitEnv` splits it back using the **manifest's** portable names, never guessed from key
names. `reconcileRestoreSecrets` stays a pure function — the new source arrives as an **argument**.
Manifest → **schema 2** + `portable_secret_env_vars` (NAMES only; the manifest is 0644).
**Precedence: the UNIT WINS.** Not "newest wins". The unit's secrets are captured in the same run as the
dumps beside them (`runVolumeDumps` → `captureAllRecoveryUnits`), so the unit's value is the one that
matches the data about to be restored; the guest's is merely the most recent. A rotated data key does not
decrypt data encrypted with the old one, and a rotated DB password does not match the hash in the
restored data directory. Pinned in both directions — an undefined precedence between two sources of a
decryption key is a data-loss bug waiting for its first disagreement.
**The fail-closed gate is unchanged and still fail-closed:** a data key in NEITHER source refuses
outright. D5 makes it normally present; "normally" is not a reason to soften a gate.
**Three comments that asserted invariants D5 makes false were corrected, not left to read as settled**
(`CaptureRecoveryUnit` "NEVER writes a secret value", `RestoreFromRecoveryUnit` "no secret is read from
the unit", `appbackup/paths.go` + `appdata.go` "secret-free"), and the O4 WARN that claimed "stored data
is unaffected" for every non-data-key secret now says what is true — finding 2 disproves it for DB
passwords.
**Backward compatible.** A schema-1 unit carries no secrets and still restores from the guest; the next
capture rewrites it (the app.yaml checksum changes). No existing backup changes, no data moves, and the
escrow / whole-guest / offsite tiers are untouched in code — the offsite copy simply carries the secrets
inside the unit it already pushed, encrypted at rest under the customer-owned restic password.
**Tests** — `TestRestoreFromRecoveryUnitWithGuestAbsent` is D5's claim as a test rather than a
description; plus fail-closed-with-both-sources-absent, precedence both directions, the schema-1
no-regression case, `readUnitEnv` splitting, and the wrong-outcome check that the withheld class appears
NOWHERE in the unit. Fixtures come from a unit written by the **real** `CaptureRecoveryUnit`, so the two
sides meet at real bytes. Seam: `Manager.stackProvider` only. **Four red-proofs, each verified to have
landed:** drop the portable merge → the consequence test fails; neuter the gate → 4 failures; flip
precedence → the unit-wins test fails; widen the class to `type: password` → the boundary test fails.
**R-120's gate does not apply to this task** — it sits in `hub/internal/web/configs.go`
`handleSetArtifacts`, the golden-**vouch** form, and never runs on a controller image deploy. Re-baking
the golden is a follow-on so that FRESH installs get D5; it is not a prerequisite here.
### v0.187.0 — R-108: network storage may not host an app's data namespace (2026-07-30) — MinAgent 0.113.0 (unchanged)
**This is D5's precondition, and it is now met.** D5 moves app secrets into the local recovery unit so
Tier-1/Tier-2 restore stop needing the guest; that is safe only once no browsing surface can reach the
backup tree. One could.
**The chain, confirmed at source end to end.** `namespaceRoot(drivePath)` returns any non-system drive
path AS-IS (`internal/backup/backup.go:262`), so an app's namespace root IS its `HDD_PATH`. Its recovery
unit therefore lands at `<HDD_PATH>/backups/primary/<stack>/` (`appbackup.RecoveryUnitPath`). Put an app
on a NAS and that directory sits inside the share, which FileBrowser binds **whole** — share ROOT,
`:rslave`, `download: true`. Live on demo-hp, the asymmetry visible in one glance:
- /mnt/felhom-drives/nvme-1tb/userdata:/srv/nvme-1tb <- drive: userdata-SCOPED
- /mnt/felhom-drives/Felhom-Share:/srv/Felhom-Share:rslave <- share: ROOT
**Why the bind was NOT narrowed** (this was the real finding, and it inverted the fix). The share-root
`:rslave` bind is **load-bearing**, not an oversight: a Phase-0 probe (2026-07-22) proved an in-container
access through it wakes the idle automount trigger, so narrowing it breaks NAS access itself. And
scoping is not even definable — apps on a share store at `<share>/<app>`, there is no `userdata/` layer,
and creating one would write Felhom's directory convention onto a customer's own NAS, which R-67
forbids outright. So the browsing surface cannot be narrowed, and the backup tree must therefore never
be placed under it. **Operator ruling 2026-07-30: refuse the placement, keep the browse bind.** Tier 2
already refuses network targets for this same class of reason (`F-6C-1`); this closes the PRIMARY
namespace, which was the last way a `backups/` tree could appear inside a share-root bind.
**Nothing is stranded.** Verified across all six hub customers including Peti: zero apps on network
storage. demo-hp's `Felhom-Share` holds only the customer's own files (no `backups/`). R-67's browse
capability is untouched — same bind, same `:rslave`, byte-identical.
**FIVE surfaces, not the four the register named.** `settings.RefuseAsAppNamespace` is the single
predicate all of them consult:
1. **the deploy POST** (`internal/api/router.go`) — **this is the boundary.** The R-108 row says "the
deploy dropdown has no `IsNetwork()` filter", which understates it: the dropdown is a UI list, and
this endpoint accepts whatever `HDD_PATH` a caller supplies, with `DeployStack` validating only that
it EXISTS (`os.Stat`, `internal/stacks/deploy.go`). Filtering the list alone would have left the
surface open.
2. **per-app migrate targets** (`internal/web/handlers.go`) — dropped from the offered list.
3. **`handleStorageMigrateApp`** — refused before `MigrateApp`, so no job starts.
4. **`handleStorageDecommission` mode=migrate, the TARGET** — **not in the register.** The existing
`refuseNetworkLifecycle` guards `req.Where`, the SOURCE; the target was unchecked, so a whole
namespace could be decommissioned ONTO a NAS. Found by enumerating the set rather than trusting the
four that were named.
5. **the FileBrowser bind** — deliberately unchanged, and now pinned by a test so it cannot drift.
**FAIL CLOSED, and the non-obvious part is why this is a function and not an `IsNetwork()` call:**
`/mnt/felhom-drives` holds BOTH kinds in-guest (`.../hdd_1` is a local drive, `.../Felhom-Share` is a
NAS), so a path prefix cannot classify — `Kind` is the only discriminator and it exists only on a
REGISTERED path. An unregistered path under that root is therefore un-classifiable, and un-classifiable
refuses. Every share is registered under that root by construction, so the network set is completely
covered without touching drives.
**UI (§5): marked, not hidden.** A registered NAS stays in the deploy dropdown, `disabled`, labelled
`(hálózati tárhely — alkalmazáshoz nem választható)`, and never pre-selected even when it is the
registry default. A share the customer registered themselves, silently missing from the list they
expect it in, reads as a bug and generates a support question; present-with-a-reason answers it in
place. Follows the existing `(nem elérhető)` disabled-option precedent.
Tests: 9 new, all asserting the **non-effect**. The refusal tests run against a Server with a
deliberately **nil `stackMgr`**, so a guard that fails to fire reaches the mutation and PANICS rather
than passing quietly. They assert no job id, no `started` flag, no `MigratedTo` written, and — for
decommission — that the source was NOT soft-marked. Fixtures are demo-hp's real two-class storage set
(both paths under the same mount root, which is the trap). Red-proofs: 4, each mutation asserted to
have landed first — drop either migrate guard → panic; break fail-closed → 3 tests; userdata-scope the
share → the R-67 regression guard fires, quoting the broken bind.
Suite rc=0, 27 packages, 0 FAIL; `go vet` rc=0; `template_id_gate.py` + `emoji_gate.py` both OK.
- `internal/settings/settings.go` — `RefuseAsAppNamespace` + the two Hungarian refusal reasons.
- `internal/api/router.go` — the deploy-POST refusal.
- `internal/web/storage_handlers.go` — `refuseAppNamespaceTarget`; wired into migrate-app + decommission.
- `internal/web/handlers.go` — migrate-target list filter; `DeployStoragePath.NotAllowed`.
- `internal/web/templates/deploy.html` — disabled option + reason; no pre-select of a disabled default.
### v0.186.0 — R-114 + R-112: tell the truth about the backup target, then show it (2026-07-29) — MinAgent 0.113.0 (unchanged)
Two defects E-2d found on a real box, fixed in this order deliberately: the message is corrected
BEFORE it is put on screen, because switching on a banner that lies is worse than a silent one.
**R-114 — the third state.** `resolveBackupTargetState` had two outcomes: a disk claims the target
(healthy), or nothing does (degraded, "the backup is on the system disk"). The state *configured, and
its drive is gone* has no branch, so it fell into the second and inherited both its message and its
offer. Live payload, target detached: `degraded:true, target:"felhom-backup"` **plus** the system-disk
copy — false, the backup was on a drive that had vanished — **plus** `offer_path` naming that same
vanished drive as the remedy (`felhom.eu` `audits/E2D-fresh-vm-2026-07-29.md` §5.3).
New `BackupTargetState.TargetAbsent` discriminates. `Degraded` keeps its meaning ("is there a
problem") so the wire contract is unchanged for every consumer; `TargetAbsent` answers "which
problem", because the two have OPPOSITE remedies — attach any second drive, versus reconnect *that*
one. Its copy is routed through `degradedMessageFor`, so there is still exactly one place that
decides what a customer reads. **No offer in this state**, suppressed on the branch itself rather
than left to `firstOfferableDrive`'s `Disconnected` skip: that flag is set by the agent-side gate in
another repo (R-113), and this state must be correct independently of it. Belt here, braces there.
The absent copy is **verbatim** the hub's `customerMessages["backup_target_absent"]`, so the banner
and the email tell one story. It now lives in two repos with nothing binding them but a test —
filed as a drift risk, not solved.
**R-112 — the state finally has a consumer.** The endpoint was byte-correct and **nothing in the
product ever asked for it**: templates fetch 18 distinct `/api/storage/*` endpoints, and
`backup-target[/assign]` were the only two with zero references. Server-rendered on the backups page
now (`backupsHandler` → `backupTargetView` → `backups.html`), following the existing
`SingleCopyWarning` banner pattern — not a 19th JS fetch, because a banner that needs JavaScript to
appear is one more thing that can silently not happen. `backupTargetView` returns **nil** for healthy
and unknown, so those render nothing at all: no badge, no reassurance. The offer control POSTs to the
existing assign endpoint behind the standard inline confirm, never auto-submits, and surfaces
`restart_required` honestly instead of adding a self-restart.
**Seam test (Scenario E)** drives `backupsHandler` over httptest and asserts the RENDERED HTML —
handler → view → resolver → template. Deleting the one line that sets `data["BackupTarget"]`
reproduces the R-112 state and fails every render assertion. Tests 326 → 338 (+12) in
`internal/web`; three red-proofs run and reverted.
**MinAgent unchanged at 0.113.0.** R-114 reads `BackupTarget`/`MountPath`/`GuestPath`/`Role`, none of
which R-113 altered — it changed `BoundUnderParent`, which this code does not read. So demo-hp
(agent 0.113.0) is not held.
**NOT LIVE-VALIDATED.** Scenario C cannot occur on a healthy box; Session C proves it.
### v0.185.1 — E-2: the offer endpoints were mounted where nothing routed to them (2026-07-29)
Registered as `/api/backup-target` inside `ServeStorageAPI`'s switch — which `main.go` mounts ONLY at
`/api/storage/`. Live result: `{"ok":false,"error":"endpoint not found"}` while every unit test
passed, because the tests called the handlers directly and never travelled the mount. Caught by the
first live call, which is exactly why the live call is part of the procedure.
Moved to `/api/storage/backup-target[/assign]`. `TestBackupTargetRoutesLiveUnderTheStorageAPIMount`
now asserts the dispatcher's own source contains both paths, so a handler that nothing routes to
fails the suite — the repo's seam-wiring rule ("a feature is not shipped until its entry point is
reachable"), applied to a route rather than a button.
### v0.185.0 — E-2 Parts 3+4: the offer, and the honest degraded state (2026-07-29) — MinAgent 0.113.0
The half that makes the rest work. A degraded backup target recorded only in config is the
silent-degradation pattern this arc has spent a week removing.
**Part 3 — the offer.** `POST /api/backup-target/assign` moves the target via the agent's
`POST /backup/target`. It is the **only** writer of the role: registration does not set it, the
drive-gate does not, no scheduler does. Declining is simply not calling it.
The agent returns `restart_required` rather than restarting itself, and the reason is E-1's own
mistake: restarting with a backup in flight cancels the wait and records a spurious tier failure for
a backup that actually succeeded. The restart belongs to whoever can re-check in-flight work
immediately beforehand.
**Part 4 — visibility.** `GET /api/backup-target` returns the state and, when degraded, the
Hungarian copy:
> „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd,
> lemezhiba ellen nem. Csatlakoztass egy második meghajtót a teljes védelemhez."
FACT → CONSEQUENCE → REMEDY, pinned by a test: a customer told only the fact cannot act on it.
**Healthy renders NOTHING** — no badge, no reassurance, no tonal change. `degradedMessageFor` is the
single decision point so there is exactly one place that could start decorating a working box.
Red-proofed: adding „A rendszermentés védett…" to the healthy branch fails Scenario E.
**UNKNOWN is not degraded.** An unreachable or pre-R-82 agent means we could not ask, which is not
evidence of degradation — the absence-read-as-a-value mistake R-88 Part 2 closed.
**A hollow test caught by its own red-proof.** `TestUnknownStateRendersNothing` originally used
`{Known:false}` with `Degraded` left false, so it passed even with the `!Known` guard deleted — the
second condition covered for it. The fixture is now `{Known:false, Degraded:true}`, which fails
properly when the guard goes. The red-proof is what exposed it; without it the test would have been
decoration.
State is derived from the AGENT, never from our own intent flag: on the two boxes migrated by hand
in E-1 the intent was never recorded while the drive really is the target.
### v0.184.1 — E-2b keying fix: the backup-target branch was unreachable (2026-07-29)
**Caught before deploy by tracing, not by a failure — and the 0.184.0 image is therefore superseded
and must not be shipped.**
`ReconcileDriveGates` resolves the target as `isTarget[a.Path]`, and `a.Path` is the **registered**
`StoragePath` — for an external drive that is the GUEST path `/mnt/felhom-drives/<name>`, not the
agent's host `MountPath` (`/mnt/<name>`) that `/disks` reports. `driveTargetByPath` keyed the map on
`MountPath` alone, so the lookup never matched: **every absent drive, the target included, fell
through to the generic `storage_disconnected`.**
The alarm would have looked wired, passed its own unit tests, shipped, and been silently wrong on
exactly the drive it exists for — the same defect class E-2b was opened to fix, one level down.
Now keyed under BOTH paths, mirroring `planDriveGates`, which already registers `present[]` under
`GuestPath` and `MountPath` for the same reason.
Red-proofed: reverting to MountPath-only keying fails with *"the backup target is not resolvable by
its GUEST path — the gate passes a.Path (the registered StoragePath), so the backup-target branch
would never fire"*.
### v0.184.0 — E-2b + Part 5: the drive-absent alarm that was never wired (2026-07-29) — MinAgent 0.112.0
**`NotifyStorageDisconnected` and `NotifyStorageReconnected` were defined and called from NOWHERE.**
Registered in `allowedEventTypes`, in `DefaultEnabledEvents`, and given a Hungarian message on the
hub — and never invoked. A drive going absent produced apps stopped, a `[WARN]` log and a UI badge,
then **silence on every channel**. Verified against the gitignored-`cmd/` trap with a positive
control. Fifth instance of this class in the project, found by E-2's Phase 0 rather than by a
failure.
A drive that is *only* a backup target has no apps to stop, so it was silent twice over.
`ReconcileDriveGates` now calls both halves. When the absent drive is the **whole-guest backup
target** it raises the more specific `backup_target_absent` (error) instead — never both; two mails
for one event trains people to ignore the channel — and recovers as `backup_target_restored` (info,
the existing pairing-gated pattern; `severityNotifies` is NOT widened). The recovery must mirror the
alarm's choice or the operator cannot match them.
**Which drive is the target comes from the AGENT** (`/disks` `backup_target`, agent ≥ 0.112.0), not
from our own `StoragePath.BackupTarget`: that field is customer INTENT, and on the two boxes migrated
by hand in E-1 the intent was never recorded while the drive really is the target. An older agent
omits the field → false → the generic disconnect alarm, never a wrong one.
Before this, an absent backup target had **no prompt signal at all**: the tier stays DUE
(`targetStoragePresent` checks name presence, never reachability), so the only evidence was the
tier's own failure at its next due cycle — up to ~24 h on the daily local tier. The R-100 shape.
**Tests** observe the WIRE, not a mock, because the failure class is "nothing arrives": a real
`Notifier` posts to an `httptest` hub and the test asserts the event type and severity that actually
went out. A typo in the type string is not cosmetic — the hub 400s it and the event vanishes.
### UNRELEASED — E-2 Part 1: the backup-target role (foundation; NOT yet wired to a UI)
**Status: foundation only. No version bump — nothing customer-visible changes yet.** The field is
written by `SetBackupTarget` and read by `BackupTargetPath`/`BackupTargetAssigned`, and by nothing
else. **The offer UI (Part 3), the degraded banner (Part 4) and the absent-target signal's
controller half (Part 5) are NOT in this commit** — tracked as E-2 in `OPEN-ITEMS.md` so this cannot
become a sixth "seam built but never wired" (the fifth, `NotifyStorageDisconnected`, was found by
E-2's own Phase 0 and is one of the things still to wire).
`StoragePath` gains `BackupTarget bool` — the sibling role to `Schedulable`/`IsDefault`/`Kind`,
marking the drive the whole-guest vzdump is written to.
**It is INTENT, not truth.** The authority is the agent's `backup.local_backup_target`; this records
what the customer ASSIGNED so the controller can render the state, notice the drive going absent,
and detect drift. Truth comes from the agent's `GET /backup/tiers`.
Invariants, each pinned by a test asserting the CONSEQUENCE rather than the mechanism:
- **A drive never acquires the role by appearing.** Registration does not set it; only an explicit
customer choice through `SetBackupTarget` does. Red-proofed: adding auto-elevation to
`AddStoragePath` fails `TestRegisteringDrivesNeverAssignsTheBackupTarget` with
`registering drives assigned the backup target "/mnt/hdd_1"`.
- **Exactly one carrier** — assigning moves the role rather than duplicating it.
- **Sticky** — a new, bigger, faster drive appearing does not steal an assigned target.
- **An absent target stays assigned.** Clearing on disconnect would be a silent retarget by
omission: the box would read "no target configured" instead of "your target drive is missing".
- **A network share is refused** — the role exists to survive a LOCAL disk failure, and a remote,
credential-bound share mounted at its own root (R-108) is a different risk model.
Attributes may suggest and may refuse the absurd; they may never select. The reference hardware
settles it: demo-felhom's backup drive is an external **USB HDD**, and **both** demo boxes' drives
report `removable=0` — a transport rule would disqualify the reference drive, a removable rule would
find no candidate at all.
Green gate: `go build` + `go vet` + `go test ./internal/{settings,web,quiesce}` all rc=0, run
separately from the commit.
### v0.183.0 — C9-F1 + C9-F2: a restore that restored nothing, and a crash loop nobody saw (2026-07-28)
Both are the same shape — the system reporting healthy while the customer is not — and both were
found by Campaign 9 on live hardware.
**C9-F1 (HIGH).** Tier-2 writes TWO things on every run: the capture legs (`hdd/`, `userdata/`) and,
always, a full `recovery-unit/` — the app's DB dumps and named-volume tarballs. `RestoreTier2Files`
reads **only the two legs** (`tier2_restore.go:101-104`) and has never opened `recovery-unit/`. For an
app whose data lives entirely in named volumes that is its ENTIRE dataset, so pressing
„Fájlok visszaállítása" stopped the app, restored 0 files, restarted it, and reported
„Nincs hiányzó fájl — minden fájl megvan a helyén." — at the exact moment the customer pressed it
BECAUSE files were missing, while 156 MB of BookStack's data sat unread in the same copy.
**Phase 0 enumerated all 53 catalog templates** (cross-checked against both demo boxes' actual copies):
**43 apps** have no readable subtree at all — the restore is a guaranteed no-op for them, forever —
**9** have file legs but never their database or volumes, and 1 is stateless. Four apps (`plex`,
`jellyfin`, `emby`, `navidrome`) are in the 43 only because their single bind is a `:ro` media mount,
which `ClassifyBinds` correctly excludes.
Fixed on the honesty axis (completeness is filed as C9-F1b, see below):
- a **pre-flight coverage check refuses UP FRONT** — no op begun, and the app is **not stopped**;
- the refusal **names the action that works** instead of dead-ending 81% of the catalog:
„Ennek az alkalmazásnak az adatai nem ebből a másolatból állíthatók vissza — az alkalmazás nem állt
le. Használd a Visszaállítás indítása gombot a Biztonsági mentés → Visszaállítás oldalon.";
- where the restore DOES run it now claims only what it **examined** —
„Minden vizsgált fájl megvan a helyén." — plus, whenever a unit is present,
„Az alkalmazás adatbázisa és belső kötetei nem tartoznak ebbe a visszaállításba." That second string
closes the QUIET half: immich's 1.3 GB Postgres unit is not covered, so the old blanket sentence was
a clean bill of health over data the operation never opened.
New seam: `Manager.Tier2RestoreCoverage` + `Tier2Coverage{Legs, HasUnit}`, computed from the RECORDED
copy on disk rather than the catalog, so an app whose template changed is judged by what it actually has.
**C9-F2 (HIGH).** `IsDownState` excludes `restarting` as "self-recovering", but with the catalog's
standard `restart: unless-stopped` Docker retries forever — so a crash loop was counted as working.
Campaign 9 watched docmost loop for nine minutes (restartcount 18) while F-OBS's heartbeat printed
„180 scans since boot, 4 deployed app(s) evaluated, **0 currently down**". No banner, no
`app_start_failed`, no email, no hub event, indefinitely.
`StateRestarting` is deliberately **NOT** added to `IsDownState` — that would alarm on every deploy and
update fleet-wide, the over-correction F-A1 nearly cost us. Instead a sustained restarting run becomes
down after `crashLoopAfter = 5m`, justified against three numbers already in this codebase: the deploy
flow's **120 s** health timeout, Mealie's **60 s** `start_period` (the slowest catalog healthcheck), and
R-97b's **180 s** quiesce grace — which the threshold must exceed so the two windows compose into one
bounded delay instead of leaving a gap. Docker's own backoff caps at 60 s, so a real crash loop
registers ≥4 attempts inside the window. New `Stack.RestartingSince` (not persisted, same reasoning as
the R-88 breaker) + `Stack.CrashLooping(now)`, used by BOTH the alarm and the dashboard counter — which
previously counted `restarting` as running, contradicting the alarm on the same screen.
Red-proofs, all observed: crash-loop term removed → A fails; **StateRestarting naively added to
IsDownState → B fails** („every deploy and update would page the operator"); quiesce term removed →
C fails; coverage guard removed → D fails with the app STOPPED; guard made unconditional → E fails
(the paperless regression guard); old blanket message restored → F fails.
**Filed, not fixed:** **C9-F1b** (route class-B apps to the Tier-1 unit restore — it puts a destructive
operation behind a button reached via a non-destructive one, so the confirm copy has to carry that
difference) and **C9-F4** (`backups/secondary/<stack>/recovery-unit/` is written by every Tier-2 run and
read by NOTHING — `RecoveryUnitPath` resolves to `backups/primary/`, so the second local copy that
exists precisely for drive loss is unreachable by any customer action).
### v0.182.0 — R-101 + F-DIAG: the customer must not be told a failed backup is a copy (2026-07-28)
**R-101.** `Tier2LastRun` is the ATTEMPT clock — `recordTier2Failure` writes it too — and it was
rendered as „Legutóbbi másolat" in the **restore confirm dialog**. That is misinformation at a decision
point, not an alarm bug: the restore it guards fills in MISSING files without touching existing ones,
so a customer whose Tier-2 had been failing was told a copy existed from last night, restored, and
silently received **older** files while believing they were recent.
`CrossDriveBackup` gains `LastSuccess` (same rule and shape as the offsite anchor) plus
`SuccessTracked`, which distinguishes "this row predates the anchor" from "this row has one and it is
empty". Without that marker the two are indistinguishable and **all 7 Tier-2 rows on the fleet** would
have flipped to „Még nincs sikeres másolat" on deploy. Legacy rows are migrated truthfully on first
touch: a row whose last known state was `ok` adopts that time; a row whose last state was `error` seeds
nothing, because the old data evidences no success.
Shipped strings: „Legutóbbi sikeres másolat: {dátum}" · „…Figyelem: a legutóbbi mentési kísérlet nem
sikerült, ezért a visszaállított fájlok ennél régebbiek lehetnek." · „Utolsó sikeres: {relatív}" ·
„Még nincs sikeres másolat" with the restore replaced by „Még nincs sikeres másolat, amiből vissza
lehetne állítani." The dialog also stops printing a raw UTC RFC3339 stamp — new `fmtTimeStr` renders
Budapest-local `2026-07-25 03:30`.
**Part 2 — the copy-site hazard, and it was in the path.** The three `record*` helpers each built a
WHOLE `CrossDriveBackup` literal with a helper re-applying exactly two fields; everything else was
zeroed on every status write. Adding `LastSuccess` to that shape would have had `recordTier2Failure`
**clear** it — the mirror image of the defect, firing on the first failure. Replaced with
`tier2Update`, which copies the existing row and overlays the outcome: **safe by construction**, a new
field carries over unless deliberately overwritten. Sweep: `SetTier2Preference` mutates in place
(safe); `SetCrossDriveConfig(name, nil)` is a deliberate delete.
**F-DIAG.** The offsite failure notification was `"…: " + err.Error()` — one string for every cause AND
a raw passthrough. `ClassifyOffsiteFailure` now returns quota / orphaned / no_repo / no_units /
transport / **unknown** (unclassifiable says so rather than being folded into a neighbour), each with
its own Hungarian message.
**The secrets half caught a bug in my own first attempt.** The initial sanitiser regex-matched
`sftp:…` and `user@host` and looked complete; its own test caught it leaking on
`ssh: connect to host <host> port 23: Connection refused`, a bare hostname in neither shape. It now
redacts the target's **actual** host/user/repo-path literally, with the regex kept only as a backstop —
guessing at what a secret looks like fails exactly where it matters.
Red-proofs, all observed failing: dialog back on the attempt clock → `the dialog does not name the last
SUCCESSFUL copy`; gate the restore on `LastRun` → `a tier that has NEVER succeeded still offers a
restore`; make the caution unconditional → `a HEALTHY tier shows the failed-attempt caution`; clear the
anchor on failure → `a FAILED run wiped the success anchor`; raw sanitiser → `the repo reference reached
the message ("sftp:" leaked)`.
### v0.181.0 — R-100: record the last SUCCESS, not just the last attempt (2026-07-28)
The producer half of R-100. `OffboxTarget` gains **`LastSuccess`** (RFC3339), carried to the hub on the
offsite report as `last_success`. The hub's staleness verdict counts from it (hub v0.80.0).
**Why a new field rather than reading `LastStatus`.** `LastRun` is written unconditionally at the end
of every run, failures included — it records an **attempt**. "How long since `LastRun`" therefore
answers "how long since we last TRIED", which is not the question a freshness verdict asks. The
alternative — "`LastStatus == error` ⇒ stale" — turns every transient blip into an immediate alarm,
which is the F-A1 noise failure mode. Anchoring on last success tolerates one bad night and catches a
persistent one, using the threshold that already exists.
The rule is a pure function, `offboxAnchorAfterRun(prev, at, runErr)`, called unconditionally beside
the `LastRun` write. Both directions are bugs if got wrong and both are pinned:
- a failure must not **advance** it → or the original defect survives;
- a failure must not **clear** it → or one bad night makes an established tier read as never-succeeded
(the mirror-image over-correction, and on the hub side the newborn-box path).
**Two silent-wipe sites found and closed**, both of the "seam built but never wired" shape — the field
exists, the writer sets it, and an unrelated routine path zeroes it:
- `offboxConfigHandler` rebuilds the target from the form and copies runtime status field by field, so
an ordinary settings save (edit the host, edit the path) would have erased the anchor;
- `ApplyOffsiteTarget` does the same on a hub re-apply — an established tier reset to "never
succeeded" every time the hub re-pushed its descriptor.
Neither would have surfaced until the hub's verdict changed, days later.
**A hollow test of my own, caught by red-proofing it.** The first version of
`TestOffboxLastSuccess_OnlyAdvancesOnSuccess` re-implemented the rule in a local closure: mutating the
production code left it **green**. That is what the extraction to `offboxAnchorAfterRun` is for — the
test now calls the real rule, and the red-proof bites.
Red-proofs, all observed failing: drop the `runErr` guard → `a FAILED run advanced LastSuccess to
"2026-07-21T02:15:00Z" — that is the R-100 defect in mirror image`; always return `prev` → `a
successful run did not advance the anchor`; drop the wire field → `OffboxReportStatus dropped
LastSuccess — the hub would degrade forever on a controller that has it`; drop the handler
preservation → `a settings save erased LastSuccess`.
### v0.180.0 — F-OBS: the dead-app check gets a positive observable (2026-07-28)
On a default `logging.level: info` box there was **no way to tell whether `deadapp-check` had run**.
Its per-cycle scheduler line goes through `Scheduler.dbg()`, which is gated on `s.debug` — so on an
info-level box the line is never *produced*, not merely filtered, and therefore cannot reach the
always-DEBUG ring either. A 30 s interval also puts the job on the scheduler's quiet path
(`quiet := job.Interval <= 30*time.Second`).
So "no alarms" was indistinguishable from "the detector never ran" — the exact fallacy this project
now has a standing rule against, and it directly undermines confidence in the **F-CRIT-1** fix in the
field: that fix's whole value is that a genuinely dead app now alarms, and an operator had no way to
confirm the thing that alarms is alive.
**A periodic summary, not a line per run.** At 30 s a per-run line is 2880 lines/day, which is
precisely why the original author chose silence — so a fix that floods is not a fix. Every 20th scan
(≈10 minutes) emits one INFO carrying the scan count, how many deployed apps were evaluated, and how
many are currently down. An operator can answer "is it running, and what does it see?" from a default
box, and a STALLED detector shows up as the heartbeat stopping.
10 minutes is chosen to stay useful as a liveness signal: it is well inside the 180 s alarm grace this
check feeds, and a test pins the cadence so nobody can widen it to hours and quietly make the
observable useless again.
### Also
Corrected the comment claiming the quiesce unquiesce is "guaranteed by defer". Campaign 8 fault 10
established that a SIGKILL runs no deferred function — the guarantee is the crash MARKER plus
`Recover()`, which brought the stacks back 1 s after restart. The `defer` covers only the graceful
exits.
Files: `cmd/controller/main.go`, `internal/quiesce/quiesce.go` (comment),
`cmd/controller/deadapp_observable_test.go` (new).
### v0.179.0 — F-CRIT-1 + F-A1: one alarm that never fired, one that fired wrongly (2026-07-28)
Both Campaign 8 findings live in `internal/quiesce` and its `classifyRunStates` consumer, and both
are "the alarm is wrong" — one missing, one spurious. Fixed together, one pass over the same code.
#### F-CRIT-1 — an app that failed to restart after a quiesce NEVER alarmed. **Two independent causes; either alone kept it dead.**
*Cause 1 — the outcome was thrown away.* `restartAll` returned nothing; a failed `StartStack` was
logged and dropped on the spot, so no caller could learn a customer's app had not come back. It now
**returns the stacks that failed**, and both call sites (the cycle's unquiesce and crash recovery)
record the result.
*Cause 2 — a documented invariant the quiesce path had made false.* `classifyRunStates` whitelists
`StateStopped` because v0.164.0 (correctly) refused to alarm on deliberate user stops, resting on
I1: *"StateStopped means deployed, deliberately stopped by the user."* **The quiesce loop stops
stacks by the same `docker compose down` path**, so a stack it stopped and then failed to restart is
also `StateStopped` — byte-identical to a user stop on the Docker side — and was whitelisted into
total silence. Campaign 8 watched a customer app sit dead indefinitely with no banner, no event and
no email while the dead-app scanner ran over it 11 times.
No state test can separate the two; they *are* the same state. The distinguishing fact is that the
loop **tried to restart it and could not**, which it now reports via `Loop.FailedRestarts()`. That
set is the only thing that lifts the whitelist, so genuine user stops stay silent (pinned by
`TestClassifyRunStates_UserStopStillSilent`).
*Why the existing tests missed it:* R-97b's Scenario F asserted that **suppression expires**. It
never asserted that **an alarm follows**. The suppression lifted correctly and the whitelist ate the
alarm one layer down — a green, red-proofed suite over a production path broken two ways.
#### F-A1 — a correct refusal reported as a failure
HTTP 409 from `POST /backup` is the agent's R-85 single-flight gate refusing while a restore-test
holds it. The start path had no 409 branch, so it called `noteTierFailure`: the R-88 breaker armed
and `whole_guest_backup_failed` was emailed. Campaign 8 saw it on both boxes in the same minute.
At real cadences a ~12-minute restore-test against a daily backup collides roughly once per 420
guest-days — about **every 4 days on a 100-guest fleet, forever** — which trains the operator to
ignore the alarm and quietly undoes R-97a.
409 is now **contention, not failure**: `agentapi` returns a typed `*StatusError` on POST, the
adapter maps 409 → `quiesce.ErrTierBusy` (the same seam that maps 404 → `ErrTiersUnsupported`), and
the loop defers instead of failing. No breaker, no event, no email; the tier stays **DUE**.
**Two traps this deliberately avoids.** *Silence:* "just ignore 409" would let a wedged
restore-test block backups forever with nobody told — so contention lasting past
`contentionAlarmAfter` (**3h**) raises its own signal, headlined **BLOCKED**, not FAILED. The bound
is set by the agent's own ceiling, not taste: its PBS restore-test task is capped at 120 minutes, so
contention outliving that is a stuck gate rather than a busy one; 3h adds margin and is 15× the
longest contention actually observed (12m01s). *App thrash:* removing the failure treatment also
removes the breaker's deferral, which had been (accidentally) preventing a re-quiesce every 5
minutes. Without a replacement the customer's apps would be stopped and restarted on **every poll**
for the whole restore-test — worse than the bug. A contended tier is therefore dropped from the due
set **before anything stops**, on a `contentionRetryAfter` of 15m (the longest observed restore-test
is 12m01s; the agent's local restore-test wait is 10m).
#### Comments corrected (three of the six)
`classifyRunStates`' I1 now states what `StateStopped` actually means and names the quiesce path;
`quiesce.go`'s "would record a spurious failure" says that this was not hypothetical until now; and
the agent's `inflight.go` "a caller that cannot acquire DEFERS" records that this was true of the
restore-test caller and not the backup caller. A standing rule was added to both copies of
`CLAUDE.md`: **a comment asserting an invariant needs a test pinning it, or it is a wish.**
Files: `internal/quiesce/{quiesce.go,suppress.go,contention.go (new)}`,
`internal/agentapi/client.go`, `cmd/controller/main.go`, plus new tests
`internal/quiesce/{failed_restart_test.go,contention_test.go}` and
`cmd/controller/failed_restart_classify_test.go`. No wire/contract change; no agent behaviour change.
### v0.178.0 — R-88 Part 2 (controller) + R-97c comment fix (2026-07-27) — **MinAgent: 0.105.0** for the age_state semantics
**The safety valve now needs a licence.** `scheduledRunAllowed` fired on ANY nil age — "no recorded
backup yet, never withhold the first one". With agent v0.105.0 the age carries a STATE, and only a
POSITIVE claim licenses the bypass:
| `age_state` | licenses the valve? | why |
|---|---|---|
| `absent` | **yes** | the agent looked; there is genuinely nothing there |
| `unknown` | **no** | unreadable storage — this is the whole fix |
| `known` | n/a | a real age; the age comparison decides |
| *(empty)* | **yes** | pre-v0.105.0 agent — see below |
**A missing field means LEGACY, not unknown, and that is deliberate.** Reading an old agent's silence
as "unknown" looks safer and regresses Scenario D: the valve would stop firing on every un-upgraded
box, so a genuinely new box would never take its first backup outside its window and nobody would
notice for weeks. Preserving the KNOWN behaviour is correct; the MinAgent floor drives the upgrade.
The degrade is logged **once** per process, the `logTierDegradeOnce` shape. An unrecognised FUTURE
value also maps to legacy — a newer agent inventing a fourth state must not inherit "unknown"
semantics from a controller that has never heard of it.
**Caught while doing it, and worth naming:** `TieredBackend` is satisfied by a RUNTIME type assertion
in `resolveDueTiers`, so when `DueFor`'s signature changed the whole repo still built and vetted
clean while `quiesceBackend` silently stopped satisfying the interface — which would have degraded
every box to the untargeted single-tier path, losing R-82's multi-tier backups entirely, with no
error anywhere. `TestQuiesceBackendSatisfiesTieredBackend` is now the compile-time witness. Sixth
instance of the inert-seam class.
**R-97c follow-through:** the comment in `internal/notify` claiming these event types are
operator-only "because they have no customerMessages entry" was **wrong** and is corrected — the hub
falls back to the raw message when the entry is missing, and the only customer gate is
`prefs.EnabledEvents`. Enforcement is hub-side `operatorOnlyEvents` (hub >= v0.79.0).
Unchanged: the R-88 Part 1 breaker and its timings, the window bounds, `dropBackedOffTiers`, and
`TriggerNow` (still ungated by everything).
Tests +7 (6 age-state + 1 interface witness); 27 packages ok. Red-proofs observed for Scenarios A, B and C.
### v0.177.0 — R-97: a failing backup is HEARD, and stops blaming the apps (2026-07-27) — MinAgent unchanged; requires hub >= v0.78.0
**R-97a — the whole-guest tier had no route to the hub.** `internal/quiesce` did not import
`internal/notify` at all, so on 2026-07-27 three failed whole-guest backups and twelve app-stack
stop/starts produced **zero** events. `NotifyBackupFailed` existed and the hub allowlisted
`backup_failed`; only the wiring was missing — the inert-seam shape this project has now hit five times.
This got MORE urgent when R-88 shipped, not less. Before the breaker a failing backup retried every
5 minutes: harmful, but loud enough to notice. Now it backs off to 4h and goes quiet, leaving the
hub's deadline monitor as the only signal — **~26h for local, ~8 days for PBS**, a full cycle of the
weekly tier. This trades that delay for an immediate one.
`quiesce.TierNotifier` is a seam, not an import (same reason `windowStartFn` is injected), wired by
the init-only `SetTierNotifier` because main.go builds the notifier *after* the loop. It is
**edge-triggered**: `BackupFailed` fires when the breaker ARMS — the first failure of a run, never
the retries behind it — and `BackupRecovered` on `recordSuccess`'s existing bool, so an operator told
a tier broke is also told it healed.
**New OPERATOR-ONLY event types**, `whole_guest_backup_failed` / `_recovered` (hub v0.78.0).
Deliberately NOT `backup_failed`: that type carries a customer-facing Hungarian template **and** sits
in demo-felhom's live `enabled_events`, so reusing it would have emailed the CUSTOMER
„A biztonsági mentés sikertelen" while the backup was still retrying. The tier travels in
`WholeGuestBackupDetails.Tier`, which is load-bearing — the hub keys its per-tier operator cooldown
on it, so `local` failing is not swallowed by `felhom-pbs` having failed within the hour.
**R-97b — stop telling the customer their app is broken when WE stopped it.** During the loop the only
customer-visible output was `app_start_failed — „Telepített alkalmazás nem fut: BookStack"`: customer
channel, Hungarian, during an outage the backup system itself caused, with no indication why.
**v0.164.0's filter does not cover this.** That predicate is state-based
(`IsDownState(st.State) && st.State != StateStopped`) and suppresses *deliberately stopped* apps.
BookStack alarmed because the third cycle caught it **mid-restart** — starting, or up but not yet
healthy — which is not `StateStopped`. No state classification can tell "restarting because a backup
stopped me" from "restarting because I keep crashing"; the distinguishing fact is that *we* stopped
it, and we know we did. So the fix is a **suppression window keyed to the cycle**, consumed at the
same single derivation point (`classifyRunStates`) that already computes both the banner dead-list
and the notifier Down-set — still one place.
**The grace window is 180 s**, derived rather than picked round: the deploy flow already allows
**120 s** for a stack to come up healthy, and the slowest catalog healthcheck start_period is Mealie's
**60 s**, after which a couple of check intervals must still elapse. It **expires** — an app that
genuinely fails to come back alarms on the first scan after the window closes. Permanent suppression
would trade a loud false alarm for a silent real one, which is R-88's Scenario D in a new costume.
Tests +9 (8 quiesce + 1 wiring reachability). Red-proofs observed for Scenarios C, E and F.
### v0.176.0 — R-88 Part 1: a failing backup stops re-quiescing (2026-07-27)
**The apps were being stopped and restarted every five minutes for a backup that could not
succeed.** Observed live on demo-felhom 2026-07-27: three full quiesce cycles at 09:02:57, 09:07:58
and 09:12:57 Budapest — each stopping and restarting all four customer app stacks
(`bookstack calibre-web docmost immich`, ~19 s down per cycle) against a PBS tier that was
unreachable. It stopped after three only because PBS came back, **not** because anything gave up:
`internal/quiesce` had no consecutive-failure counter, no backoff and no circuit breaker of any
kind, and the driver is a plain 5-minute ticker. Had the outage lasted, so would the loop.
**The failure breaker** (`internal/quiesce/breaker.go`). Consecutive failures are tracked **per
target**; a tier inside its backoff is dropped from the due set **before any stack is stopped** —
the gate is on the QUIESCE, not the backup, because the harm was never the failing backup but the
outage taken to attempt it. Backoff is `15m → 30m → 1h → 2h → 4h`, then 4h forever.
The cap is picked against two real constants rather than taste: 4h sits well inside the shortest
tier cadence (local = 24h), so a recovered tier still gets several attempts within its own cadence;
and it equals the width of the backup window gate `[W+2h, W+6h)`, so a tier at maximum backoff still
gets at least one attempt inside any given night's window instead of stepping over it.
Deliberately bounded in four ways, each with a test:
- **Never permanent.** The cap bounds the retry INTERVAL; it never stops retrying. A latched breaker
is a silent backup outage — strictly worse than the loop, which at least announced itself.
- **Never global.** One broken tier cannot suppress a healthy one.
- **Never gates `TriggerNow`.** A human pressing „Mentés most" is not deferred by a scheduler's
safety net. Manual runs still RECORD their outcome, so a manual success clears the backoff.
- **`stillRunning` is not a failure.** A first full offsite snapshot legitimately runs for hours.
State is **in-memory on purpose** — a restart forgets the backoff and re-attempts, which is the
cheap direction to fail; persisting it could carry a stale "this tier is broken" verdict across the
very restart that fixed it.
**The invariant, written where it will be read** (`scheduledRunAllowed`). A missing value means
UNKNOWN — not zero, not "never". Only a POSITIVE determination of "never backed up" may fire the
safety valve. This is the **fourth** instance of the same class (hub v0.12.0, hub v0.73.0, R-81, and
this), so the comment names all four and `TestContract_NeverBackedUp_RunsOutsideTheWindow` pins the
half that a careless fix would break.
**NOT fixed here, and deliberately so — R-88 Part 2 (agent-side).** The nil branch still fires the
valve, because the controller *cannot tell the two apart*: the agent's `/backup/due` returns
byte-identical responses for "the storage read errored" and "there has genuinely never been a
backup" — same `Due: true`, same `Reason: "no successful backup recorded yet"`, same nil `AgeSecs`.
Root cause is `localapi/server.go`'s `newestArchiveOn`, whose comment promises errors "degrade to
unknown, never to no-backup" while its `(time.Time, bool)` signature cannot represent unknown.
Splitting them needs a wire change plus a compat rule in both directions → its own task. Until then
the breaker bounds the damage: an unknown-driven cycle may still run once outside the window, but it
can no longer repeat.
Tests +11 (7 breaker, 4 contract). Red-proofs observed for Scenarios A, D and F.
### v0.175.0 — R-82: a tier that overruns the quiesce bound defers the rest (2026-07-26)
Operator ruling 2026-07-26: *"let the first backup run as long as needed; other backups shouldn't
start until finished."*
A first FULL offsite snapshot legitimately runs for **hours** — far past `max_quiesce`. When that
bound elapses the app resumes (correct, and unchanged), but `quiesceAndPollTiers` then moved on and
started the NEXT tier while the first was still uploading. That is now a `break`: the remaining tiers
are deferred to a later poll.
Why it matters: vzdump still holds the guest lock, so the second start would be **refused by the
agent (409, v0.99.0)** or fail on the lock — and a failed backup never satisfies a cadence, so the
tier would stay permanently due and retry into the same wall every poll.
`pollTier` now returns `(phase, stillRunning, err)`; `stillRunning` means the bound elapsed with the
backup still going. Nothing else changed — the app still resumes exactly once, on the same guard.
**Tests:** `TestTierOverrunsQuiesceBound_RemainingTiersDeferred`. Red-proof observed: dropping the
`break` starts the second tier and the test fails with
`the second tier MUST NOT start while the first is still running; started=[local felhom-pbs]`.
Restored; full suite green.
### v0.174.0 — R-82 Slice B: one quiesce window, two backup tiers (2026-07-26)
**MinAgent UNCHANGED — deliberately.** This release degrades gracefully against ANY older agent; it
does not require v0.97.0. Against a pre-R-82 agent it uses the untargeted single-tier path exactly as
before, logs the degrade once, and **still takes the backup**.
The agent gained per-target backup tiers in v0.97.0 ("local daily + PBS weekly"). The **controller**
owns quiescing, so the multi-tier schedule has to be reconciled here: on the weekly night both tiers
come due at once, and two quiesce cycles would mean **two app outages for one night's work** —
undoing the entire argument for weekly-over-daily.
### The dedup rule (specified, not emergent)
| local due | PBS due | result |
|---|---|---|
| yes | no | one quiesce, local backup |
| no | yes | one quiesce, PBS backup |
| **yes** | **yes** | **ONE quiesce window, BOTH backups inside it — never two cycles** |
| no | no | no quiesce |
### Added
- **`quiesce.TieredBackend`** (optional extension to `Backend`) + `quiesce.BackupTier`,
`ErrTiersUnsupported`. A backend that does not implement it — or whose `Tiers` returns
`ErrTiersUnsupported` — drives the pre-R-82 single-tier path unchanged.
- **`agentapi` per-tier client**: `BackupTiers`, `BackupDueFor`, `StartBackupFor`,
`BackupStatusFor` (`internal/agentapi/backup_tiers.go`). `targetQuery("")` yields an EMPTY suffix,
so an untargeted call hits the untargeted route byte-for-byte.
- **`Loop.resolveDueTiers`** — the dedup rule in one place, returning due tiers in AGENT ORDER.
- **`Loop.quiesceAndPollTiers` + `pollTier`** — one marker, one stop, N sequential backups, one
resume, tail polled to completion.
### Capability detection
`GET /backup/tiers` 404 ⇒ pre-R-82 agent. This is the project's documented ROUTE-PROBE mechanism
(`internal/agentapi/features.go`: "a route that shipped together with the coupled semantics either
answers (2xx ⇒ supported) or 404s"). It is **not** registered in the `featureProbes` table on
purpose: that table answers a yes/no at a UI entry point, whereas the loop needs the tier LIST
itself, so a table row would be a second probe of the same route for no gain. The degrade is logged
**exactly once per process** — once because it is a steady state during a rollout, never zero times
because a silent degrade is indistinguishable from multi-tier working.
### Two decisions worth stating plainly
**The app stays quiesced until the LAST tier snapshots.** Resuming after tier 1's snapshot would
leave the following tier capturing a RUNNING app — losing app-consistency on exactly the DR tier we
most want it on. **Consequence, user-visible:** on the both-due night downtime is
*(first tier's full backup)* + *(last tier's snapshot)*, not one snapshot. Tiers must therefore run
**fast-first**: vzdump holds a guest lock so they are necessarily sequential, and the agent
advertises primary (local) first — local-then-PBS makes downtime ≈ local backup + PBS snapshot,
whereas the reverse would be ≈ PBS backup + local snapshot, far worse.
**A manual "Mentés most" covers EVERY tier**, in one window, due-ness ignored. A manual run that
silently skipped the DR tier would be the same applied-and-empty fault in a different costume.
### Resilience (unchanged guarantees, extended per tier)
- Marker written BEFORE anything stops; unquiesce guaranteed by `defer` and fires **exactly once**
no matter which tier fails; a crash between two backups leaves the marker and `Recover()` restarts
the stacks at startup.
- One tier failing to START does not prevent the other tier's backup, and the app still resumes once.
- One tier's due-check erroring does not drop the other tier's backup.
- An agent advertising ZERO tiers falls back to the untargeted path — never "nothing to do".
- The window gate's safety valve now evaluates the OLDEST (most overdue) due tier, so a stale DR
tier cannot be starved by a fresher local one (`oldestAge`; a never-backed-up tier wins outright).
### Tests
+11 in `internal/quiesce/tiers_test.go`; full suite green. Red-proofs observed and restored:
- **#3 both-due night** — a per-tier cycle instead of one window fails with
`want EXACTLY 1 stop and 1 start, got stops=2 starts=2`. The COUNT is the assertion; asserting
only "both backups ran" would pass against a double-quiesce implementation.
- **#2 new controller ↔ old agent** — treating `ErrTiersUnsupported` as "nothing due" fails with
`OLD AGENT: a backup MUST still be taken via the untargeted path; got started=[]`. The hollow
version of this test asserts only "no error", which passes while silently skipping the backup.
### v0.173.0 — R-77: endpoint-drift detection, samba protected-set gate, channel log honesty (2026-07-26)
Source: `felhom.eu/documentation/audits/DIAG-agent-channel-2026-07-26.md`.
**Operational repair first (Part 0).** Both production controllers had been dialling their pre-island
LAN address since 2026-07-25 12:44 — the island migration rewrote `bootstrap.json` and
`controller.yaml` was never updated. `local_api.endpoint` corrected to `169.254.253.1:8443` on
demo-felhom and demo-hp (backups at `controller.yaml.pre-r77.bak`); **fingerprint and token agreed on
both boxes**, so only the address moved. Channel healthy since: zero `[channel]` lines and zero
`agent_channel_*` hub events after restart.
**Endpoint-drift detection — DETECT AND NAME, never write** (`bootstrap.DetectEndpointDrift`). When
`controller.yaml` and `bootstrap.json` both carry a complete `local_api` block and their endpoints
disagree, the controller emits one ERROR naming **both values and both paths**, raises a **new,
dedicated event type `local_api_endpoint_drift`** (error severity — drift never self-heals), and
shows its own Hungarian banner *above* the channel banner, because drift is the CAUSE and
"agent unreachable" the symptom. It **does not reconcile the files**: the mirror-image failure —
clobbering a correct `controller.yaml` from a stale `bootstrap.json` — is just as bad, fleet-wide.
That authority ruling is **R-78**. Fail-safe to silence on an absent/unparseable/incomplete bootstrap
(an unprovisioned guest is not drifted) and on an empty endpoint (that is `ensureLocalAPI`'s
fill-if-missing path, untouched). The fingerprint is compared and reported as a **boolean only**; the
token is never compared, logged or exposed.
**Hub allowlist (`felhom.eu` hub v0.74.0) — required, not optional.** `allowedEventTypes` 400s an
unknown `event_type`, so without the one-line entry the new alert would have been silently inert —
the exact seam-wiring failure this project has hit four times. Shipped with the controller.
**Samba protected-set gate.** `EffectiveProtected` now requires `smb.Enabled && smb.UserSet`,
mirroring **both** of `reconcileSambaAt`'s early returns. Sharing enabled without a household
password means the controller deliberately does not deploy samba, yet the health monitor reported
`fail` for it — demo-hp reported `health=fail` to the hub from the moment sharing was switched on.
**The doc comment was corrected in the same change**: it claimed "detection and deployment agree in
both directions" while citing only `!smb.Enabled`, an assertion that became false when the
`!smb.UserSet` return was added — a comment documenting a guarantee the code no longer provides is
how the bug returns. Not over-suppressed: sharing on **with** a password and a dead container still
alarms. `TestEffectiveProtectedTracksSharingToggle` was updated — its old fixture asserted the buggy
behaviour.
**Channel log honesty.** The debounce branch seeded an unseeded state to `"up"`, so a **born-down**
channel logged `up->down:<reason>` and `orUnseeded` was dead code. On 2026-07-25 that implied a
working channel degrading when neither controller had *ever* reached its agent, and it misdirected
the first read of the incident. The placeholder is now `stateUnconfirmed`, rendered `unseeded`.
**Logging only** — the placeholder is still matched in the re-arm condition, so F2 born-down alerting
is byte-for-byte unchanged; the Scenario-F test asserts sink call **count and arguments**, not just
the string, and all nine pre-existing channelhealth tests still pass.
Tests 951 → 959, all green. Three red-proofs (A, E, F) recorded in REPORT.md — Scenario A in **both**
failure directions: no-detection, and the auto-correcting variant that trips the byte-identical
assertion. **MinAgent unchanged; felhom-agent untouched** (DIAG refuted H1 — the island is healthy).
### v0.172.0 — R-75: canonical import root, catalog-derived skeleton, import surfaces (2026-07-26)
Spike: `felhom.eu/documentation/audits/SPIKE-catalog-data-paths-2026-07-26.md`.
**The drop-zone is now ONE canonical location on the system drive.** New `${IMPORT_PATH}` =
`<system namespace root>/userdata/import`, injected at BOTH compose-env builders
(`withUserdataPath` → `withPathVars`, `deploy.go` + `manager.go`) — the initial-deploy path missing
`USERDATA_PATH` once bound a bogus root-owned dir at the container root, and `IMPORT_PATH` has the
identical failure mode. It is derived from the SYSTEM drive, never from `HDD_PATH`, and has **no
per-drive fallback**: an unresolvable root leaves the variable UNSET so compose fails loudly instead
of quietly building a second, non-functional drop-zone. *Operator ruling, overriding the spike's
Fork-1 recommendation:* each drop-zone app has exactly one ingest bind, so a per-drive `import/`
would put a folder that LOOKS like a drop-zone on every drive while only one works — and since
import paths are `class: excluded`, files stranded in a dead one are never backed up either.
**Third `BindRoot` + the whole-block regression it prevents.** `RootImport` / `${IMPORT_PATH}` in
`composeVarRoots`, an `Import []BindSpec` list in `BackupSpec`, and `ValidateBackupSpec` /
`ClassifyBinds` extended. This is load-bearing: `ValidateBackupSpec` rejects an entry matching no
compose bind and the rejection is WHOLE-BLOCK, so moving paperless's ingest bind while leaving
`userdata: import/paperless` in place would have discarded the entire block — taking
`hdd: appdata/paperless/media class: mandatory` with it and silently degrading the customer's
document originals to legacy handling. `TestScenarioB_*` is the gate.
**Exhaustive-root audit — `resolveAbs` was the sharp one.** An import bind resolved against `hddPath`
would name a directory on the WRONG DRIVE. `resolveAbs`, `structuralGuard`, `ComputeCaptureSet` and
`ComputeFabBuckets` now take `importRoot` explicitly (compile-forced at all 4 call sites), and an
unresolvable root is refused LOUDLY into `Skipped` (`reasonNoImportRoot`) rather than joined onto "".
`GetImportRoot()` added to both provider interfaces + both adapters. `fabplan`/`tier2DestRel`/
`export.go`/`appbackup_bridge.go` audited and recorded in REPORT.md.
**Catalog-derived skeleton, deterministic by construction.** `UserdataSkeleton()` →
`UserdataSkeletonCarry()` (the v0.171.0 list verbatim, retained forever) + `BuildUserdataSkeleton()`,
which merges it with `DeriveUserdataDirs(stacksDir)` and **sorts**. The carry-list makes zero-removals
true by construction — `documents` is implied by no catalog app yet exists on both demo boxes — and
doubles as the fresh-box floor. The sort is not tidiness: the spike measured the naive map-order
derivation at **20 distinct outputs from 20 identical runs**, and `fbNeedsRecreate` force-recreates on
any byte difference across ~14 `SyncFileBrowserMounts` call sites — a fleet-wide FileBrowser restart
loop. `TestScenarioC_SkeletonDeterminism` pins 20/20. The catalog sync is deliberately **still not**
wired to `SyncFileBrowserMounts`. The canonical import root is excluded from per-app migration
(`appDataSkipSet`) so it never moves with an app.
**One authoritative compose parser.** `ParseComposeUserdataMounts` is now a thin resolver over
`ParseComposeClassifiableBinds`. The classifier won because it is the richer of the two byte-identical
scanners (it keeps the root and the `:ro` flag). One deliberate behaviour drop, recorded not hidden:
the old textual replace also accepted a LITERAL absolute path under `userdataPath`; no catalog
template has ever used that form and such a compose would be pinned to one machine's drive layout.
The deploy belt now handles both roots, gated differently — the drive-absent gate applies to the app's
data drive and must NOT suppress a system-drive import dir.
**Surfaces.** FileBrowser gains a separate `/srv/beolvasas` bind + a „Beolvasás" sidebar source
(separate, not nested — a nested source is indexed twice). New app-page block **„Hova tegyem a
fájlokat?"** for DEPLOYED apps declaring `data_paths`, with a deep link built from the shipped
Quantum router template, `url.PathEscape` per segment (**never `QueryEscape`** — it encodes space as
`+`, a literal plus in a path), the system-drive free space on import rows, and a **class-driven**
consequence line so the UI can never promise a backup the engines do not make. Copy does not promise
one click: a cold deep link goes through the FileBrowser login.
**`data_paths:` annotation** (`stacks.Metadata.DataPaths`) — role + Hungarian label over paths that
must ALREADY exist as compose binds; it can never declare one. Fork-3 asymmetry, deliberate: a
malformed PATH is a whole-block reject (data handling; reuses `ValidateBackupSpec`'s refusal set via
the extracted `appbackup.ValidateRelPath` — no second validator), an unknown ROLE fails OPEN with one
WARN (presentation; the `Lifecycle` precedent). Catalog: paperless-ngx, calibre-web, romm.
**System-owned import share.** `SMBShare.System`; a `beolvasas` share auto-created when sharing is
ENABLED (never before — deploying an app must not put SMB on the household LAN), `Offsite: false`
because the data is `class: excluded`. Deletion refused **server-side at both the handler and the
store**, and the button omitted in the template — three checks proving different things (the v0.70.1
ghost-delete lesson: a render gate is not enforcement, a handler test is not reachability). The share
is written directly rather than through `sharingResolvePath`: that guard validates CUSTOMER-supplied
picker paths, and the system drive is deliberately not a registered StoragePath.
**A latent 500 caught on the way:** the sharing template's row struct was function-local, so adding
`{{if .System}}` would have failed at render for every share. `ShareRow` is now package-level and the
render test constructs the exact type the handler passes.
**Caught during the live legs and fixed in the same version (two things):**
1. The carry-list initially kept `import`, `import/paperless` and `import/calibre`, so the skeleton
would RE-CREATE a per-drive drop-zone on every drive forever — the exact dead lookalike this arc
removes, and one that is never backed up. Dropped from the carry-list. This is not a removal:
nothing deletes the dirs an existing box has (both demo boxes' old drop-zones were verified to hold
**zero files** first); they stop being maintained and stop appearing on fresh boxes.
`TestSkeletonNeverCreatesAPerDriveDropZone` pins it, and `TestUserdataSkeleton_List` was updated to
assert their absence.
2. `EnsureImportRoot` ensured only
the leaf, so `MkdirAll`'s intermediates left `<sysNS>/userdata` at `755 root:root` — the one userdata
root on the box outside the 2775/gid-1000 convention. Both the parent and the import dir now carry it
(`TestEnsureImportRoot_ParentCarriesTheConvention`).
Tests 915 → 951, all green. Red-proofs recorded in REPORT.md for Scenario B (classification),
C (determinism) and E (server-side share refusal). No destructive filesystem call was added anywhere
in this arc. **MinAgent unchanged.**
### v0.171.0 — Disk-health card: device-model label (pairs with agent v0.95.0) (2026-07-25)
`agentapi.SmartSummary` gains `ModelName` (mirrors the agent v0.95.0 `model_name`); the "Lemezek
állapota" card row label now prefers the device model ("TOSHIBA MQ04ABF100") over the raw storage
name/UUID, falling back to Name (+ speed hint) on an older agent or a modelless disk. With agent
v0.95.0 the system SSD and the USB drive now carry real SMART, so the card shows real verdicts
(Rendben) with human labels instead of "Nincs adat" on a raw UUID. Additive; old-agent payloads render
exactly as before. Test `TestDiskDisplayLabel_PrefersModel` (red-proof: drop the fallback → A4 fails).
### v0.170.0 — Root → Indítópult; gofmt normalization; stale-note fix (2026-07-25)
- **`/` is now the Indítópult** (operator ruling, reversing the v0.163.0 landing choice). `GET /` 302s
to `/launcher` (ONE canonical URL per page — the launcher body is never served at `/`); the
Vezérlőpult keeps its own URL **`/dashboard`** and its nav slot. Nav: Indítópult active on
`/launcher`, Vezérlőpult `href="/dashboard"` active there — never both. Post-login (default `/`) and
the mobile-topbar logo (`/`) both flow through the redirect to the launcher; the login redirect target
is unchanged. Tests: the 302 (target + status), `/dashboard` 200, nav hrefs/active; red-proof: fold
`/` back into the dashboard case → the 302 test fails. Two dashboard-card tests repointed `/`→`/dashboard`.
- **gofmt normalization** shipped as a **separate, style-only prior commit** (`gofmt -w` across the
controller tree, **46 files**, `gofmt -l` now empty) — disarms the formatting landmine where a
targeted edit + an accidental `gofmt -w` swept ~46 unrelated files. Pure formatting (whitespace +
optional-semicolon removal in reflowed inline closures); one doc comment reworded to avoid gofmt's
Go-1.19 `''`→curly-quote doc-comment substitution.
- Repo `CLAUDE.md`: corrected the stale "vacation — agent DOWN at a remote site" note — felhom-pve is
back on the home LAN and the agent is up at `192.168.0.162:8443` (Tailscale alias still available).
### v0.169.1 — Disk-health card: exclude logical/network storage (2026-07-24)
Live QA follow-up to v0.169.0: the agent defaults SMART to UNKNOWN on non-physical targets (PBS,
LVM-thin), so they appeared in the "Lemezek állapota" card as spurious "Nincs adat" rows.
`isPhysicalDisk` now excludes `pbs`/`lvmthin`/`nfs`/`cifs` by type (applies to both the card and the
6h check). Test strengthened: a PBS/LVM fixture carrying UNKNOWN SMART must still be excluded.
### v0.169.0 — Disk-health card + degradation notification ("Lemezek állapota") (2026-07-24)
Consumes the agent's new `smart` payload field (agent **v0.94.0**); **MinAgent floor unchanged** — the
feature detects by payload presence (nil → "Nincs adat", never alarms). Pairs with the hub allowlist
bump (adds `disk_health_degraded`). No new smartctl load anywhere — the agent serializes
already-computed SMART; the controller only reads it.
- **`agentapi`:** `SmartSummary` extended to the full counter set (SATA reallocated/pending/offline +
NVMe critical/media/percentage_used + power-on-hours); `DiskInfo` gains `Smart *SmartSummary`; new
pure `DiskVerdictFor(*SmartSummary) DiskVerdict` (the SINGLE source of truth for card + check) with
`Label()` (Rendben / Figyelmeztetés / Hiba / Nincs adat) + `DegradedAttributes`. Mapping: FAILING →
Hiba; PASSED with any of reallocated>0 / pending>0 / offline_uncorrectable>0 / critical_warning>0 /
media_errors>0 / percentage_used ≥ 90 → Figyelmeztetés; PASSED clean → Rendben; nil/UNKNOWN/empty →
Nincs adat (never alarms).
- **Dashboard "Lemezek állapota" card:** one row per PHYSICAL disk (label + colored verdict chip +
temperature). Fed by a **60 s in-process TTL cache** around `/disks` so dashboard refresh-spam cannot
smartctl-storm the host. An unreachable agent renders "Nincs adat" — the page never blocks.
- **6-hourly `disk-health-check`:** compares each physical disk's verdict against an in-memory baseline
and emits `disk_health_degraded` **only on a degradation** (verdict worsened). First run baselines
silently; recovery/improvement notifies nothing; **UNKNOWN is excluded both directions** (a transient
UNKNOWN blip never fires and never erases history); multiple attributes on one disk → ONE event.
Severity: warn (Figyelmeztetés) / critical (Hiba). The hub applies its own per-event-type cooldown.
- **Deliberately no global alert banner** (CONTEXT ruling) — the card + email carry it; banner fatigue
is a real cost. Not wired into the dead-app/alert-banner machinery. Controller restart re-baselines
silently (accepted, consistent with the health-change pattern).
Tests: verdict table (+ ≥90 boundary red-proof); notifier emit (type/severity/subject); check
first-run-silent (red-proof: disable the guard → first run notifies), degradation-once, recovery-silent,
UNKNOWN-excluded, FAILING→critical, nil-smart card graceful, TTL cache.
### v0.168.0 — Customer-configurable backup window ("Mentési időablak") (2026-07-24)
No agent coupling; MinAgent unchanged (the disk-tier gate is controller-side; the agent's cadence-based
`/backup/due` is untouched). New pure package `internal/backupwindow`; touches scheduler, settings,
quiesce, the backup page, and main.go wiring.
**One setting drives every nightly leg.** A single customer control — **"Mentési időablak kezdete"**
(default = the effective DB-dump time, historically "02:30") — from which every leg derives at FIXED,
never-stored offsets, so misordering is impossible: DB dump at **W**, tier-2 mirror at **W+60m**,
off-box at **W+105m** (wrap-safe across midnight). Precedence: settings > controller.yaml
`db_dump_schedule` > "02:30".
- **Scheduler seam `UpdateDaily(name, timeStr) bool`** (+ a per-daily-job buffered `resched` channel and
a new select case in `runDailyJob`): a saved window fans out to all three legs and takes effect at the
next scheduling pass **without a restart**. Unknown/non-daily name or invalid time → WARN + false, job
untouched.
- **Disk-tier (whole-guest PBS/vzdump) window gate.** The quiesce loop's scheduled cycles now run only
inside **[W+2h, W+6h)** (wrap-safe, Europe/Budapest wall-clock), with a **safety valve**: if the newest
successful backup is older than cadence+24h (or none exists), the cycle runs regardless of the window —
a box powered on only outside its window never starves. Gate denials log at DEBUG with the window.
**Manual triggers ("Mentés most" / `TriggerNow`) are NEVER gated** (they bypass `runOnce`). The
`Backend.Due` seam now also returns the backup age (from the agent's own `/backup/due` answer) for the
valve; the agent, its cadence, and `/backup/due` semantics are unchanged.
- **Backup page (Áttekintés):** a compact "Mentési időablak" card — time input (value = effective
window) + "Mentés" button, and the derived rows (adatbázis-mentés / helyi másolat / távoli mentés
times, and the "teljes rendszermentés kb. W+2h–W+6h között" line). POST `/backups/window` validates →
saves → `UpdateDaily`×3 → PRG redirect with a Hungarian flash. Behind RequireAuth + CsrfProtect like
its siblings.
- Derived leg/gate times are **computed, never persisted**; no per-leg settings; the offsets are not
exposed in the UI.
Tests (5 groups, all red-proofed): `LegTimes`/`GateWindow` incl. midnight wrap + invalid-rejected;
`EffectiveWindow` precedence table; `UpdateDaily` mutate+signal + unknown/invalid + goroutine consumes
the reschedule; `scheduledRunAllowed` truth table (inside/outside/valve/wrap/nil-age) + Loop integration
(defer outside / run inside / valve runs / manual never gated); handler valid-save + invalid-rejected.
### v0.167.1 — Center the sidebar logo (2026-07-24)
CSS one-liner + test. `.sidebar-logo` gains `margin: 0 auto` so the 140px logo is horizontally
centered within the header instead of left-aligned — applies to both the desktop sidebar and the
mobile drawer (same element). Pin: `TestSidebarLogo_Centered` (red-proof verified).
### v0.167.0 — Outlined logo + favicon (Part 4, the v0.166.0 gated follow-up) (2026-07-24)
No agent coupling; MinAgent unchanged. Embedded-asset constants + one test only — no backend, no
routes, no template/CSS behavior change.
Completes Part 4 that v0.166.0 deferred at the §3a gate. Viktor pushed the text-outlined
`website/assets/logo.svg` to felhom.eu `main` (`be9edb4`): the wordmark is now **17 real `<path>`
glyphs** (Inkscape Object→Path) instead of live `<text>` with `font-family:'Vremena Grotesk'`/`'M+ 2c'`.
Under `<img>` secure static mode only locally-installed fonts resolve, so the old constants rendered
the wordmark in a fallback font on every device without those fonts — now fixed.
- **`FelhomLogoSVG`** body replaced with the outlined master. Inkscape left behind **2 empty `<text/>`
shells + font-* style leftovers on the paths** (inert, but they carried the font names); these were
stripped via a DOM pass (lxml) — **glyphs untouched, no text-to-path conversion done by CC**. Also
dropped the editor-only `<sodipodi:namedview>`. `viewBox` **unchanged** (`0 0 645.30703 408.36403`);
full palette preserved (white glyphs `#ffffff`, blue `.eu` `#008ddf`, navy `#051343`, all 14
gradients, the cloud/house/server/lock artwork).
- **`FelhomFaviconSVG`** vestigial empty `<text>` nodes + their `font-family` removed (cloud icon only;
`viewBox` unchanged `0 0 437.307 296.36403`, 11 paths).
- Both constants now contain **zero `<text` and zero `font-family`**. Served at their existing paths by
the unchanged handlers (the hub-synced-file-first fallback logic is untouched).
- Test: `TestLogoSVG_NoLiveText` (both constants free of `<text`/`font-family`) — was written as the
gated red-proof in v0.166.0 (FAILED against the old constants), now committed green.
### v0.166.0 — Mobile nav drawer + sidebar cleanup + versioned logo/favicon URLs (2026-07-24)
No agent coupling; MinAgent unchanged. Templates (layout/login/icons) + CSS + one login-handler data
key + tests only — no backend logic, no routes, no settings, no dependency changes. Desktop (>768px)
is unchanged.
**Mobile navigation rework (Option A — off-canvas drawer).** The `@media(max-width:768px)` block
predated the v0.146.0 accordion: it flattened `.nav-links` into a horizontal `overflow-x` strip, and
because the accordion's nested sub-lists share the `.nav-links` class, sub-items laid out horizontally
inside an `overflow:hidden` grid row — everything past the first sub-item was clipped. The strip is
**deleted** (not patched — Option C was rejected) and replaced by:
- a sticky **top bar** (`.mobile-topbar`, logo → `/`, single hamburger `.nav-burger` with
`aria-expanded`/`aria-controls="sidebar"`, new `#i-menu` icon);
- the existing vertical sidebar reused as an **off-canvas left drawer** (`.js .sidebar`,
`transform:translateX(-100%)`→`is-open`), a `.nav-backdrop`, body scroll-lock (`body.nav-open`);
drawer JS toggles on burger, closes on backdrop click or Escape. **The accordion handler is
untouched** — it works identically inside the drawer (all sub-items stack vertically, nothing clipped).
- a **no-JS fallback**: `<html class="no-js">` (swapped to `js` by an early head script); when JS is
off the sidebar renders static inline above the content and the burger is hidden, so no destination
dead-ends.
- z-index ladder: top bar 800 < backdrop 900 < drawer 950 < `.modal-overlay` 1000 (modals stay on top);
`height:100dvh`; drawer transition disabled under `prefers-reduced-motion`.
**Sidebar customer-name removed.** The `<span class="customer-name">` and its dead CSS rule are gone
from the sidebar header (logo only). `{{.CustomerName}}` stays in base data and on the **login page**
subtitle (identifies the box owner).
**Cache-bust on logo/favicon.** `/static/felhom-logo.svg` and `/static/favicon.svg` now carry
`?v={{.Version}}` (sidebar logo, head favicon, login logo) — Cloudflare edge-caches `/static/*` for 4h,
so unversioned asset URLs kept serving the previous release's copy after a deploy (the 0.126.1 CSS
failure mode). `renderLogin` now passes `Version`.
**Part 4 (outlined-logo swap) GATED OUT — not shipped.** The §3a precondition failed: live felhom.eu
`main` (`be9edb44`) still serves a `website/assets/logo.svg` with live `<text>`/`font-family` (the
text-outlined master is Viktor's manual Inkscape push, still pending). The `FelhomLogoSVG` /
`FelhomFaviconSVG` constants are therefore **unchanged** and still contain live `<text>` — the wordmark
renders in a fallback font under `<img>` secure static mode until the outlined asset lands and Part 4
ships. Only the `?v=` cache-bust portion of the logo work is in this release.
- Tests (5 new, all through the real layout/CSS): topbar+drawer markup, CSS strip-removed/drawer-present
(scoped to the 768px block), sidebar-no-customer-name, login-customer-name-kept, versioned asset URLs.
RED-proofs recorded pre-change (strip present, no drawer, customer-name present, no `?v=`, and the
gated logo-constant proof). `nav_accordion_test.go` invariants pass unchanged. Focus-trap on the
drawer deliberately omitted (page navigations reset state). Android drawer feel + desktop
pixel-parity are Viktor's visual acceptance step.
### v0.165.1 — Native "Megosztás…" button in the share modal (Web Share API) (2026-07-24)
No agent coupling; MinAgent unchanged. Template JS + tests only — no backend, no routes, no
settings, no dependency changes.
The "Indítópult megosztása" modal gains a **"Megosztás…"** button that opens the OS share sheet via
`navigator.share` (Messenger / WhatsApp / email / anything installed), sending the share **title +
text + URL only**. Feature-detected: the button is `display:none` in the markup and revealed only
when `navigator.share` exists; the universal **"Link másolása"** stays as the fallback and is never
demoted. A user cancel (`AbortError`) is silent; any other rejection falls back to `copyShareLink()`
so the user still keeps the link on the clipboard.
- **The QR is deliberately NOT attached** to the share payload (no Web Share Level-2 `files:`):
file-share support is narrow and several targets drop the URL when handed file+URL, leaving an
unscannable QR picture in a chat. The QR's job — physical cross-device scanning — is already served
by the modal image (mobile long-press covers "send the picture" with zero code).
- Share copy (user-to-user, deliberately conjugation-free): title `Indítópult — <domain>`, text
"Az otthoni alkalmazások egy helyen.".
- Tests: Group A (button hidden-by-default + feature-detect reveal + title/text/url-only payload,
no `files:`) + Group B (AbortError-silent + non-abort fallback to copy); 2 red-proofs verified red.
### v0.165.0 — Indítópult megosztása: guest launcher via capability URL (2026-07-24)
No agent coupling; MinAgent unchanged. New dependency: `github.com/skip2/go-qrcode`
(v0.0.0-20200617195104-da1b6568686e, MIT, pure Go, zero transitive deps) for the modal QR code.
The admin launcher gains an **"Indítópult megosztása"** button that mints a **capability URL**
(`https://<host>/s/<token>`, 160-bit token) serving a standalone, read-only guest launcher — same
tiles, opens apps in new tabs — with **no accounts and no admin session**. The link grants
**information only, zero control**: app names + public URLs; every privilege stays behind each app's
own auth and the controller admin password.
- **Capability-URL serving.** `/s/<token>` is added to the RequireAuth pre-auth allowlist (AFTER the
claim-gate block, so the claim gate stays supreme) and exempted from session CSRF (guests carry
their own pre-auth HMAC CSRF, like the claim POST). Token comparison is `subtle.ConstantTimeCompare`;
an empty stored token matches nothing, so a wrong/disabled token is **byte-identical to the mux
default 404** — nothing distinguishes it from an unknown route. Guest responses set `X-Robots-Tag:
noindex, nofollow`, `Referrer-Policy: no-referrer`, `Cache-Control: no-store`.
- **Optional per-share password.** A SEPARATE credential — its own bcrypt hash
(`settings.LauncherSharePasswordHash`, never the admin hash), its own per-IP 5/1-min attempt map
(never the admin login map). Passing it once mints a signed cookie = HMAC-SHA256 over
`token|passwordHash` (keyed with the persisted, box-scoped `web.session_secret`), so **rotating the
token OR changing the password invalidates every outstanding cookie** with zero bookkeeping.
- **Modal (admin):** copy-link, a QR code (`/launcher/share/qr.png`, ~256px, admin-authed),
"Jelszó beállítása/törlése", "Új link készítése" (rotation), "Megosztás kikapcsolása". POSTs under
`/launcher/share/*` ride the normal admin session + session CSRF.
- **Guest state labels ride the v0.164.0 ruling:** `StateStopped` ⇒ "A tulajdonos leállította";
any other non-clickable state ⇒ "Átmenetileg nem elérhető"; guests never see internal state
vocabulary (stopped/exited/degraded/unhealthy). Clickable ⇔ operational AND its public route is
published (`isOperationalState && !routeUnpublished`), so a guest tap never dead-ends on a 404.
- **Token is a secret:** never logged (the ServeHTTP debug line and the 404 WARN redact `/s/` paths
to `/s/<redacted>`), never written to CHANGELOG/REPORT/CONTEXT, constant-time comparison only.
- **Refactors:** `launcherApps()` extracted from `launcherHandler` (shared with the guest handler);
the tile visual extracted into a `launch_tile` partial (single markup source for admin + guest);
`isOperationalState` promoted to a package predicate (single source for the funcmap + guest rule).
- New files: `internal/web/share.go` (pure core), `internal/web/share_handlers.go` (HTTP surface),
`internal/web/share_test.go` (Groups A–G + 3 red-proofs verified red), templates
`launcher_shared.html` + `launcher_share_password.html`.
- Design rulings (CONTEXT): member accounts are superseded by this capability-URL model;
per-member tile visibility is parked under the SSO arc.
### v0.164.0 — Deliberately stopped apps no longer alarm (banner + email) (2026-07-24)
No agent coupling; MinAgent unchanged. Operator finding on 9201: stopping an app via the UI
(Leállítás) raised the global warning banner "Telepített alkalmazás nem fut: … (stopped)" on every
page — including the launcher, where the tile already shows the greyed state — and fired the
`app_start_failed` notification event on the running→down transition. A deliberate user action is not
a fault; it must not alarm the user anywhere. Genuine faults keep alerting exactly as before.
- **The fix is a one-line filter at the single fix-3 derivation point.** `scanDeployedAppRunStates`
(cmd/controller/main.go) is the only place both the banner dead-list and the notifier Down-set are
computed. Its pure core was extracted to `classifyRunStates([]stacks.Stack)` (testable without a
live Manager), and the down predicate changed from `stacks.IsDownState(st.State)` to
`stacks.IsDownState(st.State) && st.State != stacks.StateStopped`. `StateStopped` is therefore
suppressed from BOTH surfaces: no banner on any page (launcher included) and `Down=false` fed to
the notifier ⇒ no `app_start_failed` event and a clean transition tracker.
- **Why `StateStopped` ⇒ deliberate (two invariants, recorded at the seam and in CONTEXT.md):**
(I1) the UI stop path `Manager.StopStack` runs `docker compose down` → containers are removed, and
a deployed stack with zero containers aggregates to `StateStopped` (refreshStatusLocked). (I2) the
P2 restart-policy census (2026-07-21, 53 templates / 78 services) found every catalog service on
`unless-stopped`, so a crashing app never comes to rest at `stopped` — faults surface as
`restarting` / `unhealthy` / `exited` / `degraded`. **If either invariant changes, revisit this
suppression.** An out-of-band `docker compose stop` leaves containers present → `StateExited` →
still alerts (out-of-band tampering is reportable — acceptable).
- **`IsDownState` deliberately UNCHANGED** — other callers (e.g. `CommittedMemory`, bootrecon) rely
on stopped counting as down. The suppression lives ONLY at the scan; no template, funcmap,
notifier, dashboard-counter, or Hungarian-copy change. The launcher tile still shows greyed +
"Leállítva"; the monitoring page and dashboard RunningCount/StoppedCount are unchanged (factual
display is not an alarm). A pre-existing banner self-clears on the next health cycle (state-based).
- **Tests +4** (notify 3→4, main 4→7): Group A — `classifyRunStates` over [running, stopped, exited,
degraded] yields dead={exited,degraded} and Down flags {false,false,true,true} (red-proof: revert
the filter → both assertions fail, verified). Group B — fault parity: exited+degraded both in the
dead list, both Down=true, raw state string carried through. Group C — stop→start→crash drives
`NotifyAppStartFailures` to exactly ONE event for the crash and zero for the stop (red-proof: mark
the stop Down=true → the zero-for-stop assertion fails, verified). Plus a skip test for
deploying/undeployed.
### v0.163.1 — Launcher polish: monogram reveal-on-failure + placeholder on every icon surface (2026-07-24)
No agent coupling; MinAgent unchanged. Two live findings from the v0.163.0 operator browser pass on 9201.
- **Monogram bled through every tile.** The launcher rendered `.launch-mono` unconditionally UNDER
the logo `<img>`; app logos are white monochrome SVGs with transparent backgrounds, so the big
white letter showed through the glyph gaps on EVERY tile. The monogram is now hidden by default
(`.launch-mono { display: none }`) and revealed ONLY when the img chain fails — the final `onerror`
step adds `.launch-tile--noimg` to the tile, which flips the monogram back on. Applies to both the
operational `<a>` and the stopped `<div>` branch.
- **Placeholder reached only the canonical row.** The `/static/app-placeholder.svg` default landed in
`app_list_row` only; four more sanctioned app-logo `onerror` chains still dead-ended in
hidden/none for logo-less apps (observed: Docmost with no icon on Biztonsági mentés →
Alkalmazások). Every app-logo surface now follows one grammar — **SVG → PNG → placeholder** (infra
rows → infra icon): `backups_apps.html` (the allowlisted aligned row), `stacks.html` (the
`data-fallback` is now always present: infra → `infra-logo.svg`, else `app-placeholder.svg`),
`app_info.html` (hero logo only — **screenshots deliberately still vanish on error**),
`deploy.html` (keeps its `.LogoURL`/`.LogoPNGURL` data source). No handler/funcmap changes.
### v0.163.0 — Indítópult (app launcher page) + universal app placeholder icon (2026-07-24)
No agent coupling; MinAgent unchanged. Adds a customer-facing **Indítópult** launcher grid and a
generic fallback icon for logo-less apps on every list surface.
**Indítópult (`/launcher`, new FIRST sidebar item, above Vezérlőpult):**
- A grid of large tappable tiles, one per openable deployed app. The rule is intentionally the same
one the „Megnyitás" button already uses: a tile exists **⟺** the stack has a subdomain (env
`SUBDOMAIN` > `.felhom.yml` subdomain > `protectedStackSubdomains`). The controller's own stack is
excluded by name. `/` still lands on the Vezérlőpult — the launcher is an ADDITIONAL page.
- Tiles are colored rounded squares: a deterministic per-app color (FNV-1a of the slug → HSL hue,
fixed S/L tuned for the dark theme), overridable with an optional `.felhom.yml` `brand_color`
(`#rgb`/`#rrggbb`; an invalid value silently falls back to the slug color). The existing white
monochrome logo renders on top; a logo-less app reveals the **monogram** initial underneath
(multibyte-safe — „Óra" → „Ó").
- Operational apps are a real `<a target="_blank" rel="noopener">` to the public URL (with
`open_path`); stopped/exited/degraded apps render a **greyed, unclickable** tile with the honest
Hungarian state badge — never a dead link. Empty state: „Még nincs telepített alkalmazás." + a link
to `/stacks`.
- New template funcs `tileColor` (returns a `template.CSS` — validated/computed in Go, because the
html/template CSS filter mangles a legitimate `hsl()` from a func pipeline) and `initial`.
**Universal app placeholder icon:**
- New embedded `AppPlaceholderSVG` (2×2 rounded-square app-grid glyph), served at
`/static/app-placeholder.svg`. The canonical `app_list_row` now DEFAULTS its fallback to it, so a
catalog app with a missing logo shows a generic placeholder on every list surface instead of the
old `visibility:hidden` dead-end. Infra rows still override with `/static/infra-logo.svg`.
- Design ruling: the felhom brand mark is NEVER an app placeholder (brand = platform identity only).
**Refactor (in-scope, single reason):** the subdomain-map assembly that lived inline twice
(dashboard + Alkalmazások) is extracted to `Server.subdomainMap`; both call sites plus the launcher
now share it (byte-for-byte priority unchanged).
**Metadata:** `stacks.Metadata` gains `BrandColor` (`brand_color`, omitempty). No catalog app sets it
yet (curation is a parked follow-up).
### v0.162.0 — R-71(a): the apply-bridge waits for the dust to settle (settle-gate) (2026-07-24)
No agent coupling; MinAgent unchanged. Origin:
`felhom.eu/documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md` — the day-0 race. A fresh box
boots below the operator floor (ISO 0.153.0 < floor 0.156.0), the apply-bridge consumes the
single-use offsite password, then ~35 s later the managed auto-floor update replaces the container
mid-install → the new process finds no installed key → consume → **404** → offsite dead until an
operator Re-issue. This recurs on **every** fresh onboarding whose ISO floor lags the managed floor;
demo-felhom escaped by timing alone. The v1.25.0 golden≥floor build gate PREVENTS the trigger for
fresh installs; R-71c (hub) HEALS a burn after the fact; this (a) removes the SYSTEMATIC trigger for
every restart shape.
**The change (ordering only — the bridge's consume/install/persist internals, the 404-no-oracle
contract, and the Consumer are UNTOUCHED; R-71(b) stays rejected-by-design):**
- New seam `offsiteapply.SettleProvider.SettleState() (version, floor string, updateRunning,
floorKnown bool)` — a thin adapter (`SettleFunc`) over the self-updater's OWN knowledge in main.go
(`GetFloor()`/`IsUpdateRunning()`); the bridge never fetches the floor a second way.
- `Bridge.AwaitSettle` polls every 10 s (bounds: 90 s floor-knowledge sub-bound, 5 min overall)
BEFORE the 3-minute Reconcile context is created (the deferral never eats the reconcile budget).
Releases: `updateRunning` → wait (the swap supersedes us); `floorKnown && version<floor` → wait
(auto-floor update imminent — do NOT burn the password); `floorKnown && at/above floor` → **GO on
the first poll, zero sleep** (the B′ invariant); `!floorKnown` past 90 s → GO+WARN (a hub that
can't serve a floor can't serve a consume → no burn risk); overall bound → GO+WARN (R-71c is the
belt). `ReconcileWhenSettled` runs the gate then Reconcile.
- The bridge goroutine MOVED in main.go to after the self-updater is constructed (so the adapter can
read it). Wired ONLY when an updater exists — with no update mechanism there is no floor-update to
race, so the bridge reconciles immediately (`Settle` nil = old behavior).
**Finding (cited in the sub-bound rationale):** the floor is in-memory (report-ACK-derived), NOT
persisted — so on any restart it is unknown until the first report ACK. The startup report fires ~5 s
after boot and `SetFloor` runs synchronously in its ACK handler, so the floor is normally known in
~5–10 s (≤~45 s across the 3×15 s report retries); the 90 s sub-bound is headroom over that.
**Tests (injectable clock, no real sleeps; fake SettleState + recorded Consumer):** A below-floor
defers then GOes at floor (one consume); B update-running defers then GOes; C floor-unknown GOes at
the sub-bound + WARN; D perpetually-below GOes at the overall bound + WARN; E (B′) at-floor GOes on
poll 1 with zero wait; plus nil-provider immediate-reconcile and cancelled-gate-skips-reconcile.
**Four red-proofs, all observed FAIL then restored:** remove the gate → below-floor consumes
immediately (0 deferral sleeps); remove the updateRunning branch → mid-swap box GOes immediately;
drop the sub-bound → floor-unknown drags 90 s→5 m; drop the overall bound → perpetual below-floor
loops forever (20 s test timeout). The deferral paths ship **unit-proven + red-proofed, NOT
live-fired** — their precondition is now structurally prevented by the v1.25.0 build gate, which is
the point: **gate prevents, (a) defers, (c) heals.** Live-observable leg = the B′ first-poll GO line
on both above-floor demo boxes.
### v0.161.0 — R-70: the hub-managed offsite empty state tells the truth (2026-07-23)
No agent coupling; MinAgent unchanged. Origin: `felhom.eu/documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`
— a box stuck pre-apply (burned credential) rendered the SAME „Még nincs beállítva távoli mentési
cél." empty state as a box that was never provisioned, and the „igényelhető szolgáltatás" card
offered to order a service that was already ordered. The ambiguity hid a dead offsite tier for
2 days on demo-hp.
**The change (XS):** `backupsOffboxData` exposes `OffsiteHubEnabled` (= `cfg.Offsite.Enabled`,
the hub descriptor in controller.yaml). On Távoli mentés, when hub-managed offsite is enabled but
no `offbox` target exists yet, BOTH empty surfaces switch to the truth: „Felhom offsite tárhely
kiépítve — a beállítás automatikus, folyamatban. Ha egy napon belül nem áll be, jelezd az
üzemeltetőnek." (the status card AND the target empty-state line). Without a hub-managed offsite,
today's copy is byte-identical; a configured target renders the status block as before. The
own-NAS setup button is untouched.
Render tests per branch of the gate (the v0.70.1 template-gate lesson): hub-enabled+no-offbox →
banner + old copy asserted GONE; not-enabled → old copy asserted intact; configured → no banner.
All design-v2 template gates green (`docker_run_volume_path_gate` stays red on the pre-existing
R-29 allowlist item, untouched by this change). Hub-side sibling: felhom-hub v0.72.0 (delivery-state
detector + stuck event + R-71c self-heal).
### v0.160.0 — R-67: the NAS share appears in FileBrowser (2026-07-22)
No agent coupling; MinAgent unchanged. Origin: the R-64 pairing drill — the share said „Elérhető"
and the customer had no way to BROWSE it; FileBrowser synced drives only.
**Phase-0 probe (GO, demo-hp, 2026-07-22):** with the Felhom-Share automount confirmed IDLE (autofs
trigger in /proc/mounts, no cifs mount), `docker run --rm -v …/Felhom-Share:/probe:rslave alpine ls
/probe` listed the real share content and left cifs mounted — an in-container access through an
rslave bind DOES wake the idle trigger, one namespace further than the spike's in-guest proof. The
design shipped as specified, no fallback fork needed.
**The change:** `syncFileBrowserMounts`' path loop is extracted into the pure
`buildFileBrowserPaths` (deps injected: mount probe / FS classifier / skeleton fn / logger), which
now returns BOTH the mount lines and the config source set so the two can never disagree. Network
shares get their own branch:
- bind = the share ROOT, `…/<name>:/srv/<name>:rslave` — `:rslave` is load-bearing (host-side
automount wake / idle-unmount events propagate into the running container);
- NO `EnsureUserdataSkeleton`, no userdata scoping — nothing is ever written toward the NAS;
- the drive-absent gate does NOT apply (an idle automount is healthy and would be skipped
forever); the gate is the `stub` classifier verdict instead — **the data-safety wrong case**:
exposing a local stub dir lets a customer upload files the real mount will later shadow, so a
stub share is excluded from mounts AND sources this pass with a WARN. autofs / network /
unknown / nil-classifier all include (fail open).
- Drive behavior is byte-identical (tested: the drive line with a share present equals the
drives-only render; drives always stay in the source list as before).
- NAS add-success (`runNetAdd` done) and remove (`handleNetStorageRemove`) now trigger
`SyncFileBrowserMounts()`; removal drops the source + mount on the next sync (F2 change
detection forces the recreate).
Tests: `filebrowser_network_test.go` scenarios A–D. Red-proofs recorded in REPORT.md: A (network
routed through the drive branch → the skeleton-call assertion fails with the NAS path recorded)
and B (stub gate dropped → the stub share leaks into mounts + sources).
### v0.159.0 — R-66: the box's own address becomes visible (2026-07-22)
No agent coupling; MinAgent unchanged. Controller-only, three XS legs with one theme: **the box
must be able to tell you where it is.** Origin: the Felhom↔Felhom NAS pairing drill — the serving
box's IP was findable only as a hint line buried on the OTHER box's Megosztás page, and the add
form's failure for a NetBIOS name („FELHOM") taught nothing.
**Leg A — „Hálózat" card** on Beállítások → Rendszer (between „Verzió és frissítés" and „Szerver
memória"): Helyi cím (LAN), Hálózati név (`\\<SMBServerName>`, rendered ONLY while Megosztás is
enabled — the NetBIOS name exists only while samba runs), Átjáró, and a muted footer asking the
customer to read the page aloud during remote troubleshooting. Everything is live-computed per
render and stored nowhere (S-5); an unavailable value renders „—" („nem állapítható meg").
**Leg B — `network` section in the Debug system dump** (`GET /api/debug/dump`): guest interfaces
(veth*/docker*/br-* plumbing skipped), default route + gateway + source interface, DNS servers
from the guest's resolv.conf, and the SAME `lan_address` value Leg A shows so a support session
can cross-check the two. Best-effort per item — a failed read yields that item's error string in
place, never aborts the dump.
**Leg C — the NetBIOS trap gets named**: helper text under the NAS add form's Szerver field, plus
one hint line appended to an `unreachable`-class add failure when the submitted server is a
single-label non-IP name („Tipp: a(z) »FELHOM« Windows-hálózati névnek tűnik…"). The detection is
purely lexical (`looksLikeFlatNetworkName`: non-empty, no dot, not `net.ParseIP`-able) — no
NetBIOS/mDNS resolution is attempted anywhere, and the agent's probe/taxonomy is untouched.
**The one design decision worth recording:** the spec sketched the gateway as a `/proc/net/route`
read, but the controller runs on a docker BRIDGE — every in-process answer (own routes, own
resolv.conf = 127.0.0.11, `net.Interfaces` = 172.x) is the S-2 wrong-kind-of-true trap that
already burned the setup wizard. All guest-net reads therefore go through the ONE guest-netns door
this process has: a docker-exec into the host-networked felhom-samba container
(`internal/stacks/guestnet.go`, single `guestNetExecFn` seam). Accepted consequence, by S-5's own
logic: with Megosztás off the door is closed and the card shows „—" rather than a plausible wrong
172.x answer.
Tests: `guestnet_test.go` (pure parsers pinned: default route, interface merge, resolv.conf;
fail-quiet contracts; B1 best-effort with a scripted per-argv exec fake) + `network_card_test.go`
(A1 all rows, A2 name-row absent when sharing off, A3 „—" fallback, per-render freshness counter,
B1 dump shape with in-place error, C1/C2/C3 hint lexicon). Red-proofs run and recorded in
REPORT.md: A2 (enabled-gate dropped → `\\FELHOM` rendered while sharing is off → FAIL) and C2
(lexical check inverted → the hint nags an IP user → FAIL).
### v0.158.1 — fix: the lifecycle methods broke every app detail page (2026-07-21)
**Defect shipped in v0.158.0 and caught live within the hour. `/apps/<slug>` returned HTTP 500 for
EVERY app**, not just withdrawn ones.
`EffectiveLifecycle` / `CanInstall` / `IsAbandoned` were declared with POINTER receivers.
`appDetailHandler` puts `data["Meta"] = found.Meta` — a `stacks.Metadata` VALUE inside a
`map[string]interface{}` — and html/template cannot call a pointer-receiver method on a
non-addressable value. So `{{if .Meta.IsAbandoned}}` failed at RENDER time:
```
executing "app_info" at <.Meta.IsAbandoned>: can't evaluate field IsAbandoned in type interface {}
```
Switched to value receivers, with the reason recorded at the declaration so it is not "tidied" back.
**Why the tests missed it, which is the more useful lesson:** it compiles, `go vet` is silent, and
every v0.158.0 test passed — because none of them rendered `app_info`. The catalog-page tests
exercised the funcmap route (`lifecycleBadge .Meta`), which takes a value and works either way. A
template method call is only ever checked when the template actually runs.
Added `TestAppInfoRendersForEveryLifecycle`, which renders the real `app_info` template through the
production tree with the handler's exact data shape — `"Meta"` as a VALUE in a
`map[string]interface{}`, deliberately not a pointer, because the pointer is what hides the bug.
Red-proof: restoring the pointer receiver reproduces the 500 for every lifecycle value including the
empty one.
### v0.158.0 — apps get a lifecycle: available / hidden / abandoned (2026-07-21)
No agent coupling; MinAgent unchanged.
Until now the catalog knew only two states: a template is present, or it is gone. "Gone" is not a
usable way to withdraw an app, because **it orphans every customer already running it** — their app
gets flagged `Elavult` and offered a Törlés button, for software that works fine. That is what the
short-lived `retired/` directory move (2026-07-21, same day) would have done, and it is why this
replaces it.
`.felhom.yml` gains an optional top-level `lifecycle:`:
- **`available`** — the default. Absent or empty means this, so all 52 existing templates are
unchanged.
- **`hidden`** — not offered for new installs. Nothing is shown to anyone already running it; "we
stopped offering this" is not their problem.
- **`abandoned`** — not offered for new installs, AND every box already running it carries a
permanent „Nem karbantartott" badge plus a notice on the app page: *„Az alkalmazás fejlesztője
felhagyott a fejlesztéssel. A telepített verzió továbbra is használható, de frissítések és
biztonsági javítások már nem érkeznek hozzá."*
**A deployed instance keeps full function in every state.** Lifecycle governs what is OFFERED, never
what runs.
- **The deploy gate is server-side and fail-closed** (`api.deployStack`, before any mutation), with
the ruled Hungarian refusal „Ez az alkalmazás jelenleg nem telepíthető." Hiding a button is not a
gate — a stale link, a bookmarked deploy form or a direct POST must all be refused. A second check
in `stacks.DeployStack` covers any future caller that does not route through the API.
- **The unknown-value posture is fail-OPEN, deliberately, and it is the opposite of the gate's.** An
unrecognised value degrades to `available` with one WARN. A typo — or a state added in a later
catalog than this controller understands — must never silently pull a working app out of every
customer's catalog. The gate that actually protects installation reads the same
`EffectiveLifecycle`, so the two can never disagree.
- **Orphan detection is untouched, and that is asserted.** Withdrawn templates stay in the catalog
tree; `getCatalogTemplateSlugs` never looks at lifecycle. A red-proof adds that filter and shows
the abandoned app immediately reading as an orphan.
- **Badge plumbing is generic**: `MetaBadge` + the `meta_badge` partial + a `lifecycleBadge` funcmap
entry. R-56's difficulty labels are meant to be a sibling funcmap function returning the same type
— no new markup, no new CSS.
- **plant-it returns to `templates/`** as the first `abandoned` app, so the mechanism is proven on
the case that motivated it. Its compose is deliberately left as-is: the app is not installable, and
rewriting it would imply it is.
**Red-proofs, all four run:** removing the API gate → the wiring test reports the gate INERT;
dropping the `Deployed ||` clause from the catalog filter → a customer's running app vanishes from
their own Alkalmazások page; removing the badge line → the abandoned app renders unmarked; making
orphan detection lifecycle-aware → `catalog set = map[bookstack:true]`, the two withdrawn apps read
as orphans. The wiring test walks the AST, not `strings.Contains`, because a commented-out call
still contains the string; it also asserts the gate precedes `DeployStack`.
### v0.157.1 — anchor the `controller` .gitignore entry (2026-07-21)
Tooling only; no behaviour change, no rebuild needed.
`controller/.gitignore` carried a bare `controller`, which git matches against DIRECTORIES as well
as files — so it also matched `cmd/controller/`. Two opposite failure modes came out of that, and
both manufacture inert seams: ripgrep silently skipped `cmd/controller/main.go`, so a search for a
setter's caller returned nothing and read as "this is unused" (a false no-caller reading has already
been recorded once); and genuinely-new files under `cmd/controller/` needed `git add -f` or were
never committed at all. Anchored to `/controller` + `/controller.exe`, which still ignores the built
binary at the module root — verified both ways.
### v0.157.0 — the boot bind gate honours a customer's Stop (R-55) (2026-07-21)
**Your Stop now means Stop across a guest reboot for drive-backed apps too** — the guarantee R-52
already gave every other app. Found by STOP-1's R-52 leg on 2026-07-21, which was designed to prove
the opposite: immich, stopped from the UI seconds earlier, came back running after the reboot.
The boot bind gate (`internal/web/intermediary.go`) keyed its recreate on
`Deployed && HDD_PATH && drive-present` alone. `Deployed` is a deploy-lifecycle flag — it stays true
across a Stop — so the gate had no way to tell "the guest went down under this app" from "the
customer switched this off", and it resurrected both. R-52 was never implicated: its own gate behaved
exactly as specified (immich, at zero containers, was never a candidate for it). The gate simply
reaches every drive-backed app first.
**The fix is R-52's own predicate, translated.** `shouldRecreateOnBoot` now also requires
`len(Stack.Containers) > 0` (from `docker ps -a`, so `Exited` containers count):
- containers EXIST but are down → the guest went down under the app; docker's records survive the
reboot → boot orphan → recreate, as before.
- ZERO containers → a UI Stop is `compose down`, which REMOVES the containers → deliberate → leave it.
**What deliberately did NOT change: container STATE is still not a filter.** That is the original
design's load-bearing part — a `State != stopped` filter misses an app that simply hasn't been
auto-restarted yet after the boot, or is stuck `Exited` on a create-time bind failure with
`RestartCount=0`. `hasContainers` is a different question ("does docker still have records of it")
and, unlike liveness, it survives a reboot as a statement of intent. `TestShouldRecreateOnBoot` now
pins both axes at once — they pull in opposite directions, which is the whole difficulty of this gate.
- **Ordering trap, handled:** the evidence is sampled into the `bootStack` snapshot BEFORE any
recreate runs, because `recreate` calls `StopStack` (`compose down`) and so destroys the very
signal the decision needs.
- **The drive-absent gate is not regressed.** Apps it stopped are also at zero containers, so this
path now skips them — correctly: they are recorded in `StoragePath.StoppedStacks` and restarted by
`ReconcileDriveGates`' `Return` branch, which runs on the same `driveGateLoop` tick.
- **Honoured Stops are observable.** `leftStopped` is counted and logged separately from `skipped`
at INFO (`… left stopped — zero containers means the customer stopped them on purpose`).
Conflating them would have fired a WARN about a missing drive bind for an app behaving exactly as
asked, and a silent correct path is how an inert seam hides.
- **Red-proof (run):** dropping `hasContainers` from the predicate makes
`TestRecreateDriveBackedApps_HonoursCustomerStop` fail with `recreated=[romm immich]` — the live
defect, by name.
### v0.156.0 — a dead primary alerts (R-51); a boot orphan restarts itself (R-52) (2026-07-21)
**No new agent coupling — MinAgent stays 0.90.0.** Two independent failures from the same live
audit, both unattended-resilience holes: the box was broken and nobody was told, then the box could
have fixed itself and did not.
**R-51 — a multi-container app whose MAIN container is dead now counts as down.** On 2026-07-20
`immich-server` sat `Exited` for **18 hours** with the app 100 % unreachable, and the box produced no
dead-app banner and no `app_start_failed` event — while single-container Calibre-Web, down for the
same reason, alerted in 90 seconds (AUDIT-vacation-remote-ops-2026-07-20 F4).
The defect was one branch in `aggregateState`: a stack with *some* members running and *some* stopped
returned `StateRunning` — "partial" — and `IsDownState` (correctly) does not treat running as down.
So the alarm never had anything to fire on. *(The ROADMAP row's diagnosis — "aggregation classifies
such a stack `unhealthy`" — is wrong at the source; corrected in the row.)*
- New `StateDegraded`. The mixed branch now asks each DOWN member for its restart policy: a member
docker is supposed to keep running (`always` / `unless-stopped`) makes the stack **degraded**, a
finished one-shot (`no` / `on-failure`) leaves it running. `IsDownState` gains `degraded` and
**nothing else** — the `unhealthy` / `restarting` / `paused` / `unknown` exclusions are byte-
identical, because folding `unhealthy` into down is what fix-3 removed the flapping by not doing.
- An **unreadable** policy counts as supervised (fail-CLOSED), the opposite of the IsDownState
fail-open rule and for a different reason: there the *state* is ambiguous, here a member is known
dead and only the excuse is missing. The P2 census backs it — all 53 catalog templates / 78
services are `unless-stopped`, and zero one-shot containers exist today.
- The policy read is one `docker inspect` per down member of a *mixed* stack, cached per
container+state and pruned to the live container set, so the 10 s refresh does not grow a docker
call per container.
- Everything that asks "are there live containers here" learns the state too: quiesce
(`RunningAppStacks`), delete's stop-first guard, the export stop-first guard, telemetry, health
probes. Everything that asks "is this app working" counts it as down: the dashboard counter, the
stopped filter, the dead-app banner and the alarm. UI: „Részlegesen leállt", warn colour, and the
URL is flagged unpublished (Traefik 404s when the routed member is the dead one).
**R-52 — an app the boot left behind now gets exactly one recovery.** The same shutdown left immich
and calibre-web `Exited` while ten sibling containers came back; the controller *reported* them for
18 hours and never started them (F5).
- New `internal/bootrecon`: one bounded sweep at startup — at most 2 attempts, 30 s apart, then it
stops and the alarm owns the problem. **Never a restart loop.**
- **A deliberate Stop survives a reboot.** The UI's Stop is `compose down`, which REMOVES the
containers; an interrupted boot leaves them behind as `Exited`. So the boot-orphan signature is
"deployed, has containers, and they are down", and a zero-container stack is never touched.
- The whole sweep (5 s settle + one 30 s gap) fits inside the 90 s `deadAppBootGrace`, so a
successful recovery never alerts and a failed one alerts honestly. A test asserts that arithmetic
rather than leaving it to a comment.
**Seam discipline (the reason both features have a wiring test).** Two inert-seam defects shipped in
the two days before this: controller v0.154.0 and agent v0.91.0, both a correct component with green
tests and no production caller. So the boot sweep is asserted from `package main` — including an AST
walk proving `func main()` actually contains the `go runBootReconcile(...)`. That test was written
first as a `strings.Contains` and **its own red-proof passed it**, because a commented-out call still
contains the string. Comments are not callers; the AST version fails as it should.
Red-proofs (all run, all failed on the pre-fix shape, all restored): the mix branch reverted to
`return StateRunning` → the immich fixture and both production-path tests fail with `"running"`; the
boot hook commented out → the wiring test fails; the zero-container gate dropped → the user-stopped
app is started, which is the one thing R-52 must never do.
### v0.155.0 — the restore wizard read the wrong "is something running" flag (2026-07-21)
**No new agent coupling — MinAgent stays 0.90.0.** Fixes a defect shipped in v0.154.0 and found by
the operator on the first live click-through, plus the dead phase-strip label from the same release.
**The bug.** `backup.Manager` carries two different booleans and v0.154.0 read the wrong one:
| flag | read by | set by | covers the verification restore? |
|---|---|---|---|
| `running` | `IsRunning()` | `acquireRunning()`, **inside** the goroutine | **no — `RestoreOffboxScratch` never acquires it at all** |
| `opRunning` | `RestoreStatus()` | `BeginRestoreOp()`, in the handler, synchronously | yes, all four offsite actions |
The wizard sourced `OpRunning` from `IsRunning()`. For „Ellenőrzés" and the full-restore preparation
— the wizard's two most-used actions, and the long ones, since they stream from restic — that flag is
false for the *entire* operation. So the execution step was unreachable: the page kept offering all
three intents with live buttons while a restore was downloading, and the progress banner (which polls
the op status) contradicted the phase strip on the same screen. Any button pressed there would have
been refused by the handler — which is exactly the "offering a control guaranteed to fail" dishonesty
R-48 exists to remove.
**The fix** is one line of behaviour behind a named seam: `restoreOpInFlight(st)` takes the
`RestoreOpStatus` the handler already reads once, and its doc comment states which flag is which and
why. The handler now takes a single `RestoreStatus()` read, so the strip, the suppression decision
and the running-op name can no longer disagree with each other.
**Why the v0.154.0 tests missed it.** The Scenario-E table proved `deriveWizardStep` behaves
correctly *given* `OpRunning=true`; nothing proved the handler ever computes `true`. Hollow at exactly
that seam. `TestRestoreOpInFlight_UsesDisplayFlagNotConcurrencyFlag` now drives a real `Manager`
through `BeginRestoreOp` and asserts the wizard suppresses every form — red-proofed against the
v0.154.0 shape.
**„Eredmény" is now reachable.** The fourth phase label never lit up in v0.154.0. The strip's
highlight is now its own derived value (`Phase`), separate from `Step`: a finished restore is back on
the intent step — everything is available again — while the strip rightly reads „Eredmény" and an
outcome card shows the result. Bounded by `restoreResultWindow` (10 min) so a stale result cannot
claim to be fresh, and bound to the app, so a finished bookstack restore does not light up immich's
page with bookstack's message. The card survives a reload, which the redirect flash does not.
### v0.154.0 — one restore entry per app, and the intent is a described choice (2026-07-21)
Closes **R-48**. **No new agent coupling — MinAgent stays 0.90.0.** This is a UI-layer change:
`internal/backup`, `internal/appbackup` and `internal/selfupdate` are untouched, and the release adds
**no mutation endpoint** — every action still posts to the `/backup/offbox/*` handler it always did,
with the same field names and the same gates.
**The defect.** The „Ellenőrző visszaállítás a távoli tárolóból" list rendered up to five inline
`<form style="display:inline">` blocks per app row: verify, prepare, the revealed size-gated commit,
the missing-only merge, and the true reconstitution. Two of them sat next to each other as sibling
buttons —
- „Helyreállítás az élő adatok közé (csak a hiányzó fájlok)" — additive; **cannot** bring deleted
content back, and
- „Teljes visszaállítás (fájlok + adatbázis)" — the real restore
— and the difference between them is whether the customer's data comes back at all. This is not
theoretical: it caused the round-2 incident. An operator who had *read the source* pressed the
missing-only button, and the controller log shows `/backup/offbox/reconstitute` was never hit
(`felhom.eu/documentation/audits/DIAG-immich-restore-round2-2026-07-19.md`, finding 1). The second
half of the trap was that the decisive „Teljes visszaállítás indítása" appeared **only after**
„…előkészítése" had been pressed, with nothing signposting that a second step existed or that the
first one had done nothing to live data.
**The rule this establishes,** worth stating once and applying past this page: *two adjacent controls
whose difference is "your data comes back" vs "your data cannot come back" must never be
distinguishable only by layout.*
**The change.** Each app row on `/backups/restore` now carries exactly **one** control —
„Visszaállítás…" — linking to a per-app wizard at `GET /backups/restore/app?name=<app>`, built on the
`backups_escrow.html` precedent:
- **Three intent CARDS**, each with its own consequence sentence rather than a label alone:
ellenőrzés külön mappába (live data untouched) · hiányzó fájlok visszahozása (additive, no
database, deleted content does not reappear) · teljes visszaállítás (files + database, danger
styling, the R-43 double-confirm carried over **verbatim** with its pair-honesty facts).
- **A visible phase strip** — Előkészítés · Megerősítés · Végrehajtás · Eredmény — so the sequence is
legible before the first click instead of after it.
- **Server-derived steps.** `deriveWizardStep` is a pure function of (op running, size-gate flash,
scratch ready); the step is never accepted from the request. Precedence is strict: a running op
outranks a stale `?full_prep=` in the URL, so no commit button can reappear mid-restore.
- **Mutation forms are suppressed server-side while any op runs** — the backup manager's
single-flight is process-wide, so a restore for app X now suppresses app Y's controls instead of
offering a button guaranteed to 409.
- **No JavaScript requirement.** Every step is a real form POST and the server renders the next one.
**Redirect retargeting.** The app-scoped `/backup/offbox/{restore,place,reconstitute}` outcomes now
land back on the wizard the customer acted from rather than on the list. Fixing that surfaced a
latent bug in `offboxRedirectTo`, which hardcoded `"?"` when appending the flash — against a target
that already carries a query (`?name=<app>`) that would have buried the flash inside the `name`
value. The separator is now chosen.
**Deliberately NOT in scope:** the shares (`_shares`) entry, the local restore panel and the .fab
block are untouched; the R-45 job registry is still its own item — the wizard polls the two existing
status surfaces as-is.
### v0.153.0 — the database replay no longer races the application, on BOTH restore paths (2026-07-20)
Closes **R-47**. **No new agent coupling — MinAgent stays 0.90.0.** Nothing in this release talks to
the host agent; the whole change is inside the controller's own compose orchestration.
**The defect, measured to the second.** On 2026-07-19 the offsite reconstitution was run deliberately
and correctly (`felhom.eu/documentation/audits/DIAG-immich-restore-round2-2026-07-19.md`, finding
**H4**). It executed its designed sequence — safety dump, stop, start, replay — and the replay
aborted:
```
10:58:25 controller: replaying DB dump into immich-postgres
10:58:33 immich-server: "Reindexing clip_index" -> "Reindexed clip_index" <- the app recreates it
10:58:35 controller: ERROR relation "clip_index" already exists - exit status 3
```
The replay needs a running database container, so the code started the WHOLE stack first. That gave
immich-server an eight-second window in which to rebuild the very schema objects the dump was about
to create, and under `ON_ERROR_STOP=1` the collision aborted the script. The photos came back anyway
**by accident**: `pg_dump` emits COPY data before CREATE INDEX, so the abort landed after the rows. A
collision earlier in the script would have left a genuinely half-restored database and reported it
identically. The operation reported failure and immich then reported schema drift.
**The fix: a DB-only window.** After the files are placed, only the stack's database service(s) come
up; the dump is replayed into them with the application still stopped; the rest of the stack starts
only once the replay has exited 0. Nothing about the replay itself changed — `--clean --if-exists`
and `ON_ERROR_STOP=1` were always correct. The bug was the window, not the flags.
**This was a class defect and both paths carried it.** The local `RestoreFromRecoveryUnit` had the
same start-then-replay shape, hidden inside `RecreateStackFromUnit` (which ended in a full
`compose up -d`). Fixing only the offsite path would have left the identical race one button away.
Both are re-sequenced here.
**What changed**
- `appbackup.DBServiceNames(composePath)` names the compose SERVICE(s) whose `image:` identifies a
database — `docker compose up -d` takes service names, not container names. It is a yaml.v3
`services:` map parse, deliberately not a line scan: immich's real template carries top-level
`immich_ml_cache:` and `immich_postgres_data:` volume keys that sit at exactly the indentation a
service name does.
- The image heuristic that `DiscoverDatabases` had inline is extracted to `dbTypeForImage` and shared
by both. That sharing is what makes the safety argument hold: a `.sql` dump can only exist because
discovery matched the running container's image, and the compose `image:` value IS that image
string — so "a dump exists" and "a service can be named" are answered by one predicate.
- `stacks.Manager.StartStackServices(name, services)` runs the scoped `up -d`. It **refuses an empty
service list**: an argument-less `up -d` is a full start, which is precisely the behaviour the
window exists to avoid, and a silent fall-through would have reintroduced the race at the one call
site that most needs it not to.
- `RedeployFromEnv` is split. Its persist half is now `PersistUnitRedeployConfig` (app.yaml, locked
fields, in-memory flags — starting nothing); `RedeployFromEnv` is that plus its unchanged
up-and-report tail, so its public behaviour is byte-identical. The split is what lets the restore
path put the DB-only window between persisting the definition and starting the app.
- `StackDataProvider.RecreateStackFromUnit` becomes `RecreateStackDefinitionFromUnit` (files +
persist, no start), and gains `StartStackServices`. The rename is deliberate: the old name promised
less than the method did, and the hidden `up -d` inside it is what carried the defect on the local
path.
**Fail-closed, both paths.** If a `.sql` dump exists but no database service can be identified in the
compose, the restore **refuses before the first mutation** — no stop, no file overwrite, no volume
restore. The alternative would be to start everything and replay into the race. Given the shared
predicate this should be structurally unreachable; it is the belt for template drift, not an expected
path.
**Every exit from the window still starts the app.** A failed replay, or a failed DB-only start, is
surfaced as before — but a best-effort full `StartStack` runs first. The DB-only state is a
deliberate half-started one, and leaving a customer with a running database and no application would
turn a failed restore into an outage.
**Tests.** 19 new (Groups A–G): ordering plus **state-at-replay-time** on both paths (a recording
provider captures whether the full stack was up at the moment the import fired — asserting "no error"
would have passed on the pre-fix shape, which is how this shipped), the no-DB negatives, the
zero-mutation fail-closed effects, the replay-failure bring-up, the compose-parser decoys built from
the catalog's real immich template, and the empty-list refusal. Three companion red-proofs run and
reverted: the pre-fix full start on the offsite path, the pre-fix full start on the local path, and
deletion of both fail-closed gates — each failing on the intended assertion. 23/23 packages green.
**Live-validated on the demo box, 2026-07-20 (operator present).** Endpoint-level, against the SAME
snapshot (`49e7cb46`) that aborted in round 2:
```
15:39:42 [stacks] Stopping stack: immich
15:39:43 [stacks] Starting stack immich services only: [immich-postgres]
15:39:43 [backup] Restore immich: replaying DB dump into immich-postgres (postgres)
15:40:03 [backup] Restore immich: replayed 1 DB dump(s) <- rc-0, no "already exists"
15:40:03 [stacks] Starting stack: immich
15:40:27 [offbox] reconstituted immich: 6 file(s) placed, 1 DB dump(s) replayed, skewed=false
```
The operation reported **success** (round 2 reported failure); immich's own DatabaseService logged
**`No schema drift detected`** twice, where round 2 left it reporting drift; 11 assets `active`, all
four containers healthy, 231 `public` indexes. Details in `REPORT.md` §4b.
**Golden 0.153.0 baked + published the same day** (`build-golden.sh` v2.1.0, from the vacation site
after a registry-reachability probe). First golden carrying **all four** infra images — the list came
from `--print-infra-images` on the 0.153.0 binary itself, so the historical 3-image fallback never
fired and `felhom-samba:1.1.0` is baked. Upload 201, anonymous GET byte-matches, ranged 206.
```
GOLDEN_VERSION=0.153.0
GOLDEN_SHA256=15fdd191f3c660a60dc8651111053dd84281aeebc6c4c0f9ecdd3a87cb45a9d0
```
**Still outstanding:** the two password-gated hub saves (Day-0 manifest Golden → 0.153.0, then floor
→ v0.153.0 **last**; Agent 0.90.1 / MinAgent 0.90.0 unchanged), Viktor's C6 customer-restore UI run,
and the immich timeline screenshot — all of which need the operator UI or a browser.
### v0.152.0 — Megosztás on a Mac: mDNS in the image, and the page stops giving Mac users a dead form (2026-07-20)
Closes **S-3** of `felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md`, and fixes a copy
defect v0.151.0 shipped the same day. **Pairs with felhom-samba 1.1.0** — the pin in
`infra.SambaImage` moves with it, so `Images()` and the golden bake follow automatically.
**The finding that redirected the fix — macOS asks, gets a correct answer, and ignores it.** The
first theory was that modern macOS no longer does NetBIOS. A packet capture on the box disproved
that: on a bare `smb://FELHOM` the Mac broadcasts a well-formed NBNS query for `FELHOM<20>` (the
File Server Service suffix — exactly right for SMB), and nmbd answers in 140 microseconds with a
textbook positive response — flags `0x8580` (response, authoritative, RCODE=0), ANCOUNT 1, unique
B-node, the correct address. **macOS never opens a TCP connection.** Sixteen seconds later the same
Mac connected through `smb://FELHOM.local` on the first try. NetBIOS on macOS feeds legacy browsing,
not `smb://` URL resolution — so no change on our side can ever make the bare name work there, and
nmbd is not the thing that was broken. (nmbd answers twice per broadcast, because it holds
`0.0.0.0:137`, `<ip>:137` and `<bcast>:137` and a broadcast lands on two of them. Standard Samba;
investigated and dismissed — a duplicated correct answer is still a correct answer.)
**felhom-samba 1.1.0 — avahi + dbus, so the Mac has a mechanism at all.** The image's discovery set
was Windows-only: nmbd for flat-name resolution, wsdd for Explorer's Network view, and nothing
whatsoever for Bonjour. It now runs avahi, with `avahi-daemon.conf` and an `_smb._tcp` service file
**templated from `FELHOM_SERVER_NAME` in the entrypoint** — renaming the server in the UI
re-advertises under the new name, where a baked name would leave the box answering to something the
customer can no longer see anywhere. A static service file rather than smbd's own `multicast dns
register`: it needs no line in `smb.conf` (bind-mounted READ-ONLY, owned by the controller's
renderer) and it lets us publish `_device-info._tcp` for a sensible Finder icon. Both new daemons
are non-fatal on failure — sharing over an address must not become an outage because a discovery
daemon did not come up. Proven live from the operator's Mac before the image was built, then the
built image smoke-tested with all five daemons up and avahi registered as `<NAME>.local`.
**The page no longer tells Mac users to do the one thing that cannot work.** v0.151.0's connect card
offered `smb://<NÉV>` for Mac. That is precisely the dead form. It is now `smb://<NÉV>.local`; the
Windows line stays the flat `\\<NÉV>`, which nmbd serves correctly and which this release must not
disturb. Red-proofed: reverting the template to the bare name turns
`TestSharingConnectCard_MacLineIsDotLocalNotBareName` red on both the missing `.local` and the
present bare form, for two different configured names — and the same test asserts the Windows line
neither disappears nor wrongly gains `.local`.
**NOT claimed: automatic Finder-sidebar discovery.** The `_smb._tcp` record is published and answers
browse queries on the wire, but the test Mac's sidebar stayed empty — it had no Network/Bonjour
section shown at all, which is a Finder Settings toggle rather than something the box controls. This
is recorded as OPEN in the DIAG, deliberately not as a shipped feature.
**Two test bugs surfaced and fixed, neither a production defect.** `TestRenderSambaCompose` asserted
the literal tag `felhom-samba:1.0.0`, so a routine image bump read as a renderer regression; it now
derives from `SambaImage` and separately asserts what actually matters — that the tag is explicit and
never `:latest`. And `TestFabUpload_GCAndIdleTimeout` raced: `expireIdleUpload` nils the slot,
releases the mutex, and only then closes and unlinks the `.part`, so "the slot is free" does not yet
mean "the file is gone" — the test stat-ed immediately and passed only by luck. It failed in the full
package while passing in isolation once this release's new render tests made the `web` package
heavier. Now it waits for the outcome it asserts, on the same 3 s deadline; red-proofed by removing
the unlink from production, which still fails it.
### v0.151.0 — the Megosztás page stops reloading, and says how to connect (2026-07-20)
Closes **S-1**, **S-2**, **S-5** and the core of **S-4** from
`felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md`.
**S-1 — `/sharing` reload-looped about once a second, for every customer with sharing enabled.**
`GET /sharing/status` carries two things that mean different things to the client: `phase` (the
ensure JOB — the page answers a terminal `running` with a one-shot `location.reload()`, because the
„Állapot" badge is server-rendered) and `running` (the service LEVEL, straight from the liveness
probe). v0.147.0 coerced `idle`→`running` on the PHASE channel so that a missing job could never
contradict a live container. That duty was real, but it belongs to — and was already discharged by —
the `running` field beside it; on the phase channel the same value reads as a fresh success edge. The
poll's `tick()` runs synchronously at script end, so the FIRST poll of every steady-state page load
reported a terminal job that had never run, scheduled a reload 1.2s later, and the new page did it
again. The coercion is gone: no job, no edge. The defensive intent it was written for is now pinned
by its own named regression test on the `running` field.
**S-4 (core) — a REAL bring-up is now reported exactly once.** Without this the loop would return
after every future image update: the finished job outlives the reload it triggered, so the next page
load found `phase:"running"` waiting for it. `consumeIfRunning` serves a terminal `running` once and
clears it — and only while the single-flight slot is free, since the job goroutine sets the phase
before its deferred `release()` and eating it in that window would lose the success the customer is
waiting on. `failed` and `needs_password` stay sticky (their client path stops the timer and shows a
card with NO reload, so stickiness is informative and cannot loop), and in-flight phases are never
consumed. Accepted cost, stated rather than hidden: with two tabs open during a bring-up only the
first gets the success banner — both still show the true state, which comes from the level channel.
The unified async-job feedback layer remains the ROADMAP item; this is the minimal contract fix.
**S-2 + S-5 — the page now names both ways in.** It had only ever shown the configured NetBIOS name,
so a customer whose network fails to resolve it had no fallback but a guess — and the guess that
produced the diagnosis was the Proxmox HOST's address, which never ran smbd. New
„Csatlakozás a megosztáshoz" card: the Windows form, the Mac form, and the direct `smb://<IP>`.
The address comes from `stacks.SambaLANAddress()`, which reads the guest's netns through the SAMBA
container (`network_mode: host`) — the controller is on a docker bridge and would answer `172.x`,
the same trap `setup.DetectLocalIPs` needs `HOST_IP` for. Reading it there also makes it the right
kind of true: it is the address smbd is bound to, not merely one the box owns. **Derived per render
and cached nowhere** — the guest holds it by DHCP, so a stored copy eventually misdirects people
(S-5) — and an underivable address omits the line, because no address beats a wrong address.
`sharing.html`'s `<script>` block is byte-identical to v0.150.0: both fixes are server-side, so the
client contract is proven fixed rather than papered over. `infra.SambaHostInterface` replaces the
third `eth0` literal (smb.conf, `FELHOM_IFACE`, and now the address read must name the same nic).
Red-proofed three ways — reinstating the coercion, deleting the serve-once clear, and memoizing the
derived address each turn the corresponding test red. 23/23 packages green.
### v0.150.0 — green gate restored + the export link stops leaking the CSRF token (2026-07-20)
**F7 / R-53 — `app_export.html` built the app's public URL from the CSRF token.** The line read
`var domain = '<subdomain>.{{$.CSRFToken}}'`, so the „Megnyitás" link was wrong for every app with a
subdomain and a session CSRF token was written into a URL (history, referrers, logs). Two-part fix:
the template token becomes `{{$.Domain}}`, and `exportPageHandler` supplies `Domain` — that handler
builds its own data map instead of going through `baseData`, which is where every other page gets
the key, so the template had nothing to read. The page's real CSRF path (the `csrfH()` helper
reading the meta tag) is correct and untouched. Render tests assert the joined `<sub>.<domain>` and
that the token appears nowhere on that line; red-proofed against the pre-fix template.
**The 7 red `internal/backup` tests are green again — no behaviour change.** `TestTier2V2_*` and
`TestSharesTier2*` had been failing on DooPlex since before v0.149.0. Root cause is environmental,
one class for all seven: Tier-2's off-drive guard asks `system.SamePhysicalDevice` (st_dev equality)
whether a candidate target is really a *second* disk, and on a host where every `t.TempDir()` lands
on one filesystem the fixture's "two drives" are indistinguishable — so the guard correctly refused
the target and the tests could never reach their subject. The failure message said so outright:
`nincs másik fizikai meghajtó`.
Fixed with one behaviour-preserving seam in the package's existing style: a nil-defaulted
`Manager.samePhysicalDevice` field plus a `sameDevice` wrapper, with the seven call sites routed
through it. **Nil → `system.SamePhysicalDevice`, so production is byte-for-byte unchanged**; only
the two test fixtures install a fake that models one drive per directory subtree. No assertion was
weakened, no test skipped, renamed or deleted, and every one of the seven was mutation-proved: the
defect each guards was re-introduced one at a time and each test failed, including the notifier
test's own documented red-proof (`_shares` reaching Hungarian copy).
### v0.149.0 — the dashboard tells the truth about the last backup (2026-07-20)
Closes **F3** from `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md`.
The dashboard's backup card claimed **„Utolsó mentés: Még nem futott"** on every box, forever —
including boxes with dumps on disk and `crossdrive_completed` / `db_dump_completed` events already
recorded in the hub. It was not a backup failure; it was a lie in the view layer.
`dashboard.html` branches the row on `{{if .BackupStatus}}` and reads `.Success` / `.LastRun` from
it, but `dashboardHandler` never put `BackupStatus` in the template data. The key was always
missing, so the `{{if}}` arm was unreachable and the `{{else}}` — "never ran" — rendered
unconditionally. The neighbouring „Adatbázisok: N mentve" row kept working because it reads
`DBDumpStatus`, which *was* passed; that is exactly the contradiction the audit caught on the live
box (a card reporting "never ran" directly above "2 mentve").
The fix is the one-line pass-through the template always expected:
`data["BackupStatus"] = fullStatus.LastDBDump`. `*DBDumpStatus` nil/non-nil maps exactly onto the
template's branch, so a genuinely fresh box still reads „Még nem futott" honestly and no zero-value
timestamp is ever fabricated. No template change, no new view-model, and "utolsó mentés" keeps its
existing meaning (the last DB-dump run, consistent with the backups page's DB section).
Tests (`internal/web/dashboard_backup_card_test.go`) drive the **real handler** through
`ServeHTTP` rather than the template alone, so they bite on the handler wiring: a planted dump file
on the app's drive must surface as its own timestamp; a box with no dump must still say „Még nem
futott" and must not render `0001-01-01`; a failed run must render „Sikertelen". Red-proofed —
deleting the new assignment fails the first of those.
### v0.148.0 — coherent snapshot pairs + an offsite restore that actually restores (2026-07-19)
Closes **R-43** and **R-44**, the two findings from `DIAG-immich-restore-2026-07-19`. The short
version of that diagnosis: Viktor deleted 11 immich photos to test offsite restore, both restore
runs flashed success, and the photos stayed gone. Two independent defects, both fixed here.
**R-43 — no offsite path could restore a database.** All three offsite buttons were file-only.
The two „visszaállítás" actions staged into a scratch folder and never touched postgres; the
place-to-live action merged only files MISSING from the live tree and never replayed a dump. For a
DB-indexed app — most of the catalog — that combination cannot bring content back: the bytes
return and the app still cannot see them, because its index lives in the database. The dump was
faithfully carried INTO every snapshot and could never be replayed OUT of one.
New: **„Teljes visszaállítás (fájlok + adatbázis)"** (`ReconstituteFromOffsite`,
`/backup/offbox/reconstitute`). Safety dump → stop → files overwritten to the snapshot's version →
start → the snapshot's own dump replayed → health wait. Two invariants:
- **Nothing is ever deleted.** The full-restore copier is `rsync -a` with NO `--ignore-existing`
(a changed file becomes the snapshot's version) and NO `--delete` (a file created after the
snapshot survives as an extra). A restore that silently removed newer work would be a data-loss
event wearing a recovery button's label.
- **The undo exists before the act.** A `pre-restore-` dump of the live database is written and
verified on disk BEFORE anything is stopped, overwritten or replayed; if it cannot be taken the
whole operation refuses with zero changes. The safety dumps live in the app's own unit and
appear in `ListDumpFiles` — an undo the customer cannot see is not much of one.
The replay reads the SCRATCH unit, not the live one: the live recovery unit is still never
overwritten (it is the local restore path's source), so replaying from it would replay the current
database back over itself and restore nothing.
**R-44 — a manual push shipped an unrefreshed dump.** `RunOffboxBackup` went straight to the
restic push; dumps came only from the separate 02:30 local run, so a manual push at any other hour
shipped a dump up to ~24h old. On 2026-07-19 that dump was taken four hours before the customer's
account existed and probed to `asset: 0 / user: 0 / album: 0` — a 52MB file whose entire bulk was
immich's shipped geodata tables. Size and table count both called it healthy.
Every offsite run — **manual and nightly** — now refreshes the dumps and recovery units FIRST, then
captures. Order is the mechanism: the gap can only ADD files the DB does not reference yet (a
harmless orphan blob), never remove one it does, so the file set is always a superset of what the
restored DB points at. This also makes the nightly ordering structural instead of a coincidence of
two scheduler entries at 02:30 and 04:15. Each unit manifest carries the run's `offsite_run_id` +
`dumps_at`, so a snapshot's coherence is verifiable at restore time rather than assumed.
**Honesty surfaces** (warn-level, never gates — a false positive that blocked a restore would be
worse than the skew it guards against):
- A pre-v0.148 snapshot has no stamp → the confirm says „Az adatbázis-mentés régebbi (<ts>) — a
fájlok és az adatbázis eltérő időpontból származnak." It still restores.
- `ValidateDump` gained a content sniff: a structurally valid dump whose accounts table has zero
rows raises „A mentett adatbázis üresnek tűnik". Exact table-name matching, deliberately — a
substring match on "user" would flag `user_metadata` / `album_user` / `user_audit` on every
healthy single-user box and turn the signal into noise.
- The completion flash states an OUTCOME, not a mechanism: „A(z) X: N fájl és az adatbázis
visszaállítva (mentés: <ts>) — az alkalmazás újraindult." A no-database app says so explicitly
rather than borrowing the confident sentence.
- The old missing-only button now says what it does NOT do: „Adatbázist nem állít vissza — törölt
tartalom ettől nem jelenik meg újra."
A new `dump` progress phase („Adatbázisok mentése a pillanatképhez…") names the pre-phase, which on
a large database dominates the early wall clock and would otherwise read as a hang.
Tests: 11 new, with **5 red-proofs run and reverted** — replay removed (0 replayed), capture moved
before the dump (`[capture dump]`), both undo guards removed (no refusal), substring table matching
(join tables mistaken for accounts), buffer-exceeding rows uncounted (a wide row sniffed as empty).
Two of those red-proofs found real test weaknesses rather than confirming strength: the first undo
mutation was caught by a second guard, and the first table-matching test did not discriminate
between the two matchers at all — both tests were rewritten to the cases that actually separate
them.
**NOT in this slice:** the catalog-wide invariant check (stays on R-41), nightly cadence, retention,
quota math, tier-2, and the v0.147.x progress semantics beyond the one added phase line. The
missing-only place button's own zero-file flash is also untouched — that is the v0.147 feedback arc's
item, not R-43/R-44.
### docs — the workflow moved to DooPlex-local execution (2026-07-19)
**Docs only, no version bump, no code change.** Claude Code now runs on DooPlex
(192.168.0.180, Debian 13, `kisfenyo`) instead of the Windows workstation, working directly in
`/mnt/5_hdd/felhom.eu/git/felhom-controller`. Builds are local commands; felhom-pve is one SSH hop.
- `CLAUDE.md` — environment/access table rewritten (local DooPlex + `ssh felhom-pve`, no `$SSH`
variable), build/deploy commands de-SSH'd, workspace-root pointer now
`/mnt/5_hdd/felhom.eu/git/CLAUDE.md`.
- **New clean-tree gate** in the build section: `git status --porcelain` empty AND `HEAD` ==
`origin/main` before any build — because the CC working tree IS the tree `build.sh` builds from.
An unpushed change does not exist.
- `claude-in-chrome` is **not available** on DooPlex — endpoint-level validation ("invoke the exact
endpoint the UI invokes") is now the stated standard method; strict UI coverage is a manual pass.
- Windows knowledge is preserved, not deleted: a "Legacy: Windows workstation" note in `CLAUDE.md`
and `RUNBOOK-e2e-live-drive.md`, and `docs/vscode-ssh-fix.md` carries a LEGACY banner.
Historical Windows references in past CHANGELOG/REPORT entries are left untouched — history is
history. Platform-genuine mentions (the `_other.go` dev stubs, `\\FELHOM` shares in Windows
Explorer, the Windows-grep multibyte rationale in the gate scripts) are unchanged.
### v0.147.3 — 4c follow-up 3: the run does not end with the last app (2026-07-19)
Third real run, third thing only a live run could show. The per-app legs finished in ~15 seconds;
the remaining **40 of the 57-second run** was the shares leg and `forget --prune` — during which the
card sat frozen on „calibre-web — 8 / 8 fájl". The same frozen-looking silence 4c exists to remove,
relocated to the end of the run.
Progress now carries a **phase**. The post-app stages announce themselves („Megosztott mappák
mentése…", „Karbantartás: régi mentések rendezése a távoli tárolón…") and the app-scoped counters are
cleared when a phase starts, so the last app's finished numbers are never shown against work that is
no longer about that app. Starting the next app clears the phase again. Pinned by a test.
### v0.147.2 — 4c follow-up 2: when NO counter can move, say what is being worked on (2026-07-19)
The v0.147.1 file-count fallback fixed the incremental case but not the one the demo box actually
hits. Watching a second real run: **bookstack** reported clean byte progress (100%, 154.0 MB, 7/7
files — the byte path works), while **immich** sat at `files_done 1 of 46`, `bytes_done 0`, for 42
seconds. restic 0.14 only counts a file into `bytes_done`/`files_done` when it **completes**, so an
app dominated by a single large archive (immich's ~430MB volume tar) freezes *both* counters. No
percentage can move in that window.
So stop trying to. restic keeps reporting `current_files` and `seconds_elapsed` throughout; the card
now shows the file being processed and the elapsed time — „Mentés: immich — 1 / 46 fájl (430.2 MB) ·
feldolgozás alatt: immich_upload.tar · 42 mp". „Working on this file for 42 seconds" is a completely
different message from „0%", and it is the honest one.
The last known `current_files` value persists across ticks that omit it (restic does not send it
every time, and blanking the label every other second is its own flicker), and switching app clears
it so one app's file is never shown against another. Both pinned by tests, along with the real
42-second status line shape.
### v0.147.1 — 4c follow-up: the progress bar must move on an INCREMENTAL run (2026-07-19)
Found by watching the v0.147.0 card during a real manual run on the demo box, which is the only way
this was ever going to surface.
**The observation.** A 430MB immich push reported `0%` for 40+ seconds and then completed. The
parser was not broken — restic was genuinely reporting no transferred bytes. On an incremental run
where nothing changed, restic transfers nothing: `bytes_done` is `omitempty` on restic's side, so it
is not even present in the JSON, and `percent_done` stays 0 for the whole run. Confirmed against the
real schema by capturing `backup --dry-run --json` output from restic 0.14.0 in the controller image
(the version comment in `offbox_progress.go` now quotes those captured lines verbatim).
**Why it mattered.** A byte-only progress bar is indistinguishable from a hang in the COMMON case —
the incremental run — which is precisely the silence 4c set out to remove. Shipping it would have
replaced "no feedback" with "feedback that says 0% and looks stuck".
- `files_done` / `total_files` are now parsed and published alongside the byte counters. They move on
an incremental run even when bytes do not.
- The card prefers bytes when bytes are moving; otherwise it drives the bar from files and says
„N / M fájl ellenőrizve"; only before restic knows a total does it say „a mentendő adatok
felmérése…".
- `parseResticStatus` now returns a struct rather than four positional values, and a new test pins
the real incremental-run line shape (bytes absent, files climbing) so a future refactor cannot
quietly drop the file counters and restore the stuck bar.
### v0.147.0 — feedback slice 1: pressing a button says something (2026-07-19)
Green: `go build ./... && go vet ./... && go test ./...` all pass (23 packages);
`template_id_gate` + `emoji_gate` + `native_confirm_gate` + `offbox_rename_gate` +
`app_row_dedup_gate` + `mojibake_gate` all PASS. (`docker_run_volume_path_gate` fails on
`internal/appexport/estimate.go:179` — **pre-existing on HEAD, untouched by this release**; verified
by stashing this work and re-running.)
**The systemic complaint, twice in one evening: you press a button and nothing happens.** No
progress, no ETA, no named result. This slice fixes the three worst offenders using the two patterns
already in the codebase (the deploy 3-step panel and the storage-init status poll). It deliberately
does **not** introduce a feedback framework — that is a ROADMAP item ("unified async-job feedback"),
because three targeted cards are worth shipping tonight and a framework is not.
- **4a — a verification restore now names its result.** The completion flash said the app had been
restored „ellenőrző mappába a meghajtón" — *which* folder, on *which* drive, was invisible, so the
customer could not go and look at the thing they had just asked for. It now carries the **full
path**. The restore page gained a **„Meglévő ellenőrző másolatok"** listing (app · size · date ·
path) — until now nothing anywhere showed what these restores had accumulated, so they piled up
and the only way to find them was SSH — each with a double-confirmed **„Másolat törlése"**.
- That delete is the **only** delete this release adds, so it names a STACK, never a path: the
Manager resolves the name inside a `backups/offsite-restore` root it computed itself and refuses
anything landing outside. Red-proofed — neutralise the name guard and `stack: ""` resolves to the
offsite-restore ROOT and takes every copy with it. Every refusal is asserted as a **non-effect**
(the neighbouring copy and the live data are still on disk afterwards).
- `backups/offsite-restore` was open-coded in three places; it now has one home
(`offsiteRestoreRootFor`), and a test pins the path in the flash to the path in the listing so
the customer can never be told about a directory the page cannot show or remove.
- **4b — Megosztás enable shows what it is waiting for.** Enabling sharing ran `ReconcileSamba()`
**synchronously inside the POST handler**. On a box whose golden had not baked `felhom-samba` that
is `compose up -d` pulling ~100MB from a private registry: minutes of an apparently-hung form post,
then „Beállítás mentve." whether or not anything had come up. Now detached + polled, with a card
that distinguishes **„képfájl letöltése"** (image genuinely absent — the multi-minute case) from
**„indítás"** (already baked — seconds). The distinction is decided *before* the work starts,
because afterwards the image is always present and the card could never truthfully say „letöltés".
- Success is **probed, not inferred**: `compose up -d` exits 0 on a crash-loop, so the terminal
state is container liveness. `nil` from reconcile also covers "deliberately deployed nothing
because there is no household password yet", which now gets its own message instead of a card
spinning forever.
- The **password** form starts the same job — with `UserSet` false reconcile deploys nothing, so on
a fresh box *that*, not the enable toggle, is where the pull actually happens.
- **4c — „Távoli mentés most" streams real progress.** restic was already reporting bytes and
percentages; the runner seam used `CombinedOutput()` and threw them away. The manual run now passes
`--json`, scans stdout line-by-line, and the page shows **total bytes, percent and the app
currently being pushed**. Before the scan finishes it says „a mentendő adatok felmérése…" rather
than pinning a bar at 0%, which reads as stuck.
- **Manual only.** The nightly run stays silent and its output format is untouched — pinned by a
test that fails if the scheduled path ever passes `--json` or publishes progress.
- The poll now **arms unconditionally**. It used to start only if the page already rendered „Fut…",
which loses a race the manual trigger always runs: the POST redirects and the page renders before
the detached goroutine writes `LastStatus=running`, so the poll never armed and the customer
watched a static page during the very run they had just started.
- Red-proofed twice, both confirmed: break the parser → the percent assertion fails; drop the
wiring → the `--json` assertion fails. The `--json` stream is tail-bounded (40 lines) so a large
backup does not buffer megabytes of status spam for error diagnosis.
- **Golden/controller infra-image drift closed at the source (supports the agent-side change).**
`infra.Images()` derives the list from the existing pins, `--print-infra-images` prints it, and the
golden bake now asks the controller binary it is about to bake instead of carrying its own copy.
The copy had already drifted: `felhom-samba` was never added to it, so the golden baked 3 of 4 —
which is *why* enabling Megosztás pulled at runtime. A test parses the const block out of the
source and fails if a pin is added without reaching `Images()`.
**Live-validated** on demo guest 9201 through the real UI. **No floor change** — Viktor decides floor
timing.
### v0.146.0 — nav polish: styled scrollbars + collapsible sidebar groups (2026-07-18)
UI-only; no behavioural or backup/restore surface touched. Green:
`go build ./... && go vet ./... && go test ./...` all pass; `template_id_gate` + `emoji_gate` +
`native_confirm_gate` + `offbox_rename_gate` + `mojibake_gate` + `app_row_dedup_gate` all PASS.
- **Scrollbars (`style.css`).** The platform default is a light, chunky bar that reads as a bright
stripe against the navy and competes with the content it is scrolling. Now thin and hairline
coloured: `scrollbar-width: thin` + `scrollbar-color` for Firefox, `::-webkit-scrollbar` (8px,
thumb `--line`, hover `--text-3`, `--radius`) for WebKit/Blink — **both** declared, because
neither alone covers the browsers customers actually use. The two surfaces that really scroll take
their own panel background as the track (`.sidebar` → `--bg-2`, `html` → `--bg-0`) so the gutter
never shows through as a lighter channel. Tokens only, no raw hexes.
- **Collapsible nav groups (`layout.html` + `style.css`, vanilla JS — no framework).** Tárhely,
Biztonsági mentés and Megosztás are now accordions with a chevron indicator; **exactly one is open
at a time**, and clicking the open one closes it. Groups without sub-items (Vezérlőpult,
Alkalmazások, Rendszermonitor, Debug) are untouched plain links. Hungarian labels unchanged.
- **The header is a real `<button>`**, so keyboard and assistive-tech reachability come for free
rather than being simulated with `tabindex`/`role` on a div. It carries `aria-expanded` +
`aria-controls`, and a `:focus-visible` outline.
- **Nothing became unreachable when the header stopped being a link:** every group's own landing
page is *also* its first sub-item (`/storage` → Meghajtók, `/backups` → Áttekintés, `/sharing` →
Hálózati megosztás). This was checked before the conversion, not assumed.
- **Progressive enhancement:** the group containing the active page is rendered open
**server-side** (`.is-open`), so the correct group is open before any JS executes and stays open
if JS never runs. The listener only handles clicks.
- **No layout jump:** the collapse animates `grid-template-rows: 0fr → 1fr` (with `min-height: 0`
+ `overflow: hidden` on the sub-list) rather than `max-height`. That animates to the content's
REAL height, so there is no magic number to drift when a group gains or loses an item — the
specific way a `max-height` accordion rots. The toggle reserves its 3px active border as
`transparent` so becoming active adds no width shift. Transitions are .18s and both the collapse
and the chevron rotation are disabled under `prefers-reduced-motion: reduce`.
**Note (unchanged, pre-existing):** `docker_run_volume_path_gate.py` still fails on
`internal/appexport/estimate.go:179`. That is ROADMAP **R-29**, it is unrelated to this change, and
it was verified to fail identically on the untouched tree — deliberately not bundled here, per
R-29's own "do not bundle (a) into an unrelated feature commit".
### v0.145.0 — R-7b: share data enters the live backup runs (Model B′) + samba liveness (2026-07-18)
**The „Felhőmentés" toggle on the Megosztás page is now true.** Before this release a customer could
switch a share to „Felhőmentés: bekapcsolva" and the page would render exactly that while the files
dropped on it were in **no backup at all** — `backup.RunTier2` short-circuits on `os.Stat(unitDir)`
before the classification seam, and the offsite runner enumerates `GetOffboxApps()`. A share-only
infra stack has neither a recovery unit nor an offbox toggle, so it fell through both engines. R-7b
closes that with a **sibling shares source** in each tier.
**Model B′ (Viktor's ruling, 2026-07-18) and its invariant.** Share data enters the runs through
NEW, ADDITIVE job/leg code that reuses the proven primitives — the tier-2 mirror seam, the restic
wrappers, the soft-quota/enlargement gate, the status recorders — while leaving **every per-app engine
path byte-identical**. Not Model A (a synthetic recovery unit breaks on multi-drive shares and wraps
1 KB of JSON in dump machinery) and not engine-loop surgery. The invariant is enforced by test, in
both tiers, with red-proofs.
- **Payload (`internal/backup/shares_payload.go`, new):** a staging dir holding
`_shares-manifest.json` (the share definitions, sorted → **byte-deterministic** for an unchanged
registry, so a no-op run gives the mirror nothing to rewrite) plus a **best-effort** `passdb.tar`
captured from the samba container. A restore therefore returns the files, the share configuration
AND the SMB password hash — not just bytes on a disk. The credential copy is SECRET-BEARING: 0600,
never logged at INFO, never in a report or a committed file. A down container degrades to
manifest-only and KEEPS any previously captured copy (a stale credential beats none for DR).
- **Tier-2 shares job (`internal/backup/tier2_shares.go`, new):** runs after the per-stack loop in the
same orchestrator run. Shares are grouped **by source drive** — a household's shares can span disks
and each group needs its own cross-drive target — into
`backups/secondary/_shares/<sourceDriveKey>/<share>` with the payload at `_payload/` and the layout
marker written **LAST**. Reuses `selectTier2TargetFrom` (a narrow source-drive seam extracted from
`selectTier2Target`; the headroom math is untouched), `tier2ReconcileRoots` (a pure extraction),
`tier2SafeRemove` and the `recordTier2*` helpers.
- **Offsite shares leg (`internal/backup/offbox_shares.go`, new):** ONE additional
`restic backup --tag felhom-offbox --tag _shares` carrying the manifest staging dir plus every
MANDATORY share folder, hooked in AFTER the per-app loop and BEFORE retention — so
`forget --group-by host,tags` covers the `_shares` group with **no flag change**. Same enlargement
arithmetic as the per-app gate. **Degradation contract:** a quota-blocked push falls back to the
MANIFEST ONLY, never to nothing — definitions protection must not regress because the files stopped
fitting.
- **Restore „Megosztások" (`internal/backup/shares_restore.go`, new):** siblings of the per-app
scratch/place pair. Files are merged **missing-only** (never overwriting) and every destination is
**prefix-asserted** against registered LIVE storage roots — a snapshot is untrusted layout input, so
a path that no longer sits under a live root is refused rather than created. Definitions merge with
**existing-wins** (a restore must never silently flip a live share's settings; skipped ones are
named in the flash). Then `ReconcileSamba` re-renders smb.conf, and the credential goes back into
the named volume best-effort. Routes `POST /backup/shares/{restore,place}`.
- **Samba liveness (the fold-in):** `monitor.EffectiveProtected` gains a settings-backed dynamic
extra, so a dead sharing service raises the same protected-container issue → alert → Hungarian
degradation e-mail as a dead traefik — but only while sharing is ON. It watches the **container**
name (`infra.SambaContainerName`), which is deliberately NOT the stack name. **Finding: the
alert/e-mail pipeline needed no further change** and no new event type is introduced, so the
`allowedEventTypes` gotcha does not apply.
- **UI truth-up:** the Megosztás page states per-tier status (2. mentés / távoli mentés, amber only on
deviation) and links to the restore page. The reserved `_shares` key is mapped to „Megosztások" at
the notification and Hungarian-prose boundaries ONLY — the persisted `EnlargedBlocked` set, the
restic tag and the dest path keep the raw key, because templates index by it.
- **RESERVED-NAME FINDING (the task's assumption was false):** `settings.nbNameRe` begins with
`[A-Za-z0-9_]`, so „_shares" **was an accepted share name** — the underscore namespace was not in
fact reserved. `ValidateSMBShareName` now refuses a leading underscore (on ADD only, so existing
shares are never retroactively invalidated), and `RunAllTier2`/`RunOffboxBackup` additionally skip a
`_shares` STACK loudly as defense in depth.
- **`infra.SambaContainerName` / `SambaPassdbVolume` / `SambaPassdbMount`** become the single source of
truth for the samba container identity — the compose renderer interpolates them, and stacks, backup
and monitor all read them instead of repeating string literals.
- **Bug found by test:** `shareSourceDrive` returned a slash-normalised path, which made the target
selector's source-drive equality check miss — a share group could have targeted its own source
drive (a same-disk copy pretending to be tier 2). Fixed; POSIX-only in effect, but real.
- **Tests:** 24 new/extended cases in `internal/backup` + 2 in `internal/monitor`. **Six red-proofs
run and reverted, all fired:** (1) shares leg appending into the app's argv → B′ isolation FAILS;
(2) mandatory→offsite mapping inverted → Scenarios A+B FAIL; (3) manifest-only degradation dropped →
Scenario C FAILS; (4) prefix-assert removed → place-guard traversal FAILS; (5) dynamic samba extra
removed → Scenario E enabled-case FAILS; (6) shares destBase dropping the reserved segment → tier-2
isolation FAILS.
### v0.144.0 — „Megosztás": LAN SMB file sharing (R-7 slice 1) (2026-07-18)
The customer turns on network sharing, sets ONE household SMB password, and exports folders. The box
appears in Windows Explorer's Network view as `\\FELHOM`; opening a share and writing to it works, and
every SMB write lands as uid:gid 1000 so apps and both backup tiers see consistent ownership. SMB is
an **embedded controller feature**, not a catalog app — it needs host networking (the R-6 spike
verdict), its config is a generated share list, and its roots ride the backup classification.
- **New infra image `felhom-samba:1.0.0`** (`controller/infra-images/samba/`, built by
`controller/scripts/build-samba-image.sh`): pinned alpine 3.21 (`sha256:48b0309c…`) + smbd + **nmbd**
+ wsdd + tini. Dumb by design — `smb.conf` is bind-mounted READ-ONLY, nothing is templated inside,
no name/password is baked, passdb lives on a named volume. nmbd is REQUIRED alongside wsdd: the R-6
spike proved wsdd-only leaves the box *visible* but the Explorer double-click fails `0x80070035`
(no flat-name resolution). Anonymous pull verified from the guest.
- **Settings (`internal/settings/smb.go`, new):** `SMBSettings{Enabled, ServerName, UserSet}` +
`SMBShare{Name, Path, ReadOnly, Offsite, CreatedAt}` registry with NetBIOS-safe validation
(≤15 chars, no slash/dot) and case-insensitive name-collision refusal. **The household SMB password
is NEVER persisted** — only the `UserSet` boolean.
- **Renderers (`internal/infra/samba.go`, new):** pure `RenderSambaConfig` (hardened global block:
`server min protocol = SMB2`, `bind interfaces only = yes`, `interfaces = lo eth0`,
`disable netbios = no`, `map to guest = never`, per-share force-user block) and
`RenderSambaCompose` (`network_mode: host`, pinned image, config `:ro`, passdb volume, one bind per
share — `:ro` for read-only shares as defence in depth beside smb.conf). Exact smb.conf golden test.
- **Lifecycle (`internal/stacks/samba.go`, new):** `ensureSamba` joins `EnsureBaseStack` after
filebrowser, gated on `SMB.Enabled` (the cloudflared conditional-deploy precedent); `ReconcileSamba`
runs after every mutation. Idempotent — unchanged config + running container performs **zero**
compose calls. Config writes are atomic (tmp+fsync+rename). The password is applied via
`smbpasswd` on **STDIN** (never argv, never logged). Disable = `compose down`; the passdb volume and
every shared folder are KEPT. A share on a disconnected/decommissioned drive is rendered ABSENT from
smb.conf (never export a dead mountpoint) while its config is retained.
- **Protection:** `samba` is protected in CODE (`config.alwaysProtectedStacks`) because
`cfg.Stacks.Protected` comes from the golden-generated controller.yaml and predates it. This also
makes the app-backup loops correctly skip it (it is infrastructure, not a customer app).
- **UI (`internal/web/sharing_handlers.go` + `templates/sharing.html`, new):** a new top-nav category
**„Megosztás"** → **„Hálózati megosztás"**. Enable/server-name card, household password, shares table
(Név · Mappa · Írásvédett · Felhőmentés · Törlés — "a mappa és a fájlok megmaradnak"), and a create
flow (new folder under `<storage>/shares/` or an existing folder via the browse modal).
- **Picker security:** every customer-supplied path goes through `sharingResolvePath` — absolute →
`EvalSymlinks` → containment in a registered LIVE storage root → deny-listed system subtree →
is-a-directory. Refusals are **uniform** so the picker can never act as a filesystem oracle. The
deny-list is DERIVED from `stacks.ProtectedHDDPaths` (provably a subset, so it can only shrink,
never drift); the drive root is an exact-match denial so user-data folders under it stay shareable.
`sharingResolveStorageRoot` is a separate, strictly tighter check for the new-folder parent.
- **Backup classification [R4] (`internal/stacks/samba_classify.go`, new):** `ClassifiedBinds("samba")`
resolves from the shares registry instead of catalog metadata. Felhőmentés ON → `mandatory`
(offsite + tier-2); OFF → `optional` (tier-2 only); smb.conf/passdb never classified. Verified
through the real `ComputeCaptureSet` tier filter including the negative. **Zero backup-engine edits.**
- **KNOWN GAP (reported design fork, not improvised):** making that seam correct does NOT by itself put
share data into a live tier-2/offsite RUN. `backup.RunTier2` short-circuits on `os.Stat(unitDir)`
before it ever calls `GetStackClassifiedBinds`, and the offsite runner enumerates
`settings.GetOffboxApps()` — both are recovery-unit shaped, which a share-only infra stack has not.
Teaching them about one is more than an enumeration tweak, so per the task's STOP clause it is
reported rather than improvised. See `REPORT.md`.
- Live-validated end-to-end on demo guest 9201 through the REAL endpoints (curl against the exact
routes the UI posts to; the UI is password-gated so no browser leg): enable → password → create both
share kinds → guard refusals (appdata/backups//etc/drive-root all uniform 400) → smb.conf + container
+ `:ro` bind verified on the box → Windows 11 workstation: `Test-NetConnection 445` True, nbtstat
`FELHOM <00>/<03>/<20> Registered`, `ping FELHOM` resolves, SMB write/read byte-compare PASS, and a
**write to the read-only share refused with no effect**. SMB-written files land as `1000:1000`.
**Explorer leg PASSED (Viktor, 2026-07-18):** both shares open from the Network view; an
interactive Explorer save landed owned 1000:1000 and a write into the read-only share was refused,
folder left empty. Slice 1 is fully PROVEN-LIVE.
### Build infra — build root relocated (2026-07-18)
- `controller/build.sh`: `REPO_DIR` + `WEBSITE_ASSETS_DIR` repointed `/home/kisfenyo/…` →
`/mnt/5_hdd/felhom.eu/…`. All felhom working dirs on the DooPlex build server (180) were hard-moved
off the (filling) SSD to `/mnt/5_hdd/felhom.eu/`. No controller code/behavior change; no version bump.
### v0.143.0 — guest RAM resize UI (R-24) — MinAgent: 0.90.0 (2026-07-17)
The customer sees the guest's current memory + the allowed range on the **Rendszer** settings page and
resizes it. The controller only proxies + maps the agent's machine `code` to Hungarian — the AGENT
(felhom-agent v0.90.0) enforces every bound and applies the change live (no reboot). Memory only.
- **agentapi (`internal/agentapi/client.go`):** `GuestMemory(ctx)` (GET /guest/memory) and
`ResizeMemory(ctx, mb)` (POST /guest/memory). A ruled 412 refusal surfaces `*MemoryRefusedError`
carrying the machine code (below_min/above_max/below_usage_floor) + fresh bounds; a pre-0.90 agent
404s → the typed `*StatusError{404}` (the capability signal).
- **Capability (`internal/agentapi/features.go`):** `FeatureGuestMemoryResize` + `featureMinAgent`
**0.90.0** + a `featureProbes` row (GET /guest/memory is the probe; the probe type-asserts the one
method it needs, so the shared `SupportProber`/`netAgent` stay untouched). Per the publish-train
convention (this CHANGELOG declares MinAgent; the Supports gate sits at the handler entry point).
- **UI (`internal/web/system_memory_handlers.go`, new; `templates/settings_system.html`):** a "Szerver
memória (RAM)" card shows current/used memory + the `2048 MB – {max} MB` range; a number input
(step 256) + "Átméretezés". `POST /api/system/memory/resize` → capability gate (SupportUnknown passes)
→ agent → POST-response flash. A JS confirm fires ONLY on a shrink. Code→Hungarian map: success
"A memória átméretezése megtörtént: X MB → Y MB."; below_usage_floor "…túl közel van a jelenlegi
felhasználáshoz (N MB). Állíts le néhány alkalmazást…"; below_min / above_max; agent-outdated → the
control is not offered + a "rendszerfrissítése szükséges" note; agent-unreachable → the value falls
back to the guest's own `/proc/meminfo`, control disabled, honest note (the page never 500s). The
agent's English message is never shown raw.
- **Ripple (no code):** lxcfs updates the guest `/proc/meminfo` live, so the deploy-page memory math
follows a resize automatically.
- **Tests:** agentapi (GuestMemory decode, 404-typed, ResizeMemory success + refusal-code, capability
table 0.89→No / 0.90→Yes / probe 404→No / probe ok→Yes) + web handler (success maps to Hungarian;
below_usage_floor maps + the agent English never leaks; agent_outdated gate refuses with ResizeMemory
never called). Gates: template_id + emoji OK; `go build/vet/test` all pass.
### v0.142.0 — offsite repo continuity: orphaned-repo guard (A) + run-status auto-refresh (C) (2026-07-17)
Closes the reinstall-orphaned-repo incident class (`memory` offbox-repo-orphaned-2026-07-17): a
recreated data volume minted a new repo passphrase; the offsite repo, keyed under the old one, became
unreadable and surfaced only as a raw nightly `wrong password or no key found`. Green:
`go build ./... && go vet ./... && go test ./...` + template/emoji/native-confirm gates. Pairs with hub
v0.60.0 (Part B escrow retention).
- **Part A — orphaned-repo guard.** `ensureOffboxRepo` now CLASSIFIES the `restic cat config` failure
(`classifyResticProbe`, the exact 07-17 stderr): `wrong password or no key found` → **ORPHANED**;
no-repo → init; other (network/SFTP-auth) → existing error handling. An orphaned repo persists
`OffboxTarget.RepoState="orphaned"` and, ONLY on the transition, pushes `offbox_repo_orphaned`
(never nightly-spam — scheduled runs then SKIP). The remote page shows a calm Hungarian card
(exception color) explaining the remote holds backups under a previous, no-longer-available key —
NOT the raw restic banner. **Reset (move-aside, never delete):** an UNCLAIMED box auto-resets on
detection (Scenario B); a CLAIMED box gets an explicit reveal-then-confirm reset (Scenario C) →
`mv <repo> <repo>.orphaned-<date>` (collision-suffixed) + `restic init` + `offbox_repo_reset` event.
`internal/backup/offbox.go` (+ `ErrOffboxOrphaned`, `ResetOrphanedRepo`, an ssh-exec seam),
`web/offbox_handlers.go` (`/backup/offbox/reset` + the orphaned-run refuse), `backups_remote.html`.
Red-proof `TestOffbox_OrphanDetection_Claimed` (pre-fix = the incident: raw error, no state → FAIL)
+ `TestOffbox_OrphanDetection_UnclaimedAutoReset` + `TestOffbox_ConfirmedReset`.
- **Part C — run-status auto-refresh.** New `GET /backup/offbox/status` (JSON) + a poll on
`backups_remote.html`: while a run shows "Fut…" the page polls and flips to Rendben/Hiba + fresh
numbers without a manual reload; polling stops at the terminal state. Test `TestOffboxStatusHandler`.
### v0.141.0 — N100 polish: initialize-to-usable (F6) + Vissza back-routes (F7) (2026-07-17)
Closes two `VALIDATION-n100-baremetal-2026-07-16.md` findings. Green:
`go build ./... && go vet ./... && go test ./...` all pass; template/emoji/native-confirm gates pass.
- **F6 (MEDIUM) — drive "initialize" now ends in a USABLE (mounted+registered) drive, disconnect-safe.**
Root cause (fork verdict, source-grounded): the format→mount→register orchestration
(`internal/web/storage_handlers.go` `runStorageInit`) ran on the REQUEST context; a closed tab /
lost connection cancelled it after `FormatDisk` (the agent's mkfs continues detached, returns
`errFormatClientGone`), so the mount+register leg was aborted — device formatted but
unmounted/unregistered (the N100-observed state). The chain must reach `SyncFileBrowserMounts`
(controller-only), so it stays controller-side — **no agent change**. Fix: `POST /api/storage/init`
starts a DETACHED single-flight job (`internal/web/storage_init_job.go`, the `netAddState` shape) on
`context.Background()`; `runStorageInit` gains a nil-safe phase callback (formatting → mounting →
registering). The wizard polls the new `GET /api/storage/init/status` and renders the 3-step
progress (`storage_init.html`); the confirm/refuse verdicts surface through the same poll. Register
is the LAST step (marker-last, Scenario B) and every prior step is idempotent (`AddStoragePath`
dedups) → a crash leaves at most an unregistered orphan, never a broken/duplicate registration.
Red-proof `TestStorageInit_DetachedSurvivesClientDisconnect` (pre-fix: cancelled-ctx chain fails at
mount, NOT registered → FAIL; fixed: detached job registers exactly once). **Deeper half (found on
the live leg — a 64 GB USB):** a slow mkfs outruns the agentapi client's 15 s `Timeout`; the agent
runs it DETACHED and records the job, so `runStorageInit` now POLLS the agent's
`GET /disks/format/status` (new `agentapi.Client.FormatStatus`) to the terminal outcome on a client
timeout, then continues to mount+register (the F6 root-cause's "mkfs continues detached; poll the
status"). Test `TestStorageInit_PollsAgentFormatStatusOnTimeout` (timeout→done registers;
timeout→failed surfaces the error, no register).
- **F7 (LOW) — the "Vissza" (Back) anchor on `/storage/init` and `/storage/attach` now routes to
`/storage`** (was `/settings`). The init success link also points to `/storage` (where the new
drive appears). Test `TestStorageWizardBackAnchors_PointToStorage`.
### v0.140.0 — Direction-2 immediate-sync: hub→box wait channel client (2026-07-16)
The other half of the immediacy arc (Direction 1 = v0.139.0 box→hub trigger). An operator action on
the hub now reaches the box in **seconds** instead of on the next ~15-min cycle. Pairs with hub
v0.58.0 (the `GET /api/v1/wait` endpoint + the in-memory operator-intent notifier). Grounding:
`felhom.eu/documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md` (option b).
- **`internal/report/waiter.go` (new) `report.Waiter`:** holds a hanging authenticated GET against
the hub's `/api/v1/wait?gen=N` (reusing the SAME hub URL + key as the pusher — no new config
keys). Its own `http.Client` has **no overall Timeout** (a held GET must stay open for the hub's
~240 s hold) with sane connect/TLS/`ResponseHeaderTimeout` deadlines; a per-request context bounds
a black-holed connection. On a completion whose generation **differs** from the last seen, it fires
the v0.139.0 `report.Trigger` — and NOTHING else; the fired report's ACK delivers config/escrow/
claim/floor through the UNCHANGED machinery (this adds zero delivery logic). Behaviors:
- **First observation records, never fires** (the startup report already covered current state) —
prevents a spurious echo report on every process start / config-refresh restart. Red-proof:
disable the baseline branch → `TestWaiter_FirstObservationRecordsNoFire` fires 1 (run-fail-reverted).
- **Same-generation timeout fires nothing** (the hub's hold elapsed) — not interval-shortening.
- Heartbeat newlines tolerated; the body is read only for its `{"gen":N}` line (contentless wake).
- Any error — transport, a **404 from a hub that predates the endpoint**, or a malformed body —
backs off exponentially (5 s → 5 min, reset on success) with ONE WARN per state change, and the
15-min cycle keeps reconciling. Exits promptly on context cancel (even mid-hold).
- **`cmd/controller/main.go`:** the Waiter is constructed + started right beside the Direction-1
trigger, gated on the SAME `hubPusher != nil && cfg.Hub.Enabled` condition (strict no-op when hub
reporting is off). One INFO line on start.
- **Copy soften (Viktor-approved):** `backups_remote.html` + `backups_escrow.html` — "ez általában
néhány **másodperc**, legfeljebb 15 perc" (was "néhány perc"). The 15-min bound stays — it is the
honest worst case when both the wait and the immediate push fail. Escrow grace window unchanged.
- **Coupling (soft):** immediacy needs hub ≥ v0.58.0; against an older hub the wait 404s and the box
degrades cleanly to the 15-min cycle. No agent coupling, no `MinAgent`.
### v0.139.0 — immediate out-of-cycle hub report on user actions (Direction 1) (2026-07-16)
Viktor's ruling: a user action with hub-side effects must round-trip in seconds, not minutes. One
generic, debounced out-of-cycle report trigger now sits on the proven outbound push channel; the
15-min `hub-report` cycle is untouched and stays the reconciliation backbone. Headline UX win: the
v0.138.0 escrow "megerősítésre vár" card collapses from ~14 min to seconds (the blob is already
uploaded at claim time — the immediate report's ACK hash-match flips `pending→escrowed` via the
unchanged `EscrowAutoConfirmer`). No hub change; no UI copy change ("legfeljebb 15 perc" stays the
honest worst case for a failed immediate push).
- **`internal/report/trigger.go` (new) `report.Trigger`:** buffered-1 signal channel + single
worker (shape: hub `wgsync/reconciler.go`). Non-blocking `Fire()`; worker = quiet window 2 s
(burst coalescing) → drain → min spacing 15 s → ONE full-report fire. Coalesce-and-eventually-
fire (trailing edge): a burst yields ≤ 1 + ceil(burst/15 s) pushes and the last state always
reaches the hub — deliberately NOT the `internal/sync` refuse-debounce (a refused fire would
lose the update until the next cycle). No retries of its own (the Pusher owns 3×5 s); a fire
error logs one WARN and degrades to the cycle. Exits on context cancel.
- **`cmd/controller/main.go`:** ONE canonical fire closure (`BuildReport` + `Claimed` +
`hubPusher.Push`), constructed only when `hubPusher != nil && cfg.Hub.Enabled`; replaces the
raw per-fire goroutine behind `apiRouter.SetReportPushTrigger` (the v0.70.0 geo seam — kept,
now debounced) and feeds the new `webServer.SetReportTrigger`.
- **`internal/api/router.go`:** deploy / remove / delete endpoints now call the existing
`reportPushNow()` after success (geo save/sync already did) — all backed by the trigger.
- **`internal/web` seam + call sites (`server.go` `SetReportTrigger`/`reportTriggerNow`,
nil-safe, fired only AFTER a successful local commit):** escrow recovery-code claim
(`escrow_handlers.go`), notification-prefs save + app-email toggle (`handlers.go`), offsite
target config + per-app offsite toggle (`offbox_handlers.go`), customer claim completion
(`claim.go`). `hub.enabled: false` → seams stay nil → strict no-op.
- Tests: `internal/report/trigger_test.go` (single-fire exactly-once, Fire() non-blocking,
burst-coalescing ceiling + trailing edge with red-proof, fire-error isolation with red-proof,
prompt cancel exit), `web/report_trigger_seam_test.go` (fires-after-commit-only through the
real offbox toggle handler + nil-seam no-op), `api/report_trigger_nilsafe_test.go`.
### v0.138.0 — escrow "awaiting hub confirmation" waiting state (2026-07-16)
Closes the customer-zero (N100) UX gap: after a completed escrow ceremony the Távoli mentés page kept
showing the yellow **"Helyreállítási kód szükséges"** card for ~15 minutes, until the next hub-report
ACK flipped `pending→escrowed`. **Phase-0 diagnosis (read-only) = verdict A (report-cycle lag)**: on the
demo box the ceremony completed `16:13:39` and the very next `hub-report` ACK at `16:27:58` auto-confirmed
via hash-match (`d517ce7f…`), `escrow_state:"escrowed"` — nothing was broken; the wait simply had no UI
feedback. (Hub stale-clear-on-upload — Hypothesis B — was verified to already exist: `SaveHostEscrow`'s
`ON CONFLICT` sets `stale_at = NULL`, so **no hub change was needed or made**.)
- **`settings.OffboxTarget.CeremonyCompletedAt`** (new, `ceremony_completed_at`, RFC3339) — stamped on a
successful recovery-code **claim** (`web/escrow_handlers.go`, only while still pending; best-effort, a
stamp failure never fails the claim) and **zeroed** on the `pending→escrowed` flip (the auto-confirmer
`Flip` closure in `cmd/controller/main.go` + the deprecated manual confirm in `web/offbox_handlers.go`).
Persisted → survives a controller restart mid-wait.
- **`web/handlers.go` `offboxCeremonyWaitState` + `escrowCeremonyGraceWindow` (35m):** classifies the
wait — *awaiting* (stamped, within the window) vs *timed out* (stamped, past two report cycles + slack).
Both fall back to the plain pending CTA when escrowed, unstamped, or the stamp is unparseable.
- **`backups_remote.html`:** one new escrow-card branch ahead of the existing chain — an **info (blue)**
"Helyreállítási kód létrehozva … megerősítésre vár, legfeljebb 15 perc" card, degrading to a **warn**
"A megerősítés nem érkezett meg …" + re-ceremony CTA past the window. The existing pending/stale
(Scenario F)/escrowed branches are untouched.
- **`backups_escrow.html`:** the wizard's final "Befejezés" step gains a **"Mi történik ezután?"** note so
the customer expects the interim card on the page they land on.
- Test: `web/escrow_wait_state_test.go` (truth table + mutual-exclusion invariant; red-proof recorded in
REPORT). No scheduler/agent/endpoint changes.
### v0.137.0 — empty-email notification save guard (data-loss fix) (2026-07-15)
Fixes a silent alert-delivery wipe demonstrated on the demo customer on 2026-07-15: saving the
Értesítések form with a **blank e-mail box while events were still enabled** dropped an empty
`Email` into the prefs AND pushed it to the hub (`SyncPreferences`), overwriting the customer's
provisioning-seeded alert address — the "Kedves Ügyfél!" delivery path went dark until it was
restored by hand in 6D (P3-DELIVERY).
- **`web/handlers.go` `settingsNotificationsHandler`:** after computing the trimmed email + enabled
events, a guard refuses the save when `email == "" && len(enabledEvents) > 0` — it returns
**before** `SetNotificationPrefs` and **before** any hub sync, re-rendering the page with a
Hungarian error ("Adj meg egy értesítési e-mail címet – …") and repainting the just-submitted
checkboxes (an overlay on `notificationsPageData`'s `NotificationPrefs`, render-only). Enabled
events with no address is a purely destructive state reachable only via the bug. The legitimate
**empty-email + ZERO events** clear-all still proceeds (the empty hub push is correct there).
- **Deliberately NOT** an HTML `required` attr on the input — `required` is unconditional and would
block the legitimate clear-all case; the server-side guard is the correct, precisely-conditional
floor. `SyncPreferences` / the hub side / the seed-migration are untouched.
- **Tests (`web/notifications_guard_test.go`):** guard-fires (stored email survives — the wipe is
prevented; red-proofed: remove the guard → the email is wiped to `""`), legitimate clear-all
proceeds, normal save persists. Real temp-file `Settings` (non-hollow: asserts stored state).
### v0.136.0 — `.fab` exclusion scoping: classes in the manual export (Task 4) (2026-07-15)
Task 4 — the `.fab` column of the matrix (architecture §2; the SQ5 exclusion-scoping verdict + Viktor
ruling #1 + R1-C). **The SQ6 over-capture is fixed for classified apps:** a manual export no longer
drags every sibling app's content along in the userdata root tar. Mechanically unchanged from
v0.130.0 — ONE exclude-scoped userdata-root tar + per-mount skip — so the **manifest stays v1**
(basename keying, the `userdata` fallback), the **import side is untouched**, and old controllers
import new bundles correctly (they just extract a root tar containing less). **Legacy (no-block) apps
export byte-identically to v0.130.0** (the SQ5 safety net).
- **`appbackup.ComputeFabBuckets` (new, pure):** the same resolution + structural guards + equal-Abs
collapse as `ComputeCaptureSet` (shared pipeline extracted, not duplicated), but bucketed by class
(mandatory / optional / excluded), guards over ALL classes (a traversal path is never plannable —
opt-in or not), and NO cross-bucket containment dedup (a mandatory child inside an excluded parent
stays independently addressable).
- **The export plan (`appexport/fabplan.go`):** `ExportRequest` gains `DeselectOptional` +
`OptInExcluded`; `computeFabPlan` resolves the selection (mandatory forced-in — a **server-side
floor** ignores a client trying to deselect a mandatory path; optional default-in; excluded
default-out) into `{SkipMounts, SkipUserdataTar, UserdataExcludeRels}`. The userdata root tar is
**skipped entirely** when no userdata bind is selected (radarr → state-only, the SQ6 headline); else
the exclude list is the **topmost dirs neither ancestor nor descendant of a selected relpath**
(R1-C — the exact `tier2Reconcile` keep-rule). An HDD mount matching no classified bind is KEPT
(fail toward capture). `tarDirectoryExcluding` skips excluded subtrees in the walk (`tarDirectory`
is now a thin wrapper).
- **Estimate split (additive):** `ExportEstimate` gains `HasClassification` + `BaseBytes` +
per-path `MandatoryItems`/`OptionalItems`/`ExcludedItems` (each with its du size) — existing fields
and the fits gate are unchanged. Both estimate pipelines surface it (shared `EstimateExport`).
- **UI (`app_export.html`, classified apps only):** locked mandatory list, pre-selected optional
checkboxes, a collapsed excluded opt-in behind the two-number warning
("Alap mentés: ~X. A kihagyott, nagy méretű tartalommal együtt: ~Y.") + the FileBrowser pointer;
totals recompute client-side per toggle. Selections POST through BOTH start endpoints (the
two-call-site).
- **Tests:** `ComputeFabBuckets` bucket/guard/containment tests; `computeFabPlan` scenarios A–F +
§8 edge cases; `tarDirectoryExcluding` FS-level; export-level bundle tests (exclude-scoped tar,
legacy full-root, all-excluded no-tar); the two-call-site bundle test across both start pipelines.
All 6 §10 red-proofs verified. **CAMPAIGN-6D Accept #1** (the ≥1 GiB `.fab` full circle) now runs
against this final capture shape.
### v0.135.0 — Tier-2 engine rework: class-driven legs, v2 layout, NAS-target exclusion (Task 3b) (2026-07-15)
Task 3b of the backup-classification-redesign arc — the tier-2 column of the matrix (architecture
§2/§8). Behavior-changing but bounded: every destructive write lands ONLY under
`backups/secondary/<stack>/` (fully-derived data), asserted in code.
- **Class-driven appdata leg (`tier2_capture.go`, new):** for a classified app the tier-2 legs are
the Task-3-core `TierSecondary` set — per-bind mandatory + optional HDD/userdata paths (paperless's
copy legitimately SHRINKS as `export` drops out). Legacy apps keep a byte-identical capture set (the
resolver appdata dir(s)) mapped into the same layout. Skipped/missing **mandatory** paths are loud
gaps (English log + the app's Hungarian cross-drive `LastWarning`), mirroring the offsite pattern.
- **v2 relpath-mirroring layout:** `backups/secondary/<stack>/` = `.felhom-tier2-layout` marker (content
"2", written **LAST**) + `recovery-unit/` + `hdd/<relpath>/` + `userdata/<relpath>/`. N>1 appdata dirs
and nested binds are represented natively — the v0.131.0 flat-appdata **N>1 refusal is gone**
(`errTier2MultiDir`/`tier2AppDataName` deleted). Restore is position-derivable.
- **Migration = delete-and-rebuild** (marker absent → remove the old flat `appdata/`; `recovery-unit/`
is layout-identical, untouched) + a **reconcile** step that prunes dest dirs a bind no longer covers
(a removed/re-classed bind stops occupying the secondary drive within one run). All `os.RemoveAll`
goes through `tier2SafeRemove`, which refuses any target not strictly under `backups/secondary/`.
- **SSD fallback is an enforced STATE-ONLY tier (§2.2):** headroom is decided on unit + mandatory; the
SSD carries unit + mandatory only, optional legs skipped with an honest Hungarian reason.
- **NAS-target exclusion (F-6C-1):** `selectTier2Target` never selects a NETWORK storage path — pinned
OR auto (metadata-only `IsNetwork()`, no fs probing). NAS-only ⇒ the honest reason
("Hálózati tároló nem lehet a 2. mentés célja…"). Prevents the rsync `-og`-under-root_squash
silently-wrong-owner restore.
- **Restore reads v2 only:** a marker gate refuses a pre-v2 copy ("A 2. mentés régi formátumú…");
the reader merges the `hdd/` and `userdata/` subtrees missing-only into live (N>1 native).
- **Part 0 — prefs seed fix:** the 3a-fix un-disableable checkbox is fixed — `offbox_enlarge_blocked`
is now a ONE-TIME persisted seed at settings Load (`OffboxEnlargeNoticeSeeded`), not a getter append,
so a customer's later opt-out **sticks**. **Part 0.5:** the offsite restore scratch now prefers a
LOCAL path over a network one (a squashed scratch would feed `PlaceOffsiteRestore` wrong-owner files).
- **Tests:** the v2 suite (`tier2_v2_test.go`: A–H + reconcile keep/remove + the safe-remove boundary
proof), Part 0 seed tests (idempotent + opt-out-sticks), Part 0.5 scratch-preference tests; obsolete
v1 flat-layout / N>1-refusal tests removed. All 10 §10 red-proofs verified (mutate → fail → revert).
### v0.134.1 — Placement hardening (F-3a-1..4) + enlarge-blocked notification delivery chain (Task 3a-fix) (2026-07-15)
Follow-up hardening of the (not-yet-live) place-to-live flow from v0.134.0, plus the controller side
of the `offbox_enlarge_blocked` notification delivery chain (paired with hub v0.55.0). No new
architecture.
- **F-3a-1a (`offbox_restore.go` `PlaceOffsiteRestore`):** the live target now resolves via the RAW
`GetStackHDDPath` (mirrors `offboxCaptureSet`), NOT `AppNamespaceRoot` whose `systemDataPath`
fallback would have merged userdata onto the SSD system namespace. Empty HDD ⇒ undeployed ⇒ refuse
(`"a(z) %s nincs telepítve — előbb állítsd helyre az alkalmazást, utána az adatokat"`).
- **F-3a-1b:** placement headroom gate — refuse before any copy if `offboxFree(liveNs) <
offboxSize(scratch)` (a missing-only merge copies at most the scratch size).
- **F-3a-4:** the src-existence check is now a `os.Stat` PRE-PASS over EVERY placement before the
first copy — an incomplete scratch (e.g. a unit-only restore) refuses with ZERO copies, making the
"no partial writes" guarantee true (was interleaved: the unit could be placed before the refusal).
- **F-3a-3 (`mapOffsiteRestorePaths`):** the escape check tightened to `!HasPrefix(p, oldNs+"/")` so
the namespace root itself (`p == oldNs`) is refused instead of mapping to a junk placement.
- **F-3a-2:** on FULL success the scratch is removed best-effort (logged); `OffboxFullScratchReady`
then turns false so the place button disappears. A FAILED placement keeps the scratch for a retry.
- **Delivery chain (controller side):** `settings.DefaultEnabledEvents` gains
`offbox_enlarge_blocked`; `GetNotificationPrefs` append-if-absent migration surfaces it enabled for
existing customers (idempotent — a customer couldn't have disabled a type that didn't exist), which
the startup prefs sync (`main.go:782`) carries to the hub; the settings-page checkbox
(`"Távoli mentés — tárhelykeret-figyelmeztetés"`) + the handler's single-event slice both gain it
(missing either half re-opens the checkbox-drop trap). Hub v0.55.0 allowlists the event; NO
`customerMessages` entry (the raw dynamic two-number message must survive — `templates.go:129`).
- **Tests:** +5 placement (`offbox_place_test.go`: undeployed/headroom/incomplete-pre-pass/lifecycle
+ mapping namespace-root) + 3 settings (`notif_migration_test.go`: default-contains, migration
idempotency, no-duplicate). All 6 controller §10 red-proofs verified (mutate → fail → revert).
### v0.134.0 — Offsite tier policy engine: mandatory userdata, raw-data quota, restore rework (Task 3a) (2026-07-14)
Task 3a of the backup-classification-redesign arc — the FIRST behavior-changing task
(`felhom.eu/documentation/architecture/07-backup-architecture.md` §2/§6/§7/§9; restic mechanisms
proven in `SPIKE-restic-snapshot-shape-2026-07-14.md`). Offsite pushes now carry each app's
**mandatory** userdata, quota is measured as real Storage Box fill, retention survives the shape
change, and restore is reworked off the rootfs. **BEHAVIOR CHANGE.**
- **Multi-path snapshot (§6):** each toggled app's offsite push is now ONE restic snapshot =
recovery unit + the app's TierOffsite mandatory capture set (Task 3-core `ComputeCaptureSet`).
Optional/excluded never ship offsite. Legacy (no block) / undeployed apps stay **unit-only**,
byte-identical to v0.133.0 (the SQ5 cost guard). New `offbox_capture.go`.
- **Loud capture gaps (SP-3.4):** restic 0.14.0 does NOT error on a missing source path (exit 0,
silent partial snapshot), so a structurally-refused or on-disk-missing MANDATORY path is detected
BEFORE invocation (guard `Skipped` list + `os.Stat` filter) and surfaced in the English log **and**
the Hungarian `LastWarning`. A restic exit code never proves a path was captured.
- **Quota = raw-data (§9, SP-1):** `offboxRecordStats` now runs `stats --mode raw-data --json`
(actual deduplicated+compressed repo bytes) instead of the modeless restore-size that multiplied
by the retained-snapshot count. **The displayed remote-backup size drops once after deploy** — it
now reflects the customer's true Storage Box fill.
- **Pre-push enlargement gate (§9, ruling #1):** before an app's enlarged push, if last-known
raw-data repo bytes + the mandatory-set `du` estimate would cross the soft quota, the ENLARGEMENT
is blocked (config+DB unit-only push still proceeds — never a protection regression), the app is
recorded in `OffboxTarget.EnlargedBlocked`, `LastWarning` names it, and an **edge-triggered**
notification fires once per new block (`offbox_enlarge_blocked` event, warning severity). A
per-app "config+DB only" note renders on /backups/remote.
- **Retention grouping (§6, SP-2):** both `forget` call sites gain `--group-by host,tags` so an
app's old unit-only-shape snapshots share a group with its enlarged shape and age out naturally
(the default host,paths grouping would strand old-shape snapshots in a permanently-retained group).
- **Restore rework (§7, F-A1):** new `offbox_restore.go`. Scratch moves off the ~8 GB guest rootfs
to a data drive (`<nsRoot>/backups/offsite-restore/<app>`) behind a headroom gate (full needs
size×1.1, unit-only a 2 GiB floor; ID-first `snapshots latest --tag` → `stats <ID>`; size-unknown
fails closed). `RestoreOffboxScratch(full)` — unit-only DEFAULT via `--include <absolute-unit-path>`
(SP-3.2), full is a size-first two-step. `PlaceOffsiteRestore` places a completed full scratch into
live locations via a missing-only merge (`rsync -a --ignore-existing`, never `--delete`), unit only
if the live unit is absent; the pure `mapOffsiteRestorePaths` refuses the whole placement on no-unit
/ escape / reserved-zone. Legacy rootfs scratch is cleaned best-effort. The old `RestoreOffbox`
(whole-snapshot to an explicit dest) is retained for existing callers.
- **UI (Hungarian):** /backups/restore offers unit-only ("Visszaállítás ellenőrzéshez"), full
two-step ("Teljes visszaállítás előkészítése" → "…indítása (~méret)"), and place-to-live
("Helyreállítás az élő adatok közé (csak a hiányzó fájlok)"); /backups/remote shows the per-app
quota-blocked note. New route `POST /backup/offbox/place`.
- **Settings:** `OffboxTarget.EnlargedBlocked []string` (replaced each OK run; preserved across a
config edit). **HUB FLAG:** the `offbox_enlarge_blocked` event needs adding to the hub's
`allowedEventTypes` + `customerMessages` for delivery — until then the in-dashboard `LastWarning`
and the /backups/remote note carry the message (see REPORT §flags).
- **Tests:** +13 in `internal/backup/offbox_3a_test.go` (Scenarios A–G + all-excluded, raw-data,
both forget sites, restore argv, size-unknown refusal, scratch cleanup, place-to-live mapping);
all 10 §10 red-proofs verified (mutation → fail → revert). No tier-2 / .fab / hub / agent changes.
### v0.133.0 — Capture-set computation (INERT; Task 3-core) (2026-07-14)
Task 3-core of the backup-classification-redesign arc
(`felhom.eu/documentation/architecture/07-backup-architecture.md` §3; matrix §2; SP-1/2/3 spike
verdicts landed in `SPIKE-restic-snapshot-shape-2026-07-14.md`). Ships the **pure capture-set
computation** in the `appbackup` leaf package — **deliberately INERT**: NO backup tier changes
behavior. 3a (offsite policy engine) and 3b (tier-2 rework) are the consumers.
- **`ComputeCaptureSet(binds, hasClassification, tier, hddPath) CaptureSet`
(`internal/appbackup/captureset.go`, new):** turns an app's `ClassifiedBinds` into a tier-filtered,
structurally-guarded, containment-deduped absolute path set. Fixed pipeline (§8): legacy
short-circuit → tier filter → structural guards → equal-Abs collapse (mandatory > optional) →
containment dedup (keep ancestor) → sort by Abs. `CaptureSet{HasClassification, Paths []CapturePath,
Skipped []SkippedPath}`; each `CapturePath` carries `{Abs, Root, RelPath, Class}`.
- **Tier columns (§2):** `TierOffsite` = mandatory only (optional never ships offsite); `TierSecondary`
= mandatory + optional; **excluded** is silently dropped at every tier (never in Paths, never in
Skipped). A **legacy** app (`hasClassification=false`) resolves NOTHING — `{HasClassification:false}`,
nil Paths/Skipped — so the engines' no-block branch stays byte-identical (the SQ5 cost-regression
guard: an unmigrated app never resolves a bind into an automatic tier).
- **Structural guards (security-shaped, load-bearing):** the compose parser path.Cleans but does NOT
reject `..`, and `ValidateBackupSpec` vets only *spec* entries, so an unlisted writable
`${HDD_PATH}/../x` bind arrives classed **mandatory**. Guards (run after the tier filter) move
traversal (`..` segment / absolute), bare **HDD** drive-root (`""` — would nest `<hddPath>/backups`),
and reserved `backups/` zone captures into `Skipped` with distinct English reasons (a skipped
mandatory = a capture GAP the engines log loudly). Bare **userdata** root is allowed
(`<hddPath>/userdata`). Segment-wise `..` detection (a legit `a..b` dir passes).
- **`CrossAppOverlaps(map[app]CaptureSet) []Overlap`:** pure §4.2 advisory — same absolute path in ≥2
apps' Paths (exact-Abs only; cross-app *containment* is legitimate and does NOT report). WARN wiring
is deferred to 3a/3b by design — no log call sites here.
- **Purity:** no `os`/`exec`/`filepath`/logging; **slash algebra** (`path.Join`/`path.Clean`)
throughout — resolved paths are in-container Linux paths, and `filepath` on the Windows test host
would flip separators and break containment prefix checks.
- **Docs alignment:** architecture §3 sketch updated to the as-built API (`UnitOnly` → `HasClassification`,
`Skipped` added), felhom.eu commit `8d85da7`.
- **Tests (all green):** `internal/appbackup/captureset_test.go` (Groups A–F: per-tier split, legacy
inertness, excluded-invisible, structural guards + legit `a..b`, containment/collapse/determinism,
cross-app overlap) + the F-S3 no-seam wiring test `internal/stacks/captureset_wiring_test.go`
(Group G, real Manager → ClassifiedBinds → ComputeCaptureSet end-to-end). All 6 §10 red-proofs
verified (mutation → fail → revert). **No behavior change; no UI; no engine edits.**
### v0.132.0 — Backup classification: schema + parser + pure classifier (INERT; Task 2) (2026-07-14)
Task 2 of the backup-classification-redesign arc
(`felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md`). Ships the
referential-coupling classification as **DATA + PARSER + PURE CLASSIFIER — deliberately inert**: NO
backup tier changes behavior. Task 3 (tier policy engine) and Task 4 (manual `.fab` UI) are the
consumers; today only a validation log pass touches it.
- **Schema (`internal/appbackup/classify.go`, new):** `BackupSpec`/`BindSpec` model the `.felhom.yml`
`backup:` block (`userdata:`/`hdd:` lists of `{path, class}`); `ComposeBind` is a `${VAR}`-relative
host bind carrying the `:ro` flag. Classes: **mandatory** (COUPLED — restore-without is
broken-not-empty, SQ3), **optional** (DECOUPLED-precious), **excluded** (DECOUPLED-bulk/transient).
- **Pure classifier `ClassifyBinds`:** the SQ5 **two-level default** — an explicit block entry ALWAYS
wins (an explicit `optional` on immich's `:ro` external library beats the reader default); an
unlisted **writable** bind defaults **mandatory** (the C6B-F1 capture direction, never silent-drop);
an unlisted **`:ro`** bind defaults **excluded** (reader rule). Returns `hasClassification` — **a
nil spec (no block) → every bind is `legacy` with NO class semantics**, so an unmigrated app's
behavior is byte-identical.
- **Validation `ValidateBackupSpec` (whole-block-reject):** ANY defect — unknown/empty class (a typoed
`clas:` key leaves `""`), empty/absolute/`..`/backslash/non-clean path, duplicate `(root, path)`, or
an entry matching NO compose bind (a stale/typoed path must not silently shift the real bind onto
the mandatory default) — rejects the ENTIRE block with the first defect named. Never partial.
- **Parser `ParseComposeClassifiableBinds` (`internal/stacks`):** copies the
`ParseComposeUserdataMounts` scanner shape but stays in `${VAR}`-relative space and preserves `:ro`
(why it does NOT reuse `ParseComposeHDDMounts`, which resolves absolutes and drops the mode). Deduped
on `(root, relpath)`, first-occurrence `:ro` wins; short-syntax only.
- **Integration:** `Metadata` gains `Backup *appbackup.BackupSpec`; `LoadMetadata` is the SINGLE
validation choke point (catalog listing, deployed-stack scan, and git-sync all flow through it, so a
bad catalog push logs `[ERROR] ... backup block rejected in <dir>: <reason>` within one sync cycle
and the app degrades to legacy). `stacks.Manager.ClassifiedBinds` + a new
`StackDataProvider.GetStackClassifiedBinds` seam (delegated by `stackAdapter`, nil-stubbed in every
fake) exist so Task 3 consumes a **wired, end-to-end-tested** seam — the F-S3 lesson that wiring is
where typos hide.
INERT by design: offsite/tier-2/`.fab`/deploy/sync are byte-identical — proven by the entire
pre-existing test suite staying green with **zero test-logic edits** (only mandated nil-stub methods
added to fakes). Recovery units already carry `.felhom.yml` and git-sync already whitelists it, so the
block propagates to deployed stacks + units with zero plumbing changes; no recovery-unit SchemaVersion
bump. The 13 catalog `backup:` blocks ship in the same `app-catalog-felhom.eu` change (this controller
must be live first so the parser validates them on first sync). +14 test functions (Groups A–E,
incl. a 9-case validation table); red-proofs RP-1..RP-4 all
confirmed (validation, explicit-beats-ro precedence, capture-default direction, LoadMetadata→validate
wiring). Controller-only; no MinAgent/hub coupling.
### v0.131.0 — F-S2 + F-S3: compose-derived appdata dir resolution (paperless-ngx → appdata/paperless) (2026-07-14)
The controller assumed an app's HDD appdata dir is always `appdata/<stackName>`. paperless-ngx binds
`${HDD_PATH}/appdata/paperless/...` — stack `paperless-ngx`, dir `paperless` — so every consumer that
keyed by stack name silently missed it via a stat-and-skip. One canonical resolver
(`appbackup.AppDataDirNames`) now derives the real dir name(s) from the app's compose `${HDD_PATH}`
binds, and all consumers use it. Task 1 of the backup-classification-redesign arc
(`felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md`), deliberately independent
of the classification schema.
- **F-S2 (spike-proven live) — tier-2 backup/info/restore.** `RunTier2` now mirrors the resolved
`appdata/<name>` dir, so paperless documents get their off-drive copy (previously: the appdata leg's
`os.Stat` gate skipped `appdata/paperless-ngx`, which never existed — **silent, no copy**).
`Tier2Info`'s size + the SSD-headroom guard use the resolved dir. `RestoreTier2Files` targets the
resolved live dir (was restoring into a wrong/empty `appdata/paperless-ngx`). A `[WARN]` now fires
when the compose DECLARES an appdata dir but it is absent on disk (the silence that hid F-S2). Every
rsync leg goes through a new `tier2Mirror` seam (prod behavior unchanged).
- **F-S3 (NEW, found this session) — scope="app" migration.** `migrate.go` keyed all six per-app
appdata legs (collision check, source-size, copy, verify, cleanup, skip-set) by stack name. For
**scope="app"** there is no merge walk, so migrating paperless-ngx copied nothing, "verified"
vacuously, flipped `HDD_PATH`, and the app came up with an **empty media dir**. (scope="all" was
saved by the merge walk — data safe, accounting off.) All six legs now loop the resolved name(s);
the copy leg WARNs on a missing declared dir.
- **Multi-dir refusal (defensive; no catalog app hits it today).** An app resolving to N>1 distinct
appdata dirs is refused loudly by tier-2 backup/info (honest `no_target` status +
`"az alkalmazáshoz több adatkönyvtár tartozik — a 2. mentés jelenleg alkalmazásonként egy könyvtárat
támogat"`) and tier-2 file-restore (refused BEFORE the app is stopped). Migrate supports N naturally.
This limitation is lifted by the tier-policy engine (Task 3).
- **Display.** The storage-detail page sums the resolved appdata dir(s), so paperless-ngx shows a
non-empty size.
- **Truth repair.** The v0.130.0 entry below states "The scheduled/tier-2 backup path was NOT affected
(it copies the felhom-data namespace wholesale)" — **that sentence is false** and is left in place
only as the historical record it corrects here: tier-2 copies the recovery unit + the resolved
`appdata/<name>` dir(s) ONLY (never the userdata tree — F-S1, unaddressed here — and never the
namespace wholesale). The `main.go` export-adapter comment that repeated the claim is fixed in code.
Scope guards: destination layout unchanged (`<destBase>/appdata` stays flat); no userdata copying at
any tier (F-S1 is the classification redesign's, not this task's); `ExportDataMounts` / `.fab` /
offbox untouched. Tests: +9 (resolver table incl. dedupe/foreign-drive/whole-root; RunTier2
paperless/legacy/multi-dir; Tier2Info + restore refusals + resolved live dir; scope="app" paperless
migration). Red-proofs RP-1..RP-5 all confirmed (resolver, RunTier2 leg, restore dst, migrate copy
leg, N>1 guard). Controller-only; no agent/hub coupling; MinAgent unchanged.
### v0.130.0 — CRITICAL C6B-F1: hollow .fab export (three compounding defects) + C6B-F2 share-removal guard (2026-07-14)
CAMPAIGN-6B surfaced that `.fab` export produced a **config-only, data-free bundle** for 12/13
`needs_hdd` catalog apps, reported success, and passed the v0.125.0 anti-hollow guard (live: sonarr,
4.17 GB / 7 files → a 2308-byte bundle). Cross-box/fresh restore = silent total data loss. The
scheduled/tier-2 backup path was NOT affected (it copies the felhom-data namespace wholesale) and is
untouched. Controller-only.
- **C6B-F1 cause 2 (discovery, `cmd/controller/main.go` exportAdapter + new `stacks.ExportDataMounts`):**
the export adapter resolved only `${HDD_PATH}` binds (`ParseComposeHDDMounts`), never the standard
`${USERDATA_PATH}` convention (`<HDD_PATH>/userdata`, injected at deploy) → 0 mounts → "no HDD
mounts — skipping". Fix: `stacks.ExportDataMounts` unions the HDD binds with the **userdata ROOT**
(one mount, basename `userdata`) when the compose binds `${USERDATA_PATH}`. Root-not-per-bind is
deliberate: the manifest keys HDD tars by basename and the import maps a basename to a resolved
mount or `<HDD_PATH>/<basename>` — `userdata` round-trips through the UNTOUCHED import exactly,
while a per-bind `userdata/media/tv` would base to `tv` and restore to the wrong place (the task's
literal per-bind union + namespaced tar names would have required import changes, which the task
forbade — deviation documented in REPORT). Containment dedupe both directions. The backup-side
`stackAdapter` is intentionally unchanged. Also fixes the estimate's `data=0 B` for these apps.
- **C6B-F1 cause 1 (either/or, `appexport/export.go` executeExport):** `needs_hdd` apps never ran
`exportVolumeData`, silently dropping named volumes (sonarr_config = the whole app DB). Export is
now ADDITIVE (HDD data AND volumes); `EstimateExport` counts both so fits-on-dest stays honest.
- **C6B-F1 cause 3 (guard, `appexport/export.go` assertBundleDataComplete):** the claimed-tar checks
pass trivially on 0 claims. New assertion: a `needs_hdd` manifest with neither HDD data nor volume
data fails the job ("a mentés nem tartalmaz alkalmazásadatot…") — a future discovery gap can never
again ship a silent hollow bundle.
- **§8 latent collision (`appexport/export.go` exportHDDData):** two mounts sharing a basename used
to silently overwrite the first tar; now a loud Hungarian failure (basename-keyed manifests cannot
round-trip a collision; renaming would break the import mapping). `exportHDDData` returns error.
- **C6B-F2 (share-removal guard, `web/netstorage_handlers.go`):** `POST /api/storage/netstorage/remove`
now refuses (409, names the apps) while a DEPLOYED stack's HDD_PATH is on the share — the live
event removed campaign6 under a running sonarr and the agent's tolerated stop steps deleted the
unit files under the busy mount, leaving an unreapable orphaned autofs mount until host reboot.
The remove handler resolves the agent via the netAgent seam. **Residual (out of scope, flagged for
a felhom-agent task):** the agent-side tolerate-and-continue stop in `RemoveNetworkMount`.
Tests (non-hollow, four red-proofs run→fail→revert): `stacks/export_mounts_test.go` (6 — union,
HDD-direct regression, mixed, covering-root, literal-userdata dedupe, empty; red-proof: pre-fix
HDD-only behavior fails 3), `appexport/export_additive_test.go` (5 — scenario A both-tars bundle,
scenario E volume-strand fails loud for needs_hdd, §8 collision loud-fail, scenario A' round-trip
placement to `<HDD_PATH>/userdata`, scenario D zero-data refusal; red-proofs: either/or revert fails
A, collision-check removal fails the collision test, guard removal fails D),
`web/netstorage_remove_guard_test.go` (2 — refused-while-deployed + proceeds-without; red-proof:
disabled guard returns the live `removed:true`).
### v0.129.0 — CAMPAIGN-4 fixes: rate-limiter key (F-B) + volume-blind estimate (F-A) + no-op claim status (F-C) (2026-07-14)
Three controller-side fixes from CAMPAIGN-4 (2026-07-13). Controller-only.
- **F-B (MED, security):** the login/escrow-reauth rate-limiter keyed on `r.RemoteAddr` (which is
`IP:PORT`) whenever `X-Forwarded-For` was absent, so every fresh direct connection from one host
got a distinct ephemeral port → a distinct key → the failed-attempt counter never accrued. A
direct-to-controller path (LAN/guest, bypassing the traefik/CF proxy) therefore had **no
brute-force protection**. Fix: a single shared `clientIP(r)` helper (XFF first-hop, else
`net.SplitHostPort(RemoteAddr)` host, else raw) — replaces the former `requestIP` and the
duplicated inline derivation in `handleLogin`, so the escrow re-auth limiter shares the exact same
fixed key. Accepted limitation (out of scope, commented): XFF is attacker-controlled on a direct
path — the fix closes the port-in-key bug, not XFF trust. Red-proof: revert `clientIP` to raw
`RemoteAddr` → the distinct-ports scenario stops limiting and escrow re-auth stays 401 not 429.
- **F-A (MED, honesty):** the export size-estimate's volume branch `du`'d the raw host mountpoint
from `docker volume inspect`, which is not mounted inside the containerized controller → returned
0, so a >1 GB volume-only app reported `data_size_bytes:0` / "3.6 KB" / `fits_on_dest:true`. Fix:
a `volumeSizer` seam whose real impl reads the size from a **container view** (`docker run --rm -v
<vol>:/vol:ro alpine du -sb /vol` — the same named-volume pattern the export path uses; never a
controller-host path, the v0.125.0 strand class). A failed read now sets `size_unknown` and
**forces `fits_on_dest:false`** (never renders as "fits") with the human string "ismeretlen méret".
The export pre-flight hard-aborts only on a KNOWN doesn't-fit (an unmeasured size no longer blocks
the export — the tar stream + destination FS surface a real ENOSPC). HDD-path branch unchanged.
Red-proof: revert the estimate to the host-path read → the >1 GiB scenario reads 0.
- **F-C (LOW-MED, correctness):** a no-op escrow claim (agent `phase:none` → HTTP 404) fell through
`escrowClaimAPIHandler` to a generic **502**. Fix: relay the agent's 404 as a clean 404 ("Nincs
aktív helyreállítási folyamat…") and 409 as 409; 410 (void) and a genuinely-unreachable agent
(status 0 → real bad gateway) are unchanged. Red-proof: remove the 404 mapping → the no-ceremony
claim returns 502.
Tests (non-hollow, all red-proofed): `ratelimit_ip_test.go` (F-B: 6 — direct-distinct-ports,
stable-XFF, rotating-XFF, escrow-reauth-shared-key, success-clears, clientIP unit),
`estimate_volsize_test.go` (F-A: 3 — real-not-zero, failure-never-fits, HDD-unchanged),
`TestEscrowClaim_ProxySemantics` +3 subtests (F-C: 404-not-502, 409, unreachable-stays-502). Live:
Alpine busybox `du -sb` verified supported (prod-valid).
### v0.128.1 — USB drives never show the rotational class hint (2026-07-13, ruling F5)
`storage.html` `classTag(d)`: `if(d.type==='usb') return '';` ahead of the class branches —
covers both render sites (card badges + metarow, the latter already guards on `d.class`).
Rationale: only Observe-sourced drives ever carried `class`, so the demo's legacy PVE
`dir:`-backed USB drives showed "lassú" while registry-sourced drives never did — a misleading
inconsistency, and the card already carries the USB type tag. Non-USB storages (e.g. a future
internal SATA data drive) keep the hint. **The hub-report `ClassHint` field is UNCHANGED**
(documented hint; UI-only suppression). Pinned by `TestStorageTemplate_USBClassBadgeSuppressed`
(red-proof: guard removed → FAIL "classTag USB guard missing"). Part 1 of the demo
storage-hygiene task (Part 2 = host-side `pvesm remove` of the two legacy dir storages —
operational, no repo change). Task was numbered v0.127.3 pre-sequencing; ships as v0.128.1
(0.127.3 + 0.128.0 already taken).
### v0.128.0 — browser .fab upload on the import page: chunked, tunnel-proof (2026-07-13)
The Restore/import flow no longer requires copying `.fab` files to `{tároló}/exports/` by hand
(the FileBrowser step alpha testers stumble on): `/import` now has a drag-and-drop/file-picker
upload zone. The binding constraint is the Cloudflare tunnel's request-body cap — **step-0 probe
on the REAL tunnel (2026-07-13): 120 MiB POST → edge HTTP 413 from `Server: cloudflare` before
the origin saw it; 80 MiB → passed to the origin (302 /login)** — so the client slices the file
(`File.slice`, 64 MiB chunks, strictly sequential) and the server appends each chunk to a
`.part-<random>` file in the DEFAULT drive's exports dir via `io.Copy` (no RAM proportional to
file size), then finalize fsyncs + atomically renames. The existing bundle scan + validation +
import pipeline take over untouched — the landing dir is exactly what `isValidExportPath` and
`ScanForBundles` already cover.
- **Endpoints** (inside `ServeExportAPI` — inherits the main.go `RequireAuth(CsrfProtect(...))`
mount, nothing added at the mux): `POST /api/export/upload/{init,chunk,finalize,abort}`.
Single-flight (second init → 409). Init sanitizes the filename to a `[A-Za-z0-9._ -]` base
name with a mandatory `.fab` suffix and gates on free space (declared size + 1 GiB margin,
Hungarian error with both numbers). Chunk offset MUST equal bytes received (mismatch → 409 +
`received_bytes` so the client re-syncs one step); per-request body cap 96 MiB. Finalize
requires the exact declared size (mismatch → 422, `.part` deleted) and lands collisions on the
lowest-free `"name (N).fab"` (a re-run finds its own prior `(N)` — never `(1)(1)`).
- **No client-side sha256 — deliberate:** WebCrypto can't stream-hash multi-GB files; the `.fab`
format self-validates at import. Transport integrity = sequential offsets + exact final size +
the format's own validation.
- **Crash-safety:** upload state is in-memory (a restart loses the `.part`; the browser
re-uploads). Startup GC removes `*.part-*` in every registered drive's exports dir; an upload
idle ≥15 min is aborted server-side.
- **UI** (`app_import.html`): upload zone above the bundle list ("Fájl kiválasztása" / húzza ide),
progress "Feltöltés: {pct}% ({done} / {total} GB)" + "Megszakítás"; on success the page reloads
(the existing scan renders the new row). One retry per chunk on network error, re-synced from
the 409 echo.
- **Reuse:** `appexport.DiskFree` exported (was `diskFree`) for the space gate via the
`web.uploadDiskFree` test seam. Scenarios §7 A–F tested; red-proofs run for the traversal
sanitize, the out-of-order append and the collision overwrite (all FAILED pre-fix as required).
### v0.127.3 — reveal copy states the shown code is ALREADY the live one (2026-07-13, Viktor)
The supersede happens at upload, inside the ceremony job — BEFORE the code is ever displayed.
The reveal warning now says so explicitly: "Ez mostantól az élő helyreállítási kód — a korábbi
kód érvényét vesztette. Mentse el most: a kód többé nem jeleníthető meg." (was only the
"nem jeleníthető meg" line). Pre-generation cancel already existed (the "Mégsem" next to
"Kód létrehozása" — nothing runs until the primary button); the render test now pins it. The
typed-back step stays (proof-of-capture friction + transcription-error catch; Viktor briefed).
### v0.127.2 — wizard code hide is MANUAL-only (2026-07-13, Viktor's live finding on v0.127.1)
The v0.127.1 blur-on-verify auto-blurred the code the moment a verification input got focus —
which made typing the two words HARDER, since the customer types them from the screen. Reversed:
the code stays visible; the "Elrejtés"/"Megjelenítés" toggle appears with the reveal and is
manual-only (still useful for a screen share). Render test asserts no `onfocus` auto-blur remains.
### v0.127.1 — escrow wizard polish: CTA visibility + Hungarian preflight + typed-back highlight + blur-on-verify (2026-07-13)
Four findings from Viktor's first supervised wizard passes (drill + demo). Presentation-only —
no endpoint, agent call, or state change; §10 red-proofs N/A.
- **Escrowed-card CTA** (`backups_remote.html`): "Új helyreállítási kód készítése" is a real
`btn btn-sm btn-outline` secondary button (was an inline link inside the muted hint — nearly
invisible). Outline, NOT primary: escrowed is healthy, the CTA is available-not-urgent (the
stale variant stays primary). Render-tested.
- **Hungarian preflight details** (`backups_escrow.html` `pfDetail`): the agent's operator-English
`detail` strings no longer leak into the customer UI. OK rows keep only VALUE details (storage
id, age path; `staged_secret` → "előkészítve"); boolean-OK rows (dr_tier/hub_upload/sudo_grant)
render no detail; not-OK rows get the Hungarian explanation + the raw agent detail as a muted
diagnostic span; unknown ids fall back to the raw detail (never blank a failure);
`staged_secret` keeps its informational dot. Render test pins the not-staged copy + asserts the
three English literals never appear in the page source.
- **Typed-back highlight**: the code renders as span-per-word (createElement + textContent +
createTextNode ONLY — R still never flows through innerHTML; no parse context = no injection
surface); the two verification words get `var(--warn)` + 600 weight so they're findable on
paper. `verifyIdx` is now chosen BEFORE rendering. `finishWizard`'s `textContent=''` clears the
spans (child-node replacement — verified).
- **Blur-on-verify** (Part 4, optional — implemented; trivial to strike): first focus on either
verification input blurs the code (`filter: blur(6px)`, inline-style toggle — no style.css
change, no cache-bust) + a `btn-ghost` "Megjelenítés"/"Elrejtés" toggle. The typed-back now
exercises the WRITTEN copy, not screen transcription. R stays in the JS closure.
JS behavior (spans/blur) is review-covered; the visual leg awaits Viktor's next login. All four
UI gates green.
### v0.127.0 — customer-facing escrow ceremony wizard + stale-blob re-check (2026-07-13) — MinAgent: 0.88.0 (wizard only; everything else unchanged)
The missing friend-alpha piece: the recovery-code ceremony moves from operator-SSH to a
customer-driveable wizard (`/backup/escrow`). R is displayed EXACTLY ONCE in the browser
(one-shot claim, typed-back confirm); operator ruling F1 2026-07-13 accepts the single CF-tunnel
transit (same trust class as the claim code — threat model in felhom.eu
RUNBOOK-escrow-ceremony.md). Mechanics validated by SPIKE-controller-escrow-2026-07-13.
- **Wizard** (`templates/backups_escrow.html` + `web/escrow_handlers.go`): preflight checklist →
warning copy (re-ceremony adds the supersede warning) → password re-auth (rides the LOGIN rate
limiter) → run (poll 2 s) → one-shot reveal ("Ez a kód többé nem jeleníthető meg.") →
typed-back (two random words, client-side only — R never leaves the page's JS scope; no
copy-to-clipboard by design) → finish. Void/expired → the honest "újra nem kérhető le" state.
Page + claim response `Cache-Control: no-store`; R is NEVER templated server-side, logged, or
persisted.
- **Start-handler order (load-bearing):** re-auth → **re-stage-first** (offbox configured →
`PushOffboxPasswordForEscrow`; failure ABORTS — a ceremony without the staged secret mints the
forbidden hash-less blob) → agent version gate (`AgentVersion()` ≥ 0.88.0, header absent =
older, fail-closed) → trigger. Every refusal exits with the agent untouched (seam-asserted).
- **agentapi** (`agentapi/escrow.go`): `EscrowPreflight` / `EscrowCeremonyStart` /
`EscrowCeremonyStatus` / `EscrowCeremonyClaim` (status-aware; 410 = void; claim body never
logged) over the existing envelope helpers.
- **Scenario F — stale-blob re-check** (`report/escrow_confirm.go`): `Reconcile` no longer
early-returns on non-pending; an ESCROWED box compares the ACK hash every cycle — mismatch OR
a present blob with an EMPTY hash (the spike's hash-less supersession) sets an in-memory stale
flag (surfaced on the Távoli mentés card: "A letétben lévő helyreállítási csomag nem fedi a
jelenlegi távoli mentési jelszót") + ONE warn per distinct hub hash (`warnedHash` reuse;
hash-less dedupes under a sentinel). State NEVER flips; runs NEVER block; a matching hash (or
a fresh auto-confirm) clears the flag. NOTE: the live demo's legacy hash-less blob will show
this warning honestly — the wizard is the fix.
- **Card rework** (`templates/backups_remote.html`): the deprecated manual-confirm BUTTON is gone
(the endpoint stays for legacy blobs); states: pending → "Helyreállítási kód szükséges" + CTA;
escrowed+stale → warning + "Új helyreállítási kód készítése"; escrowed clean → secondary link;
agent < 0.88.0 → "az ügynök frissítése szükséges" note, no CTA.
- Tests: call-order (stage BEFORE trigger, from pending AND escrowed), Scenario C no-stage,
security gates (wrong password 401 + rate-limit counter, 429 lockout, passwordless 403, stage
failure 502 pre-trigger, old agent 409, busy 409 — agent seam call-count 0 in each), claim
proxy no-store + 410, §8 stale truth table incl. dedupe + clear, template render states. §10
red-proofs demonstrated (felhom.eu REPORT).
### docs — controller.yaml.example: hub api_key literal scrubbed (2026-07-13)
The example carried the REAL hub global bearer key (the `manifests/hub.yaml` committed literal,
rotation-flagged in two publish runbooks). Replaced with a placeholder — real deployments get a
hub-issued per-customer key baked by configgen; the example was never a live consumer. Part of
the hub v0.53.0 bearer de-git (felhom.eu); the value itself dies with the supervised rotation
(documentation/runbooks/secrets.md §"Operator/global bearer key" in felhom.eu). No code change,
no version bump.
### v0.126.4 — edge-safe error statuses + the native-alert ban (2026-07-13)
Two defects surfaced by the agent-0.87.0 wizard leg's decommission attempt (the M1 refusal —
correct policy — reached the operator as a JSON SyntaxError popup):
- **502/504 never leave the origin:** Cloudflare replaces origin 502/504 bodies with its own
HTML error page, so every `writeDiskJSON(StatusBadGateway…)` refusal/error rendered as
"<!DOCTYPE … is not valid JSON" in the browser. `writeDiskJSON` now maps 502/504 → 500 at the
single choke point (JSON body crosses the edge intact); the M1 last-usable-drive refusal
became the typed `errLastUsableDrive` sentinel → **409** (policy verdict, not gateway
failure). Unit tests + red-proofs for both.
- **Native `alert()` banned** (the F-11 OS-modal class, now complete): the decommission error
path's `alert()` froze browser automation exactly as F-11 predicted. All 29 native `alert(`
calls across 5 templates swept to the existing `showAlert` modal (layout.html);
`native_confirm_gate.py` extended to ban `alert(` alongside confirm/prompt.
### v0.126.3 — storage wizard on a CLAIMED box: the init/attach POST no longer dies on CSRF (2026-07-13)
First live hit during the agent-0.87.0 drill wizard leg: /api/storage/init → "CSRF token missing
or invalid" (log: token mismatch). Root cause: `storageWizardPageHandler` rendered via raw
`render()` instead of `executeTemplate()`, so /storage/init + /storage/attach shipped an EMPTY
csrf-meta token — and the wizard's fetch() posts that token. LATENT until the claim arc: an
unclaimed box skips CsrfProtect entirely, so the wizard had never run against a password-gated
box before. Fix: executeTemplate (CSRF auto-injection); regression test renders both wizard
pages with a real session and asserts the meta carries the SESSION token (red-proven: swap back
to render() → both cases fail on the empty meta).
### v0.126.2 — stylesheet cache-bust (2026-07-13)
0.126.1 live QA: Cloudflare edge-caches `/static/style.css` for 4h (`Cf-Cache-Status: HIT`), so
every controller release kept serving the PREVIOUS release's CSS to customers until TTL. The
stylesheet link now carries `?v={{.Version}}` — busts automatically on every release (the same
gotcha class as the hub v0.47.0 /style.css finding, now extinct on the controller too).
### v0.126.1 — .form-input/.form-row finally have CSS (2026-07-13)
Live QA on 0.126.0 (drill box) showed the .fab password field STILL browser-default: the
`.form-input`/`.form-row` classes used across the backups/import/offbox templates had NO
backing rule in style.css at all (only the `.form-control` twin was styled) — the actual root
cause of the operator's "unstyled clipped placeholder" screenshot. Added the missing rules
(same visual spec as `.form-control`; label-over-field rows). Presentation only.
### v0.126.0 — UI uniformity bundle: shared app-list rows, infra-app metadata, restore-form polish, mojibake gate (2026-07-13)
Operator review (2026-07-13 screenshots): app lists looked designed three different ways across
four+ surfaces; infra stacks rendered as bare names; the .fab password input clipped its
placeholder; the zero-toggle run-warning went stale. Controller-only, presentation-layer — NO
backup/engine/toggle behavior change (render tests assert the action markup is untouched).
- **Part A — ONE row grammar, four surfaces:** `templates/app_row.html` defines
`app_list_row`/`app_list_row_end` (the layout_start/_end idiom): icon + name (+ optional
one-line secondary) left, caller action block right; compact 44px row. Applied to the Távoli
mentés toggle list, the Visszaállítás restore-to-verify + .fab lists and the dashboard
Telepített alkalmazások (state edge + data-href preserved); the Alkalmazások collapsed
headers ALIGNED to the grammar (icon+name left, status dot moved right before the chevron;
expander untouched — the one allowlisted aligned copy). funcmap: `dict` + `appHref`;
`OffboxAppRow`/`AppBackupRow` gain `Slug`. Old `.stack-card`/`.storage-path-item` row CSS+markup
retired. **NEW gate `scripts/app_row_dedup_gate.py`** — row markup single-sourced (red-proven:
pasted an old row block back → exit 1).
- **Part B — infra stacks carry identity:** `inframeta.go` static display-only map —
cloudflared → „Cloudflare Tunnel", traefik → „Traefik", filebrowser → „FileBrowser", each with
curated Hungarian description; dashboard rows + app cards show icon + description + the
existing Védett chip. Generic embedded `/static/infra-logo.svg` as icon fallback. filebrowser
is the ONLY Linked stack (files.<domain> Megnyitás); guarded WRONG outcome — no customer link
for cloudflared/traefik (render test counts exactly one https:// link; red-proven by flipping
Linked on cloudflared → FAIL).
- **Part C — restore-form polish:** the .fab encryption input is a standard form field —
placeholder „Opcionális jelszó" + helper under the field („Üresen hagyva a csomag titkosítás
nélkül készül."); the import-page bundle-password input picks up `.form-input` (was a bare
browser default).
- **Part D — mojibake fixed-by-construction:** byte-level sweep found ZERO double-encoded
literals in the committed source — the live „Tárhely"-class text on the import page is the
felhom-usb DRIVE-LABEL DATA (settings.json), repaired via the label-edit UI during live
validation. **NEW gate `scripts/mojibake_gate.py`** (Python per the multibyte rule): all
templates + Go sources must strict-UTF-8-decode and contain none of Ã Â Ă ă ˘ ˇ; allowlist
ZERO. Red-proven (reintroduced „Tárhely" → exit 1 naming file:line).
- **Part E — the stale zero-toggle line tells the truth:** `offboxWarningDisplay` display pick
(no state mutation): a persisted „nincs mentésre jelölt alkalmazás" run-warning is replaced by
„A kijelölés módosult az utolsó futás óta — a következő távoli mentés már tartalmazza." once
≥1 app is toggled (rendered neutral — reassurance, not deviation); 0 toggled keeps the
v0.123.0 line verbatim; quota/partial warnings pass through. Red-proven at unit + render level.
- **Gate housekeeping (first commit):** `scripts/backups_split_move_check.py` retired — the
one-shot v0.124.0 migration gate served its purpose (it pinned the split to verbatim moves vs
df7ad37); this release legitimately rewrites those blocks onto the shared row partial. All
other template gates stay mandatory (template_id, emoji, native_confirm, offbox_rename,
docker_run_volume_path + the two new ones).
### v0.125.0 — .fab volume export/import: containerized path-strand data loss FIXED (IA finding 1, HIGH) (2026-07-13) — MinAgent: 0.81.0
Both .fab volume legs streamed via `docker run -v <controller-temp-path>` host mounts — correct
on bare metal, silently wrong under the golden containerized deployment (the daemon resolves the
`-v` host side against the GUEST filesystem): the export's tar stranded host-side while the
bundle shipped an EMPTY `data/volumes` and reported SUCCESS; the import then wiped the app's
volumes and populated them from host-side emptiness. Live-hit on demo ActualBudget (v0.124.0
validation). No .fab format change — but note the ASYMMETRY: **any bundle exported by a
containerized controller ≤0.124.0 is suspect (hollow volume data) — re-export**; the new
import-side guard refuses such bundles loudly instead of destroying the app.
- **`docker cp` tar-streaming both legs** (`appexport/export.go` + `restore.go`, new `dockerExec`
seam): a stopped helper container pins the volume (`docker create -v <vol>:/vol alpine true`),
the tar streams over the docker API (`docker cp <cid>:/vol/. -` out; `docker cp - <cid>:/vol`
in) — ZERO shared paths, correct in both deployment shapes. §3 live probe proved content,
subdirs, symlinks, empty files and uid/gid round-trip. Helpers are ALWAYS force-removed, error
paths included (test-asserted); 10-min/volume timeouts + truncated stderr preserved.
- **Export can no longer lie** (scenario B): a failed volume export is FATAL (was WARN+continue);
`assertBundleDataComplete` refuses to package any bundle whose manifest claims a tar that is
missing/empty (volumes AND HDD subdirs — the HDD leg also stopped pre-claiming subdirs before
the tar succeeds). Live-proven: an engine-invalid volume name failed the export naming the
volume, no bundle staged.
- **Import validates BEFORE it destroys** (scenario C): `validateBundleData` refuses a
claimed-but-absent/empty data tar in step 0 — before the app is stopped and before any volume
is removed (the pre-fix order wiped first and discovered later); the refusal names the hollow
≤0.124.0-exporter cause and states the app is untouched. `restoreVolumeData`/`restoreHDDData`
missing-tar soft-skips became hard errors (defense in depth).
- **The class is extinct** (scenario D): `scripts/docker_run_volume_path_gate.py` — every `"-v"`
argument in non-test Go code must be allowlisted with its WHY; the Tier-1/2 volume dump/restore
entries are documented host-visible (registered-drive namespace paths under the golden
deployment's identical `/mnt` + `/opt/docker` binds), the rest are named-volume/flag usages.
- Red-proofs: assertion removed → hollow-success test fails; pre-flight disabled → the
zero-destruction assertions fail (`removedVolumes=1`); a violating `-v` line → gate exits 1.
- **Live §13**: supervised repeat of the exact failed leg on demo — export → download (bundle
now carries the 66048-byte volume tar) → drive placement → import → **volume fingerprint
byte-identical** (`ec8ea6cb…` before == after), app healthy, zero leaked helpers.
- **NOTE for the next publish train:** the golden floor must not advance past 0.124.0 without
this fix; floor may advance to 0.125.0 now that it validates.
### v0.124.0 — backups IA restructure: four sub-pages, Felhom-offsite status card, .fab browser download (2026-07-13) — MinAgent: 0.81.0
The nine-section backups page split into four sub-pages (operator review: customers got lost);
plus the offsite status card and the .fab portability exit. Controller-only; floor untouched.
Operator decisions recorded in CONTEXT: single active offsite destination stands; the status
card never changes state; .fab is portability, NOT a backup tier.
- **IA split** (`backups{,_remote,_apps,_restore}.html` + `backups_shared.html` partials):
**Áttekintés** `/backups` (storage overview, Rendszermentés, stat cards, single-copy warning),
**Távoli mentés** `/backups/remote` (status card + toggles + manual-target form,
`#offbox-section` anchor), **Alkalmazások** `/backups/apps` (schedule, Adatbázisok, per-app
1./2./3. rows), **Visszaállítás** `/backups/restore` (restore panel, the RELOCATED offbox
restore-to-verify, the .fab loop). Sections MOVED verbatim — `scripts/backups_split_move_check.py`
compares all 15 blocks against the v0.123.0 baseline (red-proven). Shared data builder
extracted (`backupsCommonData` + `backupsOffboxData`) — no duplicated computation. Old links
survive: `/backups` = Áttekintés; tier-3 row actions → `/backups/remote#offbox-section`;
offbox/restore/tier2 flash redirects + the tier2-config back-link retargeted per page.
- **Felhom-offsite status card** (top of Távoli mentés; local data only, DISPLAY-ONLY — no form,
no button, unit-enforced): (1) no target → "Felhom offsite tárhely — igényelhető szolgáltatás.
…Érdeklődj az üzemeltetőnél."; (2) applied + zero toggled → "Aktív — nincs kijelölt
alkalmazás" (+ the v0.123.0 hint below); (3) applied + toggled → no card (the status block is
the state). States 2 and 3 live-proven on the drill box.
- **.fab browser download** (`handler_export_download.go`): the EXISTING async export pipeline
with dest = a staging dir under the data dir (`fab-downloads/`; same producer → byte-identical
bundle), estimate shown BEFORE start, then a guarded streaming exit
(`GET /api/export/download?file=` — basename-shape + dir-containment guard, red-proven against
prefix-only matching; `io.Copy`, `Content-Disposition: attachment`, post-stream removal, 1h TTL
sweep on startup + each start). Batch UI downloads apps ONE AT A TIME (no mega-zip). Import
stays drive-scan; portability copy on the section ("Hordozható pillanatfelvétel…").
Unit round-trip: real export → import → content equality; a corrupted bundle is REFUSED
(gzip CRC — there is no per-file checksum; documented).
- **FINDINGS from the live §13 run** (both pre-existing, recorded for follow-up tasks):
**(HIGH)** containerized-controller .fab export of Docker-VOLUME apps strands the volume tar on
the GUEST host (`docker run -v <container-tmp>:/out` resolves against the host FS) → the bundle
ships an EMPTY `data/volumes`, export reports success, import brings the app up EMPTY. HDD-data
apps unaffected. Live-hit on demo (ActualBudget; data restored from the stranded tar by hand).
**(MEDIUM)** agent-side: on a legacy-boot PVE with LVM root, `SystemDisks` resolves NO raw
system disk → `sysKnown=false` → every disk classified system → the drive wizard can never
offer a candidate (drill box; hot-added disk invisible).
### v0.123.0 — polish batch: F-15 instant reset codes, F-11 inline confirms, Tier-3 "Távoli mentés" rename, zero-toggle honesty (2026-07-13) — MinAgent: 0.81.0
Four independent fixes from the take-two drill + operator review. Live-validated on the drill box
(qm 300 guest 9201) and demo 9201; hub v0.52.0 is the F-15 counterpart (an old hub's bare
reset-request response is a clean no-op — no coupling gate needed).
- **F-15 instant reset codes** (`internal/web/claim.go`): `requestHubResetCode` now parses the
hub's reset-request RESPONSE (`{claim: {code_hash, generation, issued_at}}`, hub ≥0.52.0) and
applies it through the SAME generation-guarded consumer as the report ACK
(`report.ClaimSync.Reconcile`) — the emailed code works the moment it lands instead of after
the next ACK (~15 min; the take-two "Hibás vagy lejárt kód" failure). Replay/older-generation
responses can never downgrade the active hash (guard reused, not reimplemented). Live re-run of
the exact failure path: code applied 1 s after the request, accepted on first try.
- **F-11 inline confirms** (`layout.html` + templates): every native `confirm()` (an OS-modal
that freezes browser automation) replaced by the LIGHT inline two-step — `felhomConfirm(el, q,
onYes)` swaps the trigger in place to "kérdés + Igen/Mégse"; form buttons opt in via
`data-confirm="…"` (submitted with `requestSubmit`, so formaction/name-value survive).
Converted: offbox restore-to-verify, tier2 restore, whole-guest backup, app data migrate,
debug simulate-disconnect + DR trigger, deploy stale-data delete (keeps its DOUBLE
acknowledgement, chained inline). New gate `scripts/native_confirm_gate.py` (zero native
confirm/prompt in templates; red-proven).
- **Tier-3 rename** (`backups.html`, `offbox_handlers.go`, `backup/offbox.go`): customer-facing
"NAS-mentés" branding → **"Távoli mentés"** (the productized target is the Storage Box; the
tier concept is offsite). Manual-target form generalized to any SFTP target ("Cél címe (IP
vagy hosztnév — NAS vagy SFTP-kiszolgáló)", "Tároló útvonala a célgépen"). The "Hálózati
tárhely" NAS network-storage feature keeps its device-truthful wording (different feature).
Python sweep (multibyte rule) + committed gate `scripts/offbox_rename_gate.py` (red-proven).
- **Zero-toggle honesty** (`backup/offbox.go`, `handlers.go`, `backups.html`): a configured +
escrowed offbox with ZERO toggled apps shows "Nincs távoli mentésre jelölt alkalmazás — jelölj
ki legalább egyet." on the toggle list, and a run in that state reports "Sikeres — nincs
mentésre jelölt alkalmazás" via LastWarning instead of bare success (red-proven unit test).
- **Dev-box fix**: `atomicPromoteTar` fsyncs via an O_RDWR handle — read-only fsync is refused on
Windows, which kept the two F7 atomic-dump tests permanently red on the dev box (Linux
behavior unchanged).
### v0.122.0 — customer-claim password gate (closes DRILL-day0-vm F-4/F-5) (2026-07-12) — MinAgent: 0.81.0
The customer sets + OWNS the dashboard password; the old "no password → open dashboard" is gone.
An unclaimed box (hub-delivered claim-code hash present, no password) serves ONLY the claim page
— every other route answers the claim page (302 → `/claim`) or `401` (API), so a Day-0 box is
never open on the public internet (closes F-4; F-5's unauthenticated geo toggle closes with it).
Requires the hub's v0.50.0 claim engine (code generation + email + ACK/config delivery).
- **`internal/web/claim.go`** — the gate + pages. `claimGateActive()` (no password + code hash +
not claimed), `effectiveClaimCode()` (ACK-cached settings beats the config bake by generation),
the claim page (`GET /claim`), submit (`POST /claim`: verify code → set own password → claimed
→ consume generation → session), and "kérj új kódot / Elfelejtett jelszó" (`POST
/claim/request-new-code` → hub `reset-request`). Code checks: bcrypt match AND generation not
yet consumed (single-use) AND ≤ 72 h old. Per-source + global brute-force limiter (5 tries →
15-min lockout, fake-clock tested); a lockout raises the allowlisted `claim_lockout` event.
Pre-auth CSRF is an HMAC over `web.session_secret` (closes the CTRL-007 bare-double-submit
weakness), min password length 12.
- **Gate wiring** (`auth.go`, `csrf.go`, `server.go`, `cmd`): the gate sits atop `RequireAuth`; a
SET password disables it entirely (password auth wins — claimed boxes never regress). `/claim*`
+ `/static/*` stay reachable pre-auth (the code is the strong factor). Legacy-open (no password,
no hash) passes through with a red transition banner (`layout.html`) until the hub delivers a
hash. Login page gains an "Elfelejtett jelszó" link.
- **`internal/report/claim_sync.go`** — caches the ACK's `claim` {hash, generation} into
settings.json IDEMPOTENTLY BY GENERATION (offsite-descriptor one-way shape: newer generation
advances; same/older/nil never rewrites, a hub outage never clears). The report carries
`claimed` (set-only hub-side). `config.web.claim_code_*` baked by the hub gates from first boot.
- **`internal/settings`** — `Claimed` (set-only), `ClaimCode*` cache, `ClaimConsumedGeneration`
(single-use). **`--print-reset-code`** root escape hatch: prints a one-time local code (a
generation above cached/baked/consumed), the same gate consumes it.
- Tests: gate-coverage signature test (every route → claim/401, a deploy POST mutates nothing) +
happy-path/reuse-refused/expired/lockout+window-reopen; four §10 red-proofs proven
(mutate→FAIL→revert): gate skip-line, single-use generation (hub + controller), reset non-DoS,
rate-limiter.
### v0.121.0 — backups page truth pass (dead sections removed, real Tier-3 state, SQLite-honest DB) (2026-07-12) — MinAgent: 0.81.0
Pure UI/data-plumbing on `/backups`; no backup-engine behavior change, no agent-API change, MinAgent
UNCHANGED (0.81.0). The live demo page (v0.120.0) contradicted itself — a dead "Részletek" card
claimed "Nincs 2. szintű mentés konfigurálva" while six Tier-2 runs showed above it; every per-app
"3. mentés" row said "Hamarosan — B2/S3/SFTP" while the off-box (Storage Box) tier was live with 6
snapshots; and an embedded-SQLite-only box rendered "Adatbázis mentve: 0" / "Nem található adatbázis
mentés." as if backups were failing.
- **Dead "Részletek" card removed** (operator-approved as redundant — per-app rows + the Adatbázisok
section already carry the truth). This deletes the last references to the never-set template fields
`Tier2DriveGroups` and `ResticPassword`, the `restic-pw` element, and the `toggleTier` /
`toggleResticPw` / `copyResticPw` JS.
- **Per-app "3. mentés" row now shows real off-box state** — a four-state row driven by a new pure
`tier3State` (configured → toggle → escrow precedence): `unconfigured` ("Nincs beállítva" +
Beállítás link), `off` ("Kikapcsolva" + Bekapcsolás link → `#offbox-section`), `escrow_pending`
("Kulcsletétre vár" — never a false success while the fork-4 escrow gate holds), `active` (status
badge from the global off-box `LastStatus`, `restic → <host>`, relative last-run). The "hamarosan"
placeholder is gone everywhere.
- **SQLite-honest DB messaging** — a new pure `dbSectionState(discovered, dumps)` picks
`dumps` / `pending` / `embedded`. Embedded-only boxes now render "–" + "beágyazott DB-k a
kötetmentésben" on the stat card and an explanation ("…beágyazott adatbázist használnak (pl.
SQLite)…") instead of a bare "0" / "Nem található adatbázis mentés."; a discovered-but-not-yet-dumped
box shows "…az első ütemezett mentés éjjel fut le."
- **Three dead/raw display fields fixed** — `Tier1LastRun`/`Tier1LastStatus` (previously never
assigned) are now populated from `ListRestorePoints` (newest recovery-unit artifact time; correct
per-drive resolution — no fabricated time for unit-less apps); Tier-1 and Tier-2 last-run labels
now render via `timeAgoStr` (relative time) instead of raw RFC3339 (the restore confirm() dialog
keeps the precise timestamp by design).
- **Terminology split** — the off-box section is retitled "Távoli mentés (3. mentés) — titkosított,
offsite" (with an `#offbox-section` anchor); the whole-guest PBS stat card is relabeled "Távoli
rendszermentés" so two different features no longer share one customer-facing name on one page.
- **Deploy page** — the app backup card gains a "Mentési beállítások →" link to the tier-2 config panel.
- **Tests:** +9 in `internal/web` (pure `dbSectionState`/`tier3State` truth tables; `buildAppBackupRows`
off-box mapping + escrow-pending precedence + Tier-1-from-restore-points wiring; template renders for
all four Tier-3 states, embedded/pending DB messaging, and relative-time formatting). Four companion
red-proofs run→fail→revert (re-insert "hamarosan"; hardcode OffboxEnabled=false; dbSectionState
ignores discovered; drop the Tier1LastRun assignment).
### v0.120.0 — dead-app alerting (fix-3) + debug-ring revision (fix-6) — CLOSES CAMPAIGN-3 (2026-07-12) — MinAgent: 0.81.0
The last CAMPAIGN-3 findings (`felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md`). MinAgent
UNCHANGED (0.81.0). Pairs with hub v0.48.0 (accepts the new `app_start_failed` event).
- **fix-3 (MED) — a dead deployed app is LOUD, not silent.** The campaign's CWA sat dead 4 h with no
signal; F11 then produced 4 silently-dead NAS apps per reboot. A new `deadapp-check` job (every 30 s)
scans `stackMgr.GetStacks()`: a DEPLOYED app whose containers are `stopped`/`exited`
(`stacks.IsDownState`; `created`/`dead` map to `stopped` — the F11 dead-at-boot case) raises a
state-based WARN dashboard banner ("Telepített alkalmazás nem fut: <app>"; grouped above 3 to survive
a reboot storm) that SELF-CLEARS the moment the app runs again, AND fires an `app_start_failed` hub
event ONCE per running→down transition (`Notifier.NotifyAppStartFailures` tracks per-app state — the
hub owns the real cooldown; the controller adds no timer and does not spam). A 90 s boot grace skips
the controller's own startup settle so apps that legitimately take 30–60 s to come up don't
false-alarm; after the grace an app that never came up STILL fires (the whole point).
- **fix-6 (MED) — the post-incident window survives.** The 1000-entry ring wrapped in ~6.5 min under
the campaign's load and died on every restart. Three changes: **(a) cap 1000→5000** (viewer +
`Entries`/handler display cap raised to match — a larger ring is useless if unreadable); **(b)
periodic-noise policy** — the every-cycle scheduler "job finished" + `refreshStatusLocked` success
lines are demoted to a new `[TRACE]` level the ring DROPS at write-time (failures/transitions are
never TRACE, so nothing is lost); **(c) spill persistence** — `LogBuffer.SpillTo`/`LoadFrom`
atomically (tmp+rename, JSON-lines) spill the ring to `<DataDir>/debug-ring.log` on the SSD state dir
(NEVER a NAS path) every 30 s and on clean shutdown, loading it back on boot so a restart / container
recreation preserves the pre-restart window. Corruption-safe (a truncated line is skipped, never
fatal).
- **Live-validated (demo 9201 + hub):** fix-3 — `docker stop seerr` → the dashboard banner
"Telepített alkalmazás nem fut: Jellyseerr (stopped)" appeared AND the hub received exactly ONE
`app_start_failed` event across 3 down-cycles (anti-spam); `docker start` → banner self-cleared.
fix-6 — the ring showed 0 periodic-spam lines; a controller restart PRESERVED the pre-restart window
(oldest entry unchanged across the restart; 63 KB spill on the persistent SSD volume). Tests incl.
the fix-3 silent-regression + one-event-per-transition red-proofs, the fix-6 TRACE-drop-keeps-failure
+ corrupt-spill-safe red-proofs, all green.
### v0.119.0 — storage-health coherence (F8) + mapped_uid validation (F4) (2026-07-12) — MinAgent: 0.81.0
Fixes CAMPAIGN-3 (`felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md`) storage-UI findings.
MinAgent UNCHANGED (0.81.0) — controller-only; the §3 design fork took the recommended option **B**
(reuse the shipped v0.117.0 classifier), so no agent change.
- **F8 (MED) — one classification, two surfaces.** The share row's health used to come only from the
agent's SERVER-LEVEL TCP dial (`server:2049/445`), which stays green when a *single* export is
`exportfs -u`'d — so the row showed benign "Készenlét" while the stacks/dashboard already showed the
stub reality. `networkStorageItems` now FUSES the agent view with the consuming-namespace
classification (`fuseNetHealth` → the same `system.ClassifyPathFS` the stacks stub badge reads): a new
`stub` health state wins over a benign idle/ok when the namespace sees local disk at `Where`; a
whole-server `unreachable` still wins over stub; autofs-healthy / network / inconclusive `unknown`
leave the agent health untouched (never manufacture a fault, never force-mount an idle trigger). The
row badge for `stub` = "Hibás — az alkalmazások nem a NAS-t látják". The row and the stacks/dashboard
badge now derive from ONE classification and can never contradict.
- **F4 (LOW) — mapped_uid/gid validated at the door.** `handleNetStorageAdd` range-checks the container
uid/gid (1..65533) after the `<=0` default, BEFORE the job starts. Out of range → an immediate,
friendly Hungarian 400 ("Az alkalmazás felhasználói azonosítója (uid) érvénytelen…"), nothing
installed — the campaign's `mapped_uid:101000` (a host-side mapped value) previously slipped past the
controller and failed only at the agent with a raw `agent_error`.
- **Live-validated (demo 9201):** F8 — `exportfs -u` while idle + drop-mount → the share row flipped to
`stub`/"Hibás — az alkalmazások nem a NAS-t látják" AND the stacks stub badge showed (4), the two
surfaces AGREE; re-export → row cleared to `ok`/"Elérhető" (healthy idle NOT downgraded). F4 —
`mapped_uid:101000` → 400 + friendly message, registry unchanged; `mapped_uid:1000` passed the range
check. Tests incl. the F8 fusion companion (revert → row idle → fail), the autofs-not-stub guard, and
the F4 boundary (65533 pass / 65534 fail), all green.
### v0.118.0 — backup integrity: atomic volume dumps (F7) + no single-copy (F6) + stale-primary sweep (F5) (2026-07-12) — MinAgent: 0.81.0
Fixes CAMPAIGN-3 (`felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md`) backup findings. MinAgent
UNCHANGED (0.81.0) — all changes are controller-local; no new agent API consumed.
- **F7 (HIGH) — atomic volume dumps.** `backup.DumpAppVolumes` now writes the tar to `<vol>.tar.tmp`,
fsyncs it, and only atomically `os.Rename`s it over the restore point on success — the same
crash-safe pattern the DB-dump path already uses (`appbackup/dbdump.go` DumpOne), extended with a
best-effort directory fsync. Before this, tar wrote the `.tar` IN PLACE, so a mid-write NFS cut left
a 0-byte tar REPLACING the last good dump (tier-1 restore is replace-semantics → an empty volume).
Now any tar error / timeout / dead-NFS EIO removes ONLY the `.tmp`; the last good `.tar` is
byte-untouched. The `.tar.tmp` name (ends `.tmp`, not `.tar`) is invisible to the
restore-point/stale scans; orphan `.tar.tmp` from a killed run is swept. New `tarVolume` test seam.
- **F6 (LOW) — no single-copy backups.** Volume-only apps (no HDD_PATH, backups on sys_drive) now flow
through the tier-2 cross-drive copy (`RunAllTier2` no longer skips non-HDD apps) — a second copy on
the secondary drive (the 3-2-1 intent). Their restore-point drive label is no longer blank (clear
"Belső SSD (rendszer)"). A single-drive box (no off-drive target) surfaces an HONEST
`SingleCopyWarning` banner on the backup page instead of implying a 3-2-1 guarantee it cannot keep.
- **F5 (LOW) — stale primary-dir sweep.** After each backup cycle, `pruneStalePrimaryDirs` removes an
orphaned `backups/primary/<app>` dir an app left on an OLD drive when its HDD_PATH moved (invisible
disk residue). LOAD-BEARING GUARDS: removes only when the app is deployed AND its current namespace
root differs from the dir's drive; NEVER touches the app's current-drive dir (the live restore
point) or an undeployed app's dir; only ever operates strictly under a `backups/primary/` prefix.
- **Part 4 (operator fork) — backup-target locality: option A (keep locality), document-only.** NAS
apps' tier-1 artifacts stay beside the data on the NAS; tier-2's cross-drive copy is the off-NAS
leg. Documented plainly (backup feature doc) so the NAS-outage window is never a surprise; no code
change (option B, retarget-to-local, was not selected).
- **Live-validated (demo 9201):** F7 money-shot — a mid-write `exportfs -u` during a volume dump left
all 5 nas-media volume tars BYTE-IDENTICAL (sha unchanged), no 0-byte, no leftover `.tar.tmp`, run
`success:false`; next run produced fresh good tars. F6 — actualbudget/seerr now on
felhom-usb/secondary. F5 — a seeded stale dir on the wrong drive swept, current dirs kept. Restore
round-trip byte-identical. Tests incl. the F7 truncation red-proof + F5 guard red-proofs, all green.
### v0.117.0 — consuming-namespace NAS verification + deploy-view truth (RCA fixes 2+4) (2026-07-11) — MinAgent: 0.81.0
Controller half of the RCA fix pair (agent v0.84.0 ReassertNetworkMounts). Source:
`felhom.eu/documentation/audits/AUDIT-nas-cwa-rca-2026-07-11.md`. MinAgent UNCHANGED (0.81.0) —
every new check is controller-namespace-local; no new agent API is consumed.
- **`internal/system/fsclass*.go`** — statfs f_type classifier for THIS process's namespace:
`network` (NFS 0x6969 / CIFS 0xFF534D42 / SMB2 0xFE534D42) | `autofs` (0x0187 — the HEALTHY idle
trigger; NEVER force-mounted) | `stub` (anything local — the RCA's silent guest-reboot state) |
`unknown` (statfs error/3 s timeout — fail open). Seams: `statfsFn` + per-caller injectables.
- **Probe fstype assertion (fix 2a):** the `--netprobe` child creates the probe file FIRST (the
create legitimately triggers the automount), THEN requires a MOUNTED network fs — exit 5 →
category `not_network_fs` (new §3.2 Hungarian message), full rollback, nothing registered. A
writable local stub can never verify again. Red-proof: assertion disabled → the stub VERIFIED
(exit 0 / job phase `done`) → FAIL.
- **Deploy-time refusal (fix 2b):** `POST /api/stacks/{name}/deploy` refuses (409, Hungarian) when
`HDD_PATH` is a registered network path classifying as a stub (`Router.refuseNetworkStubDeploy`,
`classifyFSPath` seam). Idle autofs / live / unknown / local / unregistered / empty all proceed.
Red-proof: a mounted-only gate wrongly refuses the healthy idle trigger → FAIL.
- **Stub badge (fix 2c):** `networkStorageWarnings` returns (warnings, stubs); the controller-side
classification runs even when the agent is unreachable. Dashboard + stacks cards render the new
distinct badge "Hálózati tárhely hibás — az alkalmazás nem a NAS-t látja"; stub WINS over the
recoverable unreachable badge (never both); the unreachable line stays byte-identical
(template-asserted). Pure mapping core `networkStorageWarningsIn` (appsUsingPathIn pattern).
- **Deploy-view truth (fix 4, the RCA S-C symptom):** the deployed-app storage select marks
`selected` by the STORED `HDD_PATH` (`CurrentHDDPath`); a stored path absent from the schedulable
list renders an extra disabled `<path> (nem elérhető)` option; `IsDefault` selects only for NEW
deploys. Red-proof: IsDefault-only revert → the default drive shows selected → FAIL.
- Gates: template_id_gate + emoji_gate OK; full `go build/vet/test ./...` green.
### v0.116.1 — debug surface ungated from logging.level (2026-07-11) — MinAgent: 0.81.0
Live validation of v0.116.0 caught the last blind spot: `/debug` + `/api/debug/*` (and the nav
item) 404'd/hid unless `logging.level=debug` — the EXACT failure mode of the motivating incident,
still standing in front of the new always-on ring. The debug surface is now available at ANY
logging level (still session-authed via RequireAuth + CSRF); the nav link always renders.
`isDebug()` keeps gating only legacy log EMISSION sites, as designed.
### v0.116.0 — observability pass: always-on debug ring + leveled sweep + agent tab + self-log pull (2026-07-11) — MinAgent: 0.81.0
Controller half of the cross-repo observability task (agent v0.83.0 + hub v0.46.0). Motivating
incident: a live NAS-verify refusal on an `info` box showed NOTHING in the debug view — the ring
only existed at `logging.level=debug`, so the detail never existed.
- **Capture layer**: `setupLogger` now ALWAYS builds the 1000-entry `LogBuffer`; the logger is
`MultiWriter(LevelFilterWriter(stdout, logging.level), ring)` — DEBUG always reaches the ring,
stdout/docker-logs keep respecting `logging.level` exactly as before (red-proof: filter disabled →
the capture test fails on the stdout assertion). New `internal/logx` leveled helpers
(`Debugf/Infof/Warnf/Errorf`, caller-attributed via `Output(3,…)`); legacy `isDebug()` sites
untouched (observation, not refactor).
- **Report self-log pull** (`report/selftail.go`): ACK gains `controller_log_requested` (additive);
the NEXT report carries `controller_log_tail` (ring newest-kept, 128 KB, consume-once — the
v0.111.0 logtail.go shape copied exactly; red-proof: drain removed → ships every cycle → FAIL).
The app-tail wire is byte-compatible (schema test asserts steady-state omission + unchanged keys).
Serving a pull logs the customer-visible `operator log pull served` INFO (rides IN the tail).
- **Debug page agent tab**: Naplóviewer gains `Vezérlő | Ügynök` tabs; the agent tab proxies
`GET /api/debug/agent-logs` → agent `GET /debug/logs` (client `DebugLogs`, 10 s budget). A
pre-0.83 agent (typed 404 StatusError) renders "Az ügynök naplónézete az ügynök következő
frissítése után érhető el." — ok-response, no error spam, nothing else gated (S6 tested both
polarities). Template gates green.
- **Gap-fill sweep** (all new lines via logx; entry/decisions/outcome+duration/errors):
netstorage_job (start, per-phase transitions with elapsed, agent add/verify/probe verdicts,
rollback start+outcome, terminal WARN/INFO with duration), netprobe (exec start + result),
netstorage_handlers (per-check validation refusals, orphan-share WARN, capability-gate line now
carries the decision SOURCE via new `SupportsWithSource` — version vs probe vs cache), agentapi
client (per-call DEBUG method/path/status/duration + agent-version-change line; `SetLogger` wired
on the memoized client), migrate engine (run start, per-phase DEBUG, complete line with duration),
tier2/offbox (run-start INFO + previously SWALLOWED status-persist errors now WARN).
- **S7 log-sequence smoke**: a full fake NAS add at level info must leave the 8 ordered phase
markers in the ring (red-proof: dropped probe-verdict line → FAIL naming the marker).
- MinAgent: **0.81.0 unchanged** — the agent tab degrades to the notice on older agents; nothing
else is coupled. Demo-deploy only; Peti untouched (his visibility arrives with the next train).
### v0.115.0 — version-aware Supports (agent version channel) + DSM-validated guidance (2026-07-11) — MinAgent: 0.81.0
Capability detection upgrades from route-probing to explicit version comparison, riding agent
v0.82.0's `X-Felhom-Agent-Version` response header. **The probe FALLS BACK cleanly — agent 0.82 is
NOT required** (MinAgent stays 0.81.0: the coupled NAS semantics; Peti's 0.81.0 box exercises the
fallback in production).
- **agentapi**: every response path passively captures the header (`noteAgentVersion` on all four
`Do` sites — even on 404s/errors); STRICT bare-semver validation at capture (the publish-agent.sh
shape; garbage never overwrites); `Client.AgentVersion()` exposes the last-seen value.
- **features.go**: per-feature `featureMinAgent` table (`netstorage_verify: 0.81.0`) + optional
`AgentVersionReporter` on the prober. `Supports` order: version known → semver compare →
Yes/No with ZERO probe traffic; version unknown/garbage/table-gap → the v0.114.0 probe path
byte-identical (SupportCache stays probe-only). Gate/banner/UI unchanged — same three verdicts,
better source.
- **THE one comparator**: `selfupdate.ParseVersion/Version.Compare` moved verbatim to
`internal/util/version.go` (selfupdate keeps type aliases — call sites + tests byte-unchanged);
agentapi shares it (no import cycle, no second comparator).
- **DSM-validated NAS guidance** (SPIKE-nas-dsm-2026-07-11, real DSM 7.2 via virtual-dsm): the NFS
guidance gains the verified Synology steps — File Services → NFS → enable + **Maximum NFS
protocol: NFSv4.1** (the v3 default refuses our mount), NFS Permissions rule with Squash
„Map all users to admin”, the `/volume1/<mappa>` path hint; the "útmutató készül" caveat narrows
to **QNAP only** (Synology now validated end-to-end incl. SMB hardlink).
- Tests + red-proofs: version-known compares without probing (mutant: short-circuit dropped →
probes=1); garbage/absent header → exactly-one-probe fallback + cached (mutant: trusting an
unparseable header as "too old" → fails); non-reporter probers byte-unchanged; comparator table
incl. pre-release rejection + numeric-vs-lexicographic; wire-level: header wins over a routeless
agent through the real pinned client, garbage header ignored at capture.
### v0.114.0 — agent-capability gate for coupled features (2026-07-11) — MinAgent: —
Box-level backstop for the publish-train ordering discipline (incident: the 0.81/0.113 train's
9-minute controller-before-agent skew on Peti's box — RUNBOOK-publish-0.81-0.113-2026-07-11): the
controller now detects whether its agent supports a coupled feature and refuses that feature up
front, instead of failing mid-pipeline with a misleading rollback. No agent or hub changes; works
against agents 0.79–0.81 as they exist. (Retroactive note: v0.113.0's effective MinAgent was
0.81.0 for the NAS add — this release is the machinery that makes such coupling self-protecting.
Header convention from here on: coupled releases declare `MinAgent: X.Y.Z` on this line.)
- **agentapi typed status (1.1):** non-2xx GETs surface as typed `*StatusError{Path,Code}` (same
message text as the old formatted error) — the probe keys on `Code==404` via `errors.As`, never
string matching.
- **`internal/agentapi/features.go`:** `Feature`/`SupportState` + `featureProbes` table (one row:
`netstorage_verify` → `GET /netstorage/verify-status`, the route that shipped WITH the coupled
add semantics in agent v0.81.0) + `SupportCache` (TTL 5 min, Yes/No cached, Unknown NEVER cached
or refused) + `Client.Supports`. 2xx ⇒ Yes; 404 ⇒ No; transport/timeout/401/5xx ⇒ Unknown — an
agent problem is never claimed as "too old".
- **Add gate:** `handleNetStorageAdd` refuses on `SupportNo` BEFORE the single-flight claim —
HTTP 412, machine code `agent_outdated`, message "Az ügynök frissítése szükséges ehhez a
funkcióhoz — a frissítés megérkezése után próbáld újra." `SupportUnknown` passes through to the
existing agent-error paths. `remove`/`list`/health are NOT gated — old shares stay manageable.
- **Settings page:** `NetAddSupport` (yes/no/unknown, short 2 s probe budget + cache) — `no` swaps
the add form for the honest banner; the share list + remove render in every state.
- **Tests:** T1 gate refusal (job never starts, slot never claimed, zero agent calls), T2 unchanged
happy path + warm-cache NEGATIVE assertion (probe count stays 1 across two adds), T3
indeterminate-never-refuses, T4 classification incl. the string-match trap case, T5 banner
render, T6 TTL re-fire, wire-level 404-typing through the pinned client. Red-proofs RP1–RP5 run +
reverted (recorded in REPORT.md).
- Also: fixed a scheduling flake in `TestBackupTier2Restore_DoubleClickRefused` (pre-existing).
- **Docs:** publish-train rules codified at `felhom.eu/documentation/runbooks/publish-train-rules.md`
(manifest-before-floor; floor field LAST — the DB row overrides env and acts immediately;
MinAgent fleet gate; this gate as backstop).
### v0.113.0 — NAS verify-before-commit + page redesign + protocol-honest guidance (2026-07-11)
Kills the "bogus share sits at Készenlét forever" bug: `POST /api/storage/netstorage/add` now
verifies the share END-TO-END before anything is registered, and rolls everything back on failure.
Built on SPIKE-nas-verify-2026-07-11 (b57f6ca) with agent v0.81.0; live-validated A–E on demo 9201
against an isolated sim NAS.
- **Orchestration job** (`internal/web/netstorage_job.go`, the migrate.go shape): sync validation →
detached single-flight job on `context.Background()` (~150 s budget; a closed tab can't abort a
rollback) with phases `agent_add → verifying → probing → registering → done|failed`, polled on
NEW `GET /api/storage/netstorage/add/status`. Registration is the LAST step — the worst crash
outcome is an agent-side orphan, never a registered-but-broken path. Verify-lost after an agent
restart (`phase:none`) ⇒ controller rollback (Scenario F).
- **In-guest uid-1000 write probe** (`netprobe*.go` + hidden `--netprobe <dir>` re-exec mode in
main.go): `SysProcAttr.Credential{1000,1000}`, no shell; dot-file + nonce + readback + delete;
exit codes → `not_writable` (the squash trap — an export that mounts but denies uid-1000 writes
can no longer register) / `probe_io`; cleanup-fail = WARN on success, not a failure.
- **agentapi**: `AddNetStorage` result gains `verify/job_id/code`; typed `NetAddRefusedError`
(categorized sync refusals — unreachable pre-probe); new `NetVerifyStatus` (short GET, the 15 s
global client timeout is untouched — the long wait lives in the poll loop).
- **§3.2 Hungarian error map** server-side (`netAddMessage`): unreachable / nfs_export (MERGED
not-found+not-permitted — NFSv4 returns identical strings) / smb_auth / smb_share / timeout /
not_writable (the Route-A guidance with the computed uid+100000) / probe_io / generic.
- **Orphan surfacing**: any agent-configured share NOT in the registry renders as a remove-only
"Árva megosztás" row (closes the crash-window gap visibly; re-add with the same name = repair).
- **storage_network.html full redesign** on the canonical `storage_attach` pattern — kills the
`<details>/<summary>`-as-button hack and the NONEXISTENT `form-row`/`form-input` classes (the
unstyled-look root cause). SMB listed FIRST (`SMB (Synology, QNAP — a legtöbb NAS)`), NFS
two-recipe guidance (map-all-users simple recipe + full-fidelity `anonuid=<uid+100000>` with a
live computed host-id), staged poll progress (Kapcsolódás → Csatolási teszt → Írásteszt →
Regisztrálás), categorized errors + collapsible raw detail. Gates green; C8 render smoke guards
the class regression.
- Feature doc: `felhom.eu/documentation/controller/network-storage-nas.md` (authoritative).
Companion: agent v0.81.0 (retry=0, journal classifier, agent-side auto-rollback), host-install
v1.13.0 (`systemd-journal` group). Red-proof outcomes: REPORT.md.
### v0.112.0 — self-update without credentials: anonymous registry mode (2026-07-10)
Root cause (live on Peti's box): the updater piggybacked on the Git Sync credentials and REFUSED when
they were absent — but the registry serves the public package anonymously (Docker v2 token dance,
verified empirically). A fresh customer without a private catalog silently lost version discovery +
self-update for no reason. Credentials become what they were meant to be: optional, private-catalog only.
- **`queryRegistry` (internal/selfupdate):** both creds empty → anonymous mode — plain GET; on 401
parse `WWW-Authenticate` (realm + service FROM THE HEADER — never hardcoded, quoted/bare/any-order/
comma-in-quotes handled); GET the realm with `service` + `repository:<image>:pull` scope and NO
credentials; retry tags/list with the Bearer. Creds present → the BasicAuth path unchanged.
Half-configured pair → loud "hiányos registry hitelesítő adatok". A genuinely-denying registry →
"registry denied anonymous access — a private registry requires Git Sync credentials" (never the old
"credentials missing"). The registry base URL now derives from the image ref (was hardcoded host).
- **`pullImage`:** no creds → the `docker login` step is skipped entirely (docker's native anonymous
flow covers public packages); creds → login/pull/logout unchanged (token still stdin-only).
- **Settings page truthfulness:** "Verzió és frissítés" gains a mode line — "Registry: nyilvános
(hitelesítés nélkül)" vs "Registry: hitelesített"; credential-less is no longer an error state; the
Hiba row appears only on a real failure. `DryRun.PullCapable` counts anonymous as capable.
- Tests (`registry_anon_test.go`, httptest fake registry + fake CLI runner): full anonymous dance with
ZERO creds (token request auth-free, correct scope, highest semver); creds path byte-shape unchanged
(BasicAuth, no dance); both denial paths (token 401 / tags-with-Bearer 401) → the new clear error;
WWW-Authenticate parser table; pull with no creds → no login invocation recorded, pull still invoked;
creds → login/pull/logout order + stdin token; partial creds refuse everywhere. **Red-proof:** old
creds-required guard restored → all three anonymous tests FAIL with
"registry hitelesítő adatok hiányoznak" visible. Restored green.
- Pairs with hub v0.43.1 (Git Sync form hint: "Opcionális — csak privát alkalmazás-katalógushoz…").
### v0.111.0 — remote app-log diagnostics: error context + on-demand log tails (2026-07-10)
Extends the app-telemetry pipeline with what the live Peti support session lacked: readable error
context and a way to pull an app's logs WITHOUT any access to the customer box. Pairs with hub v0.43.0.
- **Error context (Part B) — `internal/metrics`:** the log scraper now attaches `LogIssue.Context` —
up to ±5 raw lines around the FIRST occurrence of each error-severity issue in the scrape window
(never on repeats; warns carry none). Caps: ≤11 lines, ≤400 chars/line (`…`), and a HARD 16KB
per-report budget enforced in `internal/report` (context dropped from the lowest-count issues
first). The scan loop was extracted into the pure `analyzeLogLines` (first unit tests for the
scanner). Additive `context` field on the report's `issues` — old hubs ignore it.
- **Sanitization (Part E) — `metrics.RedactLine`:** authoritative controller-side redaction applied
to every context + tail line before it leaves the box: `password|passwd|secret|token|api[_-]?key|
authorization|bearer` values → `[REDACTED]` (incl. `Authorization: Bearer <tok>` in one pass) +
64-hex strings → `[REDACTED-HEX64]` (repo-password shape).
- **On-demand log tails (Part D) — pull-based, ACK-flag pattern (same as escrow/config-refresh):**
the report ACK gains `log_tail_requests: [app…]`; the NEXT report ships
`log_tails: [{app, collected_at, lines[]}]` — 200 lines via the existing plumbing
(`stacks.GetLogs` compose-logs for stacks, scanner-style `docker logs` for the controller
container), ordered as emitted, ≤400 chars/line, ≤64KB/app head-truncated (newest kept),
redacted. Consume-once: drained at build; a failed push re-arms from the hub's still-pending
request. NO hub→controller push channel — the guest listens to no one.
- Tests + red-proofs (all three failed exactly as designed, then restored green): context capture
dropped → "context has 0 lines, want 11" FAIL; redaction gutted → `password=hunter2` shipped
visibly → FAIL; consume-once clear removed → "second drain = [gokapi cwa]" (tails every cycle)
→ FAIL. Plus: exact ±5 ordered window, first-occurrence-only context, warn-no-context, truncation,
budget drop order, byte-budget newest-kept, fetch-error skip, empty-ACK clears stale pending.
### v0.110.0 — offbox stale-lock self-heal (campaign C2) + crash-truthful status (C1) (2026-07-10)
Fixes the overnight campaign's HIGH finding: a crash mid-prune left a restic EXCLUSIVE lock the controller
couldn't clear, failing every subsequent offsite run until manual `restic unlock`. Root nuance from the
evidence: plain `restic unlock` (stale-only) does NOT clear it — the recreated container has a new hostname,
so restic can't verify the dead PID and won't treat the lock as stale for ~30 min.
- **C2 — `internal/backup`:** `resticStep` wraps the backup/prune/restore restic calls: on a lock error
(`repository is already locked`) it escalates to `unlock --remove-all` and **retries the step ONCE**,
justified by the ARCHITECTURAL single-writer guarantee (one controller per repo via per-customer
sub-account isolation + the in-process single-flight mutex every caller holds → no live sibling). A second
lock failure surfaces the error (never loops). Plus cheap pre-run `unlock` (stale-only) hygiene before
every run + restore. **Boundary (documented):** a DR-cloned second controller writing the same repo would
defeat the single-writer premise — operator-supervised territory.
- **C1 — `NewManager.reconcileCrashedRun`:** on startup, a persisted `LastStatus="running"` (a controller
that died mid-run) flips to `error` + the Hungarian "megszakadt futás (a vezérlő újraindult futás közben)"
— truthful after a crash; the next successful run clears it.
- Tests + red-proofs: self-heal-and-retry (A, **red-proof:** neuter the escalation → the exact campaign
failure `offbox backup rallly: exit status 1` → FAIL); persistent-lock → one `--remove-all` + one retry,
error surfaced, no loop (B); pre-run stale unlock issued every run (C); no lock → `--remove-all` never
fires (E); crash-status flip (D, **red-proof:** drop the flip → status lies "running" → FAIL).
### v0.109.1 — re-apply must preserve escrow custody + runtime status (live finding) (2026-07-10)
Found deploying v0.109.0: including `QuotaGB` in the bridge's descriptor hash triggered a one-time
re-apply on the demo — key-auth-first re-pinned cleanly (proven live, no password consumed) but
`ApplyOffsiteTarget` REPLACED the target with the freshly-built struct: the escrowed demo was **demoted to
pending** and its runtime status (last_run/size/snapshots) wiped — which would also false-trigger the new
staleness alert after re-confirming.
- `ApplyOffsiteTarget` now carries over the EXISTING target's `EscrowState` + runtime status fields on a
re-apply: EscrowState tracks the REPO PASSWORD's custody (preserved by `WriteOffboxSecrets`, never
rotated by this path), not the target coords; the status belongs to the runner. A fresh guest (no
existing target) still lands `pending`. **Companion red-proof:** dropped the EscrowState carry-over →
"a re-apply must NOT demote an escrowed target, got pending" → FAIL. Reverted.
- Demo repair: one manual confirm-escrow (the deprecated fallback — truthful: the same already-escrowed
password) restored `escrowed`; a manual run restored the runtime status.
### v0.109.0 — SLICE 4: soft-quota gate + usage bar + offsite report status (2026-07-09)
The shared-model soft quota (`quota_gb`) enforced controller-side (pairs with hub v0.41.0's
OffsiteChecker + freeze lever). No secrets anywhere in this slice — sizes/timestamps only.
- **`internal/settings`:** `OffboxTarget.QuotaGB` (mapped from the hub descriptor by the apply-bridge —
`OffboxEnabler` seam gains `quotaGB`; 0 = no soft limit, dedicated boxes are Hetzner-enforced) +
`RepoSizeBytes` (machine-readable size persisted from `restic stats` alongside the human string; a failed
stats call keeps the last-known value — stale-but-safe).
- **Pre-run soft-quota gate (`RunOffboxBackup`):** at **≥100%** NEW backup runs are refused —
`LastStatus="error"` with the Hungarian notice ("A NAS-mentés túllépte a tárhelykeretet (X/Y GB) — törölj
régi mentéseket vagy kérj nagyobb keretet."), operator alert via the existing `offboxNotify` path — but
the **retention/prune step STILL RUNS** (`offboxPruneOnly`; pruning is the customer's only way back under
quota — gating it would deadlock them) and **restore is never gated**. **Companion red-proof:** gated the
prune too → "prune MUST still run over quota, got 0" → test FAILED. Reverted. The gate is pre-run: a run
crossing 100% mid-flight finishes; the next refuses. At **≥80%** (<100%) an OK run sets the Hungarian
usage `LastWarning` ("A NAS-mentés a keret X%-át használja (A/B GB).").
- **UI:** `/backups` gains a soft-quota usage bar (used/quota + %, green/amber/red) — rendered only when
`QuotaGB > 0`. Template gates green.
- **`internal/report`:** the hub report gains `offsite:{enabled, escrow_state, last_run, last_status,
snapshot_count, repo_size_bytes, quota_gb}` (`backup.OffboxReportStatus`; absent when no offbox target —
the hub checker is nil-safe on old controllers).
- Tests: over-quota refusal (backup 0 calls, prune 1 call, Hungarian status, restore ungated); 84% warn +
bytes persisted; quota-0 no gate; report object + nil when unconfigured; bridge quota mapping.
### v0.108.0 — SLICE 3: hub-verified escrow auto-confirm (current-password hash match) (2026-07-09)
Replaces operator trust with a verified fact (pairs with agent v0.79.0 + hub v0.40.0): the report ACK now
carries `escrow:{identity_blob_present, restic_pw_sha256, created_at}` and the controller flips offbox
`EscrowState` pending→escrowed ONLY when `sha256(local repo_password) == restic_pw_sha256` — i.e. the
stored escrow provably covers the CURRENT key, not merely "a blob exists" (a stale blob would re-open the
un-recoverable-ciphertext gap fork-4 closed).
- **`internal/report`:** `PushResponse.Escrow` + `EscrowAutoConfirmer` (long-lived; runs on every ACK):
match → flip (`UpdateOffboxStatus`) + wipe the agent-staged secret (the v0.107.0 DELETE path, best-effort
loud); mismatch → stays pending + a LOUD warn naming the fix ("run the escrow ceremony"), **deduped per
distinct hash** (not per 15-min cycle); no row / NULL hash / hash-without-identity-blob / no local
password file → stays pending silently (fail-closed); non-pending → total no-op (**never un-confirms**).
**Companion red-proof:** modeled the blob-present-only check → the stale-blob and hash-less scenarios
flipped when they must not → tests FAILED. Reverted — hash-match is the load-bearing core.
- **`internal/backup`:** `HashResticPassword` (canonical: sha256 hex over the TRIMMED string — **pinned
cross-repo test vector**, same vector asserted in felhom-agent) + `Manager.OffboxRepoPasswordHash`.
- **`internal/web`:** the manual `POST /backup/offbox/confirm-escrow` is now a documented **deprecated
fallback** for legacy hash-less blobs (e.g. the demo's) — auto-confirm is primary.
- Hashes are safe to log (non-reversible over a 256-bit random secret); passwords never appear in logs.
### v0.107.0 — offsite hardening: key-auth-first bridge + staged-secret wipe on confirm (2026-07-09)
Part of the offsite-provisioning hardening bundle (pairs with hub v0.39.0 + agent v0.78.0).
- **Key-auth-first (`internal/offsiteapply`):** new `KeyAuthProber` seam (`SFTPKeyAuthProber` — probes the
ALREADY-INSTALLED key against the descriptor target, pinned to the freshly-verified known_hosts). On a
descriptor change where the existing key still authenticates, the bridge **re-pins + reconfigures WITHOUT
consuming a one-time password** — kills the stale-descriptor consume-404 loop seen twice in the live e2e,
and shrinks the re-issue blast radius to genuinely-fresh guests. The probe NEVER bypasses the fingerprint
verify (scan+verify still precedes it; a mismatch refuses before any probe). Fresh guests (no key / auth
refused) fall through to the full verify→consume→install path unchanged.
Tests + red-proofs: probe-success with a panicking consumer (drop the skip → panic → FAIL); fresh-guest
fallthrough (early-return on probe-fail → nothing applies → FAIL); mismatch now also asserts the probe
never runs on a failed identity check.
- **Staged-secret wipe (`internal/web` + `internal/agentapi`):** `WipeStagedEscrowSecret` (DELETE
`/escrow/stage-secret`, agent ≥ v0.78.0); the confirm-escrow handler wipes the agent-staged repo password
whenever `EscrowState` flips to `escrowed` — best-effort (a wipe failure logs a loud ERROR but never fails
the confirm; re-confirm retries). Closes the fork-4 hygiene gap where a confirm without a fresh ceremony
left the staged 0600 file behind (observed live in the e2e's Option-A close). Test: confirm wipes exactly
once; a failing wipe still confirms + logs "NOT wiped".
### v0.106.1 — offsite apply-bridge: ssh-copy-id -s needs ~/.ssh to exist (live finding F3) (2026-07-09)
First supervised live apply: scan+verify passed, the one-time password was consumed, then `ssh-copy-id -s`
died **locally** — SFTP mode mktemp's its batch file under `~/.ssh`, and the container image ships without
`/root/.ssh`. The fail-safe held (loud "password is spent" signal, no marker, no offbox config) and the
password never left the box, but the install could never succeed.
- `internal/offsiteapply.SSHCopyIDInstaller`: ensure `~/.ssh` (0700) exists before running `ssh-copy-id`.
- Live diagnosis (container, no secrets): with `~/.ssh` present, the pinned single-line known_hosts +
`StrictHostKeyChecking=yes` verifies cleanly and a wrong password fails as `Permission denied` (sshpass
exit 5) — the TOCTOU-hardened pin mechanics are sound end-to-end against the real box.
### v0.106.0 — offsite provisioning SLICE 2: controller apply-bridge (2026-07-09)
Pairs with hub v0.38.0. On startup the controller reconciles the hub-served `offsite:` descriptor into a
working key-only offbox target — closing the loop to a hands-off, hub-driven offsite target. (Auto-confirm =
SLICE 3; soft-quota = SLICE 4.)
- **`internal/config`:** `OffsiteConfig` (`offsite:` section) mirroring the hub descriptor
(enabled/type/host/user/port/repo_path/quota_gb/box_type/**host_fingerprint**) — deep-merged from `controller.yaml`.
- **`internal/offsiteapply` (the apply-bridge):** `Bridge.Reconcile` — idempotent (a descriptor-hash marker
at `<dataDir>/offbox/applied_marker` prevents re-consuming a spent password) and fail-safe (any step fails
→ nothing persisted, retried next cycle). Flow: **scan + VERIFY the box host key against `host_fingerprint`
(no blind TOFU)** → generate the controller keypair → **consume the one-time password**
(`POST /api/v1/offsite/consume-password/{id}`, Bearer APIKey, single-use, never logged) → install the
pubkey (`sshpass -e ssh-copy-id -p 23 -s -f`, **pinning the scanner-verified `known_hosts` with
`StrictHostKeyChecking=yes` — no `accept-new`/TOFU on the install or verify session**, so a MITM cannot
substitute a key in the gap between the scan and the install) + verify key auth → configure the offbox target →
`EscrowState="pending"` (fork-4 enable path via `Manager.ApplyOffsiteTarget`) → persist the marker LAST.
Seams (consume/scan/keygen/install/enable) so unit tests fake all I/O. A consumed-but-failed install logs a
loud "password is spent — reset on the hub" signal.
- **`internal/backup`:** `Manager.ApplyOffsiteTarget` reuses `WriteOffboxSecrets`/`SetOffboxTarget`/
`PushOffboxPasswordForEscrow` → `EscrowState="pending"`; the escrow stage-push is best-effort (agent-down ≠ apply failure).
- **`cmd/controller`:** wires the bridge (real seams — HTTP consumer, x/crypto/ssh host-key scanner, ed25519
keygen, sshpass installer) and runs `Reconcile` async at startup (non-blocking; the config-refresh restart re-runs it).
- **`Dockerfile`:** + `sshpass`.
- Tests (faked seams): apply-end-to-end (pinned known_hosts + key + pending + marker + **pw-not-logged**);
host-key mismatch → refuse **+ companion red-proof** (drop the verify → wrong key pinned → test fails);
idempotent (marker match → no re-consume); install-fail → fail-safe **+ companion red-proof** (persist
marker early → failed apply looks done → test fails).
- **NOT yet live-applied** — the supervised end-to-end (hub provisions on the new pool box → controller
consumes + installs + configures) is the next runbook, gated on the hub's new scoped `HETZNER_TOKEN`.
### v0.105.0 — fork-4: offsite password custody hand-off + atomicity gate + DR inject + DR coord (2026-07-09)
Pairs with agent v0.77.0 to make the restic-offsite repo password recoverable at DR (rides the customer-R
escrow) and forbids an un-escrowed offsite copy from existing. Validated design: custody spike `febdc56`.
- **Hand-off** (`internal/agentapi/client.go`): `StageEscrowSecret` pushes the repo password to the agent's
`POST /escrow/stage-secret` over the authenticated pinned local-API channel (value never logged). The
enable flow (`internal/web/offbox_handlers.go`) reads the 0600 password via a new
`Manager.PushOffboxPasswordForEscrow` (the handler never sees the value) and marks `EscrowState="pending"`.
- **Atomicity gate** (`internal/backup/offbox.go`): `OffboxRunnable()`/`offboxEscrowed()` — `RunOffboxBackup`
(and thus the daily scheduler + the run handler) **refuses to run until `EscrowState=="escrowed"`**, so no
un-recoverable offsite ciphertext can exist. `OffboxConfigured()` is unchanged (config/UI still work).
New `settings.OffboxTarget.EscrowState` (`""|"pending"|"escrowed"`, additive, preserved across edits).
- **Confirm + DR inject** (`internal/web`): `POST /backup/offbox/confirm-escrow` (operator, after the escrow
ceremony) → escrowed; `POST /backup/offbox/inject-password` (DR) → `Manager.InjectOffboxPassword`
pre-places a recovered 64-hex password 0600 (tmp+rename), refusing to clobber without `force` — a
subsequent `WriteOffboxSecrets` then uses it (the pre-place seam). `/backups` shows a pending-escrow
notice + confirm button.
- **DR recipe** (`internal/report/dr_recipe.go`): `DRRecipeAppHalf.OffsiteRestic *DRResticCoord`
{host,user,port,repo_path} — coordinates ONLY (the password is escrowed, the SFTP key is regenerable);
populated from `Manager.OffboxCoord()`; clears the `_NoSecrets` regex.
- Tests: atomicity (pending blocks run; confirm enables) **+ companion red-proof** (gate disabled → runs
while pending → FAIL); DR inject pre-place honored + refuse-clobber + companion (no-inject generates a
DIFFERENT password); `OffboxCoord`; agent stage endpoint (0600 + non-secret ack + cross-guest 403 +
value-not-in-log); `DRResticCoord` no-secrets. Web: run-gate + confirm + inject endpoints.
- **NOT yet live-validated** — the supervised escrow ceremony (enable→stage→escrow-create→confirm→gated run)
is the operator-run follow-up.
### v0.104.0 — off-box unit discovery (durable, deployment-independent) + no-silent-success (2026-07-09)
Fixes the off-box mis-resolution + silent-success landmine surfaced by the Storage-Box spike and pinned by
the DIAG report (`felhom.eu/documentation/audits/SPIKE-storagebox-restic-direct-2026-07-09.md`). Root cause:
`runOffboxInternal` located each toggled app's recovery unit via `RecoveryUnitPath(AppNamespaceRoot(stack),
stack)`; `AppNamespaceRoot`→`GetAppDrivePath` reads the app's **live** `app.yaml` `HDD_PATH` and returns `""`
for a not-currently-deployed app, **silently falling back to `systemDataPath`**. So a toggled-but-undeployed
app was looked for on the wrong drive → `os.Stat` failed → skipped → the run returned `nil` → status `ok`
with 0 snapshots (no alert).
- **Discovery over inference** (`internal/backup/offbox.go`): new `discoverOffboxUnit` / `offboxCandidateNSRoots`
scan the durable storage registry — every registered *schedulable, non-decommissioned* path
(`GetSchedulableStoragePaths`) ∪ `systemDataPath`, deduped by resolved nsRoot — for `backups/primary/<app>`,
independent of live deploy state. Multiple copies of the same unit (drive churn) → the **newest by manifest
`CreatedAt`** (mtime fallback) is backed up, the stale one WARN-logged. `AppNamespaceRoot` and the primary
WRITE paths (`CaptureRecoveryUnit`/dumps) are **unchanged**.
- **No silent success** (`RunOffboxBackup`): `runOffboxInternal` now returns `(backedUp, missing, err)`.
≥1 toggled but `backedUp==0` → a **hard error** (`LastStatus="error"` + `offboxNotify` fires with a non-nil
err → operator alert). A *partial* run stays `ok` but sets a new customer-visible **`OffboxTarget.LastWarning`**
(`internal/settings/settings.go`, `last_warning,omitempty`) naming the skipped apps; rendered on `/backups`
in the `--warn` style (`internal/web/templates/backups.html`), preserved across a config edit
(`internal/web/offbox_handlers.go`).
- Tests (`internal/backup/offbox_test.go`): six non-hollow cases (A discovery-on-registered-drive, B 0/N
hard-error+alert, C partial→warning, D newest-of-two-copies, E happy path, edge 0-toggled), asserting the
exact discovered `src` + `LastStatus`/`LastError`/`LastWarning` + notify-err. **Companion red-proofs run:**
(A) reverting to `AppNamespaceRoot` resolution → 0 backups → FAIL; (B) `if false` on the 0/N promotion →
silent `ok` → FAIL; both reverted.
- **NOT yet live-validated against the Storage Box** — awaiting supervised re-provision + endpoint round-trip
(box repos/creds were torn down with the spike). Unit suite fully covers the discovery + status logic.
### v0.103.0 — F-C2-1: config loader no longer corrupts a bcrypt password_hash (silent auth bug) (2026-07-07)
Fixes campaign-2 finding **F-C2-1** (`felhom.eu/documentation/tests/CAMPAIGN-2-2026-07-07.md`).
`loadAndParse` and `LoadFromBytes` ran `os.ExpandEnv` over the **entire** YAML before parse. A bcrypt
hash (`$2a$10$…`) is full of `$word` sequences, so `ExpandEnv` silently replaced each with its (usually
empty) env value — corrupting `web.password_hash` on load (proven: `$2a$10$N9qo8uL…` → `"a0"`). A silent
auth-integrity bug.
- **Fix:** removed both `os.ExpandEnv` calls (`config.go` :234 loadAndParse, :249 LoadFromBytes) — parse
the raw bytes directly. The sanctioned, typed env path (`applyEnvOverrides` → `FELHOM_WEB_PASSWORD_HASH`,
applied after parse) is unchanged; no shipped `controller.yaml` relies on file-level `${VAR}`
interpolation (only `docker-compose.yml` uses `${DOMAIN}`, which is compose-level).
- **Behavior change:** a literal `${VAR}` in a controller.yaml value is now preserved verbatim (was
expanded). No repo config depends on the old behavior.
- Tests (`config_test.go`): bcrypt hash loads byte-identical (file + bytes paths; red-proof: pre-fix
`ExpandEnv` mangles it to `"a0"` → FAIL, demonstrated + reverted); `FELHOM_WEB_PASSWORD_HASH` override
still wins; literal `${VAR}` preserved.
### v0.102.0 — async restore family: no more proxy-timeout error page on a succeeding restore (2026-07-06)
Re-adjudicates campaign **F4** (`felhom.eu/documentation/audits/RERUN-p1p3-2026-07-06.md`): all three
restore surfaces (`/backup/restore`, `/backup/tier2/restore`, `/backup/offbox/restore`) blocked the
HTTP request until completion. Through cloudflared's hard 100s cap + traefik, a real customer got an
**error page while the restore silently succeeded**; the off-box one was worse — it bounded on
`r.Context()`, so a proxy read-timeout **canceled the SFTP restore mid-flight**.
- **Async family** (mirrors the existing `offboxRunHandler` shape): each handler fast-path refuses a
concurrent op (`IsRunning()` → "Egy mentési/visszaállítási művelet már fut."), then runs the restore
in a **background goroutine** and redirects immediately with a "Visszaállítás elindult…" flash. The
offbox restore's context moved from `r.Context()` to `context.Background()+30m` (fixes the mid-flight
cancel). The restore functions' internal single-flight acquire is unchanged.
- **Op-status surface** (`internal/backup/opstatus.go`): mutex-guarded in-memory current-op + terminal
`last{op,stack,ok,message,finished_at}`, deep-copy getter; new `GET /api/backup/restore-status`
(distinct from `/backup/status`, which proxies the agent's PBS status). In-memory, lost on restart
(same precedent as notification cooldowns).
- **UI** (`backups.html`): a progress banner polls the status every 3s — neutral while running (shown
even on a fresh page load mid-op), success on completion, red **only** on failure.
- Tests: `opstatus_test.go` (begin→running→terminal, deep-copy, failure); `async_restore_test.go`
(handler returns <500ms while the restore parks in a blocking provider + op-status transitions;
double-click refused with no second launch). Red-proof: the pre-fix synchronous handler blocks the
request indefinitely (test killed at 30s) vs <500ms async.
### v0.101.0 — campaign findings F3 (sync deadline) + F2 evidence gap (agent refusal surfacing) (2026-07-06)
From the no-mercy campaign (`felhom.eu/documentation/audits/CAMPAIGN-nomercy-2026-07-06.md`).
No behavior change for the happy paths; two robustness/diagnosability fixes.
- **F3 — sync git subprocess deadline** (`internal/sync/sync.go`): `runGitInDir` had no context,
so a hung remote parked the sync goroutine in `cmd.Run()` — the `doSync` defer never ran,
`syncing` stayed true, and every manual + periodic sync was refused with "Szinkronizálás már
folyamatban" until a controller restart. Each git command now runs under
`exec.CommandContext` with a fresh per-command `gitCmdTimeout` (120s); the deadline error names
the timeout and the masked git args (no silent hang). Debounce + failed-sync-arms-debounce
unchanged. Tests: `TestRunGitInDir_CancelledContextKillsSubprocess` (red-proof: pre-fix
`exec.Command` runs to completion → FAILs), `TestTriggerSync_FailureReleasesSyncingAndAllowsRetry`.
- **F2 evidence gap — agent refusal surfacing** (`internal/agentapi/client.go`): `EjectDisk` and
`Decommission` used `c.post`, which discards a non-2xx body — so the agent's informative refusal
(`"…decommission refused (role: system)"`) was flattened to a bare `HTTP 403` (the exact campaign
evidence). Both now use `postWithStatus` + a shared `refusalError` that carries the agent's
reason (truncated ~300, no bodies/secrets) through to the controller's Hungarian error and the
UI/API response. The generic `post` and all other callers are untouched. Tests: T-D1/T-D2
(fake-agent 403 → reason surfaced; red-proof: pre-fix `c.post` yields bare `HTTP 403` → FAILs),
T-D3 success unchanged, plus an ok:false 2xx-business-refusal case.
### v0.100.0 — one-click class-C file restore from the Tier-2 copy (2026-07-05)
TASK C2 — closes drill finding **F2** (`DRILL-appdata-restore-2026-07-04.md` §4): HDD bind-mount
user files (`appdata/<stack>`) had no customer recovery path — Tier-2 protected them nightly, but
getting deleted files back was an operator copy-back by hand.
- **Engine** (`internal/backup/tier2_restore.go`): `Manager.RestoreTier2Files(stack)` — in-place,
**additive-only** restore from the RECORDED Tier-2 copy (`CrossDriveBackup.DestinationPath`, never
a fresh `selectTier2Target`). Semantics = `rsync -a --ignore-existing`: files missing live are
copied back; existing live files are NEVER overwritten (a customer edit after the last copy wins);
nothing is EVER deleted (the `rsyncMirror --delete` trap in this direction would erase every file
created since last night — the new `rsyncRestoreMissing` copies the mirror's exec shape with the
opposite-direction flags). Single-flight with backup/restore; all refusals (no copy / LastRun
empty / copy dir gone / either drive disconnected / live drive decommissioned) happen BEFORE the
stop, with customer-readable Hungarian reasons; stop → copy → start → health; copy/restart errors
surface (F17). File count from `--itemize-changes` (`>f` lines); file names never logged at INFO.
- **Endpoint + UI**: `POST /backup/tier2/restore` (`internal/web/server.go` + `handlers.go`,
backupRestoreHandler-shaped guards) + a **"Fájlok visszaállítása"** button on the healthy Tier-2
layer row (`templates/backups.html`; hidden when unconfigured / never ran / target drive
disconnected/inactive) with a confirm dialog stating the additive-only contract + last-copy time.
Zero files copied = success ("Nincs hiányzó fájl — minden fájl megvan a helyén."), not an error.
- Out of scope by design: overwrite/point-in-time restore (offbox + operator paths), per-file
selection, `recovery-unit/` (backup artifacts are not user files). Apps that index their data dir
(e.g. Nextcloud) may need a rescan before restored files appear in their own UI — noted in
`felhom.eu/documentation/controller/backup-architecture.md`.
- Tests: orchestration via a `restoreFilesCopier` seam (stop→copy→start order, src/dst contract,
refusal NON-effects: never stopped, copier never invoked), Scenario-D zero-copy success, itemize
parsing, handler guards, and an FS-level semantics test of the real rsync (LookPath-skipped where
rsync is absent). Companion red-proof: swapping the flags for `rsyncMirror`'s mirrors the backup
over live — the differing live file gets clobbered AND the live-only file gets deleted (both
assertions red; verified on the build server, reverted).
### v0.99.0 — restore-path fixes: dead restore UI + volume dumps + blank-secret redeploy (2026-07-05)
TASK C1 — fixes F1/F3/O4 from the 2026-07-04 restore drill
(`felhom.eu/documentation/audits/DRILL-appdata-restore-2026-07-04.md`). F2 (one-click in-place
class-C restore) deliberately NOT included — product-design work (C2).
- **F1 (HIGH — the restore panel was dead):** `GET /api/backup/snapshots?stack=` now exists
(`internal/api/router.go` + `backup.Manager.ListRestorePoints`, `internal/backup/restore_points.go`).
The backups.html restore panel fetched this restic-era route, got the catch-all 404, so the
snapshot dropdown never populated and "Visszaállítás indítása" could never enable. Returns the
ONE honest keep-side restore point (the current recovery unit): `time` = newest artifact mtime
(manifest / db-dumps / volume-dumps), `short_id:"helyi"`, `tier:1` always (Tier-2 copies are NOT
restorable via POST /backup/restore — never listed), `drive_label` from the storage registry.
Guards: traversal/empty → 400 (`validStackParam`), unknown stack → 404, no unit yet → `ok:true, data:[]`.
No template change needed — the JS payload contract was honoured server-side.
- **F3 — named-volume data was never backed up:** `DumpAppVolumesSafe` had no production caller.
New `runVolumeDumps` loop in `runDBDumpsInternal` (`internal/backup/backup.go`), running BEFORE
`captureAllRecoveryUnits` so manifests enumerate the fresh tars. Gate order is load-bearing:
protected-stack + has-volumes checks precede the Safe call (which stops the stack before its own
check — unconditional calls would bounce every volume-less app nightly); disconnected/decommissioned
drives skip like the DB loop. Failures land in the run summary and fail the run (no silent
partial). Zero-DB early return removed (volume-only apps still get dumps + unit refresh).
Test seam: `dumpVolumesSafe` func field (F17-style).
- **O4 — missing resettable secret redeployed blank:** the restore proceed-path now generates a
replacement credential via the catalog field's `generate` spec (`stacks.Manager.GenerateSecretForField`
→ `backup.SetSecretGenerator` seam, wired in main.go), persisted encrypted through the existing
`RecreateStackFromUnit` → `SaveAppConfig` path. Data-keys are NEVER generated (gate untouched +
generator refuses `data_key` fields); values never logged. No-generator fields keep proceeding
with an upgraded "may fail to start" WARN. Residual case documented: a restored volume tar
carrying the OLD internal credential hash may still need a manual in-DB reset.
Tests: 272 → 286 top-level test funcs (+14: api snapshots ×3, backup restore-points ×4, volume-dump gating ×3, backup secret-gen ×3, stacks secret-gen ×1);
all three fixes companion-red-proofed (hollow `[]` endpoint / removed volume gate / no-generation
each fail their test). Full `go build && go vet && go test ./...` green.
### docs — CLAUDE.md refresh: slim-down to stable orientation (2026-07-03)
No code change, no version bump. CLAUDE.md 338 → ~160 lines: full 30-package layout map (was 7);
stale bare-metal `/opt/docker` deploy steps replaced with the verified 9201 bootstrap deploy
(`/etc/felhom-controller-image` + `felhom-controller-bootstrap.service`); embedded hub build section
deleted (points to felhom.eu); deep runbooks/design/testing content moved to the new skills
(`felhom-build-deploy`, `felhom-ui-design`, `felhom-testing` — source `felhom.eu/skills/`); "Key
patterns"/"lessons" pruned to session-critical invariants (rest live in REUSE.md). Standing rule
adopted: CLAUDE.md carries no version-pinned current state — that lives in CONTEXT/CHANGELOG/REUSE.
### docs — REUSE.md introduced (2026-07-03)
Cross-repo reuse-map rollout (docs-only, no code change, no version bump). New `REUSE.md` at the
repo root: curated map of canonical helpers (62 rows), patterns, dangerous lookalikes (rsyncMirror
`--delete`, raw os.RemoveAll on drive paths, fresh agentapi.New per request…), test seams, extension
points, and observed duplication (12 clusters, NOT fixed). Every entry code-verified at file+symbol;
cited paths machine-checked by `felhom.eu/scripts/reuse_refs_check.py` (green). CLAUDE.md gains the
"See REUSE.md before writing new code" pointer + the same-commit maintenance rule.
### v0.98.3 — hide "Eltávolítás a listából" on wizard-enrolled drives (2026-07-02)
User feedback follow-up: list-removal (registry-entry delete; data + mount untouched) is only
meaningful as the undo of a MANUAL path add. On a wizard-enrolled drive (/mnt/felhom-drives/) the
resulting de-registered-but-still-agent-bound limbo is never what the customer wants — its real
lifecycle is Biztonságos leválasztás / Végleges leszerelés. New `StoragePathView.IsEnrolled`
(path-prefix check) gates the button; manually added paths keep it, and the decommissioned-branch
"Eltávolítás a rendszerből" (final cleanup) is unchanged. Endpoint untouched.
Test: enrolled card must not render the remove form, manual card must (TestListRemovalHiddenForEnrolledDrives).
### v0.98.2 — drive-card action clarity: dedupe + self-documenting labels (2026-07-02)
User feedback: two "Leválasztás" buttons per drive, and four near-synonymous labels (Letiltás /
Leválasztás / Eltávolítás / Leszerelés) for very different operations. storage.html only:
- **Dedupe:** the agent-level eject no longer renders on a card that already offers the registry
safe-disconnect (enrichCard passes `hasSafeDisconnect`, detected from the card's own
storageDisconnect button) — one detach affordance per card. Non-USB registered drives (no
safe-disconnect) and unregistered drives keep the agent eject.
- **Labels + tooltips (endpoints unchanged):** Letiltás/Engedélyezés → "Új telepítések
letiltása/engedélyezése"; registry Leválasztás → "Biztonságos leválasztás" (apps stop, safe
unplug, reconnectable); Eltávolítás → "Eltávolítás a listából" (registry-entry removal only,
data untouched); Leszerelés → "Végleges leszerelés" (permanent, optional migrate-first); agent
Törlés… → "Formázás…" (that's what it does). Every action button carries an explanatory `title`.
### v0.98.1 — drive-card spacing: enrichment rows no longer touch (2026-07-02)
User feedback: on the Meghajtók cards the agent tag row ("Felhasználói adat", "lassú", uuid) and the
agent action row ("Leválasztás", "Törlés…") rendered with zero vertical gap. The
`.drive-agent-extra` slot is now a flex column with a .6rem gap (+ .6rem top margin, hidden when
empty); the inline margin in `enrichCard` dropped in favor of the slot styles; metarow horizontal
gap tightened to .75rem. CSS + one JS-string line only (storage.html, style.css).
### v0.98.0 — storage IA follow-up: Meghajtók / Hálózati tárhely subpages (2026-07-02)
User feedback on the D1 Tárhely page: the NAS-add button rendered directly next to the local-drive
enrollment buttons ("Új meghajtó inicializálása" / "Meglévő meghajtó csatolása") — two different
storage classes confusingly interleaved. The page splits into two subpages under Tárhely:
- **`/storage` — Tárhely — Meghajtók** (storage.html): physical drive registry + unified agent view
+ migrate + wizard entry points + manual add. The enrollment buttons now live unambiguously in the
local-drive context.
- **`/storage/network` — Tárhely — Hálózati tárhely (NAS)** (storage_network.html, new): the NAS
share list ("NAS-megosztások") + add form + its JS moved verbatim (incl. its own `openDialog` copy
for the remove overlay).
- **layout.html:** the Tárhely main-nav item gains two always-visible nested sub-links (Meghajtók /
Hálózati tárhely; `.nav-links-nested` CSS); the parent stays highlighted on both subpages.
- **handlers.go / server.go:** `NetworkStoragePaths` moved from `storagePageData` into the new
`networkStoragePageData` (page key `storage-network`) + `storageNetworkPageHandler`;
`GET /storage/network` route. No `/api/storage/*` change.
- Tests updated: `/storage` must NOT render the NAS section, `/storage/network` renders it and
nothing drive-related; the section inventory + no-native-confirm scans cover the new template.
Both template gates green; `go build/vet/test ./...` green (18 pkgs).
- Live-validated on 9201: both subpages render with correct sidebar active states; the NAS add-form
toggle + `nsToggleSmb` + `openDialog` exercised on the new page.
### v0.97.0 — TASK-D1: settings split + Tárhely page (unified drive view) (2026-07-02)
The 1451-line settings monolith becomes four pages; storage is promoted to a first-class main-nav
page with a unified (registry + agent) drive view; native browser dialogs migrate to the overlay
pattern. **IA/appearance only** — no `/api/storage/*` payload, storage semantics, or agent-client
change. Four commits (d50a919, f8e18a9, cb6f04c, 622d932).
- **Routing (server.go):** new `GET /storage` (Tárhely), `GET /settings/notifications` (GET→page /
POST→save split on the same path), `GET /settings/security`; the enrollment wizards move to
`/storage/init` + `/storage/attach`, with **301** permanent redirects from the old
`/settings/storage/{init,attach}`.
- **Data builders (handlers.go):** `settingsData()` decomposed into `settingsBaseData` +
`systemPageData` / `storagePageData` / `notificationsPageData` / `securityPageData`; each GET
handler and each error-re-rendering POST handler uses exactly its page's builder + template. All
five storage action redirects now land on `/storage?storage_msg=…`.
- **Template split:** `settings.html` deleted; sections moved verbatim into `settings_system.html`
(Rendszer konfiguráció, Verzió és frissítés, Vezérlő/Kiszolgáló újraindítása),
`settings_notifications.html` (Értesítések, Alkalmazás-email), `settings_security.html` (Jelszó
módosítás, Földrajzi korlátozás, Vészhelyzeti információk — misspelled heading + section copy
accents fixed), and `storage.html`. The NAS + migrate sections (previously nested inside
`{{if .StoragePaths}}` and invisible with zero drives) are now unconditional on `/storage`.
- **Sidebar (layout.html):** Tárhely main-nav item (hard-drive icon) + a "Beállítások" group with
Rendszer / Értesítések / Biztonság és hozzáférés sub-links (active state per page); orphaned
`.sidebar-settings-link` CSS deleted, `.nav-group-label` / `.nav-links-sub` added.
- **Unified drive view (storage.html):** registry cards render server-side as before; the agent
`/api/disks` list ENRICHES each connected user-data card in place (role tag via `i-lock`, drive
class, durable-id mono line, agent-only register/eject/wipe actions) joined on mount path — one
card per drive. Two extra groups: **Rendszermeghajtók** (system/backup, read-only, lock tag, no
actions) and **Nem regisztrált meghajtók** (register action only). Agent-down → one warn note
(`Az ügynök nem elérhető…`), all registry cards still render (graceful degradation). Agent-view
helpers emit design-system `.tag` markup (no `.badge`); the 🔒 emoji is gone.
- **Overlay migration:** every native `confirm()`/`prompt()` on the four pages routes through a
light `.confirm-overlay` dialog (`openDialog`; texts verbatim) — storage remove forms,
netStorageRemove, storageMigrateAll, storageDisconnect, storageDecommission (migrate + the
type-to-confirm anyway branch preserved like-for-like), storageReEnroll, triggerUpdate,
controller/server restart, and the two geo Hungary-removal confirms. One froze a browser tab
during D0 validation; none remain.
- **D0 leftovers:** the D0 grep gate false-negatived multibyte emoji on Windows — a Python
codepoint gate (`scripts/emoji_gate.py`) found and removed **8** survivors (📁🔄🔒📦 across
backups/debug/deploy/storage; the ★ default-marker → „(alapértelmezett)"). Orphaned
`.badge-lock`/`.lock-ico` CSS deleted (grep-zero first).
- **Gates & tests (+8):** `scripts/template_id_gate.py` (JS element-ID integrity — every
`getElementById`/`querySelector('#…')` resolves in its own template; red-proven by a misplaced
function), `scripts/emoji_gate.py` (0), Go tests for the four-page render + cross-leak, the h3
section inventory (all 11 old headings accounted for), 301s, storage redirect + flash, password
inline re-render, no-native-confirm scan, agent-down warn-note, and codepoint emoji scan.
Redirect + inventory tests red-proven against pre-split code. `go build/vet/test ./...` green.
- Live-validated on 9201 via claude-in-chrome: all four pages + the 301 redirect, the unified view
(3 enriched cards with role tags + durable-ids, Rendszermeghajtók group read-only with 0 action
buttons), the Leválasztás overlay opened + cancelled (drive untouched), and a full label-rename
round-trip through the real UI (flash on /storage, renamed back). NOT live-validated: agent-down
degradation (static/unit only — the agent must not be stopped on the live host); destructive
storage ops (endpoints unchanged; the moved UI paths await a supervised session).
### v0.96.0 — TASK-D0: design system v2 re-skin (appearance only) (2026-07-02)
Full customer-UI re-skin to the approved Felhom design system v2 — navy token palette, exception-based
status color, vendored fonts/icons, flat metadata. **Appearance only:** no route/handler/IA changes;
every page keeps its URL, sections, forms and behavior. Canonical reference:
`felhom.eu/documentation/design/design-system.md`. Four commits (b073cc4, 5dc277f, f100cef, 7df061c)
+ a bug-fix (4906524).
- **Vendored assets (`internal/web/static/fonts/`, `templates/icons.html`, `embed.go`, `server.go`):**
Plus Jakarta Sans + JetBrains Mono as variable woff2 (latin + latin-ext — ő/ű), embedded and served
from `/static/fonts/` (font/woff2, immutable cache); Google Fonts `@import` removed (CDN silently
broke offline nodes). Vendored 30-icon Lucide sprite included at top of `<body>`; all emoji replaced
by sprite icons or plain text (templates AND JS-built strings).
- **Setup CSS fix (`internal/setup/handlers.go`):** `handleCSS` served a dataDir-derived filesystem
path that never exists in the container — production setup mode silently fell back to `minimalCSS`.
Now serves the embedded `web.StyleCSS()` (new accessor); minimalCSS only if the embedded read errors
(logged). `minimalCSS` retokened to v2.
- **`templates/style.css` (rewritten in place):** v2 `:root` tokens; single 2px radius; every
`box-shadow` + the bg grid overlay deleted; new components — `.meter` (3px hairline track, blue
nominal fill, neutral 70/85 ticks, warn/crit `.meter-flag` „Fogyóban a hely" / „Kritikusan kevés
hely"), `.tag` (square state chip + dot, pulse on progress, reduced-motion respected), `.metarow`,
`.panel`/`.list`/`.section-h`, boxless `.stats`, buttons (danger = crit outline until confirm),
`:focus-visible` outlines.
- **funcmap (`internal/web/funcmap.go`):** `stateColor` → `run/progress/warn/neutral/off`
(**stopped/exited is neutral, NOT red** — operator-approved exception-color change; restarting =
warn); `usageColor`/`tempColor` → `nominal/warn/crit` (thresholds unchanged); `stateLabel`
Hungarian copy untouched (byte-identity guarded by test). New `timeAgoStr` (see fix below).
- **All 19 web templates + setup templates:** bars → meters (template + JS-generated markup),
badges/pills → tags, informational pills → metarows with icons, legacy `var(--*)` names in inline
styles/JS renamed to v2 tokens, monitoring Chart.js palette (cpu `#2EA8F5`, memory `#8E7CE8`, temp
`#E0A93E`, load `#5EC4B6`; v2 tooltip/grid/tick literals), deploy 3-step progress → sprite icons,
catchall page (standalone) fully retokened with inline SVGs, login two-tone H1.
- **fix(backups) 4906524:** `OffboxTarget.LastRun` is an RFC3339 *string*; backups.html passed it to
`timeAgo` (expects `time.Time`) → GET /backups 500'd on any node where an off-box backup had ever
run. Pre-existing since v0.93.0, exposed by the D0 click-through; fixed with `timeAgoStr`.
- **Tests (+7):** §8 truth tables for stateColor/usageColor/tempColor + stateLabel guard (red-proven
vs the old funcmap), font route + `StyleCSS()` accessor, setup embedded-CSS (Scenario E, red-proven
vs the old handler). Grep gate: 34 banned patterns (old hexes, 999px, box-shadow, CDN import,
legacy class names, emoji) at **zero** in `internal/{web,setup}` (baseline: 143 hits).
- Live-validated on guest 9201 via claude-in-chrome: full click-through, no Google Fonts requests,
`document.fonts.check` true, ő/ű render in PJS latin-ext, dashboard Scenario-A assertions
(0 green fills, 0 shadows, 0 radii >2px) DOM-verified. NOT live-validated: setup wizard rendering
(unit-tested only), warn/crit meter states on real hardware (demo node healthy; unit-tested).
### v0.95.0 — enrollment wizards use the raw-device scan `/disks/candidates` (Impl-2b) (2026-07-01)
Final drive-enrollment piece: both enrollment wizards now source candidates from the agent's Impl-2a
raw-device scan instead of the `Observe()`-based `/api/disks` list — so a brand-new (non-PVE-storage)
drive is finally visible + enrollable end-to-end. The enroll flow (`runStorageInit`/`runStorageAttach`)
and the Impl-1 guarded `mkfs` are UNCHANGED; the wizards just get the right candidate list.
- **`internal/agentapi/client.go`:** `ListCandidates(ctx) (CandidatesResult, error)` → agent
`GET /disks/candidates`; types `CandidatesResult{Initialize,Attach []DiskCandidate}` +
`DiskCandidate{Device,SizeBytes,Model,FSType,DataBearing,Mountable,MountSource,DurableID}` mirroring
the agent's `candidates.go`.
- **`internal/web/agent_disk_handlers.go`:** `GET /api/disks/candidates` proxy
(`agentDiskCandidatesHandler`, copy of `agentDisksListHandler`) — passthrough, NO controller-side
filtering (the agent's unclaimed-disk filter is authoritative + fail-safe).
- **`templates/storage_init.html`:** fetch `/api/disks/candidates` → render the `initialize` list
(model/size/current-FS + a data-bearing marker); dropped the client-side "already-managed" filter
(the server list already excludes OS/enrolled/claimed disks). Data-bearing → the existing wipe-confirm.
- **`templates/storage_attach.html`:** fetch `/api/disks/candidates` → render the `attach` list
(mountable-FS disks); selecting posts the FS-bearing node + its fstype to the existing
`/api/storage/attach` (mount + bind, NO format).
- **TOCTOU:** the wizard trusts the agent's Impl-1 `Format` guard as the backstop (re-checks unclaimed at
format time), not the list's freshness — a device claimed between scan and enroll is refused.
- Tests: `agentapi` `TestListCandidates` + `_Error`. `go build/vet/test ./...` clean. Live end-to-end
raw enrollment of `/dev/sdd` validated through the real UI (see REPORT).
### v0.94.0 — pull-based config-refresh (re-pull controller.yaml + self-restart on a config change) (2026-06-30)
Config delivery is now pull-based, riding the report ACK exactly like the Phase 2 version floor — the hub
never connects into the box. This replaces the hub's retired "Push Config" (companion hub change v0.26.0)
and is the mechanism by which an operator config edit reaches a running box.
- **`internal/report/pusher.go`:** `PushResponse` gains `ConfigVersion int` (`json:"config_version"`).
0 = the hub didn't advertise it (old hub / report-only customer) → no action.
- **`internal/report/config_refresh.go` (NEW) — `ConfigRefresher.Reconcile`.** The testable reconcile
(all side effects injected): on a config_version change vs. the last-applied version, **Refresh** (re-pull
`controller.yaml`) → **Record** → **Restart**. First-ever ACK (nothing recorded) records the baseline
WITHOUT restarting (the first-boot pull already has the current config); an unchanged version is a no-op
(no restart storm); a failed pull keeps the current config and does NOT record/restart (retried next
cycle); record-before-restart so the restarted process sees it applied and doesn't loop.
- **`internal/bootstrap/bootstrap.go` — `RefreshConfig`.** Re-pulls `controller.yaml` from the hub and
rewrites it, re-merging `local_api` from the same read-only `bootstrap.json` mount (no secret stashed
elsewhere). Reuses the existing `pullWithRetry`/`mergeLocalAPI`/`writeFileAtomic`. Overwrites
`controller.yaml` (hub = source of truth); NEVER touches `settings.json`; fail-safe (any failure leaves
the current config unchanged + returns an error). NOT first-boot-gated (unlike `MaybeIngest`).
- **`internal/settings/settings.go`:** `applied_config_version` + `GetAppliedConfigVersion` /
`SetAppliedConfigVersion` (persisted so the version survives the restart).
- **`internal/api/selfrestart.go`:** exported `GracefulSelfRestart` (the unexported one now calls it) so
the main.go reconcile reuses the one graceful-restart mechanism instead of reinventing an `os.Exit`.
- **`cmd/controller/main.go`:** wires the reconcile into `OnPushResponse` beside the floor reconcile —
same report cycle, no new timer, no agent involvement. The first-boot `MaybeIngest` never-clobber is
untouched (the refresh is a separate explicit re-pull).
- Tests: `Reconcile` (change→refresh+record+restart; **same-version no-op RED-PROOF**; baseline-no-restart;
failed-pull no-record/no-restart; zero-version no-op; record-fail skips restart) + `RefreshConfig`
(re-pull overwrites + re-merges local_api; failed pull leaves config unchanged; absent bootstrap errors
without writing). `go build/vet/test ./...` green.
### v0.93.0 — NAS Part B: off-box backup target (restic-over-SFTP) (2026-06-30)
Closes the NAS arc: back the app-data tier (each off-box app's recovery unit + DB dumps + volume tars) up
to the customer's NAS as an **encrypted restic repo over SFTP** — the "1 off-site" leg of 3-2-1, distinct
from the local cross-drive rsync copy and the agent's PBS whole-CT DR. No kernel mount; restic talks SFTP
to the NAS directly. Spike-validated (SPIKE-nas-storage Q8).
- **`Dockerfile`:** restic was dropped when cross-drive migrated restic→rsync — re-added `restic` +
`openssh-client` (restic's sftp backend shells out to `ssh`); version pinned by the Debian release.
- **`internal/backup/offbox.go` (NEW):** the restic-SFTP backend + orchestration.
- **Fail-fast (the load-bearing spike Q8 lesson):** every restic call carries
`-o sftp.args=…-oConnectTimeout=10…` so a dead NAS errors in ~10 s instead of a multi-minute TCP hang.
Also `-oStrictHostKeyChecking=yes -oUserKnownHostsFile=<pinned>` (no blind TOFU) + `-oBatchMode=yes`.
- init-if-absent (idempotent — a present repo is reused, never re-init), per-app `restic backup --tag`,
`forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --prune` retention, single-flight (shares
`m.running`) + migration-guard, restic's **own exit code** checked (never pipe-swallowed), restore via
`restic restore latest --tag <app> --target <scratch>` (non-destructive).
- **Secrets:** the SSH private key + the auto-generated repo password are 0600 files in the data dir —
never logged, never in a committed/non-0600 file; the repo is encrypted so the NAS sees only ciphertext.
They ride DR via the PBS whole-CT snapshot of the rootfs (the data dir), so a rebuilt box can reach the
off-box repo (the recovery-unit/dr-recipe stay secret-free).
- **`internal/settings/settings.go`:** `OffboxTarget` (host/port/user/repo path/schedule + runtime status,
no secrets) + per-app `AppBackupPrefs.Offbox` toggle + helpers.
- **`cmd/controller/main.go`:** daily `offbox-backup` schedule (04:15) gated on enabled+configured; a
failure (incl. fail-fast dead-NAS) alerts the operator via the allowlisted `backup_failed` event.
- **UI (`backups.html`):** "Külső (NAS) mentés" section — target config (host/port/user/repo + out-of-band
SSH key + known_hosts textareas), status (last run / repo size / snapshots), per-app toggles, run-now,
restore-to-scratch.
- Tests: ConnectTimeout present in base args + **the fail-fast companion red-proof** (a fake SSH transport
hangs to the ctx deadline WITHOUT the arg, fails fast WITH it); dead-NAS run fails fast + alerts + records
status; restore round-trip byte-identical (SFTP-shaped seam); single-flight skip; repo idempotency;
secrets are 0600. `go build/vet/test ./...` green.
### v0.92.0 — NAS network storage Part A2: registry + UI + per-share health (2026-06-30)
The controller side of NAS network storage, proxying to the validated agent foundation (felhom-agent
v0.50.0 `/netstorage/*`). An operator can add a customer's NAS share and point a media app at it — all via
the UI. A NAS is a **distinct storage kind** (NOT a drive): no enroll/eject/decommission/migrate/wipe/SMART.
- **`internal/agentapi/client.go`:** `AddNetStorage`/`ListNetStorage`/`RemoveNetStorage` + the
`NetworkMountStatus` mirror (health `ok|idle|unreachable`; `idle` is benign, only `unreachable` degraded).
The SMB credential is passed STRAIGHT THROUGH to the agent (which writes the 0600 file) — **never persisted
by the controller**.
- **`internal/settings/settings.go`:** `StoragePath.Kind` discriminator (`""`/`drive` | `network`) + the
network descriptors (Protocol/Server/Export/MappedUID/MappedGID — **no password**); `IsNetwork()` +
`IsNetworkStoragePath()`; `NetworkMountRoot`.
- **`internal/web/netstorage_handlers.go` (NEW):** `POST /api/storage/netstorage/{add,remove}` +
`GET /api/storage/netstorage`; registers/deregisters a Kind=network `StoragePath`; merges the agent's live
per-share health for the UI.
- **Kind-gating (the safety centerpiece):** `refuseNetworkLifecycle` blocks the drive ops
(eject/decommission/migrate/wipe) on a network path server-side; the drive-absent **gate**
(`planDriveGates`) and the **missing-storage** surface now SKIP network paths — so an `unreachable` NAS is a
**recoverable warning**, never the drive "missing → stopped" cascade (Scenario C). `networkStorageWarnings`
drives a distinct "Hálózati tárhely nem elérhető" app-card badge.
- **UI (`settings.html`):** a "Hálózati tárhely (NAS)" section — add form (NFS/SMB, server/export, uid,
SMB creds), per-share health badges, remove. Network shares are auto-selectable as a media app's `HDD_PATH`
(they register Schedulable). Hungarian, minimal emoji.
- Tests: agentapi round-trip (creds forwarded, health states); registry Kind-gate **companion** (drive
lifecycle refuses a network path; a drive path is not gate-refused); `unreachable`≠`missing` **companion**
(the drive gate stops an absent drive but NOT a network path). `go build/vet/test ./...` green.
### v0.91.0 — F2: alert on born/persistent-down channel (not only transitions) (2026-06-29)
- **What:** closes F2 from the full-stack testrun — a channel failure present at **startup/reseed**
(e.g. the controller boots right after a leaf regen → first observation is `pin_mismatch`) was
dashboard-only, **no operator email, forever**. Now a born-down non-transient reason alerts on cycle 1.
- **`internal/channelhealth/checker.go`:** added an `alerted` flag (have we emitted for the CURRENT
down-spell?). A confirmed down that comes from up/unseeded OR changes reason re-arms (`alerted=false`)
then emits once; a steady down that already alerted does not re-fire; recovery (up) re-arms. Removed
the `prev==""` silent-seed-for-down branch (a born-down IS a real down-spell). Debounce stays intact:
a **transient** born-down (refused) still needs N≥2 (the cold-boot agent-not-yet-up race), and a
**healthy** first-obs still seeds silently.
- Tests: F2 born-down non-transient **red-proof** (one alert cycle 1) + companion showing the old
seed-silent path would not have alerted; born-down transient still debounced; recovery re-arms the
spell. Version `0.90.0 → 0.91.0`.
### v0.90.0 — Controller→agent channel health-check (periodic probe + classified operator alert) (2026-06-29)
- **What:** the next self-health slice — a ~60s scheduler job that proves the controller↔agent
local-API channel, classifies failures, and alerts the operator + dashboard on a state change.
Closes the gap the R1 pin-mismatch incident exposed (the channel was only checked once at startup and
only logged). Spike-proven: `felhom.eu/documentation/audits/SPIKE-controller-agent-channel-health-2026-06-29.md`.
- **`internal/channelhealth` (NEW):** a `Checker` over two seams — a `Probe` (the channel call) and a
`Sink` (dashboard + operator notify). Each run classifies the result into `up | down:<reason>` by the
spike Q1 error-substring map (`pin_mismatch` / `unauthorized` / `unreachable` / `timeout` /
`misconfigured` / `construction_error` / `unknown`). **Debounce:** transient reasons (refused/timeout)
require **N≥2 consecutive** down-probes before alerting — so the ~1s agent-restart socket gap (spike
Q2) does NOT page; pin/401/DNS/construction alert on the first down observation. First scheduler
observation **seeds** state without notifying (mirrors host_staleness/host_capability). A
construction error (`agentClient()` can't build — latches via `sync.Once`) is surfaced **distinctly**.
- **Probe via the PRODUCTION memoized client (`Server.ProbeAgentChannel`):** GET /storage through
`s.agentClient()` — NOT a fresh `agentapi.New` per probe. Per the spike, the memoized client
self-heals after an agent restart, reflects exactly what the disk UI sees (zero divergence), and
avoids the per-call transport leak the singleton fixed.
- **Operator alert + dashboard:** `Notifier.NotifyAgentChannelDown/Recovered` relay an **English,
operator-only** event (the customer can't act on "the agent re-keyed"; the event type isn't a
customer toggle, same as the host_* events) on a **transition** (up→down, down→up, or reason-change),
with the existing hub 1h cooldown. `AlertManager.SetAgentChannelAlert` shows a short Hungarian
banner whenever the channel is down (state-based, idempotent — a born-down channel shows even though
it's seeded silently). **No customer email.** No agent or hub change (the hub relays the new event
types generically).
- Tests: classifier per-reason, transitions (up→down once, no duplicate, recovered, reason-change
re-alerts), first-obs seed, and the **debounce red-proof** (one refused → no alert; two consecutive →
exactly one). Version `0.89.0 → 0.90.0`.
### v0.89.0 — App-email: plaintext-only listener (:2526) + split-From mapping (2026-06-29)
- **What:** closes the two relay gaps from `FINDING-app-email-rollout-2026-06-29.md` so the
opportunistic-STARTTLS clients (cal.com, nextcloud) can use the relay.
- **Gap 1 — `internal/mailrelay/server.go`:** a **third listener `:2526`** that is plaintext and does **NOT
advertise STARTTLS** (`TLSConfig` left nil ⇒ go-smtp omits the STARTTLS capability from EHLO). Clients that
opportunistically upgrade to STARTTLS and then validate the cert with no skip-verify knob (Nodemailer/Symfony
Mailer) never attempt TLS against it. Accepted posture: plaintext on the single-tenant app Docker bridge only
(never host/internet). `:2525` (STARTTLS) and `:2465` (implicit-TLS) unchanged. New config
`mail_relay.plain_no_tls_listen` (default `:2526`).
- **Gap 2 — `internal/stacks/metadata.go` + `mailenv.go`:** `SMTPMapping` gains **`tls_mode`** (`""`/`starttls`
→2525 default; `plaintext`→2526; `implicit-tls`→2465 — `smtpEnv` now picks the port from it instead of the
hardcoded 2525) and **`from_domain_var`** (split-From: when set, inject `FromVar=<local>` +
`FromDomainVar=<domain>` separately, for nextcloud's `MAIL_FROM_ADDRESS`+`MAIL_DOMAIN`; unset = the current
`<local>@<domain>`).
- **No regression:** default `tls_mode` keeps vaultwarden/gitea/rallly on 2525 and mealie's plaintext path
unchanged; the hub is untouched (it relays whatever raw MIME the shim sends).
- **Tests:** `smtpEnv` port-by-tls_mode (+ companion that plaintext≠starttls port), split-From (+ companion
single-From), and the `:2526` listener has `TLSConfig==nil` & a real EHLO showing it does NOT advertise
STARTTLS while `:2525` does.
### v0.88.0 — App-email SMTP relay: in-process shim + per-app injection (2026-06-29)
- **What:** deployed apps can now send outbound email (password resets, invites, confirmations) through one
managed path — **app → in-controller SMTP shim → hub → Resend** — with the Resend key staying hub-side.
Implements `SPIKE-smtp-app-relay-2026-06-28.md` (verdict READY). Architecture: **Shape 1**, the shim runs
**in-process inside the controller** (operator-confirmed), reusing the controller's existing hub client.
- **New `internal/mailrelay/`:** a `go-smtp` server with two listeners — `:2525` plaintext+STARTTLS and
`:2465` implicit-TLS (self-signed cert generated at boot, CN/SAN = the shim service name). Advertises AUTH
PLAIN+LOGIN and **accepts any credentials, ignoring them** (apps send none; some require the offer; a
~15-line LOGIN sasl server fills go-sasl's gap). `policy.go` validates the **From header** domain against
an allowlist (default `felhom.eu`) and rejects with a clean 5xx **before** any hub call. `forward.go`
POSTs the **raw MIME** to the hub `POST /api/v1/mail` with the controller's hub Bearer key — **single-shot**
(no retry, no spool in v1); the hub HTTP status maps to an SMTP reply (2xx→250, 4xx→451, 5xx→554) so the
app surfaces the real outcome. `lifecycle.go` starts/stops the shim at runtime so the global toggle takes
effect without a controller restart. Listeners bind to the app Docker network only — never host/internet.
- **Settings + injection:** new global **app-email** toggle (`settings.AppEmail{Enabled,FromName}`); new
`.felhom.yml` **`smtp_mapping`** block (renames host/port/security/from/from-name to an app's env keys, plus
fixed `extra` vars); per-app toggle persisted in `app.yaml` (`AppConfig.EmailEnabled`). The relay env is
injected at compose time in `stackEnv` (host=shim, port=2525, security/from per mapping) **only when**
global ON + per-app ON + the app has a mapping — derived each compose, never persisted. New
`config.MailRelayConfig` (listeners, shim host, From allowlist; kill-switch).
- **UI (Hungarian):** Settings page "Alkalmazás-email" card (global toggle + optional household From-name);
per-app "Email-küldés" toggle on the deployed app's config page (only for apps with `smtp_mapping`),
save → recreate the stack to apply.
- **Tests:** `mailrelay` (happy-path passthrough byte-equality, From-reject-before-forward + companion,
single-shot-on-hub-failure + companion, status mapping, LOGIN lifecycle, real-socket STARTTLS end-to-end);
`stacks` (mapping parse, both-toggles-on injection, per-app/global-off no-injection, no-mapping, Mealie-style
mapping, household From-name). New dep `github.com/emersion/go-smtp` v0.24.0 + go-sasl.
### v0.86.0 — Phase 2 managed updates: floor-driven auto-update (2026-06-27)
- **What:** the controller now honors an operator-enforced **minimum version** (FLOOR) delivered on the
hub report ACK and **auto-updates to the floor** when below it — the managed default (no customer
click). The customer "update to latest" button is unchanged (latest, opt-in); the floor is the
**auto-target**, never latest.
- **`internal/report/pusher.go`:** `PushResponse` gains `min_controller_version` + `latest_version`
(the pusher already parsed the ACK for `customer_blocked` — extended, not a new path). *(The task
pointed at `notify/notifier.go`'s response-discards, but the actual report sender is `pusher.go`,
which already had an `OnPushResponse` seam — used here.)*
- **`cmd/controller/main.go`:** the existing `OnPushResponse` callback now also calls
`updater.SetFloor(resp.MinControllerVersion)` + `updater.MaybeAutoUpdate()` — riding the existing
report cycle; **no new timer/endpoint**.
- **`internal/selfupdate/updater.go`:** `SetFloor`/`GetFloor` + `MaybeAutoUpdate()` which **reuses the
Phase 1 `performUpdate`** (in-guest pull → agent `SwapController` → rollback on failure) with the
**floor** as target (`initiatedBy="auto-floor"`). Strict no-op unless: floor set, current parses,
current < floor (at/above = nothing — does NOT chase latest), agent wired, no backup running, no swap
in flight, not already attempted this floor (in-memory + persisted-state guard = no flapping/storm),
and the floor is **pullable** (floor ≤ latest available; floor > latest → warn + do nothing).
- **UI (settings, Hungarian):** shows "Minimális verzió (üzemeltető): X" when a floor is set, and during
an auto-update surfaces the same restart-poll panel as the button (auto-polls `/api/health` on load).
- **Tests (`internal/selfupdate/floor_test.go`):** below-floor→floor (not latest); at/above→no-op;
no-floor inert; floor>latest→no chase + warn; no-flap (one swap across repeated reconciles); raised-floor
honored (Scenario C/E); dev/no-agent→no-op. **Companion red-proof (verified):** making `MaybeAutoUpdate`
always no-op fails the below-floor test → restored → green.
- **No agent change** (reuses Phase 1 swap). Live (demo 9201): dogfood-deployed 0.86.0 via the Phase 1
self-update (the exact endpoint the Settings button invokes), then global floor set to 0.87.0 → the box
auto-updated 0.86.0 → 0.87.0 with **no click** (`last_state.initiated_by="auto-floor"`, success);
at/above-floor produced **no second update** (no flap). Hub `controller_version`=0.87.0. See REPORT.md.
### v0.85.1 — version-only build (live self-update validation target) (2026-06-26)
- No code change vs v0.85.0. Pushed as the registry "latest" so the live e2e self-update path could be
validated via the real Settings button (demo `0.85.0 → 0.85.1`: in-guest pull → agent swap → reload).
### v0.85.0 — Self-update reworked: in-guest pull + agent swap (Phase 1) (2026-06-26)
- **Problem:** the self-update button was dead in the LXC architecture — `selfupdate/updater.go` drove
the old bare-metal flow (`docker compose -f /opt/docker/felhom-controller/docker-compose.yml up -d`),
a path that doesn't exist in the guest ("docker-compose.yml nem elérhető"). The stranded 0.77.0 demo
could detect 0.84 but not install it.
- **Fix (Phase 1):** the controller now **pulls** the target image in-guest (its existing registry token
via `docker login --password-stdin` → `docker pull` → `docker logout`, over the shared docker socket),
then **delegates the swap to the host agent** (`agentapi.Client.SwapController` → agent
`POST /controller/swap`). The agent — external to the controller container — rewrites
`/etc/felhom-controller-image`, restarts the bootstrap unit, verifies health, and **rolls back** if the
new controller doesn't come up. The controller never `docker rm`/recreates itself.
- **Removed** the dead compose path: `performUpdate`/`updateComposeFile`/`composePath` and the
`docker compose up -d` flow are gone. `DryRun` now reports `agent_reachable` + `pull_capable` instead of
`compose_writable`.
- Success/failure is detected on the **next boot** by the existing `VerifyStartup` (running version vs
target) — a rollback lands the previous version → "failed (version mismatch)". The UI button + poll
(`triggerUpdate`/`pollUntilBack`) are unchanged. **Latest-only** (no version picker — Phase 2).
- `agentapi`: new `SwapController` (202) + `SwapStatus`. `NewUpdater` takes an `AgentSwapper` (nil on an
un-provisioned guest → update unavailable) instead of a compose path.
- Tests (`internal/selfupdate/updater_test.go`): up-to-date → no pull/no agent; pull-fails → agent never
called; happy → pull then one `SwapController` with the right ref; no-agent → unavailable.
### v0.84.0 — Show an app's auto-generated initial login on its page (catalog-driven) (2026-06-26)
- **Problem:** some apps generate a random first-login password into a file at first boot (Crafty →
`/crafty/app/config/default-creds.txt`) instead of taking it from a deploy field. Customers had to
read the container logs to find it — the static `app_info.default_creds` hint can't carry a
per-install secret.
- **General, catalog-driven mechanism (not Crafty-specific):**
- `.felhom.yml` gains an optional `initial_credentials` block: `{file, format: json|regex|plain,
container?, username_key/password_key (json), username_pattern/password_pattern (regex), note}`.
- `internal/stacks/metadata.go`: new `InitialCredentials` struct + `Metadata.InitialCreds` (deep-copied
in `deepCopyStack`).
- `internal/stacks/initialcreds.go`: `ReadInitialCredentials(stack)` reads the file **live** from the
running container (`docker exec <c> cat <file>` — path passed as a single arg, no shell) and parses
it via the pure, unit-tested `parseInitialCreds` (json/regex/plain). Never persists the secret to
`app.yaml`; returns a non-Available result (card hidden) when the container is down / file missing /
parse fails. Container defaults to the stack's main container (`findProbeContainer`).
- `internal/web/handlers.go`: `appDetailHandler` populates `InitialCreds` for deployed apps with the
spec; `app_info.html` renders a "Kezdeti belépési adatok" card with username + masked password
(Megjelenítés/Másolás, value read from a hidden element — never inlined into JS), labelled clearly as
the **initial** password (stays valid only until the customer changes it in-app).
- Tests: `parseInitialCreds` json/regex/plain + error paths.
- **Security note:** this surfaces a live working credential on the app page — same exposure class as the
existing post-deploy password reveal and `default_creds` card. It relies on the dashboard being
auth-gated in production (the demo's public-unauth dashboard is a separate, pre-existing tracked issue).
- Paired with `app-catalog-felhom.eu` adding the `initial_credentials` block to crafty-controller.
### v0.83.0 — Traefik scoped serversTransport for self-signed HTTPS backends (fixes crafty 502) (2026-06-26)
- **Problem:** the crafty-controller healthcheck fix (catalog `68ce009`) un-withheld its Traefik route,
exposing a pre-existing 502 — Traefik proxied **HTTP** to Crafty's **HTTPS-only** self-signed backend
on `:8443`. Crafty is the first/only catalog app with an HTTPS backend; all others serve plain HTTP, so
Traefik's default HTTP transport works for them.
- **Fix (scoped, Option B — verification stays ON by default):** the controller now renders a Traefik
file-provider dynamic config defining a **named** `insecure-skip-verify` serversTransport
(`http.serversTransports.insecure-skip-verify.insecureSkipVerify: true`). A service opts out of backend
TLS verification only by referencing it (`serverstransport=insecure-skip-verify@file`) — no global
`insecureSkipVerify` (the rejected Option A). `insecureSkipVerify` is not settable via Docker labels in
Traefik v3, so it must live in file/static config; the matching `scheme=https` + `@file` reference
labels go on the app (catalog repo).
- `internal/infra/infra.go`: new pure `RenderServersTransports()` + exported `ServersTransportInsecure`
constant.
- `internal/stacks/infra.go`: new `ensureServersTransports(traefikDir)` writes
`dynamic/serverstransports.yml` (0644) idempotently (write-only-on-change, like `wireController`, so
the traefik file-watcher doesn't reload each self-heal tick). Called from `EnsureBaseStack` **outside**
`ensureTraefik` (which early-returns when traefik is already running) so an established node still
materializes the file on the next self-heal tick; the watcher hot-loads it (no traefik restart).
- Rationale for skip-verify: a per-container self-signed cert has no CA to verify against and the hop
never leaves the host's internal docker bridge.
- Paired with `app-catalog-felhom.eu` adding `scheme=https` + `serverstransport=insecure-skip-verify@file`
to the crafty service. Tests: `TestServersTransports` + the YAML-parse matrix.
### v0.82.0 — FileBrowser sync no longer bounces the file UI on no-op; drop dead restic binary (2026-06-24)
- **F2 — gate the FileBrowser recreate on an actual change.** `syncFileBrowserMounts` (`internal/web/handlers.go`)
previously ran `docker compose up -d --force-recreate --remove-orphans` **unconditionally**, so every
controller restart and every storage sync force-recreated the FileBrowser container even when its
`config.yaml`/compose were byte-identical — bouncing the customer's file-access UI and contradicting the
"Vezérlő újraindítása → apps keep running" promise. Now it captures the on-disk `config.yaml`+compose
**before** the writes and re-reads the **final** content **after** them (so the integrations'
`ReapplyConfigForTarget` edits are included), and recreates only when something actually changed via the
new pure helper `fbNeedsRecreate(oldCfg,newCfg,oldCompose,newCompose)`; otherwise a plain `up -d` ensures
it's running without a bounce. The restore-mode DB reset (`sourcesChanged && resetDBOnChange` → `down -v`)
is preserved and forces `changed=true` (a reset removes the container). First-ever run (empty old files)
still recreates. Unit-tested (`filebrowser_gate_test.go` `TestFbNeedsRecreate`, incl. red-proof against the
old unconditional behaviour).
- **F1 — dropped the unused `restic` binary from the image** (`controller/Dockerfile`). The disk-tier restic
work moved to the host agent; no controller code execs the binary (the only `"restic"` references are a
backup-dir-name comparison and the `Method` config string, both unaffected). Removed the `restic` apt line
and its comment. The `ResticSchedule`/`migrateResticToRsync` config+settings paths are **untouched** (still
live in the dashboard).
### v0.81.0 — retire the drive-activation banner; add a standalone "Kiszolgáló újraindítása" button (2026-06-23)
- **Removed the obsolete drive-activation banner.** In the intermediary-mount model an enrolled drive
binds **live** into the running guest (agent `disks.go` — no `pct set -mpN`, no slot, no reboot), so
the "… meghajtó aktiválásra vár / Újraindítás most (~30 mp)" banner was a relic of the old per-drive
reboot model. It was also effectively dead since v0.78 (`pendingActivationDrives` keyed `attached` by
the agent's RAW `MountPath` but compared it to the now-STABLE `sp.Path`). Removed: the
`{{if .PendingDrives}}` banner block + `window.activatePendingDrives` JS (`settings.html`), the
`data["PendingDrives"]` feed (`handlers.go`), and the dead `pendingActivationDrives` helper +
its now-unused `internal/system` import (`storage_handlers.go`).
- **Repointed the reboot endpoint to a non-storage route.** Renamed `handleStorageActivate` →
`HandleServerReboot` and split out a testable `serverReboot` core (mirrors `runStorageInit`); removed
the `/api/storage/activate` case from `ServeStorageAPI`; mounted the handler at the new
`/api/server/reboot` (same `RequireAuth` + `CsrfProtect`) in `cmd/controller/main.go`. The agent
`GuestReboot` primitive is reused unchanged. (`/api/storage/activate` now returns 404.)
*Note:* the handler is exported (`HandleServerReboot`) because `cmd/controller/main.go` wires it
cross-package — same convention as every other web handler mounted there.
- **Added the standalone "Kiszolgáló újraindítása" settings card.** A deliberate full-server (guest)
restart affordance, a sibling to the existing "Vezérlő újraindítása" controller-only restart.
New `settings-card` + `restartServer()` JS (reuses the existing `pollRestart()` loop) in
`settings.html`; posts to `/api/server/reboot`.
- **Test:** `TestHandleServerReboot_CallsGuestReboot` (`storage_handlers_test.go`) — a fake `diskAgent`
asserts `GuestReboot` is invoked exactly once and the response is 202 `{ok:true, rebooting:true}`.
`diskAgent`/`mockAgent` gained `GuestReboot`. Green gate: `go build ./... && go vet ./... && go test ./...`.
### v0.80.0 — disk card: show + act on the stable path, not the raw host mount (2026-06-23)
- Follow-up to v0.78/0.79. The storage disk card displayed the drive's **raw** host PVE mount
(`/mnt/<name>`) — which doesn't exist inside the guest — instead of the **stable** in-guest path
(`/mnt/felhom-drives/<name>`, the `guest_path`) the registry, app `HDD_PATH`, and FileBrowser use.
- It also passed the **raw** path to the Leválasztás/Törlés buttons, so those would unmount the drive
but leave its **stable** registry entry orphaned (`RemoveStoragePath` is keyed on the stable path), and
the impact warning (`/api/storage/impact?where=`) found no affected apps (HDD_PATH is the stable path).
- Fix (`settings.html`): the card sub-line + the eject/wipe buttons now use the stable path (`regKey(d)`);
the type-to-confirm name is derived from the basename so it still matches the server check; **register**
keeps posting the raw path (its agent guest-attach operates on raw). `handleStorageWipe` now maps the
registered path to raw via `agentWhere()` for the agent eject (matching `handleStorageEject`), so the
drive deregisters cleanly. Agent-facing ops are unchanged (same raw paths); only display + the
controller's own registry bookkeeping are corrected.
### v0.79.0 — disk view: key the "registered" check on the stable path (2026-06-23)
- Follow-up to v0.78.0. The storage disk-view JS (`settings.html` `regBadge`/`actions`) decided whether
a drive was registered by looking up its **raw** `mount_path` (`/mnt/<name>`) in the registry — but
since v0.78.0 the registry correctly stores the **stable** path (`/mnt/felhom-drives/<name>`), so an
enrolled, working drive showed a spurious **"Nem regisztrált"** badge + **"Regisztrálás"** button.
- Fix: new `regKey(d)` = `d.guest_path || d.mount_path` (the agent already reports the stable
`guest_path` per disk); `regBadge`/`actions` now key on it. `registerDrive()` still posts the RAW
`mount_path` (the agent operates on raw; `handleStorageRegister` maps it to stable). Display-only.
### v0.78.0 — storage register: use the STABLE intermediary path, not the raw path (2026-06-23)
- **Bug:** `handleStorageRegister` (the "Regisztrálás" action for an already-mounted, unregistered drive)
registered the **raw** `/mnt/<name>` host path verbatim, unlike its siblings `runStorageInit`/
`runStorageAttach` which map through `stablePathForName` to the **stable** intermediary path
`/mnt/felhom-drives/<name>`. The agent binds the drive at the stable path (intermediary model), so the
controller ended up watching an empty placeholder dir on the guest **rootfs** → the drive showed as on
the system drive (**"Rendszermeghajtón"**, ~31 GB), `0 connected / N disconnected`, and the
**"… meghajtó aktiválásra vár"** banner never cleared (`registerStoragePath`→`EnsureUserdataSkeleton`
even `mkdir`'d those rootfs placeholders). Surfaced after a destroy+re-provision, where surviving host
mounts make "Regisztrálás" the natural action. Diagnosis:
`felhom.eu/documentation/audits/DIAGNOSE-drive-bind-after-reprovision-2026-06-23.md`.
- **Fix** (`internal/web/storage_handlers.go` `handleStorageRegister`): register
`stablePathForName(path.Base(req.Where))`, matching init/attach. `attachIntoGuest` still receives the
**raw** path (the agent operates on raw); the success log + JSON now report the stable path (+ raw).
Test: `TestHandleStorageRegister_RegistersStablePath` (+ red-proof). No other behavior changed.
### v0.77.0 — per-app open_path for the "Megnyitás" link (2026-06-23)
- The dashboard/deploy/app-info **"Megnyitás"** (open) button was hardcoded to the bare subdomain root
`https://{sub}.{domain}` for every app. Apps whose UI isn't at `/` (e.g. Gokapi redirects `/` away;
Ghost's admin is at `/ghost/`) opened to the wrong place.
- New optional `open_path` field on `Metadata` (`internal/stacks/metadata.go`, `.felhom.yml`) appended to
the URL in all three link sites (`dashboard.html`, `deploy.html`, `app_info.html` via `.Meta.OpenPath`).
Empty = bare root (unchanged for the other 50 apps). No handler changes (all three templates already
carry `.Meta`). Catalog: `gokapi` → `/admin`, `ghost` → `/ghost/`.
### v0.76.0 — campaign-#3 hardening: settings recovery, restore-name validation, quiesce-marker quarantine (2026-06-22)
Three controller findings from chaos campaign #3, all small, all controller-side.
- **S1 [MEDIUM] — no more crash-loop on a corrupt `settings.json`.** `internal/settings/settings.go`:
`save()` now writes a last-known-good `<path>.bak` **after** the primary rename succeeds (best-effort);
`Load()` on a JSON-parse error recovers from `.bak` (re-promotes it to primary) and, failing that,
**preserves** the corrupt file as `*.corrupt-<ts>` and starts on safe defaults — never returns the
error that made `main.go` `Fatalf`/crash-loop. New `Settings.LoadWarning` surfaced as a dashboard
banner. (`main.go`'s `Fatalf` stays — now only the genuine IO-unreadable path is fatal.) Recovery is
safe: an empty `PasswordHash` falls back to `controller.yaml`, the storage registry re-discovers.
- **F2 [MEDIUM, defense-in-depth] — validate `stack_name` against path traversal.** New
`web/validate.go` `validStackName` (single segment; rejects `/`, `\`, `..`, NUL). Gated in
`backupRestoreHandler` (`handlers.go`) and `apiExportStart` (`handler_export.go`) before any
restore/export work. (Storage `where=` was already validated by `gateWhere`.)
- **S3 [LOW] — quarantine a corrupt quiesce marker.** `quiesce/quiesce.go` `readMarker` now logs a
`[WARN]` + renames a bad-JSON marker to `*.corrupt-<ts>` instead of silently dropping it (still
returns "no marker" → no recovery, the correct contract).
- Tests: T-S1a-d (settings recovery), T-F2a-c (validation + both handlers), T-S3a/b (quarantine), all
red-proofed against the pre-fix code. Agent/hub untouched.
### v0.75.0 — gate userdata MkdirAll on a live mountpoint (no writes into an absent drive) (2026-06-22)
**Bugfix — two `MkdirAll`-into-`<drive>/userdata` sites fired without checking the drive was mounted**,
producing `mkdir …/userdata: permission denied` + transient `Created` flapping during a drive-absent
window (campaign-#2 findings #2/#3). Worse than noise: writing into an unmounted mountpoint lands app
data on the guest **rootfs**, shadowed when the drive returns (data-integrity + rootfs-fill hazard).
- `internal/stacks/manager.go` — `ensureUserdataMounts` (the deploy belt) now skips when the
`HDD_PATH` drive root is an **external** path (not `sysDataPath`) that is **not a live mountpoint**;
the app is held by `planDriveGates` instead. New injectable `Manager.isMountPoint` seam (defaults to
`system.IsMountPoint`) for testability. The system/local path is never gated (it's legitimately not a
mountpoint).
- `internal/web/handlers.go` — the FileBrowser sync loop skips (and does not mount) a registered path
under `StableParentDir` that isn't a live mountpoint, via a new pure `skipFileBrowserPath` helper.
Matches `planDriveGates`' external-only rule.
- `EnsureUserdataDir`/`EnsureUserdataSkeleton`/`planDriveGates` unchanged (gated the callers).
- Tests: `TestEnsureUserdataMounts_{SkipsAbsentExternalDrive,EnsuresWhenMounted,SystemPathNeverSkipped}`
+ `TestSkipFileBrowserPath` (both red-proofed against the pre-fix code).
- **Boot-time** occurrence (docker boot-restore starting drive-backed apps before the agent mounts the
drives) is a separate cause — documented as a design note (CONTEXT.md), not changed here.
### v0.74.0 — fix the controller→agent connection leak (per-call agentapi client) (2026-06-22)
**Bugfix — agent local-API socket leak that took down the whole agent-backed feature set after ~5 days.**
`Server.agentClient()` built a fresh `agentapi.Client` (hence a fresh bare `http.Transport` with
`IdleConnTimeout:0`) on **every** call and discarded it without closing idle connections. The agent's
keep-alive left one idle ESTABLISHED socket per call to `192.168.0.162:8443`; these accumulated
(~5.8k/day, measured 206 in 47 min) until the ephemeral source-port range for that tuple exhausted →
`connect: cannot assign requested address` (EADDRNOTAVAIL), killing storage UI, host-metrics, and
whole-guest backup. (`:8006`/pveproxy was immune — the controller never dials it.) Diagnosis:
`felhom.eu/documentation/tests/unattended-test-campaign-2026-06-22-8443-diagnosis.md`.
- `internal/web/server.go` — `Server` gains `agentCli *agentapi.Client` + `agentCliErr error` +
`agentCliOnce sync.Once` (and the `agentapi` import).
- `internal/web/agent_disk_handlers.go` — `agentClient()` now memoizes the build via `agentCliOnce`
and **reuses one shared client** (cfg.LocalAPI is static per process — a config-apply self-restarts).
The empty-endpoint "not configured" guard stays OUTSIDE the Once. All 19 call sites unchanged.
- `internal/agentapi/client.go` — `New` Transport hardened: `MaxIdleConns:4`, `MaxIdleConnsPerHost:2`,
`IdleConnTimeout:90s` (was a bare Transport, `IdleConnTimeout:0`). Added optional `Client.Close()`
(CloseIdleConnections) hygiene helper.
- Tests: `TestAgentClient_ReusesSameInstance` (+ `TestAgentClient_UnconfiguredErrors`) and
`TestNew_TransportIdlePoolBounded` — both red-proofed against the pre-fix code.
- Agent, its bridge-IP bind, and firewall rules were **not** touched (controller-only fix).
Separate open item: the defense-in-depth host firewall rule scoping `:8443` to the guest bridge
subnet is still absent (pve-firewall disabled) — to be closed independently.
### v0.73.0 — DR recipe: emit the secret-free customer+apps half in the hub report (2026-06-16)
**DR recipe slice (controller half).** Additive `dr_recipe` section on the controller's hub report — the
customer + apps half of the secret-free reconstruction recipe (`SPIKE-dr-recipe-2026-06-16.md`). The hub
assembles it with the agent's storage/guest/PBS half into one customer recipe.
- `internal/report/dr_recipe.go` — `DRRecipeAppHalf{recipe_version, customer{id,display,domain}, apps[]}`
built by the pure `BuildDRRecipeAppHalf(...)` over the DEPLOYED, non-protected stacks. Per app:
`AppRecipe{catalog_ref (Meta.Slug, falls back to name), enabled, storage_bindings[]}`. Storage bindings
are parsed from the compose (`appStorageBindings`) — each `${HDD_PATH}`/`${USERDATA_PATH}` volume bind
becomes `{container_path, drive (basename of HDD_PATH), subpath}` (e.g. romm → felhom-flash:userdata/roms);
named volumes are excluded. Wired into `BuildReport`.
- **THE BOUNDARY (the emitter is the enforcement point).** v1 ships an explicit ALLOWLIST — only the three
fields above — and the emitter reads **NOTHING** from `AppConfig.Env`, so no `ENC:` value / token /
password can ride along. Allowlist, not denylist → a new field is excluded by default.
- Tests (the load-bearing no-secrets boundary test + companion): `TestBuildAppRecipe_NoSecrets` feeds an
app whose `Env` carries synthetic secrets (an `ENC:` value + a token-shaped value) and asserts the
emitted recipe contains NONE of those values and NO credential-shaped key;
`TestBuildAppRecipe_AllowlistIsLoadBearing` is the red-proof (a guard-removed shape leaks the token, the
production emitter does not); `TestAppStorageBindings` (+ `_NoHDD`) pins the compose parse; and
`TestBuildDRRecipeAppHalf` checks assemble-correctness (deployed/non-protected only) with a whole-half
secret sweep. Red-proofed live: forcing the emitter to dump `Env` makes the boundary test fail.
`recipe_version=1`, ignore-unknown on read.
### v0.72.0 — FileBrowser converges on boot-recreate (2026-06-16)
Follow-up to v0.71.0: a host-reboot test found `processGuestBootChange` recreated the drive-backed app
stacks but **never re-synced FileBrowser**, so its drive mounts went stale after a reboot (FileBrowser
binds each drive's `userdata` but is base-infra with no `HDD_PATH`, so it is not in the recreate set).
Now, **after** `pollLiveBinds` confirms the live binds and the apps are recreated, the boot-recreate path
triggers `go s.SyncFileBrowserMounts()` so FileBrowser converges against the now-live drives (the sync
runs unconditionally so FileBrowser reflects the current bind state even if no app needed recreating).
Refactored the recreate loop into a pure, testable `recreateDriveBackedApps(stacks, present, recreate,
syncFB)`. Tests: FileBrowser sync invoked once, AFTER every recreate (red-proofed companion: pre-fix path
never synced); and synced even when nothing was recreated. Pairs with felhom-agent v0.37.0's host-reboot
remount-by-UUID fix. **Live-accepted with TWO real `felhom-pve` reboots:** both fired the boot-recreate
(`live bind confirmed — recreating … → re-syncing FileBrowser mounts → FileBrowser mounts synced — 3
storage path(s)`), all drive-backed apps recovered, FileBrowser non-stale — while the agent tolerated a
`/dev/sdb`↔`/dev/sdc` letter swap by mounting each drive by UUID.
### v0.71.0 — fix guest-reboot recovery of drive-backed apps (boot-race + the agent-path blocker) (2026-06-16)
A `pct reboot` of the guest left drive-backed apps (audiobookshelf, calibre-web, immich-server,
jellyfin, komga, radarr, romm, paperless-webserver) stuck `Exited` forever. On guest boot the in-guest
dockerd auto-starts the `unless-stopped` apps **~18s before** the agent re-binds the drive under the
stable parent; the create-time volume bind fails (`mkdir /mnt/felhom-drives/<drive>/userdata:
permission denied` on the empty fail-closed placeholder) and, being a create-time failure
(`RestartCount=0`), is **never retried**. The intended recovery (`processGuestBootChange`) did not fire.
Live diagnosis pinned **three** sub-causes, fixed together (harden the existing mechanism — no parallel
one):
1. **The agent-path blocker (the live root cause).** `agentClient()` returned **"agent not configured"**
— `cfg.LocalAPI.Endpoint` was empty — so `processGuestBootChange` (and the **entire** drive gate)
bailed at its first guard, never reaching any boot-id/bind logic. `bootstrap.json` *had* a complete
`local_api` block, but `MaybeIngest` returned immediately on "already configured" (customer.id set),
so a controller.yaml seeded before `local_api` existed never got the agent path merged. **Fix:**
`MaybeIngest` now calls new **`ensureLocalAPI`** on the already-configured path — it merges `local_api`
from bootstrap.json into the existing controller.yaml when missing (no hub re-pull, config preserved),
idempotent + fail-safe.
2. **The boot-race readiness gate.** `processGuestBootChange` sampled the agent's `BoundUnderParent`
**once** during fast startup, racing the ~18s rebind, recreated nothing, and burned its boot-id
one-shot. **Fix:** it now gates on the **REAL live in-guest bind** — new `driveBindLive` checks
whether `/mnt/felhom-drives/<drive>` is an actual mountpoint in the controller's own `/mnt` (rslave)
`/proc/self/mountinfo` (true only once the agent's bind propagated, exactly when docker can recreate
the app), and new `pollLiveBinds` **waits** for it (bounded ~120s, poll 2s) before recreating via the
normal pipeline (`compose down`→`up -d`). `shouldRecreateOnBoot` is unchanged and state-independent,
so a stuck-`Exited` create-time-failure app is included.
3. **The single-shot fragility.** `processGuestBootChange` ran only once at startup; right after a guest
reboot the agent's local API can be briefly unreachable/stale, so the one attempt bailed and never
retried. **Fix:** `driveGateLoop` now runs it on every periodic tick too — idempotent (boot-id
gated), so it retries until the agent is reachable.
Apps on a drive that never goes live in the window are left to the normal gate. The host-reboot path the
earlier sweep validated is unaffected (same code path, strictly more robust); the **guest-only reboot
path** (never exercised by host-reboot sweeps) is now covered.
Tests (non-hollow, with pre-fix companions, red-proofed): `pollLiveBinds` waits through the rebind then
reports live (recreate fires) / never-live stays absent / a single early sample misses the not-yet-live
bind; `ensureLocalAPI` merges `local_api` into an already-configured controller.yaml that lacks it
(companion: pre-fix MaybeIngest left it empty) and no-ops when already present. Live-accepted with
**guest reboots ×2 AND host (`felhom-pve`) reboots ×2** — all 8 drive-backed apps recover automatically,
zero manual starts; the persisted boot-id now advances per boot (it had been frozen at the first-boot
value). Note: the agent path worked at the v0.68 acceptance (it surfaced a real bug there) and regressed
afterward — `controller.yaml` is reset to the golden's no-`local_api` baseline on each container recreate
and the old `MaybeIngest` never re-merged it; `ensureLocalAPI` closes that. The boot-race manifests on
host reboots too (not just guest), so both paths needed this fix.
### v0.70.0 — config-apply self-restart + geo-restriction UX fixes (2026-06-16)
Fixes found during live geo testing (rotating the Cloudflare API token).
- **Config-apply now self-restarts (core fix).** `POST /api/config/apply` previously wrote the new
`controller.yaml` but logged "restart needed" and left stale in-process singletons — the Cloudflare
client is built once at startup, so a rotated CF token kept 403'ing until a manual LXC restart. Now:
if the pushed config is byte-identical to the current one, do nothing (no flap on idempotent
re-push); otherwise write, respond 200 (flushed), then **gracefully self-restart** (`os.Exit(0)` after
~500ms; the container is `restart: unless-stopped`, so it comes back with fresh config). The exit is
behind an injectable `Restarter` seam (`Router.restart`/`SetRestarter`) for unit testing. Removed the
stale "restart needed" wording and the dead `OnConfigApplied` hook (Phase-1-retired infra-backup push).
- **Manual "Vezérlő újraindítása" button** on the settings page → `POST /api/selfrestart` (auth + CSRF
via the `/api/` mount) using the same helper. Confirm dialog → POST → polls `GET /` every 2s until the
controller answers → reloads. Self-serve restart without rebooting the whole guest.
- **Immediate hub report push on geo change.** A successful geo settings save and a successful manual
geo sync now fire an out-of-band, non-blocking report push (`Router.reportPushNow`), so the hub
reflects the new geo state / clears a stale `last_sync_error` within seconds instead of after the
next ~15-min cycle. (Pattern can extend to other settings later; scoped to geo handlers for now.)
- **Always report `geo_restriction`.** `BuildReport` now always populates the field (Enabled=false,
empty countries when never configured) instead of omitting it when nil — so the hub always renders
the geo section ("Inaktív" when off) rather than hiding it.
- **Country autocomplete fixed.** Root cause (diagnosed live): `filterCountries` populated the list
correctly but revealed it with `style.display = ''`; the `.geo-country-list` CSS default is
`display:none`, so clearing the inline style kept the populated dropdown hidden — no console error,
just an invisible list. Latent since the geo feature's first commit (not the hypothesised JS throw).
Fix: reveal with `display = 'block'`.
### v0.69.0 — remove dead infra-backup stubs + the unused restic-password report field (2026-06-16)
Controller half of the Phase-1 Infra Backup retirement (hub v0.12.0; see
`felhom.eu/documentation/audits/SPIKE-infra-backup-2026-06-15.md`). Pure dead-code removal — no
behaviour change (everything removed was already caller-less).
- **Removed `report.Pusher.PushInfraBackup`** — pushed the infra-backup payload to the now-removed hub
endpoint `POST /api/v1/infra-backup`. Dead since slice 8C; no callers.
- **Removed `notify.Notifier.NotifyBackupCompleted`** (the `backup_completed` event) — no callers
since whole-guest backup moved to the agent in slice 8C. The hub's backup-deadline check now reads
the agent host-report's PBS snapshots instead of this event. `NotifyBackupFailed` and the DB-dump
notifiers are untouched and still used.
- **Removed `report.BackupReport.ResticPassword`** (`json:"restic_password"`) — the live report
builder (`buildBackupReport`) has left it empty since slice 8C, but the field historically leaked
the restic password into the hub's plaintext `reports` store. Confirmed via pushed source (builder
never sets it) **and** live data (current reports carry no `restic_password`) before removal.
### v0.68.3 — fix Beállítások page endless-refresh loop after a migration (2026-06-15)
Found while live-validating the M3 migration: once any data migration finished, the **Beállítások
(settings) page reloaded itself every ~1.5 s, forever**. The migration journal keeps returning the
last completed job indefinitely (`MigrationStatus` is not cleared on `done`); the page's resume-view
IIFE called `migWatch()` for *any* returned job, and `migWatch`'s `done` branch does
`setTimeout(location.reload, 1500)`. So every load saw the persisted `done` job → watched it →
reloaded → saw it again → looped endlessly.
Fix (settings.html): the resume-view now starts the watcher **only for an in-progress job**
(`phase !== 'done' && phase !== 'aborted'`). The one-time post-completion reload still fires from the
*active* watcher started by `storageMigrateAll`, so a real migration still refreshes drive state once
when it finishes — but a stale terminal job in the journal no longer triggers the loop. Template-only
change.
### v0.68.2 — fix stack-card state-badge clipping on unhealthy apps (CSS) (2026-06-15)
The `.stack-detail-header` is a `flex` / `space-between` row holding the `.stack-title-row` (logo +
title + subdomain link + the `route-unpublished` warning) and the `.stack-state-badge`. On an
**unhealthy** app the long "⚠ URL nem elérhető – útvonal nincs publikálva" warning inflated the
title-row; because the title-row had no `min-width:0` it refused to shrink below its content, and
because the `white-space:nowrap` badge had no `flex-shrink:0` the flexbox compressed the BADGE
instead — clipping "Nem egészséges" to "Ner…". Healthy / not-deployed cards don't render that
warning, so only unhealthy cards clipped.
- `.stack-title-row` → `flex: 1; min-width: 0;` (allowed to shrink + wrap its own content).
- `.stack-state-badge` → `flex-shrink: 0;` (never compressed).
Pure CSS; no behavior change. Browser-verified on /stacks: komga's badge now reads the full "Nem
egészséges" and the warning wraps within the title column; healthy ("Fut") and not-deployed cards
unchanged.
### v0.68.1 — boot-id recreate ALL deployed drive-backed apps (state-independent) (2026-06-15)
Fix caught live in the E1 host-reboot test: `shouldRecreateOnBoot` filtered on container state
(`State != stopped`), so apps docker hadn't auto-restarted yet at the one-shot boot-id instant were
MISSED (5 apps stayed exited after a host reboot). The boot-id recreate now recreates EVERY deployed
drive-backed app whose drive is present, independent of current state (`app.yaml` deployed = should run)
— truly deterministic. Test updated.
### v0.68.0 — storage lifecycle on the intermediary model: H2/H3/M1/M3 + deterministic boot-id (2026-06-15)
Pairs with agent v0.36.0. Finishes the storage lifecycle on the new mount model.
- **Boot-id determinism (kills the 1/8 race).** `processGuestBootChange` replaces the fragile
container-uptime sample: the agent reports `guest_boot_id` (changes per guest boot, stable across a
controller-only restart), persisted in settings (`LastGuestBootID`). On a change, every deployed
drive-backed app whose drive is present and that docker brought back (`shouldRecreateOnBoot`: state not
stopped/not_deployed) is DETERMINISTICALLY recreated onto the populated path. Respects user-stop;
gate-stopped apps stay the gate's job.
- **H2 — decommission UI button** (settings.html): a "Leszerelés" button on every connected drive →
migrate-then-decommission (uses the inline target select) OR decommission-anyway (type-to-confirm the
drive name). Both modes were already server-side; the new model never touches the parent mp.
- **H3 — one-click re-enroll/reconnect.** `handleStorageReconnect` now also handles a DECOMMISSIONED
drive: clears the soft marker + schedulable, re-attaches under the parent, restarts apps (re-discovered
via `appsOnStoragePath` since decommission-anyway doesn't persist StoppedStacks). New
"Visszacsatlakoztatás" button on decommissioned drives.
- **M1 — default reassignment.** `defaultPromotionTarget` + `finalizeDecommissionWith`: decommissioning
the DEFAULT auto-promotes another schedulable drive (preferring the migrate target); if NONE exists the
decommission is BLOCKED with a clear message (never zero default).
- **M3 — userdata setgid on migrate.** The merge-walk now RE-ASSERTS 2775-setgid/gid-1000
(`EnsureUserdataDir`) on the userdata tree (`isUserdataDir`) instead of merely preserving the source
mode — so a pre-existing stale 755 target dir (e.g. import/calibre) is corrected.
- Fix: the H1 disconnect/reconnect/restart-apps JS sent `{path}` but the handler decodes `{where}`
(always 400); response keys realigned (`restarted`). New buttons use `{where}`.
Tests (non-hollow + companions): `TestShouldRecreateOnBoot` (old sample missed a healthy-stale app),
`TestDefaultPromotionTarget`, `TestIsUserdataDir`.
### v0.67.5 — gate: startup recreate waits for stack scan + handles exited apps (2026-06-15)
Adds a bounded wait for the stack scan (GetStacks is empty at NewServer time, so the recreate found no
apps) before the one-time boot-stale recreate, and recovers exited/restarting/unhealthy drive-backed
apps (not only recently-started). The deterministic guest-reboot convergence.
Refines v0.67.3's `recreateBootStaleApps`: recreate a present drive-backed app when it is boot-stale
(recently started) OR currently `exited`/`restarting`/`unhealthy` (came up wrong on the empty bind and
bailed) — the recency-only gate missed apps that had already exited. Still skips healthy long-running
apps (no bounce on a controller-only restart) and cleanly user-stopped apps.
### v0.67.3 — gate: startup recreate of boot-stale drive-backed apps (2026-06-15)
Completes guest-reboot convergence (caught in the live migration). On a guest reboot docker auto-starts
the app containers (restart:unless-stopped) potentially BEFORE the agent re-propagates the drive under
the parent, so they bind the empty fail-closed stable dir and (leaf-bind pinning) never pick up the
later propagation. `driveGateLoop` now runs a one-time `recreateBootStaleApps` at startup (the controller
restarts with the guest): for each deployed drive-backed app whose drive is NOW present
(BoundUnderParent) and whose containers started recently (a fresh boot, not a controller-only restart —
`stackStartedRecently`), it recreates the app (down+up) onto the populated path. Apps whose drive is
still absent are left to the normal stop→return→restart gate. Paired with agent v0.35.0 (the drive
re-propagation).
### v0.67.2 — gate: key "present" on BoundUnderParent (reboot convergence) (2026-06-15)
The drive-absent gate now treats a stable path as usable only when the agent reports it BOUND UNDER THE
PARENT (`BoundUnderParent`), not merely host-mounted (`State==attached`). This makes a host reboot
converge correctly: at boot the raw drive mounts early but the agent binds it under the parent slightly
later, so until then the apps' stable-path binds are empty — the gate keeps the apps stopped and
restarts (recreates) them once the bind is live. Legacy raw paths still use the host-mount signal.
### v0.67.1 — gate: only act on external drives under /mnt/felhom-drives/ (2026-06-15)
Fix (caught live on the v0.67.0 deploy): `planDriveGates` marked the internal SSD path
`/mnt/sys_drive/felhom-data` "disconnected" because the agent never reports it as a drive — which would
have blocked starting SSD-resident apps. The gate now only considers EXTERNAL drives registered under
the stable parent `/mnt/felhom-drives/<name>`; always-present SSD/system paths are skipped. Regression
case added to `TestPlanDriveGates`. (No apps were stopped — no app depended on the SSD path.)
### v0.67.0 — intermediary-mount: HDD_PATH repoint + drive-absent gate + H1 routes (2026-06-15)
Controller half of the intermediary-mount re-architecture (pairs with agent v0.34.0). Drives are now
visible in the guest ONLY at the STABLE path `/mnt/felhom-drives/<name>` (the host swaps the backing
drive underneath it; no per-drive `pct` mp, no guest reboot).
- **Repoint** (`internal/web/intermediary.go`): the registered storage path + every app's HDD_PATH +
FileBrowser source = the stable `/mnt/felhom-drives/<name>` (`stablePathForName`); the AGENT still
operates on the raw `/mnt/<name>` host mount, so controller→agent `where` is mapped back via
`agentWhere()` at the assign/attach/eject/decommission call sites. Enroll now binds-under-the-parent
BEFORE register/skeleton (the controller can only see/write the drive at the stable path post-attach).
`agentapi.DiskInfo` gains `GuestPath` + `BoundUnderParent`. New `settings.RepointStoragePath` for the
migration. FileBrowser + monitoring follow `sp.Path` automatically.
- **Drive-absent GATE**: `ReconcileDriveGates` (pure decision `planDriveGates` + executor) on a 30s loop
(`driveGateLoop`, replacing the retired slice-8C watchdog) — an ABSENT drive's apps are STOPPED +
recorded (`StoppedStacks` = the gate-stopped set, distinct from a user stop); a RETURNED drive is
re-attached under the parent and its gate-stopped apps AUTO-RESTARTED. Start-gate in `actionStack`:
refuses to start an app whose drive is disconnected/decommissioned (clear "tárhely nem elérhető"
message) — so it can't write to the empty fail-closed stable path.
- **H1 endpoints routed** (were 404): `POST /api/storage/{disconnect,reconnect,restart-apps}` →
host-side eject/reconnect (stop→agent-detach→fail-close / agent-attach→restart→clear) — no guest
reboot.
Tests (non-hollow + companions): `TestPlanDriveGates` (4 states; trivial impls fail), `TestAgentWhere`
(stable↔raw idempotent mapping), `TestRunStorageInit_Success` (agent gets RAW, registry gets STABLE).
### v0.66.2 — FileBrowser umask 002 (customer folders group-writable) (2026-06-15)
FileBrowser (uid 1000) created folders with umask 022 → mode 2755 (setgid from the parent, but
group-READ only), so a folder a customer made in FileBrowser could not be written by the content apps
in group 1000. The gtstef/filebrowser image is a single Go binary (`entrypoint ./filebrowser`) and does
NOT honor a `UMASK` env (verified live: `-e UMASK=002` leaves PID1 at 0022), so `RenderFileBrowserCompose`
(`internal/infra/infra.go`) now wraps the entrypoint:
`["sh","-c","umask 002; exec /home/filebrowser/filebrowser"]`. Customer-created folders now come out
**2775** (group-writable) so all group-1000 apps can use them. Test asserts the rendered compose carries
the wrapper. (Pre-existing pre-fix folders stay 2755 — recreated on the demo; no data.)
### v0.66.1 — fix USERDATA_PATH on first deploy (2026-06-14)
The initial deploy path (`DeployStack` → `composeExecWithEnv`) builds its compose env from the deploy
values, not from app.yaml via `stackEnv` — so v0.66.0 injected `USERDATA_PATH` only on start/redeploy,
NOT on the FIRST deploy. A freshly-deployed app resolved `${USERDATA_PATH}` to `""` and Docker bound a
bogus root-owned dir at the container root (e.g. `/media/movies`) instead of `<drive>/userdata/...`
(found live: radarr's media mount was `0:0 755` at the container root). Fix: a shared `withUserdataPath`
injector used by BOTH `stackEnv` and `composeExecWithEnv`. Regression test asserts injection on/off by
HDD_PATH presence.
### v0.66.0 — userdata layout + shared-storage ownership convention (2026-06-14)
Customer-facing `userdata/` tree (sibling of appdata/backups under each drive's felhom-data namespace)
with a shared-ownership convention so FileBrowser + content apps collaborate without permission
collisions. Spike: `felhom.eu/documentation/audits/SPIKE-userdata-layout-2026-06-14.md`. Pairs with the
app-catalog commit that repoints media mounts to `${USERDATA_PATH}`.
- **Convention helper** (`internal/appbackup/userdata.go`): `EnsureUserdataDir`/`EnsureDirOwned` =
MkdirAll → explicit `Chmod(ModeSetgid|0775)` (MkdirAll's mode is umask-masked AND drops setgid) →
chown group to `SharedContentGID` (1000). `UserdataDir`, `UserdataSkeleton` (media/{movies,tv,music,
audiobooks,books,comics,photos}, downloads, import/{paperless,calibre}, roms, documents),
`EnsureUserdataSkeleton`. Linux chown via `chownGID`/`StatGID` (`userdata_linux.go`); no-op stub
off-Linux (`userdata_other.go`).
- **USERDATA_PATH injection** (`stackEnv`, manager.go): injects `USERDATA_PATH = <HDD_PATH>/userdata`
(HDD_PATH is the namespace root) alongside HDD_PATH, so the catalog's `${USERDATA_PATH}/...` mounts
resolve.
- **Skeleton pre-create**: `registerStoragePath` + `syncFileBrowserMounts` ensure the full skeleton on
every storage path (system + additional drives) with the convention.
- **Deploy belt**: `composeExecCustomEnv` (gated on `up`) pre-creates every `${USERDATA_PATH}/...`
bind source the stack declares (`ParseComposeUserdataMounts` + `ensureUserdataMounts`) so Docker
never auto-creates a userdata dir as guest-root — covers apps not in the skeleton.
- **FileBrowser mount switch** (`syncFileBrowserMounts`): mounts `<drive>/userdata` (was `appdata`) →
`/srv/<name>`. FileBrowser runs as uid 1000 → can now create folders + upload into the 2775 setgid
userdata (fixes the permission-denied); app internals (appdata/) are no longer browsable.
- **#8 migration fix** (`migrate.go`): the non-app merge walk now preserves the SOURCE dir's full mode
(incl. setgid via `preserveDirOwnership`) + group, and `copyFile` preserves the full file mode
(`fi.Mode()`, not `.Perm()`) + group — so the ownership convention survives a whole-drive `MigrateAll`.
- Non-hollow tests: `EnsureDirOwned` produces 02775+setgid+gid (Linux companion proves a plain MkdirAll
has NO setgid); skeleton structure; `ParseComposeUserdataMounts` selectivity; deploy belt creates the
declared dirs; **migration preserves setgid+group** (Linux; mutation-proven against the pre-fix
0755/.Perm() path).
### v0.65.0 — data migration + self-serve decommission (B1+B2) (2026-06-14)
Customer-self-serve storage **migration** (move app data between drives) and **decommission** (retire
a drive), implemented trunk-based with the locked spike design
(`felhom.eu/documentation/audits/SPIKE-decommission-migration-2026-06-14.md`). Pairs with agent
v0.32.0 (the self-serve `/disks/decommission` endpoint + intent-aware re-assert). Built + deployed to
demo guest 9201. **Live decommission/migration of real data is NOT yet validated — that is the
supervised B3 session.**
- **B1 — migration engine** (`internal/stacks/migrate.go`). In-process over the controller's
`/mnt:/mnt:rslave` RW mount; crash-safe + resumable via a single journal (`<dataDir>/migration.json`).
Two entry points share one pipeline: `MigrateAll` (whole namespace — every app + a conflict-merge
walk for non-app/customer content) and `MigrateApp` (one app subtree; handles drive→drive AND
SSD→drive). Pipeline: validate → stop → copy (`rsync -a --checksum`, additive, NO `--delete`) → verify
(`rsync -ani --checksum`, zero pending) → flip+redeploy (`RedeployFromEnv`, one idempotent unit) →
cleanup. **CLEANUP is the only destructive step and is gated on every unit verified AND every app
redeployed.** Conflict-merge: skip-identical (checksum vs the target file AND its `(N)` siblings),
rename-on-differ to the lowest-free `<base>(N)<ext>`, never overwrite; idempotent (no `(1)(1)`).
Single-flight; **mutual exclusion with the backup orchestrator** (Change 3 — migration refuses while a
backup runs; the scheduled DB-dump/Tier-2 skip while a migration runs).
- **B1 UI** — `POST /api/storage/migrate` (whole-namespace), `POST /api/storage/migrate-app` (per-app),
`GET /api/storage/migrate/status` (poll). The greyed migrate-all `<span>` in settings.html is now a
real target-select + button; app_info.html gains a per-app "Áthelyezés másik tárhelyre" control; both
share a Hungarian progress panel.
- **B2b — decommission orchestration** (`handleStorageDecommission`, `POST /api/storage/decommission`).
Two choices, no partial (Change 2): **migrate-all-then-decommission** (runs `MigrateAll`; the
migration done-hook soft-marks the source + calls the agent once every app has moved) or
**decommission-anyway** (type-to-confirm; stops the apps but KEEPS their `HDD_PATH` so they show
"missing storage"). `agentapi.Decommission` added; both branches end at `SetDecommissioned` (soft
marker retained — blocks A1 resurrection) + agent `Decommission`.
- **"Hiányzó tárhely" indicator** — a deployed app whose `HDD_PATH` resolves to a decommissioned/
disconnected/absent registry path now shows a distinct warning badge on the dashboard, stacks page,
and app card (label via `GetStorageLabel`); persists until re-enroll or migrate.
- **Change 4 — re-enroll clears the marker.** `registerStoragePath` now un-retires a re-plugged
decommissioned drive (`ClearDecommissioned` + restore `Schedulable`) — previously `AddStoragePath`
deduped the re-register into a no-op and the soft marker (and the apps' missing-storage badge) would
persist forever. (`ClearDecommissioned` had zero callers before this.)
- Non-hollow tests across `internal/stacks` (engine: collision-refuse, merge dedup/idempotency,
cleanup-only-after-redeploy, verify-catches-corruption, resume, single-flight, SSD→drive, backup
exclusion), `internal/backup` (scheduled backup skipped while migrating), and `internal/web`
(finalize soft-mark+agent, re-enroll clears marker, missing-storage label). Companions for the
collision guard, cleanup gate, and Change-4 clearing were mutation-proven to fail on the pre-fix code.
### v0.64.0 — storage-lifecycle cleanups (2026-06-14)
Two settings-layer cleanups from the F9 storage-registration diagnosis
(`felhom.eu/documentation/backlog/DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md`), trunk-based on
`main`, each with table-driven tests that fail on the pre-fix code.
- **A1 — `AutoDiscoverStoragePaths` is now ADDITIVE** (`internal/settings/settings.go`). It previously bailed
early (`if len(s.StoragePaths) > 0 { return }`), so a drive a deployed app referenced but that was missing
from the registry was never picked up after first run. It now registers only the discovered paths NOT
already present, while honouring strict invariants: never removes/modifies a manually-added path; SKIPS any
path already in the registry IN ANY STATE — including a `Decommissioned` soft-marked entry — so it can't
re-add or reactivate it; never flips `IsDefault` (a newly-discovered path becomes default ONLY if the
registry currently has no default at all, and only the first such new path). NOT auto-register-on-attach —
it only picks up paths deployed apps already reference (that recommendation was rejected; manual enrollment
is by design). New `internal/settings/storage_discovery_test.go` covers it, incl. a companion test that
FAILS if the skip-by-presence guard is removed (verified: removing the guard re-adds the decommissioned
path).
- **A2 — internal-SSD label disambiguation** (`InferStorageLabel`, `internal/settings/settings.go`). A path
whose basename is the `felhom-data` namespace dir (the internal system volume, e.g.
`/mnt/sys_drive/felhom-data`) previously labelled as `Tárhely (felhom-data)`, colliding with the per-drive
felhom-data namespace. It now reads **`Belső SSD (rendszer)`**. Discriminator is `base ==
appbackup.FelhomDataDir`; Model-A user drives register their MOUNT ROOT (e.g. `/mnt/felhom-usb`), never
`.../felhom-data`, so this can't mislabel a user drive. Still overridable via `SetStorageLabel`. The demo's
already-seeded `settings.json` label for that path on guest 9201 was updated out-of-band (the seeded value
doesn't auto-change). Separate host-metrics label in `web/agent_host_metrics_handler.go` was intentionally
left untouched.
### v0.63.0 — reflect agent F9/F20-BUG2 disk fields (2026-06-14)
Pass through two new fields the host agent (v0.31.0) now returns on `/disks`, so they reach
`/api/disks` and the dashboard (the controller previously dropped them when re-marshalling the agent
response). Additive only — `agentapi.DiskInfo` gains:
- **`wipe_durable_id`** (F20-BUG2) — the device's wipe-binding id in the gate's scheme
(`byid:`/`byuuid:`), distinct from `durable_id` (`uuid:`, used for assign). A customer-confirmed
data-bearing wipe must carry THIS id; confirming with the `uuid:` id was rejected (binding_mismatch).
- **`guest_attached`** (F9) — whether the drive is actually bound into this guest (usable in-guest) vs
merely present on the host — the signal whose absence let an unattached HDD look available.
No behaviour change in the controller itself; the agent owns the fix. (Agent v0.31.0: F9 startup
bind re-assert, F20-BUG2 single wipe-id scheme, F20-BUG3 detached/restart-surviving format.)
### v0.62.0 — M18 + M19 backlog fixes (2026-06-14)
Two verified-LIVE backlog bugs (preserved fix-plans in `felhom.eu/documentation/backlog/`), implemented
trunk-based on `main`, each with a regression test that fails on the pre-fix code. Built, deployed to demo
guest 9201, and both verified live.
- **M19 — `deriveStackName` DB-container misattribution (correctness)** — commit `6bab68b`.
`deriveStackName` pure-suffix-stripped on `-` (postgres/db/mariadb/mysql/database/redis/cache), so a
stack whose slug *ends* in a role token (e.g. `my-cache`) was misattributed (stripped to `my`), filing
its DB dump under the wrong/nonexistent stack. Now threads the set of deployed stack names
(`m.knownStackNames()` ← `ListDeployedStacks`) into `DiscoverDatabases` and cross-references: use the
suffix-strip candidate if it's a known stack, else the container name if it IS a known stack (don't
strip), else the longest known stack that is a `-`/`_`-bounded prefix (handles `<stack>_postgres`,
`<stack>-1`), else the legacy strip. `nil`/empty known = legacy behaviour (appexport passes nil).
**Live:** romm-db → `romm-mariadb.sql` (correct). Table test incl. the `my-cache` case (fails pre-fix).
- **M18 — DB-dump validation re-run every cycle (performance)** — commit `f8afe5c`.
`ListDumpFiles` ran `ValidateDump` (line-by-line scan) for every dump on every ~5-min `RefreshCache`
cycle — wasted I/O+CPU on large customer dumps. `ListDumpFiles` now takes an optional
`cached(name,size,mod)` lookup; on a size+modtime match it reuses the prior result and skips
`ValidateDump`. `settings.DBValidationCache` gains `Size`+`ModTime`; `listAllDumpFiles` builds the
lookup from the persisted cache and writes back only fresh validations (cache miss) — so an unchanged
dump triggers neither a re-validation nor a `settings.json` write each cycle. `nil` cached = legacy
validate-always (back-compat). **Live:** the cache now persists `size`+`mod_time`. Tests: cache-hit
skips validate (sentinel), cache-miss validates, nil validates.
### v0.61.0 — live-drive Batch 1 (+F17) fixes (2026-06-14)
Controller-side fixes triaged in `LIVE-DRIVE-FIXSPEC-2026-06-14.md` from the 2026-06-14 live-drive
findings. Each fix has a regression test that fails on the pre-fix code. Built, deployed to demo guest
9201, and the key fixes live-verified. (F9, F20-BUG2, F20-BUG3 are the SUPERVISED agent/golden next
session — not in this batch.)
- **F17 (CRITICAL) — per-app restore now replays the captured `.sql` DB dump.** `RestoreFromRecoveryUnit`
(and the `RestoreApp` fallback) repopulated Docker volume tars but NEVER replayed the captured
`<stack>-<dbtype>.sql`, so DB-resident data did not come back. New `appbackup.ImportDump` (read-side
counterpart to `DumpOne`, reuses `DiscoveredDB`'s own discovered credentials) + `backup.reimportDBDumps`
replay the dump AFTER volume restore + stack bring-up, so the logical dump **wins** over any volume-tar
copy of the DB (operator-chosen precedence). Volume-restore and DB-import failures now **surface** (the
restore returns an error) instead of a swallowed WARN. **Live-validated** on guest 9201: a marker row
dropped after backup was restored by `/backup/restore` (log: "replayed 1 DB dump(s)"). Reuse note:
`ImportDump` lives in `appbackup` (the DB-domain package) — `appexport→appbackup` already exists so
reusing appexport's unexported copies would cycle; appbackup is the clean shared home.
- **F1 (HIGH) — guest RAM cap read from the Docker daemon; deploy guard uses committed memory.** The
controller container reported the Proxmox host's 16 GB (no lxcfs in the container; its own cgroup is
unlimited — the 2 GB cap is on the LXC ancestor), defeating the deploy memory-headroom hard-block.
`system` now sources the cap from `docker info` MemTotal (the daemon runs in the LXC → reports the
guest's real RAM; cgroup limit still preferred when present). The deploy guard now uses the controller's
own committed-app memory (sum of running mem requests) for "used" — accurate and cheap — instead of
host RSS. `/api/system/info` reports the guest cap + committed used. **Live-verified:** `total_mem_mb`
2048 (was 15771).
- **F20-BUG1 (HIGH) — `agentapi.FormatDisk` surfaces the agent's error.** A failed format (agent 502
"device is mounted", `ok:false`, `data:null`) fell through to `return out, nil`, so the web layer
reported a zero-value result as `ok:true` — a failed DESTRUCTIVE op read as success. Now returns a
non-nil error on any non-2xx/`ok:false` that is not a recognized refusal (403/needs-confirmation).
- **F5 (HIGH) — broken healthcheck → 404, two parts.** (catalog, `app-catalog-felhom.eu`) uptime-kuma's
healthcheck pointed at a v1-era `node /app/extra/healthcheck.mjs` absent in `:2`, so the container
stayed unhealthy and Traefik withheld the route (404 though running) — fixed to the v2 compiled
`extra/healthcheck` binary + 180s start_period. (dashboard) new `routeUnpublished` helper + a distinct
"URL nem elérhető – útvonal nincs publikálva" indicator on the dashboard/stacks cards for
unhealthy/restarting deployed apps (operator decision: keep gating the route, surface it distinctly).
**Live-verified:** uptime-kuma healthy → route publishes → status URL 302 (was 404).
- **F8 (LOW-MED) — `controller.yaml` persisted 0600.** It holds infra credentials (cf/hub tokens) in
plaintext; the Hub config-apply handler wrote 0644. New `writeConfig0600` enforces 0600 even on a
pre-existing 0644 file.
- **F6 (LOW) — deploy POST reports "started", not "deployed".** The deploy runs async (UI polls); the
POST now returns 202 Accepted + "Telepítés elindítva…" so API/script consumers aren't told a deploy
finished before it has.
- **F7 (LOW) — dashboard state lag.** `status-refresh` tightened 30s → 10s (cheap docker-ps refresh).
- **F4 (TRIVIAL) — `GET /api/stacks/rescan`** now returns 405 + `Allow: POST` instead of the misleading
"stack not found: rescan" fall-through.
### v0.60.0 — M25 data-race fix (backlog-Medium cleanup) (2026-06-13)
Backlog-Medium reconciliation from the 2026-06-13 BUGHUNT reconcile. M4/M5/M6 verified already FIXED
(no action). M18 (dump re-validation every 5 min — perf) and M19 (naive `deriveStackName` misattribution
— low-incidence correctness) verified LIVE but cross-package-entangled; prepared on branches
`fix/m18-dump-validation-cache` / `fix/m19-stackname-crossref` (notes + fix plan, pending review, not
deployed).
- **M25 (Server.integrationMgr data race) — FIXED.** `NewServer` launches the `SyncFileBrowserMounts`
goroutine (which reads `integrationMgr`) from the constructor, *before* `main.go` calls
`SetIntegrationManager` — so the init-only happens-before that covers the other `Set*` fields did not
hold, making it a genuine data race (reads at `handlers.go:358/360/1433` vs the unsynchronized write).
Converted the field to `atomic.Pointer[integrations.Manager]`; setter `Store`s, all readers `Load()`.
Regression test reproduces the concurrent access and is clean under `-race` (verified on the build
server); it flags on the pre-fix plain-pointer field.
### v0.59.0 — security/crash-safety fixes from the 2026-06-13 audit (2026-06-13)
Fixes the validated findings from the deep-sweep audit + BUGHUNT reconciliation
(records under `felhom.eu/documentation/audits/`). All shipped with permanent
regression tests.
- **CTRL-001 (path traversal on `.fab` import) — High.** `appexport.UnmarshalManifest`
did zero validation; the attacker-controlled `manifest.AppName` / `HDDSubdirs` /
`VolumeNames` reached `filepath.Join`+`MkdirAll`/`extractTar` (restore.go:339/606/678),
so `../..` in any escaped the stacks / HDD destination dir (arbitrary write as the
controller). New `appexport.ValidateSegment` + `validateManifestPaths`;
`UnmarshalManifest` now fails the parse on a traversal segment, with defence-in-depth
guards at the HDD-subdir and volume-name join loops. `ConfigFiles` intentionally not
validated (holds dotfiles, never used in a restore join).
- **CTRL-T2-1 (ghost-deployed stack on crash) — High.** `DeployStack` wrote `app.yaml`
`deployed:true` to disk *before* the async `docker compose up -d`; a crash during the
image-pull window left a ghost-deployed stack with no containers that the app then
refused to redeploy. The env is now persisted `deployed:false` (transitional) and
flipped to `deployed:true` by `runComposeDeploy` only after `up -d` succeeds. The
in-memory flag still goes true during the pull (no stale "Telepítés" button).
- **H10 (plaintext secret on encrypt failure) — fail-closed.** `SaveAppConfig` logged a
WARN then fell through to persist the secret in plaintext on a `crypto.Encrypt` error.
Now returns an error instead — never writes plaintext.
- **M2 (misleading lock).** `backup.Manager.SetStackProvider` was mutex-guarded while all
reads were unlocked; it is init-only (one call before any goroutine), so the lock was
removed and the contract documented. No behaviour change.
- **AGENT-001 (wrong-disk wipe race)** is fixed on the agent branch `fix/agent-001-wipe-durable-reresolve`
(PENDING REVIEW — not deployed; stored out-of-band per the supervised-merge rule).
### v0.58.0 — infra-protection prevention layer for the OS/Docker-data split (2026-06-13)
Phase 2 of the storage-split slice (Phase 1 = felhom-agent golden + provision). The OS rootfs and
Docker data are split onto separate volumes for resilience; infra (controller/traefik/cloudflared/
filebrowser) shares the one Docker data-root and is protected by **prevention, not placement**.
- **Reserved-buffer headroom guard (`internal/system/dockervol.go`):** `GetDockerVolumeHeadroom()`
measures the Docker-data volume via `statfs("/")` (the controller container's root overlay is backed
by the guest's `/var/lib/docker` volume) and computes a reserved floor `DockerVolumeReserveGB` =
`max(5 GB, 10% of total)`. Fail-open on a measurement error (the buffer is a safety net, not a
security control).
- **Deploy-time hard gate (`internal/api/router.go` `deployStack`):** a new deploy is **refused** (HTTP
507 + Hungarian message) when free space on the Docker-data volume is at/under the reserved buffer,
so customer apps can't fill the volume the infra containers depend on.
- **Deploy-page surfacing (`deploy.html`):** for a new deploy, when below the buffer the page shows a
clear Hungarian warning and **disables** the "Telepítés indítása" button (mirrors the memory-blocked
pattern) — the customer sees it before clicking; the API gate is the hard backstop.
- **Runtime monitoring (2C):** confirmed `monitor/healthcheck.go` already watches `sysInfo.DiskPercent`
= the Docker-data volume post-split (statfs `/`); warn 80% / crit 90% used trip ABOVE the 10%-free
reserved buffer, so the customer is warned before the deploy gate engages. Comment added to make the
"SSD disk" alert's target explicit.
- **Log rotation (2D):** baked into the golden's `daemon.json` (`max-size 10m`, `max-file 3`) in the
felhom-agent golden build — every guest inherits it. Per-app xfs-project-quota caps deferred.
- Tests: `DockerVolumeReserveGB` floor/scale.
### v0.57.0 — UI fixes: stable host-storage list + per-app Tier-2 config panel (2026-06-13)
Part A of the UI-fixes/storage-spike spec (Part B is a build-nothing findings report).
- **A1 — host storage list no longer reorders (item 2):** the monitoring page's `#host-storage-bars`
list (the client-side one filled from the agent's PVE-storage list — `local`, `local-lvm`,
`felhom-pbs`, `felhom-usb` with thin-pool % + temperature) reordered on every 8 s poll because the
agent enumerates `pvesm` in a non-deterministic order and the list never passed through a Go sort.
Now `enrichHostStorageTargets` (`agent_host_metrics_handler.go`) sorts the `/api/host-metrics`
response server-side (user-data → system+apps → backup → other; alphabetical by id within a tier)
and attaches a **friendly Hungarian label + one-line purpose** per entry (e.g. `local-lvm` →
"Belső SSD – rendszer és alkalmazások"). The raw PVE id is kept and shown muted — **display labels
only; PVE storage ids are never renamed** (vzdump/PBS configs reference them by name). The
monitoring JS renders the friendly label + the purpose sub-line. (Note: this is the JS-driven list,
NOT the server-rendered user-data `buildStorageBars` list that v0.56.0's 4C already sorted.)
- **A2 — per-app Tier-2 config panel (item 4):** the "2. mentés" row's **Beállítás** button used to
link to the app's deploy page, which has no backup-location setting (a dead end). New route
`GET/POST /stacks/{name}/backup` (`tier2_config_handler.go` + `tier2_config.html`) is the real
surface: it shows the current/effective off-drive target, whether it's the size-limited internal
SSD, the last-run status, and lets the customer **pin a different registered drive** or **turn
Tier 2 off**. The control is **always visible** — even when only the internal SSD qualifies (shows
"automatikus: belső SSD — csak DB/konfiguráció" + the rootfs-headroom note) and for non-HDD apps
(shows honest "already in the PBS whole-guest snapshot; the off-drive copy is supplementary"
context). The button is repointed on every "2. mentés" branch (incl. the unconfigured + disabled
states).
- Persistence: two preference fields on `settings.CrossDriveBackup` — `UserDisabled` and
`PreferredTarget` — set via `SetTier2Preference` and **preserved across the runner's status
writes** (`withTier2Prefs`). `selectTier2Target` now honors a valid pinned target (off-disk,
registered) before the auto-pick; an invalid pin silently falls back to auto. `RunTier2` skips a
customer-disabled app. Saving with Tier 2 on for an HDD app triggers an immediate run so the
result shows on return.
- Tests: `enrichHostStorageTargets` order/labels/determinism; `selectTier2Target` honors/falls-back
on a pin; status writes preserve the preference.
### v0.56.0 — Phase 4: FileBrowser scoping + deploy DB-on-SSD note + monitoring storage descriptions (2026-06-13)
Polish layer closing the slice.
- **4A FileBrowser scoping (safety):** the FileBrowser bind mount is now scoped to each drive's
`appdata/` subtree (`<drive>/appdata:/srv/<name>`) instead of the whole drive root. The recovery
units + Tier 2 copies under `backups/` are therefore **not mounted into FileBrowser at all** — the
customer browses their userdata but cannot reach (or even see) the thing that restores them. The
appdata dir is `mkdir`-ed before the bind so the source exists. (`syncFileBrowserMounts`.)
- **4B Deploy-UI communication:** the storage-selection step now states plainly (Hungarian) that the
chosen drive holds the app's **files**, while its **database runs on the fast internal SSD** and is
backed up alongside the app — so "the DB is on the SSD" stops being a surprise. (`deploy.html`.)
- **4C Monitoring storage list:** `buildStorageBars` now sorts deterministically (by path) and carries a
**purpose description** explaining the user-data drives (rendered on the monitoring "Tárolók
kapacitása" list). Note: this list is the controller's registered user-data drives only (the agent's
local/local-lvm/pbs storage is not in this registry), so the role-tier sort/`local`-vs-`local-lvm`
descriptions belong to the agent-backed storage-management page, not here.
### v0.55.0 — Phase 3: auto off-drive Tier 2 (rootfs-headroom guard, durable off-disk target) (2026-06-13)
Tier 2 = an **off-drive copy** of each HDD app's recovery unit + bulk userdata to a **different physical
disk** — the only off-drive protection browsable HDD userdata can get (PBS can't reach bind mounts).
Auto-enabled for every HDD app; the target is auto-picked and the dangerous case (the small guest
rootfs) is refused rather than filled.
- **Engine** `internal/backup/tier2.go` (`RunTier2`/`RunAllTier2`): rsync `-a --delete` of the recovery
unit (`backups/primary/<app>/`) and the app's `appdata/<app>/` to `<target>/backups/secondary/<app>/`.
restic is **not** revived — plain browsable mirror.
- **Auto target selection:** prefer another registered user-data drive on a **different physical disk**
(can hold bulk userdata); else fall back to the internal SSD for **small units only**. Off-disk is
enforced by `system.SamePhysicalDevice` (block-device identity; new exported helper, linux + stub) —
defense-in-depth re-checked before the copy.
- **Rootfs-headroom guard (the key safety):** the SSD target is the ~8 GB guest rootfs, so a size-aware
guard (`tier2FitsHeadroom`, unit-tested) **refuses** unless the unit fits while leaving a reserve free
(`max(2 GB, 20% of total)`). When nothing fits, it records an **honest** "needs a 2nd HDD" status
rather than silently doing nothing or endangering the rootfs.
- **Status + UI:** results persist via the surviving `settings.CrossDriveBackup` (rsync method, dest,
last-run/status/size). The "2. mentés" card is now **populated** (`buildAppBackupRows`): real target
("belső SSD (csak DB/konfiguráció)" vs an external drive) on success, or the honest no-off-drive-target
reason. Notifications via the surviving `NotifyCrossDrive{Completed,Failed}` hooks.
- **Scheduling + trigger:** daily `tier2-backup` job (03:30, after the DB dump); manual
`POST /api/backup/tier2`.
- Fixed a stale pre-existing test (`TestBackupCopiesOnPath`) that still used the old
`felhom-data/backups/secondary` layout — now the Model-A in-guest layout the Tier 2 copies actually use.
### v0.54.0 — Phase 2b: restore-from-recovery-unit + fail-closed data-key gate (2026-06-13)
Restore now recreates an app from its on-drive recovery unit **plus the guest's own secrets** — never
from secrets stored in the unit (there are none), and **regenerating nothing**.
- **Fail-closed data-key gate** (`reconcileRestoreSecrets`, `internal/backup/restore_unit.go` — a pure,
exhaustively unit-tested function): merges the unit's non-secret env with the secret values recovered
from the guest's live app.yaml. A missing/empty **data-encrypting key** (`data_key`) **aborts the
restore** with a clear message (a PBS whole-guest restore is required) — because regenerating it would
render stored data unreadable. A missing *resettable* secret (DB/admin password) is non-fatal (warn +
proceed; the app may need a credential reset). Secrets are recovered, never regenerated.
- **`RestoreFromRecoveryUnit`**: reads the unit manifest → recovers secrets from the guest
(`RecoverStackSecrets`) → applies the gate → restores named-volume data from the unit's tars →
recovers the app definition from the unit and redeploys with the reconstructed env (re-pulling the
pinned image). Falls back to the legacy volume-only `RestoreApp` if no unit exists. Wired into the
`/backup/restore` web handler.
- **New seams:** `StackDataProvider.RecoverStackSecrets` / `RecreateStackFromUnit` (main.go
`stackAdapter`, with the controller `encKey` for decrypting the live app.yaml); `stacks.Manager.
RedeployFromEnv` (writes app.yaml from the full env incl. locked secrets, then `compose up -d`).
- **Tests:** the gate (all recovered / data-key missing → refuse / empty data-key → refuse / resettable
missing → proceed+warn, recovered values used verbatim) and `data_key` parsing from `.felhom.yml`
(`Metadata.DataKeyEnvVars()`).
- **Live-validated on guest 9201 (AdventureLog, a real data_key app):** its recovery-unit manifest
correctly carries `data_key_env_vars: [SECRET_KEY]` (catalog→metadata→manifest flow proven live); and
with `SECRET_KEY` made unrecoverable, `POST /backup/restore` **refused** with the exact fail-closed
message ("…[SECRET_KEY] could not be recovered … a PBS whole-guest restore is required first…"),
**before any compose-up** (no side effects). The demo has no dashboard password, so the API is open
(auth + CSRF are both skipped in that mode) — this was driven via the public URL. Gate + reconciliation
+ orchestration + data_key parsing are also unit-tested.
- **One e2e not run (environment limit, not a code gap):** the full "deploy with data → restore →
confirm data decrypts" — AdventureLog's images don't fit the **8 GB guest rootfs** (the deploy hit "no
space left on device"). This is exactly the Phase 3 rootfs-headroom concern, now observed live.
Key-preservation/regenerate-nothing is covered by the gate's verbatim-recovery unit test.
### v0.53.1 — Phase 2: recovery units refresh on the periodic cache cycle (idempotent) (2026-06-13)
The recovery-unit capture now also runs from `RefreshCache` (controller startup + every 5m), not only
the daily DB dump — so a unit exists shortly after startup and stays current with config changes
(redeploy / optional-config) without a 24h wait. `CaptureRecoveryUnit` builds the captured content in
memory and **skips all writes when the unit is already current** (same config checksums + dump set +
controller version), so the periodic refresh does not thrash a spinning USB drive. Added an idempotency
test (unchanged → skip; config change → rewrite).
### v0.53.0 — Phase 2: per-app self-contained recovery unit (capture side, SECRET-FREE) (2026-06-13)
Each app's on-drive backup becomes a complete, recreatable **recovery unit** — not just DB dumps +
volume tars, but the app's *definition* too, so it can be recreated. The unit is **secret-free by
design** (decided after reading the actual hub code: the hub is deliberately zero-knowledge and holds
no app secrets; app.yaml + the encryption key live on the guest rootfs → already inside the PBS
whole-guest snapshot). Secrets/data-keys are recovered at restore from the guest's own app.yaml (live,
or via PBS) — **never stored in the unit, never regenerated**.
- **Unit layout** (rooted at the existing `backups/primary/<app>/` — no risky dump-dir migration):
`compose/` (docker-compose.yml + .felhom.yml + a **secret-stripped** app.yaml) + the existing
`db-dumps/` + `volume-dumps/` + `manifest.json`. New path helpers `RecoveryUnitPath` /
`RecoveryUnitComposePath` / `RecoveryUnitManifestPath` in `internal/appbackup/paths.go`
(`AppDBDumpPath`/`AppVolumeDumpPath` refactored onto `RecoveryUnitPath` — identical resolved paths).
- **Secret-free manifest** (`internal/backup/recovery_unit.go`): app id, display name, controller
version, timestamp, drive, namespace root, pinned **image tags** (image NOT stored — re-pulled on
restore), the **NAMES** of secret env vars (values never stored), the `data_key` env-var names, the
explicit `secret_source` note ("guest app.yaml (live) or PBS — never stored in this unit"), captured
config-file list, enumerated dumps, and sha256 checksums of the captured config.
- **Capture has no secret access:** non-secret env is plaintext in app.yaml; the capture simply excludes
the secret-named keys (plus a defensive `crypto.IsEncrypted` guard), so it reads no secret value. New
`StackDataProvider.GetStackRecoveryInfo` + `RecoveryInfo` (in `appbackup`), implemented by the main.go
`stackAdapter`; `ParseComposeImages` extracts the image pins.
- **`data_key` annotation** (`DeployField.DataKey`, `Metadata.DataKeyEnvVars()`): marks a
data-encrypting key (e.g. AdventureLog's "Titkosítási kulcs", `SECRET_KEY`) — a **fail-closed** safety
annotation for restore (refuse + warn rather than regenerate-and-corrupt), NOT a per-secret
preserve/regenerate decision. Catalog: `adventurelog/.felhom.yml` `SECRET_KEY` marked `data_key: true`.
- **Wired into the dump flow:** `RunDBDumps` refreshes every deployed app's recovery unit after the DB
dumps (best-effort per app; skips disconnected/decommissioned drives). Capture test
(`recovery_unit_test.go`) proves the unit is secret-free (a secret in the source app.yaml never
appears in the unit) and the manifest structure.
- **NOT in this increment (next):** the restore-from-unit *recreate* (re-pull + compose-up + secret
recovery from guest/PBS) and its fail-closed `data_key` gate, with live AdventureLog readable-data
validation. The README backup-paths section (stale restic/secondary) is rewritten when Tier 2 lands.
### v0.52.0 — Phase 1 GATE: deploy-side double-nest fix + path-agreement lock (2026-06-13)
Completes the Model-A double-nest reconciliation deferred in v0.48.0. v0.51.0 fixed the **backup
helper** side (`NamespaceRoot` provenance); the **deploy/compose** side still wrote one segment too
deep. On a Model-A in-guest drive the guest mount `/mnt/<drive>` already IS the host's
`<drive>/felhom-data` namespace, so the catalog templates' `${HDD_PATH}/felhom-data/appdata/<app>`
double-nested to `.../felhom-data/felhom-data/...` on disk — diverging from where the backup helpers
look (`AppDataDir(NamespaceRoot(HDD_PATH,true))`, single-nested).
- **Fix lives in the app catalog** (`app-catalog-felhom.eu`): all four HDD app templates
(`romm`, `nextcloud`, `immich`, `paperless-ngx`) changed `${HDD_PATH}/felhom-data/appdata/<app>` →
`${HDD_PATH}/appdata/<app>`. The controller passes `HDD_PATH` through verbatim and never appended
the segment, so no controller runtime change was needed. Catalog change lands via git-sync /
"Sablonok frissítése".
- **Agreement test (new):** `internal/stacks/hddpath_agreement_test.go` resolves a compose's
`${HDD_PATH}` bind mounts via the real deploy-side `ParseComposeHDDMounts` and asserts they are
byte-identical to the backup-side `AppDataDir(NamespaceRoot(HDD_PATH,true))` — no doubled
`felhom-data`, deploy and backup locked together so they cannot drift again.
- **Live migration:** existing drive-resident apps whose data sat at the doubled
`…/felhom-data/felhom-data/appdata/<app>` are migrated (stop → move → verify → redeploy) to the
single-nested path (RomM confirmed on the demo guest).
### v0.51.0 — offsite-backup UI (felhom-pbs DR) + Model-A double-nest fix (2026-06-12)
Pairs with felhom-agent v0.28.0 (whole-guest backup re-targeted to the offsite PBS tier).
**Backups page — the whole-guest backup is now shown as real DR (separate hardware).** The
"Rendszermentés" section's target label calls out the offsite tier: `backupTargetLabel` returns
**"Biztonsági szerver – külön hardver (PBS)"** for a PBS-stored backup (detected via `backupIsPBS`
on the target id / archive volid), so the customer sees the backup survives a host hardware failure.
The app-data section's **"Távoli mentés"** card stops reading "nincs beállítva": a new
`guestBackupView.Offsite` flag drives it to **"külön hardveren (PBS)"** with a ✓ when the whole-guest
backup landed on PBS. The restore-test "Visszaállítás ellenőrizve" trust signal is unchanged.
**Model-A double-nest fix — drive-resident app backups land single-nested.** Under slice-10 Model A the
host agent binds `<drive>/felhom-data` onto the guest mountpoint, so an enrolled drive's in-guest mount
IS the felhom-data namespace root (basename need not be `felhom-data`, e.g. `/mnt/felhom-usb`). The
backup path helpers were re-prepending `felhom-data`, producing `.../felhom-data/felhom-data/...` on the
host. `appbackup` path helpers now take a NAMESPACE ROOT (no internal `felhom-data` join) plus a new
`NamespaceRoot(drivePath, inGuestDrive)`; `backup.Manager.namespaceRoot`/`AppNamespaceRoot` resolve
provenance (a drive-resident app's mount is the root as-is; only the SSD-only `systemDataPath` fallback
appends `felhom-data`). All parallel constructions updated coherently so writes, deletion
(`GetStackBackupData`, `RemoveStack` backups-base + `ProtectedHDDPaths` — legacy double-nest dirs kept
protected), the wipe-warning secondary scan, and export all agree. `api.router` passes the namespace
root across the package boundary. New `appbackup` test asserts no doubled `felhom-data` segment for an
in-guest drive and exactly one for the system fallback.
### v0.50.0 — slice 10 P4: dual-role drives + backup-aware wipe warning (2026-06-12)
Pairs with felhom-agent P3 (self-heal). Establishes the dual-role MODEL + the backup-aware wipe
warning; the cross-drive backup ENGINE (restic USB1↔USB2) is a follow-on slice (needs a 2nd physical
drive to validate) and is deliberately NOT built here.
- **4A dual-role eligibility:** a user-data drive is appdata AND backup-target-eligible (it may hold
cross-drive backup copies of *other* drives) — it is not locked to a single role. Surfaced in the
drive overview's per-card purpose note ("Más meghajtók biztonsági mentési céljaként is szolgálhat").
`felhom-pbs` stays the dedicated whole-guest backup datastore (operator-signature); system/backup
roles unchanged.
- **4B backup-aware wipe/eject warning:** `handleStorageImpact` now also returns `backup_copies` — the
apps whose cross-drive (secondary) backups are stored on the drive (`backupCopiesOnPath` scans
`felhom-data/backups/secondary/<app>`, skipping the shared restic repo / `_infra`). The type-to-
confirm modal names them ("Ez a meghajtó más alkalmazások biztonsági másolatait is tárolja — a
törlés ezeket is eltávolítja"). The wipe stays **customer-confirmable** (the copies are redundant —
originals live on the source drive), not operator-signature. Forward-compatible: empty until the
cross-drive engine writes there. Test: `TestBackupCopiesOnPath`.
### v0.49.0 — slice 10 P2 activation: pending-drive detection + "Újraindítás most" (2026-06-12)
A drive enrolled into a running guest activates only at the next guest boot (the host-side live inject
is blocked on unprivileged LXC — see felhom-agent v0.26.0). Per the decision: enroll persists (no forced
reboot), and the customer activates pending drives with one batched restart.
- **Pending detection** (`pendingActivationDrives`): a registered StoragePath whose backing drive the
agent reports present+attached but which is NOT a live mount in this container → "pending activation".
- **Settings UI:** a banner ("N meghajtó aktiválásra vár") with an **"Újraindítás most (~30 mp)"** button
(one restart batches all pending drives). `POST /api/storage/activate` → `agentapi.GuestReboot` →
agent `POST /guest/reboot`. The reboot takes the controller down too, so the JS reloads after the
restart window rather than awaiting the (cut-short) response.
### v0.48.0 — slice 10 P2C: enroll passes the drive into the guest (passthrough) (2026-06-12)
Pairs with felhom-agent v0.25.0 (`POST /disks/guest-attach`) + the golden's `/mnt:rslave` controller
bind. Closes the diagnosed Branch-A gap: enrolling an external drive now makes it actually usable in
the guest, not just mounted on the host.
- **agentapi:** new `GuestAttach(where)` → `POST /disks/guest-attach` (idempotent on the agent side).
- **Enroll triggers attach:** `runStorageInit`, `runStorageAttach`, and `handleStorageRegister` call
`attachIntoGuest` after recording the StoragePath. Best-effort (logged, non-fatal) — the registration
is the durable intent; a transient attach failure is healed by P3 self-heal (next slice). Test:
`TestRunStorageInit_Success` now asserts the drive is guest-attached.
- Note: app data on these drives is written via `HDD_PATH` (the registered `/mnt/<name>`), which Model A
binds to the drive's `felhom-data` namespace — so app bytes land on the external drive, and the
controller's storage probe (os.Stat + IsMountPoint) sees a real mount → the "nem elérhető" banner
clears. (The controller's own backup-path helpers' `felhom-data` level is reconciled when app-data
backup-to-drive is wired; not P2.)
### v0.47.0 — backups page: whole-guest backup visibility + manual trigger (agent-sourced) (2026-06-12)
The backups page previously showed only the app-data (DB-dump) tier and had **zero** view of the
agent's whole-guest PBS/vzdump backup. Adds visibility + a manual trigger over the agent's existing
per-guest backup API (no agent change). Cadence/retention CONFIG stays out (hub-served policy, slice 10).
- **agentapi (2A):** `StatusResponse` gains `Backup *BackupRecord` (the agent's latest recorded
whole-guest backup — target/archive/mode/size/success/started-at); `DueResponse` gains `age_seconds`;
new `RestoreTestStatus()` → `*RestoreTestRecord` (the "verified restorable" signal, nil until one
runs). Non-hollow client tests (`backup_test.go`): parse the documented JSON + assert `StartBackup`
POSTs to `/backup`.
- **Section "Rendszermentés (teljes mentés)" (2B):** new read-only cards above the app-data section —
last whole-guest backup (time + size + **target: PBS vs Helyi (local)**, surfaced from the archive
volid/target-id), next-due (from `/backup/due` age vs cadence), restore-test result, and the running
phase. Agent-unconfigured/unreachable degrades to a note, page still renders.
- **Manual trigger "Mentés most" (2C):** **the controller owns quiescing** (confirmed: the
`quiesce.Loop` stops stacks → `POST /backup` → polls → resumes; the agent's vzdump is crash-consistent
only). The button therefore goes **through the loop**, not a bare agent call. `quiesce.Loop` gains a
mutex + `TriggerNow()` (single-flight via `TryLock` + the existing marker; `ErrBackupInProgress` on
overlap) that runs the same stop→backup→resume cycle async, bypassing the due-check. New
`POST /api/guest-backup/trigger` + `GET /api/guest-backup/status` (distinct prefix from apiRouter's
app-data `/api/backup/{run,status}` to avoid shadowing). The button warns per mode (snapshot ≈ a few
seconds' downtime on lvm-thin; stop = full downtime).
- **App-data section (2D):** the existing per-app DB-dump rows/table are now under a clear
"Alkalmazás-mentések (adatbázis + konfiguráció)" divider, distinct from the whole-guest tier above
(whole-guest = appliance restore; app-backup = granular per-app). No structural change.
- **Config (2E):** OUT OF SCOPE — whole-guest cadence/retention is hub-served policy (slice 10), so it
survives re-provision; no agent config surface added.
### v0.46.0 — fix: /backups 500 (template referenced disk-tier fields stripped in 8C) (2026-06-12)
`GET /backups` returned **HTTP 500**. Root cause (from the live log, not guessed):
`backups.html:64: executing "backups" at <.Backup.RepoStats>: can't evaluate field RepoStats in type
interface {}`. The 8C de-privileging slimmed `FullBackupStatus` to **app-data only** (DB dumps +
Docker-volume tars; the disk-tier restic/cross-drive backup moved to the host agent), but
`backups.html` still carried the full pre-8C restic UI. It referenced `.Backup.X` struct fields that no
longer exist: `RepoStats, LastBackup, ResticSchedule, NextBackup, PruneSchedule, Retention,
SnapshotHistory, LastCheckTime, LastCheckOK`. While those fields existed-but-nil, `{{if .Backup.X}}`
short-circuited safely; once the fields were *removed from the struct*, the field access itself errors →
500. (Not a panic, not a funcmap/nil-subfield issue; root-level map keys like `.PerDriveRepoStats` are
map lookups → nil on miss → safe, and the `Tier1*/Tier2*` fields are on `AppBackupRows`, still supplied.)
Fix — removed the dead disk-tier UI from `backups.html`, keeping the app-data backup view:
- Section 0 storage-stats: dropped "Mentési tároló" + "Pillanatképek" (RepoStats); kept "DB mentések".
- Section 1 cards: the status card now keys on `.Backup.LastDBDump` (was `.Backup.LastBackup`); removed
the "Tároló méret" card.
- Section 2 schedule: removed the "Restic pillanatkép" + "Karbantartás" rows and the
restic-last-backup/retention summary; kept the DB-dump schedule + a DB-dump last-run summary.
- Section 5 "Pillanatképek" (restic snapshot history): removed entirely.
- Section 6 "1. szint" tier: removed the per-drive/repo-stats + integrity rows (relabeled "(restic)" →
"(adatbázis + konfiguráció)"); kept the DB-dump-count row.
No Go change (`FullBackupStatus` was already correct); template-only. `settings.html`'s `.ResticSchedule`/
`.LastCheckTime` are unaffected — they're root-map lookups (nil-safe), not struct-field access.
### v0.45.0 — storage UX polish: deterministic order, init filter, register shortcut, system-storage clarity (2026-06-12)
Builds on v0.44.0's role-aware drive management. Pairs with felhom-agent v0.24.0 (the eject role-gate
lives at the agent — see its CHANGELOG). This release is the controller-side clarity/ordering polish.
- **Deterministic disk order (B1)** — `GET /api/disks` now sorts the agent's drive list server-side:
**user-data → system → backup** (then unrecognized), alphabetical by storage name within each tier.
The agent's storage view iterates an unordered Go map, so the list previously reordered on every
reload (CLAUDE.md lesson #3). The customer's manageable drives are now always on top, stably.
`sortDisksForView` in `agent_disk_handlers.go` + `TestSortDisksForView`.
- **Init wizard excludes mounted drives (B2)** — `storage_init.html`'s formattable filter gained
`&& !d.mount_path`, matching the attach wizard: an already-mounted drive (e.g. `felhom-usb`) no
longer appears as an "initialize" candidate. Eject it first to make it an init target.
- **Register shortcut (B3)** — a mounted, unregistered **user-data** drive now offers **Regisztrálás**
as its PRIMARY per-card action (Leválasztás/Törlés stay secondary). It records the existing mount
into the `StoragePath` registry (no format, no eject) via the new `POST /api/storage/register` →
`registerStoragePath`, then FileBrowser syncs. The natural "use this drive" intent, not "wipe it".
- **System-storage clarity (B4)** — `local` and `local-lvm` are both kept (not collapsed); each
storage card now carries a plain-Hungarian **purpose description** keyed on the agent's role/type,
the app-backing storages (`local-lvm` → "Alkalmazás-rendszer"; user-data → "Alkalmazás-adatok") are
tagged, and a one-line tiering note above the list answers "which storage do the apps use?". Pure
controller-side presentation — no agent contract change; role/type stay authoritative from the agent.
- **Eject impact (B5)** — the eject confirmation already lists, by name, the deployed apps that lose
their storage (via `/api/storage/impact`), at parity with the wipe warning — verified, no change.
### v0.44.0 — role-aware drive management: protected lockout + customer type-to-confirm wipe + drive-list restyle (2026-06-11)
The controller half of the storage-authorization redesign. The drive UI is now driven by the agent's
authoritative **role** (`system` | `backup` | `user-data`, from `GET /api/disks`): the appliance's own
system storage and the backup safety-net are visibly protected with NO destructive controls; the
customer manages their own data drives with informed consent instead of a support ticket.
- **`agentapi` client** — `DiskInfo` gains `role` + capacity (`total_bytes`/`used_bytes`/
`used_fraction`); `FormatResult` gains `role`/`needs_confirmation`/`durable_id`. `FormatDisk` now
takes `confirmed` + `durableID` and returns a new `ErrNeedsConfirmation` (user-data, awaiting the
customer's confirmation) distinct from `ErrFormatRefused` (system/backup, operator signature).
- **Role-aware overview** (`settings.html`, "Meghajtók (ügynök nézet)") — restyled from a raw
`<table>` to **cards** in the house style: prominent storage name, mono device/mount detail, badges
for class (gyors/lassú), data ("Adatot tartalmaz"), **role** (🔒 Rendszer / 🔒 Biztonsági mentés —
védett / Felhasználói adat) and registered state, plus a **capacity bar** (the monitoring page's
green→amber→red `system-bar`). Destructive controls (Leválasztás / Törlés) render **only** for
user-data drives mounted under `/mnt`. System/backup get the lock badge and no controls.
- **Type-to-confirm + name-the-apps** — a modal that (1) lists, **by name**, the deployed apps whose
data lives on the drive (`GET /api/storage/impact` → `appsUsingPath`), and (2) requires the customer
to **type the mount name** before the destructive button enables. No reflex-clickable destructive
action. Applies to both eject and wipe.
- **Customer wipe** (`POST /api/storage/wipe`) — eject (unmount + deregister) then a server-side
two-step customer-confirmed format (learn the agent's durable id, then re-submit `confirmed:true`
bound to it). The mount name is re-checked server-side. A system/backup device is refused by the
agent regardless of what the controller sends.
- **Init wizard** (`storage_init.html`) — the data-bearing path now uses the **customer-confirmation**
flow (type-to-confirm → re-submit confirmed) instead of the `felhom-opsign` instruction; the disk
selector is restyled to cards and lists only user-data targets. `storage_attach.html` likewise
restyled (cards, user-data only). No raw `<table>` remains in the storage UI.
- **Tests** — `agentapi`: blank → ok, system/backup → ErrFormatRefused (+pending op), user-data →
ErrNeedsConfirmation (+durable id), confirmed → formatted. `web`: init surfaces NeedsConfirmation and
does NO mount/register; confirmed init forwards the confirmation+durable id and proceeds; the
dependency-impact (`appsUsingPathIn`) names the right deployed apps.
Pairs with **felhom-agent v0.23.0** (the authoritative role classifier + the tiered wipe gate).
### v0.43.0 — rebuilt storage management (guided init/attach/eject on the agent disk model) (2026-06-11)
After the 8C de-privileging, the storage UI's buttons pointed at deleted routes (`/settings/storage/init`,
`/attach`, `/migrate-drive`, per-stack `/migrate`) — all 404. Everything underneath already worked (the
agent owns disk execution + the data-bearing signature gate; the controller has the `agentapi` client +
`/api/disks/*` proxies + the `StoragePath` registry). This is a controller-only UI/orchestration layer
over those.
- **Storage overview** (`settings.html`, driven by `GET /api/disks`): the agent's live disk view — name,
type, state, device, mount, class, and the **`data_bearing` badge** + registered cross-reference.
- **Guided init** (`/settings/storage/init` + `POST /api/storage/init`): pick a disk → format → resolve
the new fs UUID from the re-listed disks → assign (mount) → register the `StoragePath`. **A data-bearing
device is REFUSED** by the agent; the UI surfaces the exact `felhom-opsign -op storage_wipe -host … -durable-id …`
command and stops — **there is no force-format path** (the gate is the agent's; the controller has no
destructive authority).
- **Guided attach** (`/settings/storage/attach` + `POST /api/storage/attach`): non-destructive — resolve
the existing fs UUID → assign → register.
- **Eject** (`POST /api/storage/eject`): benign unmount (data preserved) + deregister, surfacing the
agent's dependent-guest warning.
- **`agentapi`**: `DiskInfo` gains `DurableID` (+ `FSUUID()` to strip the `uuid:` prefix — the assign
key); `FormatResult` gains `PendingOp` (+ `OpsignCommand()`), now parsed from the agent's 403 body
(the old path discarded it). Pairs with `felhom-agent` v0.22.0, which exposes `durable_id` in `/disks`.
- **Honest buttons**: init/attach are wired; migrate (drive + per-stack) is disabled "Hamarosan" — no 404s.
- **De-priv template debt (Phase 3)**: removed the dead `CrossDrive*` blocks in `deploy.html` (the "2.
mentés" form + 3 JS fns) and `backups.html` (the run buttons + 2 JS fns) — they referenced fields the
de-privileged handlers no longer provide (a `gt/eq` over a missing field 500s the page).
- Migration (controller-side rsync) is intentionally deferred to its own slice (the migrate buttons are
disabled, not dead).
- Tests: the init refusal surfaces the `pending_op`/opsign and performs **no** assign/register; success
assigns with the resolved UUID + registers the expected `StoragePath`; a template-parse test guards all
pages.
### v0.42.1 — real Let's Encrypt cert: wildcard proactive issuance via the controller route (2026-06-11)
The base-infra traefik obtained **no** real cert (acme.json empty) — both routers relied on the
websecure entrypoint-default `certResolver`, which does not trigger proactive DNS-01 issuance, so
everything ran on traefik's self-signed default (masked externally by the tunnel's `noTLSVerify`).
This blocks LAN-direct (a LAN client TLS-handshakes straight to traefik and needs the real cert).
- **`infra.RenderControllerRoute(domain, wildcardTLS)`** — the always-present controller route is now
the **wildcard-issuance anchor**: when DNS-01 ACME is configured it carries router-level
`tls.certResolver: letsencrypt` + `tls.domains: [{main: "*.<domain>", sans: ["<domain>"]}]`, so
traefik **proactively obtains `*.<domain>` + apex at startup** via Cloudflare DNS-01. Every other
router (filebrowser, future apps) then serves that one wildcard by SNI match — **no per-app
certresolver labels**, real cert ready before the first client connects. `stacks.wireController`
passes `wildcardTLS = (CFAPIToken != "" && Email != "")`.
- **Empirically established (staging on 9201):** traefik v3 issues from a **router-level** `tls.domains`
but **NOT** from the entrypoint-level `http.tls.domains` (acme.json stayed empty with the latter). The
v0.42.0 attempt (entrypoint `domains` + `TraefikData.Domain`) was reverted accordingly.
- Validated staging→prod on guest 9201 (Fake LE wildcard → real LE wildcard), then GATE: `felhom.<domain>`
+ `files.<domain>` return `200 0` (real wildcard cert, TLS verify OK) direct-to-guest from a real LAN host.
### v0.41.2 — fix controller-route auto-connect + dead dashboard cross-drive block (2026-06-11)
Two fixes found while live-validating v0.41.1 routing on guest 9201:
- **`containerOnNetwork` false-positive (v0.41.1 regression):** the membership check used
`{{index .NetworkSettings.Networks "traefik-public"}}`, whose output for an absent key is `<nil>`
(non-empty) — so `wireController` thought the controller was already attached and **skipped the
`docker network connect`**. traefik then matched the route but 502'd (backend unresolvable). Fixed by
listing the network names and matching exactly. Live: `felhom.<domain>` now reaches the controller.
- **Dead cross-drive dashboard block (pre-existing, slice-8C leftover):** `dashboard.html` still
referenced `.CrossDriveTotal/.CrossDriveConfigured/.CrossDriveFailed`, which the de-privileged
dashboard handler stopped providing — so `gt <nil> 0` **500'd the entire dashboard**. Only surfaced
now because v0.41.1 finally made the dashboard reachable. Removed the dead block (cross-drive backup
is the host agent's job since 8C).
### v0.41.1 — wire the controller dashboard into traefik (`felhom.<domain>` routing) (2026-06-11)
Completes v0.41.0: the base-infra bring-up stood up traefik/cloudflared/filebrowser but nothing routed
the **controller itself** through traefik, so `felhom.<domain>` 404'd (live-confirmed: controller on
`bridge` only, no traefik labels, empty `dynamic/`). filebrowser self-registers via Docker labels +
network membership baked into its compose; the controller can't — it's started by the golden bootstrap
*before* `traefik-public` exists, and the v2 `bootstrap.json` carries no domain (it comes from the hub
pull). So the wiring must happen post-pull.
- `infra.RenderControllerRoute(domain)` — a traefik file-provider dynamic route:
`Host(felhom.<domain>)` → `http://felhom-controller:8080` on websecure (`tls: {}` inherits the
entrypoint's default `letsencrypt` resolver when ACME is configured, else self-signed).
- `EnsureBaseStack` now calls `wireController`: writes `dynamic/controller.yml` (write-if-changed, so the
traefik file watcher doesn't reload every health tick) and `docker network connect traefik-public
felhom-controller` (idempotent — skipped when already attached) so traefik can resolve the controller
by name. Runs on first boot and every self-heal tick. The Section-G shared `/opt/docker/stacks` mount
means traefik picks up the dynamic file live.
- Diagnostic confirmed the tunnel chain was already healthy (token tunnel-id matches the DNS tunnel;
CF ingress `*.<domain> → https://traefik`); the only gap was this controller wiring.
### v0.41.0 — first-boot base-infrastructure bring-up + self-heal (+ Section-G mount fix) (2026-06-11)
Lockstep with `felhom-agent` v0.20.0 + a golden rebake. A freshly-onboarded controller came up ONLINE
but **Health = FAIL: protected containers not running — traefik, cloudflared, filebrowser**: nothing
ever deployed the base stack on a Proxmox bootstrap (it was only ever created by the bare-metal
`scripts/docker-setup.sh`), and the health loop only *detected* the gap. This release makes the
controller stand up its own base infrastructure.
- **New `internal/infra` package** — pure renderers (`//go:embed` templates lifted verbatim from
`scripts/docker-setup.sh`) for traefik (`traefik.yml` + compose + a 0600 `.env` carrying the CF DNS
token only when set), cloudflared (compose; `TUNNEL_TOKEN`), and filebrowser (compose + `config.yaml`).
**Image tags are PINNED here as the single source of truth** — `traefik:v3.6.7`,
`cloudflare/cloudflared:2026.6.0`, `gtstef/filebrowser:1.3.3-stable` (no `:latest`). The web
FileBrowser sync path now **delegates** to `infra` so the pins can never diverge.
- **`stacks.Manager.EnsureBaseStack`** (`internal/stacks/infra.go`) — creates the `traefik-public`
network, then deploys traefik → cloudflared → filebrowser under `${stacks_dir}/<name>`. **Single-flight**
(TryLock — it's fired from both first-boot and every health tick), **idempotent** (skips a stack whose
container is already running), **non-fatal** (logs, never crashes). cloudflared is deployed only when a
tunnel token is configured; filebrowser is not overwritten if its compose already exists (preserves the
storage mounts the web sync path manages).
- **Triggers** (`cmd/controller/main.go`): first-boot bring-up after stack init (goroutine, non-fatal);
self-heal calls `EnsureBaseStack` unconditionally on every `system-health` tick (decoupled from the
issue strings — safe because of the single-flight + idempotency).
- **Dynamic protected set** (`monitor.EffectiveProtected`): cloudflared counts as a protected container
only when a tunnel token is configured, so a LAN-only node doesn't report FAIL forever for a stack it
intentionally skips. Detection and the bring-up condition agree.
- **Section-G fix (in `felhom-agent` build-golden.sh):** the controller writes compose stacks under
`/opt/docker/stacks` inside its container, but the bootstrap `docker run` never bind-mounted that path,
so the guest daemon resolved every relative bind source on the guest filesystem (empty dirs) — breaking
**all** bind-mounted stacks (base infra + customer apps). Fixed with a same-path host bind
(`-v /opt/docker/stacks:/opt/docker/stacks`). Empirically confirmed on guest 9201 (probe printed
`cat: read error: Is a directory` before, `hello-from-controller` after).
- Tests: non-hollow `infra` render tests (customer params present, no `:latest` survives, both ACME/CF
branches render, `.env` 0600, rendered YAML parses), `EnsureBaseStack` single-flight, and
`EffectiveProtected`.
### v0.40.0 — bootstrap pull+merge onboarding (controller pulls its config from the hub) (2026-06-11)
Lockstep with `felhom-agent` v0.19.0. Fixes the onboarding 401: a freshly provisioned guest used to
seed a "configured" controller.yaml from the agent's **host** hub key, which the hub's `/api/v1/report`
(customer-scoped auth) rejects → the controller could never report ONLINE. Now the controller **pulls**
its full controller.yaml from the hub on first boot (the hub mints the **customer-scoped** key) and
**merges in** the per-guest `local_api` block.
#### Changed — bootstrap contract `v1 → v2` (`internal/bootstrap`)
- `SchemaV1 → SchemaV2 = "felhom.bootstrap/v2"`. `BootstrapCustomer` drops `name`/`domain`/`email` (keeps
`id`); `BootstrapHub` drops `api_key`/`host_id`, adds **`retrieval_password`** (SECRET). `local_api`
unchanged. A non-v2 schema → setup mode.
- **`MaybeIngest(configPath, cfg, logger, pull PullFunc)`** — new injected `pull` arg (decision (b): keeps
`bootstrap` from importing the heavy `internal/report` package; wired in `main.go` to `report.PullConfig`).
Flow: idempotent (configured → return, **no pull**) → parse + validate v2 → **pull** the hub config with
bounded retry (1 + 3 backoff attempts on transient `ErrPullTransient` only; auth/not-found fail fast) →
**merge** the per-guest `local_api` at the YAML-map level (preserves every hub-emitted field — assets,
CF, backup) → write 0600 atomic → reload. Fail-safe throughout: a hub outage at first boot leaves the
guest in setup mode (the manual wizard remains the fallback), never crashes.
- New sentinel **`ErrPullTransient`**; `main.go`'s pull adapter maps `report.ErrHubUnreachable` onto it
(transient/retryable) and passes auth/not-found through as permanent. Removed `configFromBootstrap`
(the host-key-seeding path) and the struct-marshal writer.
#### Tests (`internal/bootstrap`)
- Pull+merge (asserts the merged controller.yaml carries the **customer** key + identity + a preserved
unmodeled `assets.source_url` **and** the bootstrap's `local_api`, with **no host key**); idempotency
(pull **never invoked** when configured); transient-retry (N attempts then setup); permanent-no-retry;
non-v2 schema reject; missing-required reject; malformed/absent. Cross-repo render→ingest round-trip
verified against the agent's v2 renderer. `go build ./... && go test ./...` green.
### v0.39.1 — 8C orphan-template cleanup (source hygiene) (2026-06-11)
Dead-template removal — no behaviour change. Slice 8C de-privileged the controller and retired the
disk/storage/restore web handlers (`storage_handlers.go`, `handler_restore.go` and the `/api/storage/*`
+ `/api/restore/*` routes), but five HTML templates that those handlers rendered were left behind.
They have **zero** `.go` references, **zero** cross-template `{{template …}}` references, no route, and
no nav entry; the embed is a glob (`//go:embed templates/*.html templates/*.css`), so deleting them is
safe and the remaining 14 templates still embed cleanly.
#### Removed (`internal/web/templates`)
- `storage_init.html`, `storage_attach.html`, `migrate.html`, `migrate_drive.html`, `restore.html` —
orphaned pages for removed endpoints. Re-confirmed unreferenced before deletion
(`grep -rn` over `internal/`: only the templates' own `{{define}}` lines matched).
#### Noted, not changed (dead-but-harmless restic/cross-drive remnants)
- Two never-called notifier methods `NotifyCrossDriveCompleted`/`NotifyCrossDriveFailed`
(`internal/notify/notifier.go:353,359`) and a vestigial `crossdrive_failed` entry in the
notification-events list (`internal/web/handlers.go:937`) that still renders a settings toggle for an
event that can no longer fire. Plus restic config fields/comments in `config/config.go`,
`settings/settings.go`, `report/types.go`. None are live emitters — left in place, flagged for a
future dedicated cleanup.
### v0.39.0 — slice 9: host metrics in the controller (customer host-health view) (2026-06-10)
The customer-facing half of slice 9. Pairs with `felhom-agent` v0.14.0. The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read the host. The monitoring page now shows the **real Proxmox box** — CPU% + load, memory used/total, **CPU/chassis temperature** (or "n/a" when the hardware exposes none), uptime, and **per-storage capacity** (used/total bar, thin-pool fill, disk temp/wear) — proxied from the agent's new `GET /host/metrics`.
#### Added (`internal/agentapi`)
- **`Client.HostMetrics(ctx)`** — calls the agent's `GET /host/metrics` over the leaf-pinned, per-guest-token channel (same client as the 8C disk proxy) and returns `HostMetricsResponse` (host block + per-storage targets). New mirror structs `HostMetrics` (with nullable `CPUTempC`), `StorageTarget`, `ThinPoolFill`, `SmartSummary` (subset — only the fields the UI renders; unknown wire keys ignored).
#### Added (`internal/web`)
- **`ServeHostMetricsAPI`** (`agent_host_metrics_handler.go`) — a thin read-only proxy: `GET /api/host-metrics` → agent `GET /host/metrics`. Returns the `{ok,data,error}` envelope; 503 when the local API is not configured (unprovisioned guest), 502 on an agent error. Wired in `main.go` behind `RequireAuth` (GET-only → no CSRF wrapper).
- **Monitoring view** (`templates/monitoring.html`): a new **"Szerver állapota (gazdagép)"** card at the top renders the agent's host block + per-storage capacity bars (reusing the existing `system-bar`/`storage-item` styling). `cpu_temp_c: null` renders as **"n/a"** cleanly. Polls `/api/host-metrics` every **8 s** while the page is open (the host view is a live snapshot, distinct from the controller's own 60 s metric charts); shows a yellow "nem elérhető" banner when the agent is unreachable.
#### Tests
- `agentapi/host_metrics_test.go`: decodes host + storage (thin-pool, SMART temp + NVMe wear), USB drive's null SMART, and a null `cpu_temp_c` → nil pointer.
### v0.38.0 — slice 8B.2: quiesce downtime optimization (resume at `snapshotted`) (2026-06-10)
The controller half of slice 8B.2. Pairs with `felhom-agent` v0.13.0. The quiesce loop now resumes
the app at the **`snapshotted`** phase (storage snapshot taken) instead of `done` — app downtime
drops from *whole-backup* to *until-snapshot* (seconds), with no loss of app-consistency (the
snapshot froze the app-stopped state).
#### Changed (`internal/quiesce`)
- The status-poll loop **resumes (`StartStack` + clears the marker) at `snapshotted`**, then **keeps
polling to `done`/`failed`** — so a new backup isn't started until this one truly finishes, and a
post-snapshot failure is observed (the backup isn't "successful" until `done`; resuming early does
not mark it done).
- **Fallback preserved:** if `snapshotted` never arrives (stop/downgraded mode), it resumes at `done`
exactly as 8B. **Crash-safety unchanged:** marker written before stop; guaranteed unquiesce;
startup `Recover()`. A backup that fails *after* `snapshotted` is harmless — the app is already up.
#### Tests
- resume at `snapshotted` (RESUME event before `done`, marker cleared, then tracked to `done`);
stop-mode fallback (resume at `done`, no `snapshotted`); fail-after-`snapshotted` (one resume, app
stays up); the 8B crash-safety tests stay green.
### v0.37.0 — slice 8C: controller de-privileging + disk management via the agent (2026-06-10)
The in-guest controller half of slice 8C (closes slice 8). The disk-execution subsystem moves to
the host agent (`felhom-agent` v0.12.0); the controller becomes **Docker-only with no disk
privileges** and drives disk management through the agent's local API. ~12.3k LOC retired.
#### Added
- **`internal/web/agent_disk_handlers.go`** — agent-backed disk API (`ServeDiskAPI`): `GET
/api/disks` (list + data-bearing flags), `POST /api/disks/assign` (mount), `POST /api/disks/eject`
(unmount + dependent-guest warning), `POST /api/disks/format`. Thin proxies over the slice-8A
`agentapi` client (leaf-pinned, own token). **Execution is the agent's**; the UX stays here.
A data-bearing format refusal (`agentapi.ErrFormatRefused`) is surfaced as **HTTP 409 "operator
authorization required"** (the 8C invariant — the agent inspects the device; the controller's
claim is irrelevant).
- **`internal/agentapi`**: `Disks`/`AssignDisk`/`EjectDisk`/`FormatDisk` + `ErrFormatRefused`.
#### Retired (moved to the host agent / obsolete)
- **`internal/storage/`** — the entire package (scan/format/attach/migrate/safety, DriveMigrator).
- **`internal/backup/`** — restic (`ResticManager`), `crossdrive` (`CrossDriveRunner`),
`restore_drives*`, `disk_layout`, `local_infra`, `restore_scan`, the restic path helpers, and the
drive-restore `restore_app*`. **`backup.Manager` surgically split to app-data only**: kept DB
dumps, Docker-volume tars, and per-app restore; dropped restic snapshots, cross-drive, per-drive
repo stats, integrity check, snapshot history. `RestoreApp` now restores from the on-disk
volume-tar dumps (snapshot/restic restore is the agent's domain).
- **`internal/report/infra_backup*` + `infra_pull`** (kept the setup fresh-install config download as
`config_pull.go`); **`internal/setup/scanner.go`** + the wizard's drive-recovery flows (restore is
the agent's job now); **`internal/monitor/watchdog.go` + `pinger.go`** (storage watchdog →
agent; Healthchecks.io pinging → the Hub owns monitoring); **`web/storage_handlers.go` +
`handler_restore.go`** (replaced by the thin agent-backed disk API).
- Wiring dropped from `main.go` / `api/router.go` / `web/server.go`: CrossDriveRunner, DriveMigrator,
storage watchdog, infra-backup push, the restic backup scheduler jobs (kept the **db-dump** job).
#### De-privileged
- `scripts/docker-setup.sh` controller compose template: dropped `privileged: true`, the `/mnt`
rshared bind, `/sys`, `/dev`, `/etc/fstab`, `/run/udev`. The golden's bootstrap `docker run`
(felhom-agent `build-golden.sh`) was already minimal (bootstrap config + data + docker socket).
#### Tests / build
- `go build ./...` + `go test ./...` green (app-data backup / stacks / quiesce / bootstrap / agentapi
/ disk-client tests pass). The data-bearing-format refusal is proven in `agentapi` tests.
### v0.36.0 — slice 8B: app-consistent backup quiesce loop (stack-stop) (2026-06-10)
The in-guest controller half of slice 8B (doc 03 §6/§8). Pairs with `felhom-agent` v0.11.0. An
agent-initiated vzdump is crash-consistent only (an LXC has no fsfreeze); this makes app-consistency
the controller's job — it stops its app stacks around the backup so the captured state is
clean-shutdown-consistent.
#### Added
- **`internal/quiesce`** — the background quiesce loop: poll the agent's `GET /backup/due` → when
due, **quiesce** (stop deployed, non-protected, running stacks) → `POST /backup` → poll
`GET /backup/status` to `done`/`failed` → **unquiesce** (restart exactly the stacks it stopped).
- **Crash-safety (the centerpiece — a stranded-down app is worse than a crash-consistent backup):**
a persisted **marker** (atomic, `0600`) written **before** stopping anything; **guaranteed
unquiesce** (a deferred closure restarts the stacks on a backup error, a status-poll error, the
max-quiesce bound, or context cancellation); a **max-quiesce-duration** hard bound that restarts
the app no matter what (the backup continues on the agent); **crash recovery** at startup
(`Recover()` restarts stacks left stopped by a mid-quiesce crash, then clears the marker); and the
marker as a **single-flight** guard.
- **`agentapi`**: `BackupDue` / `StartBackup` / `BackupStatus` methods + a `post` helper.
- **`stacks.Manager.RunningAppStacks()`** — deployed, non-protected, currently-up stacks (protected
infra — traefik/cloudflared/felhom-controller — is never stopped), sorted for deterministic order.
- **`config.QuiesceConfig`** (`quiesce`: enabled, poll_interval, status_poll_interval,
max_quiesce_duration). Wired in `main.go`: `Recover()` at startup, then the loop goroutine, gated on
the local API being configured (a provisioned guest) + quiesce enabled.
#### Tests
- happy path (stop → backup → poll done → restart exactly those, in order; marker cleared);
**backup-start failure → stacks STILL restarted**; failed phase → restarted; **max-quiesce guard →
restarted at the bound**; **crash recovery → marker stacks restarted + cleared**; single-flight (no
second backup while a marker is active); **only the stacks we stopped are restarted** (an
already-stopped stack is never started); and **marker-written-before-stop** ordering.
### v0.35.0 — slice 8A: bootstrap.json ingestion + pinned agent local-API client (2026-06-10)
The in-guest controller half of slice 8A (doc 03 §6). Pairs with `felhom-agent` v0.10.0. No
behaviour change for an already-configured controller; adds the first-run provisioning path.
#### Added
- **`internal/bootstrap`** — first-run **`bootstrap.json` ingestion** (config-contract decision (c)).
On startup, if the controller is NOT yet configured AND the host agent's back-half attached a
`bootstrap.json` config mount, the controller **seeds `controller.yaml` from it and comes up
configured, skipping the setup wizard**. Idempotent (an existing `controller.yaml` is **never**
clobbered) and fail-safe (a malformed/absent/missing-identity/unsupported-schema bootstrap leaves
the controller in setup mode — logs, never crashes). The agent emits the stable contract; the
controller owns the translation (the two stay decoupled).
- **`internal/agentapi`** — a minimal **pinned client** for the agent's local API. It reaches the
agent over the bridge, **pinning the agent leaf-cert SHA-256** from the bootstrap (fails closed on
mismatch — `VerifyPeerCertificate` exact leaf-DER match, the same pin convention the agent uses for
the Proxmox/PBS host certs), and authenticates with the per-guest bearer token. In 8A it exercises
`GET /storage` (connectivity + the controller learning its mounts); the `/backup/due` quiesce loop
is 8B.
- **`config.LocalAPIConfig`** (`local_api`: endpoint, fingerprint, token) — seeded from the bootstrap.
- **Startup probe** — when seeded with a local-API endpoint, the controller proves the channel at
boot and logs this guest's mounts (non-fatal).
#### Tests
- bootstrap: seeds when unconfigured (reloads configured, skips setup); never clobbers a configured
controller; stays in setup on malformed / missing-identity / unsupported-schema / absent bootstrap.
- agentapi: correct pin + token reaches `/storage`; a **wrong pin fails closed**; a bad fingerprint
is rejected at construction; colon-separated fingerprints are accepted.
### docs: reflow CLAUDE.md; unify REPORT/CHANGELOG convention; add no-secrets rule (2026-06-08)
#### Changed
- **Reflowed `CLAUDE.md`** — removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged.
- **Added the uniform REPORT/CHANGELOG convention**: `CHANGELOG.md` is the cumulative log (newest on top); `REPORT.md` is overwritten with the most-recent implementation only. Added an explicit **no-secrets** rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band). Docs/meta only — no code change, no version bump.
### Repo rename — `deploy-felhom-compose` → `felhom-controller` (2026-06-08)
#### Changed
- **Gitea repo renamed** `admin/deploy-felhom-compose` → `admin/felhom-controller` (via API). Sibling rename: `admin/proxmox-controller` → `admin/felhom-agent` (docs repo, future agent code).
- **Reference rework (no functional change)**: updated every reference to the old repo name across docs and scripts — clone URLs, clone dirs (`~/git/felhom-controller`), the customer bootstrap URL in `scripts/felhom-wipe.sh` / `scripts/README.md`, `controller/build.sh`, `controller/BUILDING.md`, `controller/README.md`, `CLAUDE.md`, `CONTEXT.md`, `TASK.md`. Local working-copy dirs and the build-server source clone (`192.168.0.180:~/git/`) renamed to match.
- **Intentionally unchanged**: Go module path `gitea.dooplex.hu/admin/felhom-controller` (already matches new name), Docker image path `gitea.dooplex.hu/admin/felhom-controller` (registry is namespaced by owner, not repo), binary name `felhom-controller`. Historical CHANGELOG entries left as-is (they record what was true at the time).
### Refactor — extract app-data-backup primitives into `internal/appbackup` (no behaviour change) (2026-06-08)
#### Changed
- **New package `internal/appbackup/`**: extracted the stateless, keep-side app-data backup primitives out of `internal/backup/` — DB dump discovery/execution (`dbdump.go`: `DiscoverDatabases`, `DumpAll`, `DumpOne`, `ValidateDump`, `ListDumpFiles`), Docker-volume/app-data discovery (`appdata.go`: `StackDataProvider`, `DiscoverAppData`, `ParseComposeNamedVolumes`, `ResolveDockerVolumeNames`, `HumanizeBytes`), and keep-side path helpers (`paths.go`: `FelhomDataDir`, `PrimaryBackupPath`, `AppDBDumpPath`, `AppVolumeDumpPath`, `AppDataDir`). Pure move — logic unchanged.
- **backup/appbackup_bridge.go** (new): re-exposes the moved symbols to the `backup` package via type/const aliases and one-line function forwarders, so the still-present disk/host-side code (restic, cross-drive, drive-mount) and the both-side consumers (web, api, report) compile unchanged.
- **appexport/export.go, storage/migrate.go, storage/migrate_drive.go**: rewired to import `internal/appbackup` directly and dropped their `internal/backup` import — these keep-side consumers are now independent of the delete-side code.
- **Why**: Part-2 prerequisite for the Proxmox port. Isolating the keep-side now (as a separate green, behaviour-identical commit) means the disk/host-side code can later be removed without breaking app-data backup or `appexport`. `appbackup` has zero references to restic/cross-drive/drive-mount and does not import `backup` (no import cycle).
- **Not moved (documented coupling)**: the `*Manager` methods `RunDBDumps`/`DumpAppVolumes`/`DumpAppVolumesSafe` (share one mutex/running-flag + status state with the delete-side `RunBackup`) and `RestoreAppFromTier2` (intrinsically reads the cross-drive mirror via `copyFile`/`AppSecondaryRsyncPath`) stay on `Manager`; they delegate to `appbackup` and are left for the later re-platform step.
### v0.34.0 — Backup safety: stop-before-dump, streaming restore, health check, per-app restic, infra configs (2026-02-28)
#### Changed
- **backup/backup.go**: `DumpAppVolumesSafe()` stops stack before volume dump, restarts after — prevents inconsistent tars of live database volumes (PostgreSQL, MariaDB, SQLite)
- **backup/backup.go**: `backupDrive()` includes per-app stack config dirs instead of full StacksDir; `controller.yaml` only on system drive — reduces snapshot duplication across drives
- **backup/crossdrive.go**: `VolumeDumper` interface extended with `DumpAppVolumesSafe()`; cross-drive backup uses safe variant for pre-backup volume dumps
- **backup/restore.go**: Tier 2 DB dump copy uses streaming `copyFile()` (io.Copy + atomic rename) instead of `os.ReadFile`/`os.WriteFile` — eliminates full-file memory allocation for large dumps
- **backup/restore.go**: Post-restore health check via `waitForHealthy()` polls container state (with docker ps refresh) for up to 90s after restore
#### Added
- **backup/appdata.go**: `RefreshAndIsRunning()` on `StackDataProvider` interface for reliable post-restore state checks (forces docker ps refresh before reading state)
- **report/infra_backup.go**: `InfraStack` now includes `DockerComposeB64`, `AppYamlB64`, `FelhomYamlB64` — actual stack config files for disaster recovery (derived from `GetStackComposePath`, no signature change)
### v0.33.0 — Docker volume backup + Tier 2 restore + restore dropdown fixes (2026-02-27)
#### Added
- **backup/backup.go**: `DumpAppVolumes()` exports Docker named volumes to tar files using `docker run alpine tar`; `runVolumeDumpsInternal()` runs volume dumps for all stacks in nightly schedule (Phase 1b between DB dumps and restic); volume dump dirs included in per-drive restic snapshots
- **backup/appdata.go**: `ResolveDockerVolumeNames()` resolves full Docker volume names with project prefix (e.g., `mealie_mealie_data` instead of `mealie_data`); `GetDockerVolumes()` added to `StackDataProvider` interface; `HasVolumeData` field on `AppBackupInfo`, `HasVolumes` on `StackSummary`
- **backup/paths.go**: `AppVolumeDumpPath()` returns `<drive>/felhom-data/backups/primary/<stack>/volume-dumps/`
- **backup/restore.go**: `RestoreAppFromTier2()` restores from cross-drive rsync mirror (config, HDD data, DB dumps, Docker volumes via rsync); `restoreDockerVolumes()` populates Docker volumes from tar files after Tier 1 restore; `restoreDockerVolumesFromDir()` for Tier 2 volume restore
- **backup/crossdrive.go**: `VolumeDumper` interface + `SetVolumeDumper()` for pre-backup volume dumps; `copyStackVolumeDumps()` copies volume tars to `_volumes/` in rsync mirror
- **backup/backup.go**: `ListSnapshotsForApp()` returns snapshots only from the app's home drive primary repo
- **backup/restic.go**: `Source` field on `SnapshotInfo` ("restic" or "rsync")
- **api/router.go**: `backupSnapshots()` now accepts `?stack=` param to filter by app's home drive; appends synthetic Tier 2 entry from cross-drive config when backup succeeded
- **web/handlers.go**: `backupRestoreHandler()` routes `tier2-rsync` snapshot ID to `RestoreAppFromTier2()`
- **web/templates/backups.html**: Import from `.fab` bundle link in restore section; `data-has-volumes` attribute on restore app options; volume-aware restore type banners; "Konfig + Adatok" label for volume-backed apps
#### Fixed
- **Volume name resolution bug**: `ParseComposeNamedVolumes()` returned short names but Docker Compose V2 uses `<project>_<name>` — fixed in both backup and export adapters via `ResolveDockerVolumeNames()`
- **Double Tier 1 in restore dropdown**: snapshots from non-home drives appeared because stacks dir is in every drive's primary repo — now filtered by app's home drive via `ListSnapshotsForApp()`
### v0.32.8 — Move optional config to deploy/settings page (2026-02-27)
#### Changed
- **web/templates/deploy.html**: Optional config fields (metadata providers, API keys) now render on the deploy/settings page instead of the app info page — consistent with integrations and geo-restriction which already live there
- **web/handlers.go**: `deployHandler` now passes `OptionalConfig`, `CurrentValues`, `HasOptionalConfig` to the deploy template for deployed apps; `appDetailHandler` cleaned up to remove optional config data
- **web/templates/app_info.html**: Removed optional config section (HTML + JS) — no longer rendered here
### v0.32.7 — Fix FileBrowser config not being read on fresh deployments (2026-02-27)
#### Fixed
- **web/handlers.go**: `generateFileBrowserCompose()` now sets `FILEBROWSER_CONFIG=/home/filebrowser/config.yaml` environment variable — the `gtstef/filebrowser` image bakes in `FILEBROWSER_CONFIG=/home/filebrowser/data/config.yaml` which reads a stale initial config from the data volume instead of the controller-managed bind mount. This caused fresh deployments to show only a single "srv" source, ignore per-drive sidebar entries, and create the database outside the persistent volume (triggering the "new database was created" warning on every container recreation)
#### Changed
- **scripts/docker-setup.sh**: Initial FileBrowser compose template also includes the `FILEBROWSER_CONFIG` override for consistency
### v0.32.6 — Format empty partitions on system disk (2026-02-27)
#### Added
- **storage/scan.go**: New `FormatablePartition` struct and `FormatablePartitions` field on `ScanResult` — detects empty (no filesystem), unmounted, non-system partitions on system disks
- **storage/scan_linux.go**: New `getSystemPartitionPaths()` resolves actual system partition device paths from fstab (more granular than `getSystemDiskNames()` which returns parent disk names). `ScanDisks()` now populates `FormatablePartitions` after enrichment
- **storage/safety_linux.go**: New `IsSystemPartition()` — checks if a specific partition is a system partition (/, /boot, /boot/efi, swap) or is currently mounted; more granular than `IsSystemDisk()` which blocks the entire disk
- **web/storage_handlers.go**: Scan API response now includes `formatable_partitions` array
- **web/templates/storage_init.html**: Init wizard shows formatable system-disk partitions as a separate selectable section with info banner, conditional warning text, and hidden partitioning progress step
#### Changed
- **storage/format_linux.go**: `FormatAndMount()` now uses `IsSystemPartition()` for partition-only operations (`CreatePartition=false`) instead of `IsSystemDisk()` — allows formatting empty data partitions on the system disk while still blocking system partitions
### v0.32.5 — USB badge fix + graceful Tier2 backup on disconnected/inactive/removed destinations (2026-02-27)
#### Fixed
- **system/mounts_linux.go**: `IsUSBDevice()` and `diskModel()` now strip findmnt bind-mount suffix (`[/subdir]`) before parsing device path — fixes USB badge and disk model not showing for drives mounted via the attach wizard
- **backup/crossdrive.go**: Disconnected source/destination drives now silently skip with WARN log instead of returning error — prevents noisy error aggregation in `RunAllScheduled()` and false "failed" counts
- **web/handlers.go + backup/crossdrive.go**: Tier2 destination check now covers drives **removed** from storage (not just marked disconnected) — `IsStoragePathKnown()` detects when destination path is no longer in any registered storage, UI shows yellow "Cél meghajtó leválasztva" and scheduler skips silently
- **web/handlers.go + backup/crossdrive.go**: Tier2 destination check now also covers **inactive** (Schedulable=false) drives — `IsStoragePathSchedulable()` detects when destination drive is deactivated, UI shows yellow "Cél meghajtó inaktív" and scheduler skips silently
#### Added
- **settings/settings.go**: New `IsStoragePathKnown(path)` method — returns whether a path belongs to any registered storage (connected, disconnected, or decommissioned); paths removed entirely return false
- **settings/settings.go**: New `IsStoragePathSchedulable(path)` method — returns true only if path belongs to a registered, active (Schedulable), non-disconnected, non-decommissioned storage
- **web/handlers.go**: New `Tier2DestDisconnected` and `Tier2DestInactive` fields on `AppBackupRow` — detect when Tier2 destination is disconnected/removed/inactive, sets yellow status dot instead of green/red
- **web/templates/backups.html**: New template branches for disconnected ("Cél meghajtó leválasztva") and inactive ("Cél meghajtó inaktív") Tier2 destinations — grayed-out info, warning badge, no "Futtatás most" button
### v0.32.4 — Controller telemetry: include controller in hub app telemetry (2026-02-27)
#### Added
- **report/telemetry.go**: Include the `felhom-controller` container as a special entry in the `app_telemetry` array sent to the hub — reuses all existing hub telemetry infrastructure (memory trends, known issues, fleet aggregation) with zero hub-side changes
- **report/telemetry.go**: New `buildControllerTelemetry()` function collects controller container metrics (memory, CPU) and log scan results (warnings, errors, deduplicated issues)
### v0.32.3 — Logging cleanup: consistent tags, dedup, standardized prefixes (2026-02-26)
#### Fixed
- **All modules**: Standardized `[LEVEL] [module]` format across every log line — added missing module tags (`[stacks]`, `[backup]`, `[cloudflare]`, `[sync]`, `[scheduler]`, `[storage]`, `[monitor]`, `[metrics]`, `[report]`, `[settings]`, `[setup]`, `[api]`, `[integrations]`, `[selfupdate]`, `[assets]`, `[web]`)
- **Removed duplicate logs**: ScanStacks double completion, GetLogs INFO+DEBUG, LoadAppConfig WARN+DEBUG, copyStackDBDumps DEBUG+INFO, invalidateAllSessions INFO+DEBUG
- **Standardized stale prefixes**: `[CF]`/`[CF-DEBUG]` → `[INFO/DEBUG] [cloudflare]`, `[SYNC]` → `[LEVEL] [sync]`, `[SCHED]` → `[LEVEL] [scheduler]`, `[API]` → `[LEVEL] [api]`, `[STORAGE]` → `[storage]`, `[HEALTH]` → `[monitor]`, `[ROLLBACK]`/`[ROLLBACK-ERROR]` → `[LEVEL] [storage]`, `[DEBUG-SIM]` → `(simulation)`
- **Fixed wrong log levels**: Restic restore start WARN→INFO, ungated DEBUG lines in crossdrive/dbdump/onlyoffice/alerts removed or gated
- **Improved vague messages**: Settings SetDisconnected/SetDecommissioned now include storage path and migration target
- **Added missing logs**: `execCommand()` error, `DiscoverAppData()` completion, `BuildInfraBackup()` completion
### v0.32.2 — Comprehensive INFO/WARN/ERROR logging across all modules (2026-02-26)
#### Added
- **stacks/manager.go**: INFO logs for status refresh container/stack counts, log fetching, encryption migration, ScanStacks completion
- **stacks/deploy.go**: INFO logs for config updates, InjectMissingFields summary; ERROR logs for SaveAppConfig failures; WARN for LoadAppConfig errors
- **stacks/delete.go**: INFO log for ParseComposeHDDMounts result count
- **stacks/metadata.go**: Fixed LoadMetadata error to use `log.Printf` instead of `fmt.Fprintf(os.Stderr)`
- **backup/backup.go**: WARN for perDriveRepoStats failures; INFO for drive stats, aggregate stats, dump file count, snapshot history save
- **backup/crossdrive.go**: INFO for cross-drive backup start/completion with success/fail counts; ERROR for rsync failures; INFO for DB dump copy counts
- **backup/restic.go**: INFO for Snapshot and Check success
- **backup/dbdump.go**: INFO for DiscoverDatabases count; INFO for DumpAll start/completion
- **backup/restore_drives_linux.go**: INFO for fstab entry additions
- **backup/local_infra.go**: INFO for backup version pruning with kept/removed counts
- **cloudflare/geosync.go**: Standardized all `[GEO]` prefixed logs to `[INFO]/[WARN]/[ERROR] [cloudflare]` format
- **scheduler/scheduler.go**: Standardized all `[SCHED]` prefixed logs to `[INFO]/[WARN]/[ERROR] [scheduler]` format
- **sync/sync.go**: INFO for catalog sync start/completion; ERROR for git/network failures; WARN for file copy errors (replaced `[SYNC]` prefix)
- **report/pusher.go**: WARN for Push and InfraBackup push failures
- **report/builder.go**: INFO for BuildReport start
- **monitor/healthcheck.go**: WARN for CPU/memory/disk/temperature threshold breaches; INFO for health check result status
- **system/mounts_linux.go**: WARN for unsafe backup destinations and storage path probe failures
- **settings/settings.go**: INFO for settings load/save, storage path add/remove, disconnect/decommission, pending events; ERROR for save failures
- **storage/attach_linux.go**: INFO for disk attach start/success; ERROR for attach failures
- **storage/scan_linux.go**: INFO for disk scan start/completion with count
- **storage/format_linux.go**: INFO for format start/success; ERROR for format failures
- **storage/migrate.go**: INFO for migration start/completion; ERROR for migration failures
- **integrations/manager.go**: ERROR for integration apply failures; WARN for context build and env load failures
- **integrations/lifecycle.go**: Added `[integrations]` module tag to all logs; upgraded re-apply failure from WARN to ERROR
- **integrations/onlyoffice_filebrowser.go**: ERROR for all Apply/Revoke error paths
- **integrations/onlyoffice_nextcloud.go**: ERROR for all Apply/Revoke error paths
- **selfupdate/updater.go**: INFO for up-to-date and update-available results; INFO/ERROR for compose file updates
- **selfupdate/state.go**: INFO for state cleared
- **assets/syncer.go**: ERROR for manifest save failures (previously silent); changed sync failure log from WARN to ERROR
- **appexport/restore.go**: INFO for import start
- **web/auth.go**: INFO for logout/session invalidation/session cleanup; WARN for unauthorized API requests
- **web/server.go**: WARN for 404 Not Found on unknown routes
- **web/handlers.go**: INFO for default storage path and schedulable state changes
- **web/handler_restore.go**: INFO for restore-all initiation
- **web/handler_export.go**: ERROR for export/import start failures
- **web/storage_handlers.go**: INFO for disk disconnect/reconnect/restart-apps completion
- **api/router.go**: ERROR for stack action failures, backup snapshot listing failures, metrics query failures
#### Changed
- **stacks/healthprobe.go**: Summary log now always prints — WARN when unhealthy, INFO when all ok (was debug-only for all-ok)
- **backup/restore.go**: Changed RestoreApp start log from `[WARN]` to `[INFO] [backup]`
- **backup/restore_app_linux.go**: Changed restoreUserData/restoreDBDumps failure logs from `[WARN]` to `[ERROR]` where data loss could occur
### v0.32.1 — Comprehensive debug logging across all modules (2026-02-26)
#### Added
- **stacks/delete.go**: Debug logging for DeleteStack/RemoveStack with stack state, HDD mounts, compose output, path removal; GetStackHDDData/GetStackBackupData path scanning
- **stacks/manager.go**: Debug logging for ScanStacks per-stack discovery, refreshStatusLocked container resolution, Start/Stop/Restart pre-operation state, MigrateEncryption progress, getCatalogTemplateSlugs count
- **stacks/deploy.go**: Debug logging for UpdateStackConfig/UpdateOptionalConfig changed keys, InjectMissingFields per-stack checks, SaveAppConfig encryption counts, LoadAppConfig results
- **stacks/healthprobe.go**: Debug logging for per-target interval calculations and target collection summary
- **backup/restic.go**: `debug` field + `SetDebug()` method; debug logs for Snapshot/Prune/Check/ListSnapshots/LatestSnapshot/Stats/RestoreAppData with timing, sizes, and command details
- **backup/restore_scan.go**: Debug logging for ScanDrivesForBackups drive/app scanning with per-drive availability and backup component summary
- **backup/restore_app_linux.go**: Debug logging for RestoreAppFromBackup step timing, restoreUserData per-dir rsync, restoreDBDumps per-file copying
- **backup/restore_drives_linux.go**: Debug logging for MountDrivesFromLayout device discovery, mount strategy selection, fstab checks
- **cloudflare/geosync.go**: `debug` field + `SetDebug()` method; debug logs for Sync zone/ruleset resolution, existing/desired rule diffing, rule create/update/delete operations
- **cloudflare/waf.go**: Debug logging for GetCustomRulesetID/GetRules/GetFelhomRules counts, CreateRule/UpdateRule expression snippets
- **cloudflare/zone.go**: Debug logging for GetZoneID progressive domain lookup attempts
- **integrations/manager.go**: `debug` field + `SetDebug()` method; debug logs for Toggle validation/timing, ListForProvider counts, buildApplyContext details, ReapplyConfigForTarget per-integration progress
- **integrations/lifecycle.go**: Debug logging for OnStackStop/OnStackStart/OnStackRemove with integration counts, state checks, revoke/re-apply operations
- **integrations/onlyoffice_filebrowser.go**: Debug logging for Apply/Revoke config path, JWT secret presence, office URL
- **system/**: Package-level `DebugLogger` variable; debug logs for GetInfo timing/summary, readMemInfo/readDiskUsage/readLoadAvg/readTemperature raw values, CPU collector samples, GetDiskUsage/GetFSInfo/CheckBackupDestination/ProbeStoragePath/IsUSBDevice details
- **monitor/pinger.go**: `debug` field + `SetDebug()` method; debug logs for Ping/Fail/Start with UUIDs, send URL/attempts/response status
- **settings/settings.go**: `debug` field (json:"-") + `SetDebug()` method; debug logs for Load counts, save data size, AddStoragePath/RemoveStoragePath, SetDisconnected/SetDecommissioned, AddPendingEvent/DrainPendingEvents, SetGeoRestriction, SetIntegrationState, AutoDiscoverStoragePaths
- **scheduler**: `debug` field + `SetDebug()` method; debug logs for job registration, execution timing, daily job wait calculations
- **storage/**: Consistent `[DEBUG] [storage]` prefix; scan timing; drive migration debug logging
- **metrics/logscanner**: Debug logging for per-container scan timing, error/warning counts
- **api/router**: `debug` field + `SetDebug()` method; logs incoming API requests and handler entry points
- **selfupdate**: Expanded debug coverage with `dbg()` helper for TriggerUpdate preconditions, performUpdate step transitions, docker pull timing
- **assets/syncer**: Expanded debug coverage with `dbg()` helper for per-file hash comparison, download timing, manifest fetch details
- **web/auth.go**: Debug logging for RequireAuth middleware decisions, login attempts (IP, success/fail), session creation/cleanup
- **web/handlers.go**: Debug logging for deploy/restore/settings/storage handler entry points with key parameters
- **web/handler_restore.go**: Debug logging for restore page, status polls, restore-all execution per-app timing
- **web/storage_handlers.go**: Debug logging for all storage API operations (scan, init, migrate, disconnect, reconnect, attach, cleanup)
- **web/server.go**: Debug logging for NewServer initialization, template loading, ServeHTTP request routing
- **main.go**: Wire `SetDebug()` for settings, pinger, geoSync, integrationMgr, scheduler, apiRouter
### v0.32.0 — App export/import (.fab bundles) (2026-02-26)
#### Added
- **App export**: Per-app export to `.fab` bundles containing config, database dump, and all user data (HDD bind mounts or Docker named volumes)
- **App import**: Restore apps from `.fab` bundles — works for both existing and new apps (standalone import page)
- **Password protection**: Optional AES-256-CTR + HMAC-SHA256 encryption with scrypt key derivation for exported bundles
- **Pre-export estimation**: Size estimation with free space check before starting export
- **Export UI**: New export page accessible from app info header with drive picker, password field, stop-app checkbox, and real-time progress tracking
- **Import UI**: Standalone import page (`/import`) scans all registered storage drives for `.fab` files, shows manifest details, and handles encrypted bundles with password prompt
- **FileBrowser link**: After export, link to open the exports directory in FileBrowser
- **Bundle format**: `{appname}_{timestamp}.fab` — tar.gz internally with `manifest.json`, `config/`, `database/`, `data/` directories
- **New package**: `internal/appexport/` — export/import engine with provider adapter pattern (same as backup.StackDataProvider)
- **API endpoints**: `/api/export/estimate`, `/api/export/start`, `/api/export/status`, `/api/export/bundles`, `/api/export/manifest`, `/api/export/import`, `/api/export/import/status`
#### Changed
- **backup/appdata.go**: Exported `ParseComposeNamedVolumes` (was lowercase) for reuse by appexport package
### v0.31.7 — Infra backup retention + version picker (2026-02-26)
#### Changed
- **docker-setup.sh hub mode**: `--hub-customer` now generates a minimal `controller.yaml` (no `customer.id`) instead of installing the full hub config — this triggers the setup wizard on first run, giving the user a choice to restore from an infra backup or start fresh
- **docker-setup.sh**: Hub credentials are passed to the controller via `FELHOM_SETUP_CUSTOMER_ID` and `FELHOM_SETUP_PASSWORD` environment variables so the setup wizard auto-fills them
- **Local infra backup**: `WriteLocalInfraBackup()` now rotates previous backup into `history/` subdirectory before writing new files (keeps last 5 versions per drive)
- **Setup wizard scan results**: Table now shows app names/count, disk count, and "korábbi" badge for historical versions
#### Added
- **Setup wizard hub pre-seeding**: When deployed with `--hub-customer`, the wizard auto-detects pre-seeded credentials and auto-processes Hub API calls (no manual form entry needed)
- **Hub mode welcome page**: Shows three options instead of two — "Visszaállítás a Hub-ról" (auto-connects to Hub), "Helyi mentés keresése" (local drive scan), "Friss telepítés" (fresh config download)
- **Auto-process fallback**: If Hub auto-connect fails, the wizard clears the pre-seeded password and falls back to the manual form with the error displayed
- **Hub backup version picker**: When multiple backup versions exist on the Hub, the setup wizard shows a version picker page (date, controller version, app names, disk count) — user selects which version to restore
- **Local backup history restore**: Setup wizard can restore from historical versions found in `history/` subdirectory on local drives
- **`ReadLocalInfraHistory()`**: Scans `history/` directory for all retained backup versions with rich metadata (stack names, disk count, integrity status)
- **`ReadLocalInfraBackupFromHistory()`**: Reads a specific historical version by timestamp prefix
- **`PullRecoveryVersion()`**: Fetches a specific backup version from the Hub recovery endpoint via `?version=ID` parameter
#### Fixed
- **Bind mount write**: `atomicWriteFile()` now falls back to direct write when rename fails (fixes "device or resource busy" on Docker bind-mounted `controller.yaml`)
- **Drive mounting after restore**: Restore flow now calls `MountDrivesFromLayout()` to mount drives by UUID and add fstab entries — previously drives referenced in the infra backup were not mounted, causing "Adattároló nem elérhető" warnings
- **Post-restore redirect**: UI now polls until the controller is actually up instead of using a fixed 5-second timeout (which was too short for container restart)
- **FileBrowser DB reset scoped to restore**: `SyncFileBrowserMounts()` no longer resets the FileBrowser database volume on source changes — only the post-restore startup path (`SyncFileBrowserMountsReset`) does, preserving user accounts, permissions, and share links during normal storage operations
### v0.31.6 — UI: Brand-consistent button & card styling (2026-02-25)
#### Changed
- **Buttons**: Replaced traffic light colors (green/yellow/red) with brand-consistent palette — primary actions use blue gradient, secondary actions use ghost/outline, destructive actions show red tint on hover only (modal confirmations keep filled red)
- **Card borders**: Running apps now show a subtle blue glow instead of green top border; all other states have neutral borders
- **Status badges**: Running state badge uses brand blue instead of green
- **Button alignment**: Cards use flexbox column layout with `margin-top: auto` on actions — buttons always align to the bottom regardless of card content height
- **Dashboard cards**: Left border indicator changed from green to blue for running apps
### v0.31.5 — Fix Nextcloud-OnlyOffice callback URL + trusted_domains (2026-02-25)
#### Fixed
- **StorageUrl trailing slash**: `http://nextcloud` → `http://nextcloud/` — without trailing slash, Nextcloud's OO connector concatenates the hostname with `/apps/...` path, producing `http://nextcloudapps/...` (unresolvable hostname)
- **trusted_domains**: OO Document Server callbacks arrive with `Host: nextcloud` header; added `nextcloud` to Nextcloud's trusted_domains so these internal callbacks are not rejected
### v0.31.4 — Fix FB container not restarting + OO mixed content (2026-02-25)
#### Fixed
- `SyncFileBrowserMounts` now uses `--force-recreate` so the container always restarts and picks up config.yaml changes (bind mounts are invisible to `docker compose up`)
- OnlyOffice compose template: added Traefik `X-Forwarded-Proto=https` middleware to fix mixed content errors when OO generates `http://` URLs behind HTTPS proxy
- Nextcloud integration: added `StorageUrl=http://nextcloud` for internal file download callbacks from OO Document Server
### v0.31.3 — Fix FileBrowser integration config persistence (2026-02-25)
#### Fixed
- FileBrowser integration config (OnlyOffice URL, JWT secret) was lost after `SyncFileBrowserMounts` regenerated `config.yaml` — the async `OnStackStart` re-apply hook failed due to timing issues
- New `ReapplyConfigForTarget()` method applies integration config synchronously between config generation and container restart, ensuring it survives regen cycles
### v0.31.2 — Show FileBrowser URL on app card (2026-02-25)
#### Fixed
- Protected stacks (e.g. FileBrowser) now show their subdomain URL link on the app card — condition relaxed from `Deployed` to `Deployed OR Protected`
### v0.31.1 — Move integration & geo settings to deploy page (2026-02-25)
#### Changed
- **Integration toggles** and **geo-restriction settings** moved from app info page to deploy/settings page (user feedback: settings belong on the "Beállítások" page)
- Data wiring moved from `appDetailHandler()` to `deployHandler()` in handlers.go
### v0.31.0 — App-to-App Integration Framework (2026-02-25)
#### Added
- **Generic integration framework** (`internal/integrations/`) — Extensible system for connecting deployed apps to each other via toggle switches on the provider's app info page
- **OnlyOffice → FileBrowser integration** — Toggle enables document editing in FileBrowser by patching `config.yaml` with OnlyOffice URL and JWT secret
- **OnlyOffice → Nextcloud integration** — Toggle installs and configures the OnlyOffice connector app via `occ` CLI commands
- **Integration lifecycle hooks** — Integrations auto-suspend when provider or target stops, auto-re-enable when both are running again, permanently removed on app deletion
- **Integration API endpoints** — `GET /api/integrations/{provider}` (list), `POST /api/integrations/{provider}/{target}` (toggle)
- **Integration UI** — "Integrációk" section on app info page with toggle switches, status badges, and target availability indicators
- **`IntegrationDef`** in `.felhom.yml` metadata — Apps can declare integrations with target app slug, label, and description
- **`IntegrationState`** in `settings.json` — Persistent integration state with enabled/status/error tracking
- **SyncFileBrowserMounts re-apply** — After config regeneration (which overwrites config.yaml), active integrations are automatically re-applied
### v0.30.7 — Monitoring: Fix Memory Legend Overflow (2026-02-25)
#### Fixed
- **Memory legend overflow** — Legend items in the memory distribution chart now wrap properly instead of overflowing off-screen (`flex-wrap`, `white-space: nowrap`)
#### Improved
- **Sort by consumption** — Memory distribution bar and legend are now sorted by memory usage (descending), largest consumers first
### v0.30.6 — Telemetry: Better Log Deduplication (2026-02-25)
#### Fixed
- **ANSI escape code stripping** — Log scanner now strips ANSI color codes (e.g. `\x1b[35m`) before classifying and fingerprinting lines, preventing color codes from polluting error messages and breaking deduplication
- **Timezone offset in timestamps** — ISO timestamp regex now handles `+01:00`/`-0500` timezone offsets and optional trailing colons (fixes Vikunja-style log entries)
- **Mid-line timestamps** — Removed `^` anchor from both ISO and syslog timestamp regexes, so timestamps embedded after log-level keywords (e.g. `ERROR 2026-02-24T21:27:05`) are now stripped correctly
#### Improved
- **`cleanLine()` helper** — Consolidated ANSI + timestamp stripping into a single reusable function used by both message display and fingerprint deduplication
### v0.30.5 — Health Probe: Fast Initial Checking (2026-02-25)
#### Improved
- **Clear stale health probes on start/restart** — `StartStack` and `RestartStack` now clear the previous `HealthProbe` result, preventing stale "unhealthy" state from being re-applied by `RefreshStatus`
- **Fast 10s probing until healthy** — Stacks with no probe result (just started) or failing probes use 10-second intervals instead of waiting the full 5-minute default; reverts to normal interval once healthy
- **Scheduler frequency 1m → 10s** — Health probe scheduler runs every 10 seconds (interval logic inside `RunHealthProbes` skips stacks that don't need probing, so no extra overhead for healthy stacks)
### v0.30.4 — Deep Bug Hunt II: Concurrency, Security & Optimization (2026-02-25)
#### Fixed (Critical)
- **Watchdog mutex panic** — Wrapped `handleDisconnect` call in anonymous func with deferred re-lock to guarantee mutex re-acquisition even on panic (C1)
- **SetGeoAppOverride nil crash** — Added nil guard; passing nil override now correctly deletes the entry instead of panicking (C2)
- **SSD-only app DB restore** — `restoreDBDumps` now falls back to `app.DrivePath` when `HDDPath` is empty (C3)
#### Fixed (High)
- **Double deploy race** — Added atomic check-and-set of `Deploying` flag with `clearDeploying()` helper on all error paths (H1)
- **Delete/Remove during deploy** — Both `DeleteStack` and `RemoveStack` now reject operations while stack is deploying (H2)
- **ScanStacks overwrite** — Skips updating `Deployed`/`AppConfig` for stacks with active deploy in progress (H3)
- **FileBrowser mount race** — Added `fileBrowserMu` mutex to prevent concurrent `SyncFileBrowserMounts` calls (H5)
- **PushEvent history gap** — Added `recordHistory` calls on both success and failure paths in PushEvent goroutine (H6)
- **PushOnce silent failure** — Now returns error for non-2xx HTTP responses instead of nil (H7)
- **DB dump file corruption** — Added `tmpFile.Sync()` and `tmpFile.Close()` before rename in `DumpOne` (H8)
- **Restic retry timeout** — Creates fresh 30-minute context for retry after unlock instead of reusing near-expired original (H9)
- **Encrypt failure silent** — Added warning log when encryption fails in `SaveAppConfig` (H10)
- **Cross-backup path traversal** — Validates destination path against registered storage paths in both web and API handlers (H11)
- **deepCopyStack incomplete** — Now deep-copies `Meta.OptionalConfig`, `Meta.HealthCheck`, and `DeployField.Options` (H12)
#### Security
- **Constant-time API key** — Replaced `==` with `subtle.ConstantTimeCompare` for API key comparison, preventing timing attacks (M1)
- **Login rate limiting** — Added per-IP rate limiter (5 attempts/minute) to login handler (M8)
- **Git credential masking** — Applied `maskRepoURL()` in `runGitInDir` log output to prevent credential leakage (M23)
- **Path prefix traversal** — Fixed `storageAttachBrowseHandler` prefix check to require trailing `/`, preventing sibling directory matches (M24)
#### Concurrency & Logic
- **MigrateEncryption race** — Moved `encKey == nil` check inside the mutex lock (M5)
- **SubdomainInUse I/O under lock** — Collect stack dirs under RLock, release, then perform disk I/O outside (M4)
- **Scheduler late jobs** — Jobs registered after `Start()` now immediately get their goroutine launched (M10)
- **SQLite WAL verification** — WAL pragma now verified via `QueryRow` + `Scan` instead of silent `Exec` (M13)
- **Metrics shutdown** — `sampleContainers` now uses parent context instead of `context.Background()` for clean shutdown (M14)
- **Telemetry scan logging** — Row scan errors now logged instead of silently swallowed (M15)
- **Asset sync lock** — Refactored to hold mutex only for status updates, not during entire HTTP download (M22)
#### Optimization
- **DB dump copy** — Replaced `os.ReadFile`/`os.WriteFile` with streaming `io.Copy` via `copyFile` helper for large dumps (M16)
- **Restic stats dedup** — Per-drive stats now computed once and aggregated, eliminating duplicate restic subprocess calls (M17)
- **Infra config atomic** — `syncInfraConfig` controller.yaml copy now uses atomic write via `copyFile` (M20)
### v0.30.3 — Comprehensive Bug Hunt Fixes (2026-02-25)
#### Fixed (Critical — P0)
- **Encrypted env vars** — `UpdateStackConfig` now uses decrypted values when building compose env, preventing `ENC:...` literals in containers (C01)
- **Silent decrypt failures** — `DecryptMap` now logs warnings on decrypt failure instead of silently returning empty values (C02)
- **Deploy race condition** — `Deployed = false` flag now set inside the mutex lock in `runComposeDeploy` (C03)
- **Shared state mutation** — `GetStack`/`GetStacks` now return deep copies preventing callers from mutating cached state (C04)
- **Watchdog races** — Added per-state mutex to `pathProbeState` for thread-safe probe state access (C05)
- **Metrics double-start** — `MetricsCollector.Start()` guarded with `sync.Once` (C06)
- **Raw mount race** — `diskJobMu` now held across entire cleanup+mount+set operation (C07)
- **Encryption key race** — Added mutex to `SetEncryptionKey` (C08)
#### Fixed (High — P1)
- **Restic lock detection** — `Snapshot()` now extracts stderr from `*exec.ExitError` and checks `unlockCmd.Run()` error (H01)
- **Disconnected drives in backup** — `activeDrives()` now skips disconnected/decommissioned drives (H02)
- **Template rendering** — Buffered via `bytes.Buffer` to prevent partial HTML on error (H07)
- **Sync stop panic** — `Stop()` uses `sync.Once` for safe channel close (H08)
- **Sync race** — `syncing = true` set before releasing lock in `TriggerSync` (H09)
- **Cloudflare context** — Threaded `context.Context` through all Cloudflare API calls for cancellation support (H10)
- **Cross-drive collision** — Replaced flawed leaf-name dedup with proper `seen` map (H15)
- **CSRF bypass** — Bearer token now validated against Hub API key before skipping CSRF (H16)
- **Nil pointer** — Added nil check for `crossDriveRunner` in handlers (H17)
- **Selftest panic** — Replaced `out[:len(out)-1]` with `strings.TrimSpace` (H18)
- **Stderr goroutine** — Added `sync.WaitGroup` in `MigrateDrive` (H19)
- **UUID slice** — Guarded `uuid[:8]` with length check (H20)
- **Fstab matching** — Parse fields exactly instead of loose `strings.Contains` (H21)
- **Atomic save** — `SaveAppConfig` writes to `.tmp` then renames (H04)
- **Deploy failure** — `SaveAppConfig` on failure now includes `encKey` (H05)
- **Encryption migration** — Uses write lock instead of read lock (H03)
- **Deep copy** — `GetFullStatus` deep-copies `lastDBDump`/`lastBackup` (H11)
- **IPv6** — TCP health probe uses `net.JoinHostPort` for IPv6 compatibility
- **Backup path validation** — `RemoveStack` validates paths under expected directory (M12)
- **Updater race** — `SetBackupRunningCheck` protected by mutex (M18)
#### Fixed (Medium — P2)
- **Config env overrides** — `LoadFromBytes` now calls `applyEnvOverrides` (M05)
- **Selfupdate state** — Compose-up failure now sets `state.Status = "failed"` (M16)
- **Memory check** — `usableMB` clamped to min 0 (M22)
- **Cross-backup trigger** — Removed invalid "manual" schedule from `triggerAllCrossBackups` (M23)
- **mmcblk support** — Partition path and `stripPartition` now handle mmcblk devices (M21, L25)
- **Scheduler** — `Start()` guarded against double-start, `Stop()` acquires mutex (M14, L24)
- **Pending events** — Events restored on save failure in `DrainPendingEvents` (M03)
- **Duplicate storage** — `AddStoragePath` rejects already-registered paths (M04)
- **Setup scan** — `CleanupTempMounts` called after drive scan (H13)
- **Setup state** — `SetStep` now logs save errors (M25)
#### Fixed (Low — P3)
- **UTF-8 truncation** — `TruncateStr` now operates on runes and handles negative maxLen (L05/L06)
- **AllDone** — Returns false for empty restore plans (L14)
- **PushOnce** — Returns actual errors instead of swallowing them (L39)
- **CSRF token** — Panics on `crypto/rand.Read` failure instead of using static fallback (L40)
- **Logout** — Requires POST method (L32)
- **Server.Close** — Uses `sync.Once` to prevent double-close panic (L49)
- **Log cap** — `lines` query parameter capped at 10000 (L31)
- **Hash function** — Replaced custom `simpleHash` with `crc32.ChecksumIEEE` (L48)
- **hasPrefix** — Replaced custom implementation with `strings.HasPrefix` (L13)
- **DefaultEnabledEvents** — Copied in `GetNotificationPrefs` early return (L09)
- **Variable shadowing** — Renamed `copy` to `cp` in `SetNotificationPrefs` (L07)
#### Removed
- Dead `imageName` function in selfupdate (L02)
- Dead `detectHostIPViaRoute` function in setup (L03)
- Custom `hasPrefix` function in restore_scan (L13)
### v0.30.2 — Report geo-restriction + logo/favicon update (2026-02-25)
#### Added
- **Geo-restriction in reports** (`internal/report/`) — New `GeoRestrictionReport` struct and `geo_restriction` field in the Report JSON. Hub can now display current geo-blocking status (enabled, allowed countries, per-app overrides, sync state) on customer detail pages.
- **Favicon route** (`/static/favicon.svg`) — Separate favicon SVG served from synced assets or embedded fallback. Uses the cloud icon from `logo_favicon_2.svg`.
- **Hub Bearer auth for geo API** — `/api/geo/` routes now accept `selfUpdateAuthMiddleware` (session auth OR Hub API key), allowing the Hub to send geo-disable commands to controllers.
#### Changed
- **Logo SVG updated** (`internal/web/templates.go`) — Replaced embedded logo with the latest `logo.svg` from the website (white text variant).
- **Favicon link** — Layout and catch-all templates now reference `/static/favicon.svg` instead of the full logo.
### v0.30.1 — Geo-Restriction fix (2026-02-25)
#### Fixed
- **WAF rule creation** — Removed custom block response body from WAF rules (requires paid Cloudflare plan). Block action now uses Cloudflare's default 403 page.
### v0.30.0 — Geo-Restriction via Cloudflare WAF (2026-02-25)
#### Added
- **Geo-restriction feature** (`internal/cloudflare/`) — New package for managing Cloudflare WAF Custom Rules. Allows restricting access to apps by country using the `http_request_firewall_custom` phase. Rules are identified by `[felhom-geo]` description prefix — other WAF rules are untouched.
- **Cloudflare API client** (`internal/cloudflare/client.go`) — HTTP client with Bearer token auth for the Cloudflare v4 API. Supports zone lookup, ruleset management, and rule CRUD operations.
- **Country data** (`internal/cloudflare/countries.go`) — Embedded map of ~250 ISO 3166-1 alpha-2 country codes with Hungarian names. Includes search helpers for the UI.
- **Geo sync manager** (`internal/cloudflare/geosync.go`) — Orchestrator that diffs desired vs existing Cloudflare rules and applies changes. Runs on settings change, after app deploy/remove, and every 6 hours for verification.
- **Settings page UI** (`templates/settings.html`) — New "Földrajzi korlátozás" section with searchable country selector (autocomplete dropdown → tag chips), enable/disable toggle, per-app override summary, and sync status display. Hungary removal triggers a confirmation warning.
- **Per-app override** (`templates/app_info.html`) — Each app's detail page now has a "Földrajzi korlátozás" section (when the feature is globally enabled) to set app-specific allowed countries.
- **Geo API endpoints** (`internal/api/geo.go`) — `GET /api/geo/status`, `POST /api/geo/settings`, `POST /api/geo/sync`, `GET /api/geo/countries`, `POST/DELETE /api/stacks/{name}/geo/override`.
- **Settings model** (`internal/settings/settings.go`) — New `GeoRestriction` struct with `AllowedCountries`, `AppOverrides`, and sync state (zone ID, ruleset ID, last sync). Thread-safe getter/setter methods following existing RWMutex pattern.
#### Changed
- **Router** (`internal/api/router.go`) — Added `OnGeoRelevantChange` callback triggered after app deploy/remove to re-sync geo rules when hostnames change.
- **Main wiring** (`cmd/controller/main.go`) — Cloudflare client, geo sync manager, and scheduler job initialized when `cf_api_token` is configured. New `geoStackAdapter` provides deployed app hostnames.
#### Hub Changes
- **Config form** (`hub/internal/web/templates/config_form.html`) — Updated CF API token help text to indicate Zone WAF:Edit permission is needed for geo-restriction.
#### Notes
- The existing `cf_api_token` needs **Zone WAF:Edit** permission added (in addition to existing Zone DNS:Edit for ACME). No new token field is needed.
- Local network access is inherently unaffected — local traffic bypasses Cloudflare entirely.
- Cloudflare Free plan supports up to 5 custom rules, which is sufficient for a global rule + a few per-app overrides.
### v0.29.3 — Controller-side Health Probes (2026-02-25)
#### Added
- **HTTP/TCP health probes** (`internal/stacks/healthprobe.go`) — The controller now probes deployed apps directly over the Docker network to verify services are actually responding, not just that containers are running. Runs every minute, configurable per-app interval (default 5 min).
- **Three probe types**: `http` (any response = alive), `api` (validates status code and response body), `tcp` (port reachability). Multiple checks per app supported.
- **`.felhom.yml` healthcheck config** (`internal/stacks/metadata.go`) — New `healthcheck:` section with `interval`, `checks[]` (type, port, path, method, expect). Parsed from app catalog metadata.
- **State override** (`internal/stacks/manager.go`) — If a running container's health probe fails, the stack state is overridden to "unhealthy". Clears automatically when probe passes again.
#### Fixed
- **Vikunja healthcheck** — Removed Docker-level healthcheck (distroless image has no wget/curl). Controller-side API probe to `:3456/api/v1/info` replaces it.
### v0.29.2 — Dynamic Logo & Favicon (2026-02-25)
#### Changed
- **Logo served from synced assets** (`internal/web/server.go`) — `serveLogoHandler` now checks the Hub-synced assets directory for `felhom-logo.svg` first, falling back to the embedded SVG constant if not found. This allows logo updates via Hub without a controller rebuild.
#### Added
- **SVG favicon** (`templates/layout.html`, `templates/catchall.html`) — Added `<link rel="icon" type="image/svg+xml">` pointing to `/static/felhom-logo.svg` so browsers display the Felhom logo as a tab icon.
### v0.29.1 — Fix Git Lock File Stale After Interrupted Sync (2026-02-24)
#### Fixed
- **Stale git lock file recovery** — Catalog sync now removes stale `.git/index.lock`, `.git/shallow.lock`, and `.git/HEAD.lock` files before running `git fetch`/`git reset`. Previously, if the container was killed mid-sync, the leftover lock file would block all subsequent syncs until manual intervention.
### v0.29.0 — Encrypt Sensitive Values in app.yaml (2026-02-23)
#### Added
- **AES-256-GCM encryption for app.yaml secrets** — Sensitive deploy field values (`type: password` and `type: secret`) are now encrypted at rest in each stack's `app.yaml` using a per-node 32-byte key. Encrypted values are stored as `ENC:base64(nonce+ciphertext)`. New `internal/crypto` package provides `Encrypt`, `Decrypt`, `LoadOrCreateKey`, `DecryptMap`, and `IsEncrypted` helpers.
- **Encryption key in infra backup** — The encryption key (`encryption.key`) is included in the Hub infra backup bundle (`encryption_key_b64` field) and local drive infra backups for disaster recovery.
- **Encryption key restore** — The setup wizard's infra restore flow restores `encryption.key` from the backup bundle so encrypted app.yaml values remain readable after disaster recovery.
- **Startup migration** — On first start after upgrade, existing plaintext sensitive values in deployed stacks' `app.yaml` files are automatically encrypted in-place.
#### Changed
- **`SaveAppConfig` signature** — Now accepts `encKey []byte` and `sensitiveVars []string` parameters for encryption. All callers (deploy, update, optional config, inject missing fields, HDD path update, storage handlers) updated.
- **`LoadAppConfigDecrypted`** — New helper that loads app.yaml and transparently decrypts all `ENC:` values for docker-compose env injection and web UI display.
- **`SensitiveEnvVars`** — New exported helper that identifies sensitive env vars from `.felhom.yml` metadata (`type: password` or `type: secret` deploy fields).
- **Manager struct** — Added `encKey` field and `SetEncryptionKey()` / `MigrateEncryption()` methods.
- **Web Server struct** — Added `encKey` field and `SetEncryptionKey()` method; deploy handler decrypts values before template rendering.
### v0.28.8 — Password UX Polish (2026-02-23)
#### Fixed
- **Password fields empty after deployment** (`templates/deploy.html`) — Password-type deploy fields now read their stored value from `DeployedFieldValues` (app.yaml env) when viewing settings for an already-deployed app, instead of always using the field's `.Default` (which was empty).
- **Post-deploy credentials masked** — Passwords on the post-deploy success card are now shown as `••••••••••••` with "Megjelenítés" (reveal) and "Másolás" (copy to clipboard) buttons, instead of displaying plaintext.
#### Changed
- **Settings page: initial password hint** — Deployed password fields show a note: *"Telepítéskor beállított kezdeti jelszó — ha az alkalmazásban megváltoztattad, az itt nem frissül."* Generate button is hidden for already-deployed apps.
- **Post-deploy credential detection** — Added EMAIL to the username-detection heuristic (catches Kimai's `ADMIN_EMAIL`).
### v0.28.7 — Password Field UX (2026-02-23)
#### Changed
- **Password deploy fields: masked input with reveal & confirmation** (`templates/deploy.html`) — `type: password` fields now render as masked inputs (hidden by default) with an eye toggle button to reveal/hide. Added a "Jelszó megerősítése" confirmation field below each password input. The "Generálás" button fills both fields simultaneously. Form validation checks that both fields match before allowing deploy. Confirmation fields are only shown for new deployments.
- **App catalog: admin passwords use `type: password`** (separate repo: `app-catalog-felhom.eu`) — Changed 4 apps (Nextcloud, Grafana, Kimai, Code-server) from `type: secret` to `type: password` so users can see/edit/generate admin passwords during deployment (matching the existing Paperless-ngx pattern).
### v0.28.6 — Filebrowser Link, Appdata Paths & Log Timestamps (2026-02-23)
#### Added
- **Post-deploy credential display** (`templates/deploy.html`) — The success page now shows actual username/password values from the deploy form instead of a generic message. Reads from deploy field metadata, filtering out internal DB passwords and secret keys. Falls back to `defaultCreds` for apps without typed deploy fields.
#### Fixed
- **Filebrowser "open" link on stacks page** (`web/handlers.go`) — Protected stacks like filebrowser have no `.felhom.yml` or `app.yaml`, so the subdomain lookup found nothing. Added `protectedStackSubdomains` fallback map for programmatically managed protected stacks (filebrowser → "files"). Now shows `files.<domain> ↗` link on both the stacks page and dashboard.
- **App catalog: appdata volume paths** (separate repo: `app-catalog-felhom.eu`) — 4 compose templates (nextcloud, immich, paperless-ngx, romm) used `${HDD_PATH}/appdata/` instead of `${HDD_PATH}/felhom-data/appdata/` as designed in the v0.26.0+ storage structure. Fixed all templates. Existing deployments need redeployment or manual volume path update.
- **Debug log viewer timestamps** (`web/logbuffer.go`, `templates/debug.html`) — Naplóviewer showed relative times like "-3586mp" (negative due to timezone bug: `time.Parse` assumed UTC but `log.LstdFlags` outputs local time). Now uses `time.ParseInLocation` with `time.Local`, and displays absolute `HH:MM:SS` timestamps.
### v0.28.5 — Post-Deploy Info Card (2026-02-23)
#### Added
- **Post-deploy success page** (`web/templates/deploy.html`) — After a successful deploy, instead of auto-redirecting to the apps list, shows a rich info card with: direct app link ("Alkalmazás megnyitása ↗"), first steps from catalog metadata (with DOMAIN placeholders replaced), default credentials info, documentation link, and a link to the settings page where passwords can be revealed. Also shown for unhealthy/timeout states since apps may still be usable during initialization.
### v0.28.4 — Telemetry: Skip Stopped Apps (2026-02-23)
#### Fixed
- **Stopped apps no longer send zero-value telemetry to hub** (`report/telemetry.go`) — Previously, deployed-but-stopped apps were included in the telemetry report with all-zero memory/CPU values, which dragged down hub-side averages. Now `buildAppTelemetry` checks `isStackRunning()` and only includes apps in running, starting, unhealthy, or restarting states.
### v0.28.3 — Catch-All Page, Deploy Controls, Dashboard Open (2026-02-23)
#### Added
- **Catch-all page for stopped/undeployed apps** — When a user visits a stopped app's subdomain (e.g., `travel.demo-felhom.eu`), they now see a branded felhom page with the app name and status ("Az alkalmazás jelenleg le van állítva") instead of Traefik's raw 404. Implemented via a low-priority (1) Traefik catch-all router on the controller container + `CatchAllMiddleware` in `server.go` that intercepts non-controller hosts and renders standalone `catchall.html` without auth.
- **Start/Stop/Restart buttons on deploy settings page** — Deployed apps now show Indítás/Leállítás/Újraindítás buttons in the page header, plus a "Megnyitás ↗" link to the app's subdomain (visible when running). Previously the deploy page had no state controls.
- **"Megnyitás ↗" button on Vezérlőpult** — Running apps on the dashboard now show an open button that launches the app in a new tab. Uses the `Subdomains` map built from `app.yaml` SUBDOMAIN env with metadata fallback.
- **`findStackBySubdomain()`** helper in `server.go` — looks up stacks by subdomain, checking deployed `app.yaml` env first, then `.felhom.yml` metadata.
#### Changed
- **Subdomain links on Alkalmazások page** — Links now only shown for deployed apps (previously shown for all apps including non-deployed ones where the subdomain isn't final yet).
- **`docker-compose.yml`** — Added 6 catch-all Traefik router labels (`traefik.http.routers.catchall.*`) with `priority=1` and `certresolver=letsencrypt`.
### v0.28.2 — Async Deploy & AdventureLog Fix (2026-02-23)
#### Changed
- **Async deploy** — `DeployStack()` now runs `docker compose up -d` in a background goroutine instead of blocking the HTTP response. The deploy API returns immediately after validation + config save, so the UI switches to the progress panel instantly (previously waited 30-60s for image pulls). New `StateDeploying` container state shown while compose-up is in progress. On failure, the goroutine reverts both disk and in-memory state and stores the error in `DeployError` for the polling UI to display.
- **Deploy progress UI** — Polling now handles the `deploying` state ("Képek letöltése, konténerek indítása...") and `deploy_error` (shows error message with links to logs). Previous behavior only showed progress after compose-up completed.
#### Fixed
- **RestartStack uses `up -d` with env vars** — `RestartStack()` previously used bare `docker compose restart` which only sends SIGTERM+start without re-reading the compose file or injecting env vars from `app.yaml`. Now uses `docker compose up -d` with full env, matching `StartStack()` behavior. This ensures template changes (images, healthchecks) and env var updates are picked up on restart.
- **AdventureLog backend healthcheck** — Replaced `wget` (not available in v0.11.0 image) with `python urllib.request`. Also uses `127.0.0.1` instead of `localhost` to avoid IPv6 resolution issues.
- **AdventureLog frontend healthcheck** — Changed `localhost` → `127.0.0.1` to fix IPv6 resolution causing connection refused (Node.js only listens on IPv4).
- **AdventureLog SECRET_KEY** — Added `SECRET_KEY=${SECRET_KEY}` env var alongside `DJANGO_SECRET_KEY` for v0.11.0 compatibility (Django settings now reads `SECRET_KEY` directly).
### v0.28.1 — Telemetry Debug Section (2026-02-23)
#### Added
- **Telemetria teszt section on Debug page** — New collapsible section between "Hub & Kapcsolatok" and "Önfrissítés teszt". Click "Telemetria futtatása" to run the full telemetry collection pipeline on-demand without waiting for the 15-minute report cycle.
- **`GET /api/debug/telemetry`** — New debug endpoint in `handler_debug.go`. Invokes `GetTelemetryPreview` callback, returns per-app data: container list, memory (current/avg/peak), CPU avg, catalog limit, log error/warning counts, top issues, and overall latency. Response: `{latency_ms, app_count, total_errors, total_warnings, app_telemetry[]}`.
- **`GetTelemetryPreview` callback** added to `DebugCallbacks` struct. Wired in `main.go` debug-mode block: calls `report.BuildAppTelemetryForDebug(stackMgr, metricsStore, logger)`. Available regardless of hub configuration.
- **`report.BuildAppTelemetryForDebug()`** — Exported wrapper in `internal/report/telemetry.go` around the private `buildAppTelemetrySection()`. Allows debug endpoint access without exposing internal package details.
- **JS rendering** — `runTelemetryTest()` fetches the endpoint and shows a summary message. `renderTelemetryDetail()` builds a table with per-app rows (color-coded errors in red, warnings in yellow) and sub-rows for top issues. Includes a collapsible "Nyers JSON" section showing the exact payload that would go to the hub.
### v0.28.0 — App Telemetry & Analytics (2026-02-23)
#### Added
- **App telemetry in Hub reports** — `Report.AppTelemetry` (new field in `report/types.go`) carries per-stack memory/CPU metrics and log scan results to the Hub on every report push. Backward-compatible: old Hub versions silently ignore the new field.
- **`internal/metrics/telemetry.go`** — New `MetricsStore.GetContainerTelemetry(since)` method aggregates container memory (current/avg/peak) and CPU averages from the existing `container_metrics` SQLite table over the last 15 minutes.
- **`internal/metrics/logscanner.go`** — New `ScanContainerLogs(containerNames, since, logger)` function runs `docker logs --since=15m --tail=1000` on each non-protected deployed container. Detects errors/warnings by keyword matching, deduplicates via fingerprinting (strips timestamps, replaces 6+ digit numbers with `<N>`, hex with `<HEX>`, UUIDs with `<UUID>`). Returns `[]ContainerLogSummary` with counts and `RecentIssues` (top 10 per container).
- **`internal/report/telemetry.go`** — New `buildAppTelemetrySection()` and `buildAppTelemetry()` functions assemble per-stack `AppTelemetry` records by aggregating container-level metrics and log summaries. Only non-protected, deployed stacks are included.
#### Changed
- **`internal/report/builder.go`** — `BuildReport()` now calls `buildAppTelemetrySection()` after the stacks section, populating `r.AppTelemetry`.
- **`internal/report/types.go`** — Added `AppTelemetry []AppTelemetry` field to `Report` struct. Added new `AppTelemetry` type with fields: app_name, display_name, containers, memory metrics, catalog estimate/limit, log error/warning counts, and top issues.
### v0.27.3 — Real System Memory Everywhere (2026-02-23)
#### Changed
- **Deploy page uses real system memory** — Memory bar now shows actual `/proc/meminfo` usage instead of declared `mem_request` sums. Labels changed from "Jelenlegi foglalás" to "Jelenlegi használat". `system.GetMemoryMB()` provides real-time total and used memory.
- **Pre-start memory check uses real memory** — `actionStack("start")` in `router.go` and `DeployStack()` in `deploy.go` now check real used memory (`usedMB + newReqMB > usableMB`) instead of declared committed sums. `CommittedMemory()` kept only for soft overcommit warnings.
#### Added
- **`system.GetMemoryMB()` helper** — Lightweight function in `internal/system/info_linux.go` that returns real total and used memory from `/proc/meminfo` without the overhead of full `GetInfo()` (no disk/CPU/temp). Stub in `info_other.go` for non-Linux.
- **Monitoring page memory distribution bar** — New stacked bar on `/monitoring` showing per-container memory usage (colored segments), OS/system overhead (gray), and free memory. Built dynamically from container summary data + real-time `/api/system/info`. Color-coded legend with per-app labels.
### v0.27.2 — Comprehensive Fixes and New Labels (2026-02-23)
#### Fixed
- **Deploy error popups now copyable** — Replaced all native `alert()` calls with a custom modal (`showAlert()` in layout.html) using a `<pre>` block with `user-select:text`. Error messages can now be selected and copied. Applied across deploy.html and layout.html.
- **Manual Tier2 backup now reports to Hub** — Added `OnCrossDriveComplete` callback to `Router` (`internal/api/router.go`). Both `triggerCrossBackup` (single-app) and `triggerAllCrossBackups` (run-all) now call `pushInfraBackup()` + `writeLocalInfraBackup()` after completion, matching the automatic scheduled path.
- **Memory bar excludes stopped apps** — `CommittedMemory()` in `internal/stacks/manager.go` now skips apps with `StateStopped` or `StateExited`. Only running/starting/unhealthy apps count toward committed memory.
- **Pre-start memory check** — `actionStack("start")` in `internal/api/router.go` now validates available memory before starting a stopped app. Returns 409 Conflict with a descriptive Hungarian error if insufficient.
#### Added
- **`hungarian_ui` metadata field** — New `HungarianUI bool` field in `ResourceHints` (`internal/stacks/metadata.go`). Shows "Magyar felület" green badge on deploy, stacks, and app info pages when `hungarian_ui: true` in `.felhom.yml`.
- **USB badge on storage cards** — Settings page storage cards now show an orange "USB" badge next to Aktív/Alapértelmezett when the drive is USB-attached (using existing `IsUSB` sysfs detection).
- **`StackMemoryMB()` helper** — New method on `Manager` to get a specific stack's memory request.
#### App Catalog (app-catalog-felhom.eu)
- **AdventureLog** — Fixed image tags from `v0.12.0` (non-existent) to `v0.11.0` for both backend and frontend.
### v0.27.1 — Fix FileBrowser Mount Sync (2026-02-22)
#### Fixed
- **`internal/web/handlers.go`** — `SyncFileBrowserMounts()` was reading the domain from a `.env` file that doesn't exist in the filebrowser stack directory (domain is baked into the compose labels by `docker-setup.sh`). It always logged `[WARN] Cannot read DOMAIN from FileBrowser .env — skipping mount sync` and returned early, so storage paths were never synced to FileBrowser's config.yaml or docker-compose.yml. Fixed by using `s.cfg.Customer.Domain` directly from the controller config.
### v0.27.0 — User-Configurable App Subdomains (2026-02-22)
#### Added
- **User-configurable subdomains**: Users can now customize the subdomain (e.g., `wiki`, `cloud`, `my-notes`) for each app during deployment, instead of using a fixed value. The deploy page shows an editable text input with the default subdomain pre-filled and the base domain as a suffix (e.g., `[wiki] .demo-felhom.eu`).
- **New deploy field type `"subdomain"`** — `internal/stacks/metadata.go`, `deploy.go`: A new field type that is user-editable with a default value, validated, and locked after deployment. Changing the subdomain requires removing the app (clean install) and redeploying.
- **Subdomain validation** — `internal/stacks/deploy.go`: Three-layer validation: DNS-safe format (lowercase alphanumeric + hyphens, max 63 chars), reserved name blocklist (`felhom`, `files`, `traefik`, `api`, `www`, `mail`, `admin`, etc.), and uniqueness check across all deployed stacks.
- **Backward compatibility** — `internal/stacks/deploy.go`: `InjectMissingFields()` auto-fills `SUBDOMAIN` from the `.felhom.yml` default for existing deployed apps when templates are synced, so no manual intervention is needed.
- **`internal/web/handlers.go`** — `stacksHandler()` builds an effective subdomain lookup map (stored env → metadata fallback). `appDetailHandler()` passes `EffectiveSubdomain` to templates.
- **`internal/web/templates/deploy.html`** — New `.subdomain-input-group` widget with inline `.domain` suffix. Client-side validation enforces DNS-safe format with real-time lowercasing.
- **`internal/web/templates/stacks.html`**, **`app_info.html`** — Subdomain links now read from stored `app.yaml` env (via lookup map) instead of hardcoded metadata, showing the user's actual chosen subdomain.
#### Changed
- **`internal/stacks/deploy.go`** — `PreviewDeployValues()` domain case simplified: shows just the base domain now (subdomain is a separate field).
- **`internal/web/handlers.go`** — Deploy page domain auto-field no longer prepends `meta.Subdomain + "."`. Passes `DeployedFieldValues` for rendering stored subdomain on settings page.
#### App Catalog (app-catalog-felhom.eu)
- All 51 template `docker-compose.yml` files updated: hardcoded `{subdomain}.${DOMAIN}` replaced with `${SUBDOMAIN}.${DOMAIN}` in Traefik labels, app env vars (APP_URL, trusted domains, webhook URLs, etc.), and comments.
- All 51 `.felhom.yml` files updated: added `SUBDOMAIN` deploy field with `type: subdomain` and `default:` matching the existing `subdomain:` metadata value.
### v0.26.2 — Show Full App URL on Deploy Page (2026-02-22)
#### Fixed
- **`internal/stacks/deploy.go`** — `PreviewDeployValues()` now shows the full reachable URL (`subdomain.base_domain`) for domain-type fields instead of just the base domain. Informational only — stored env var remains the base domain.
- **`internal/web/handlers.go`** — Same fix applied to the already-deployed settings page: domain field displays `subdomain.base_domain` matching what the app card shows.
### v0.26.1 — Show Auto-Generated Values on Deploy Page (2026-02-22)
#### Changed
- **`internal/stacks/deploy.go`** — Added `PreviewDeployValues()` method: pre-generates domain and secret field values when the deploy page is loaded, so the user can see (and note down) exact values before deploying. Updated `DeployStack()` to accept pre-generated secret values from the form instead of always regenerating.
- **`internal/web/handlers.go`** — `deployHandler` now calls `PreviewDeployValues()` for non-deployed apps and populates `AutoFieldValues` (previously empty for pre-deploy).
- **`internal/web/templates/deploy.html`** — "Automatikusan generált értékek" section now shows actual values on the pre-deploy page too: domain as a readonly text input, secrets as readonly password inputs with a "Megjelenítés" reveal button. Updated section description to inform the user to note down passwords. Pre-generated secret values are submitted as hidden inputs so the same values shown to the user are saved to `app.yaml`.
### scripts — Hub Mode + FileBrowser Controller-Managed Volumes (2026-02-22)
#### `scripts/docker-setup.sh` — v6.0.0
- **Hub mode** (`--hub-customer` / `--hub-password`): downloads `controller.yaml` from Hub API early in setup, extracts `domain`, `email`, `cf_api_token`, `cf_tunnel_token` and auto-populates all infrastructure settings. Single one-liner deploys fully configured Traefik + TLS + Cloudflare Tunnel with no additional flags needed. CLI flags always override hub values.
- **`yaml_get()` helper**: strips leading whitespace before key comparison — required because Go's `yaml.v3` uses 4-space indentation.
- **`apply_hub_config()`**: called before `print_banner` in `main()` so hub-sourced values are reflected in the plan display.
- **FileBrowser initial install**: removed drive auto-discovery from `install_filebrowser()`. FileBrowser is now installed with no drive volumes and a minimal `config.yaml` with `/srv` fallback. Drive volumes are managed entirely by the controller (`SyncFileBrowserMounts()`) after storage is registered via the dashboard.
- **Bug fix**: `((found_mounts++))` → `found_mounts=$(( found_mounts + 1 ))` — `set -euo pipefail` traps post-increment when var=0 (exit code 1). Same fix applied to `step_num` in `install_filebrowser()`.
#### `scripts/felhom-wipe.sh`
- **`cleanup_scan_dir()`**: removes `/mnt/.felhom-scan/` (ephemeral DR scan directory) — called from `full` level onwards.
- **`cleanup_raw_mounts()`**: removes raw helper mount infrastructure (`/mnt/.felhom-raw/`) at `nuclear` level: unmounts bind mounts first, then raw mounts, strips fstab entries, removes empty directories. Physical drive data untouched.
- **Bug fix**: `do_soft_wipe()` used `[ -f "$f" ] && rm -f "$f" && info "..."` — with `set -euo pipefail`, when a state file doesn't exist `[ -f ]` returns 1, the whole `&&` chain returns 1, and `set -e` exits the script. Nuclear wipe was silently stopping after removing only the first two state files that existed. Fixed with `if [ -f "$f" ]; then ...; fi`.
#### `scripts/README.md`
- Hub mode quick start simplified to one-liner
- Updated installation steps table: step 7 reflects controller-managed FileBrowser volumes
- Added "Raw helper mounts" section explaining two-level mount architecture
- Updated wipe levels table for `full` (scan dir) and `nuclear` (raw mounts + scan dir)
### v0.26.0 — Storage Namespace `felhom-data/` + Test Node Wipe Script (2026-02-22)
All felhom-managed data on external drives now lives under a `felhom-data/` subdirectory, cleanly separating controller-managed data from user files. Plus a multi-level wipe script for repeatable test node cleanup.
**Key design principle:** `HDD_PATH` env var stays as the mount point (e.g., `/mnt/hdd_1`). The `felhom-data` segment is embedded in path helpers and compose templates — not in `HDD_PATH`.
#### Changed
- **`internal/backup/paths.go`** — Added `FelhomDataDir = "felhom-data"` constant. Updated 8 path functions to insert `felhom-data` between the drive root and data subdirectory:
- `PrimaryBackupPath` → `<drive>/felhom-data/backups/primary`
- `PrimaryResticRepoPath` → `<drive>/felhom-data/backups/primary/restic`
- `AppDBDumpPath` → `<drive>/felhom-data/backups/primary/<stack>/db-dumps`
- `SecondaryBackupPath` → `<drive>/felhom-data/backups/secondary`
- `AppSecondaryRsyncPath` → `<drive>/felhom-data/backups/secondary/<stack>/rsync`
- `SecondaryResticRepoPath` → `<drive>/felhom-data/backups/secondary/restic`
- `SecondaryInfraPath` → `<drive>/felhom-data/backups/secondary/_infra`
- `AppDataDir` → `<drive>/felhom-data/appdata/<stack>`
- `InfraBackupDir` **unchanged** — stays at drive root for DR scanner
- **`internal/stacks/delete.go`** — Added local `felhomDataDir = "felhom-data"` constant (cannot import `backup` due to architectural boundary). Updated `ProtectedHDDPaths()` to protect `<drive>/felhom-data`, `<drive>/felhom-data/appdata`, `<drive>/felhom-data/backups`. Fixed hardcoded paths in `GetStackBackupData()`.
- **`internal/storage/migrate_drive.go`** — Added `backup` package import. Fixed 4 issues:
- Conflict check: uses `backup.AppDataDir()` instead of hardcoded `appdata/`
- Verify step: uses `backup.AppDataDir()` instead of hardcoded `appdata/`
- rsync excludes: updated from `backups/primary/restic/` to `felhom-data/backups/primary/restic/`
- Size estimation: now scans inside `felhom-data/` namespace, skipping restic repos correctly
- **`internal/storage/migrate.go`** — Added `backup` package import. Post-migration DB dump copy now uses `backup.AppDBDumpPath()` instead of hardcoded paths.
- **`internal/web/handlers.go`** — Fixed legacy `"storage"` path in storage app detail size calculation (was dead code — path never existed); now uses `backup.AppDataDir()`.
- **`internal/storage/format_linux.go`** — Format wizard creates `felhom-data/` subdirectory instead of legacy `storage/`.
- **`internal/storage/attach_linux.go`** — Attach wizard creates `felhom-data/` subdirectory instead of legacy `storage/`.
#### Added
- **`scripts/felhom-wipe.sh`** — Test node cleanup script with 4 wipe levels:
- `soft` — Removes controller state files (settings.json, metrics.db, session/setup/update/snapshot state)
- `controller` — Soft + removes all app containers, volumes, and stack directories (skips protected stacks by default)
- `full` — `controller`-level cleanup + removes `felhom-data/` on all storage drives (also removes old-style `appdata/` and `backups/` for migration compatibility); infra containers preserved, controller restarted after cleanup
- `nuclear` — Full + removes controller.yaml, all infrastructure containers (controller, traefik, cloudflared, portainer), DR markers, and runs `docker system prune -af --volumes`
- Auto-detects paths from `controller.yaml` and `settings.json`
- Dry-run by default; requires `--yes` to execute
- Interactive confirmation prompt with `--yes` execution
#### Notes
- **Migration**: Pre-v0.26.0 restic snapshots reference old paths (without `felhom-data/`). Existing installations need data migration before upgrading.
- **App catalog**: Compose templates need separate update: `${HDD_PATH}/appdata/` → `${HDD_PATH}/felhom-data/appdata/` (tracked as separate task).
- All backup, crossdrive, and restore logic automatically picks up new paths via `paths.go` helpers — no changes needed in `backup.go`, `crossdrive.go`, or `restore.go`.
---
### v0.25.0 — Debug Page: Operator Testing & Diagnostics Dashboard (2026-02-21)
**Full debug dashboard with 8 sections for testing all controller subsystems in debug mode.**
Only available when `logging.level: "debug"` — sidebar link, page, and all `/api/debug/*` endpoints return 404 otherwise.
#### New files
- `internal/web/logbuffer.go` — Ring buffer (1000 entries) implementing `io.Writer` for capturing log output. Parses Go standard log format (with/without `Lshortfile`), extracts level/source/timestamp. Supports filtered retrieval by level and timestamp.
- `internal/web/handler_debug.go` — Debug page handler + 20 API endpoint handlers organized in 8 sections. `DebugCallbacks` struct (6 fields) for wiring main.go closures.
- `internal/web/templates/debug.html` — Full debug dashboard template with 8 collapsible sections, complete JS framework (lazy-load, polling, action buttons, log viewer with filter/auto-refresh).
#### Debug page sections
1. **Rendszer diagnosztika** — Diagnostic dump (migrated from `api/router.go`) with structured UI rendering: controller info, storage paths, deployed stacks, scheduler jobs, alerts. JSON download button.
2. **Értesítés teszt** — Send test events with configurable type/severity, view event history ring buffer (last 50 events, newest first).
3. **Mentés teszt** — Trigger individual backup phases: full backup, DB dump only, cross-drive only, restic integrity check, infrastructure backup.
4. **Tárhely teszt** — Storage watchdog status table with per-path probe state. Simulate disconnect (stops apps, marks disconnected, skips unmount) and reconnect (cleans locks, clears state). 5s auto-refresh.
5. **Hub & Kapcsolatok** — Hub report push, infra backup push, Hub/Gitea connectivity tests with latency, preference sync.
6. **Önfrissítés teszt** — Version check + dry-run (shows current/new image lines, compose writability, backup status).
7. **DR / Telepítő varázsló** — Infra backup status per drive (files, timestamps). "RESET" confirmation + infra backup pre-check before triggering setup mode via marker file.
8. **Naplóviewer** — In-memory log viewer with level filter (DEBUG/INFO/WARN/ERROR), 2s auto-refresh, color-coded entries, clear display.
#### Module additions
- `notify/notifier.go`: `PushTestEventSync()` (synchronous, returns Hub status), `GetEventHistory()` (ring buffer), `recordHistory()` for debug page.
- `backup/crossdrive.go`: `RunAllConfigured()` — runs all enabled apps ignoring schedule filter.
- `selfupdate/updater.go`: `DryRun()` — checks update availability, compose writability, backup status without performing changes.
- `monitor/watchdog.go`: `SimulateDisconnect()` / `SimulateReconnect()` with `simulatedPaths` map, `GetDebugStatus()` for per-path probe state. Watchdog `Check()` skips simulated paths.
- `setup/setup.go`: `NeedsSetup()` now checks `.needs-setup` marker file. `ClearSetupMarker()` for cleanup.
#### Routing changes
- **Mux carve-out**: `/api/debug/` routes to web server (same pattern as `/api/storage/`), with auth + CSRF.
- **Removed** `SetDebugDumpDeps()` from `api/router.go` and the `/api/debug/dump` route — dump handler migrated to `handler_debug.go` using Server's existing fields.
#### Infrastructure
- `setupLogger()` now returns `(*log.Logger, *web.LogBuffer)`. In debug mode, creates `io.MultiWriter(os.Stdout, logBuffer)` so all log output is captured from the start.
- Debug CSS: ~170 lines of styles for sections, result badges, log viewer, confirm input, danger button, spinner.
### v0.24.0 — Pre-Testing Observability (2026-02-21)
**Three features for pre-testing diagnostics: verbose debug logging, diagnostic dump endpoint, and startup self-test.**
#### Feature 1: Debug logging across all modules
All `[DEBUG]` log lines are gated behind `logging.level: "debug"` — zero overhead at `info` level.
- **New** `internal/util/strings.go`: shared `TruncateStr()` for safely truncating command output in logs.
- **Backup** (`backup.go`, `dbdump.go`, `crossdrive.go`, `restore.go`, `local_infra.go`): added `isDebug()` method and per-operation debug logging. DB dump logs container discovery, per-dump command details (passwords masked as `***`), validation results. Cross-drive logs source/dest paths, rsync results, auto-enable decisions. Restore logs step-by-step progress.
- **Storage** (`scan_linux.go`, `format_linux.go`, `attach_linux.go`, `migrate.go`): added `Logger`/`Debug` fields to request structs. Logs raw lsblk output (truncated), per-disk classification, pipeline steps for format/attach, rsync progress for migrate. Updated `*_other.go` stubs.
- **Sync** (`sync.go`): logs masked clone URLs, per-file hash comparison, post-sync hook triggers.
- **Self-update** (`updater.go`): logs registry API calls, tag parsing, version comparison, compose file edits.
- **Monitor** (`watchdog.go`): smart logging — periodic 60-probe summaries (~5 min), immediate log on unexpected failures, reconnect attempt details. (`healthcheck.go`): logs raw check values and per-check results.
- **Notify** (`notifier.go`): logs event push URL/type/response, preference sync details.
- **Report** (`pusher.go`, `builder.go`): logs payload sizes, section summaries, push responses.
- **Assets** (`syncer.go`): logs manifest fetch, per-file hash comparison, download/removal actions.
- **Setup** (`scanner.go`, `handlers.go`): logs drive scan details, hub recovery/config write operations.
#### Feature 2: Diagnostic dump endpoint (`GET /api/debug/dump`)
Returns a comprehensive JSON snapshot of all controller state. Only available when `logging.level: "debug"` — returns 404 otherwise.
- Sections: `controller` (version, uptime, config hash, PID), `storage` (per-path usage), `stacks` (deployed/running/stopped counts + list), `backup` (status, repo stats), `hub` (push status, consecutive failures), `scheduler` (all jobs with last_run/running/errors), `health` (fresh check), `notifications`, `self_update`, `alerts`.
- API router expanded with `SetDebugDumpDeps()` setter for scheduler, hub pusher, alert manager, version, and start time.
#### Feature 3: Startup self-test
- **New** `internal/selftest/selftest.go`: runs 9 diagnostic checks on boot with 5s timeout each.
- Checks: Docker socket, stacks directory, data directory (write test), system data path (mount point), storage paths (connected vs disconnected), git catalog (.felhom.yml files), Hub connectivity (/healthz), restic repos, metrics DB.
- Results logged in a clear block: `[PASS]`/`[WARN]`/`[FAIL]` per check, summary at end.
- Self-test summary (pass/warn/fail counts) sent to Hub via `NotifyControllerStarted` details map.
- Never blocks startup — purely diagnostic.
#### Constructor/signature changes
- `notify.New()`: added `debug bool` param. `NotifyControllerStarted()`: added `details map[string]interface{}` param.
- `report.NewPusher()`: added `debug bool` param. `BuildReport()`: added `logger *log.Logger` param.
- `monitor.RunHealthCheck()`: added `logger *log.Logger` param (5 call sites in main.go).
- `selfupdate.NewUpdater()`: added `debug bool` param.
- `assets.New()`: added `debug bool` param.
- `backup.NewCrossDriveRunner()`: added `debug bool` param. `WriteLocalInfraBackup()`: added `debug bool` param.
- `backup.DiscoverDatabases()`, `DumpOne()`: added `debug bool` param.
- `storage.ScanDisks()`: added `logger, debug` params. `FormatRequest`, `AttachRequest`, `MigrateRequest`: added `Logger`/`Debug` fields.
- `setup.ScanDrivesForInfraBackups()`: added `debug bool` param.
### v0.23.0 — CSRF Protection (2026-02-21)
**CSRF (Cross-Site Request Forgery) protection on all browser-facing POST endpoints — controller and hub.**
**Controller changes:**
- New `internal/web/csrf.go`: `CsrfProtect` HTTP middleware validates CSRF tokens on all state-mutating requests (POST/DELETE/PATCH).
- Reads token from `_csrf` form field or `X-CSRF-Token` request header.
- Exempt paths: `Authorization: Bearer` requests (selfupdate, config/apply hub→controller calls) — browsers cannot auto-send Bearer headers, so no CSRF risk.
- Auth-disabled mode (no password set): CSRF check is skipped entirely.
- On rejection: JSON error for `/api/` paths, HTTP 403 text for page routes.
- `internal/web/auth.go`: `session` struct gains a `csrfToken string` field. `createSession()` generates a second 32-byte random CSRF token alongside the session token. New `csrfTokenForSession(sessionToken)` method returns the CSRF token for a given session.
- `internal/web/server.go`: New `executeTemplate(w, r, name, data)` wrapper auto-injects `CSRFField` (`template.HTML` hidden input) and `CSRFToken` (raw string) into every page render data map.
- `cmd/controller/main.go`: All route registrations wrapped with `webServer.CsrfProtect(...)` middleware. Version bumped to `v0.23.0`.
- All handlers (`handlers.go`, `storage_handlers.go`, `handler_restore.go`): Switched from `s.render(w, ...)` to `s.executeTemplate(w, r, ...)`.
- All templates updated:
- `layout.html`: Added `<meta name="csrf-token">` and inline `csrfHeaders()` JS helper (returns `{'X-CSRF-Token': ...}`) in `<head>` (before page-specific scripts). Updated 4 fetch POST/DELETE calls.
- `settings.html`: Added `{{$.CSRFField}}` to 5 forms inside `{{range .StoragePaths}}` (must use `$` for outer scope inside range). Added `{{.CSRFField}}` to 3 page-level forms. Inline-label form uses `document.querySelector('meta[name="csrf-token"]').content`. Updated 5 fetch calls.
- `deploy.html`: Added `{{.CSRFField}}` to cross-backup form. Updated 3 fetch calls.
- `backups.html`: Updated 3 fetch calls. Dynamically-created restore form injects `_csrf` from meta tag.
- `storage_init.html`, `storage_attach.html`, `migrate.html`, `migrate_drive.html`, `app_info.html`, `restore.html`: All fetch calls updated.
- `storage_attach.html`: Replaced `navigator.sendBeacon()` with `fetch(..., {keepalive: true})` — `sendBeacon` cannot send custom headers, making CSRF impossible.
**Hub changes (v0.3.8):**
- `internal/web/server.go`: Replaced insecure literal `hub_session=authenticated` cookie with proper server-side session map.
- New `hubSession` struct with `csrfToken string` and `expiresAt time.Time`.
- `sessions map[string]*hubSession` + `sessionsMu sync.RWMutex` on `Server` struct.
- `handleLogin`: Generates cryptographically random 64-char hex session token + 64-char hex CSRF token. Cookie gains `SameSite=Lax` and `Secure` (when TLS) attributes. Session expires after 7 days.
- `RequireAuth`: Validates session token against map (constant-time compare), redirects to `/login` on failure.
- `CleanupSessions(ctx)`: Goroutine that purges expired sessions every hour.
- CSRF validation block at top of `ServeHTTP`: checks `X-CSRF-Token` header or `_csrf` form field on POST/DELETE/PATCH. Skips when no session cookie (Basic Auth / API path).
- `csrfToken(r)`, `csrfField(r)` helpers for template data injection.
- `internal/web/configs.go`: Added `html/template` import. All template render calls pass `CSRFField template.HTML` and/or `CSRFToken string`. `renderConfigForm` gains `r *http.Request` parameter.
- Templates updated:
- `config_form.html`: Added `{{.CSRFField}}` inside the `<form>`.
- `customer_unified.html`: Added `<meta name="csrf-token">` + inline `csrfHeaders()` in `<head>`. Added `{{.CSRFField}}` to all 5 POST forms (unblock, block, delete config, create-config, regen-password). Updated 3 JS fetch POST calls (trigger-update, push-config, pull-config).
- `cmd/hub/main.go`: Started `go webServer.CleanupSessions(ctx)` goroutine.
### v0.22.3 — Hub Asset Sync (2026-02-21)
**Hub-managed asset downloads**
- New `internal/assets` package: downloads and caches app assets (logos, screenshots) from the Hub API with SHA-256 change detection.
- Asset syncer resolves files from downloaded cache first, falls back to baked-in `/usr/share/felhom/assets/` directory.
- Config: `assets.sync_enabled: true` + `assets.sync_schedule: "05:00"` to enable daily sync.
- API: `POST /api/assets/sync` triggers on-demand sync, `GET /api/assets/status` returns sync status.
- Web server's `serveAsset()` now routes through syncer's `Resolve()` when available.
### v0.22.2 — Setup Logo Fix (2026-02-21)
- **Fix setup wizard logo**: Logo failed to load because `handleLogo()` tried to read it as a file from the filesystem, but it only exists as an embedded string constant. Now imports and serves `web.FelhomLogoSVG` directly.
### v0.22.1 — Setup Wizard Bugfixes (2026-02-21)
- **Fix setup mode detection**: Remove `demo-felhom` from `NeedsSetup()` check — only empty `customer.id` triggers setup mode. Previously the demo customer was stuck in setup mode.
- **Fix CSRF nil pointer panic**: `renderError()` was passing `nil` instead of `*http.Request` to `ensureCSRFToken()`, causing panic when rendering error pages.
- **Fix double-v version display**: Welcome page showed "vv0.22.0" — removed redundant `v` prefix from template.
- **Fix IP detection in Docker**: Setup wizard showed container bridge IP (172.x) instead of host LAN IP. Now reads `HOST_IP` env var (set by docker-setup.sh).
- **Add Hub download logging**: Log Hub config download attempts and errors for easier debugging.
- **docker-setup.sh**: Inject `HOST_IP` env var into generated docker-compose.yml.
### v0.22.0 — First-Run Setup Wizard & Local Infra Backup (2026-02-21)
Major feature release: moves ALL initial configuration and disaster recovery setup from `docker-setup.sh` into the controller itself as a web-based wizard.
**Setup Wizard (`internal/setup/`):**
- New web-based setup wizard replaces interactive CLI wizard from `docker-setup.sh`
- Dual listener: `:8080` (behind Traefik) + `:8081` (direct HTTP for LAN access before DNS is configured)
- Setup mode detection: controller enters wizard when `customer.id` is empty
- Two paths: "Restore from backup" (local drive scan + Hub recovery) and "Fresh install" (Hub download or manual config)
- Drive scanner: detects `.felhom-infra-backup/` on all connected drives, validates checksums
- Hub recovery: `GET /api/v1/recovery/{id}` with retrieval password auth — returns combined config + infra backup
- CSRF protection (cookie + hidden field) for all wizard POST endpoints
- State persistence (`setup-state.json`) survives browser crashes
- All UI text in Hungarian, uses existing dark theme CSS
- After setup: writes `controller.yaml`, creates `settings.json`, `os.Exit(0)` → Docker restart into normal mode
**Local Infra Backup (`internal/backup/local_infra.go`):**
- Writes infrastructure backup to all connected drives as `.felhom-infra-backup/backup.json` + `metadata.json`
- Schema-versioned with SHA256 checksum validation
- Runs on startup and after each nightly backup cycle
- Enables disaster recovery without Hub connectivity — any drive can bootstrap a new controller
**Hub Verification:**
- Pusher parses Hub report response for `customer_blocked` field
- Updates `hub_verified` / `hub_verified_at` in settings on each successful push
- `IsLimitedMode()` checks verification state + 7-day grace period
**Recovery Info:**
- New `internal/recovery/` package generates `recovery-info.txt` in data directory
- Settings page shows recovery info section (customer ID, Hub URL, masked retrieval password)
- Recovery file auto-regenerated on each startup when retrieval password is set
**Pending Events:**
- New `PendingEvent` type in settings with `AddPendingEvent()` / `DrainPendingEvents()`
- Events queued during setup (e.g., DR completed) are drained and pushed to Hub on first successful report push
**Config & Settings Schema:**
- `config.go`: Added `SetupListen` field (default `:8081`), `LoadPermissive()`, `Default()`
- `settings.go`: Added `hub_verified`, `hub_verified_at`, `retrieval_password`, `pending_events` fields with RWMutex accessors
**Infrastructure:**
- `docker-compose.yml`: Added port `8081:8081` mapping for setup wizard
- Removed old fresh-deployment auto-restore code from `main.go` (lines 70-141)
- Removed `restoreSettingsFromHub()` and `restorePasswordsFromHub()` helpers
### v0.21.3 — Config Apply Infra Push + Fixes (2026-02-20)
- **Push infra backup after config apply**: After a successful `POST /api/config/apply`, the controller immediately pushes an infra backup to the Hub so the config sync status updates right away.
- **Fix double "v" prefix in startup event**: "Controller elindult (vv0.21.2)" → "Controller elindult (v0.21.3)".
### v0.21.2 — Config Apply Bind Mount Fix (2026-02-20)
- **Fix config apply on Docker bind mounts**: `POST /api/config/apply` failed with "device or resource busy" because `os.Rename()` doesn't work on bind-mounted files. Now falls back to direct write when rename fails.
### v0.21.1 — Config Content Endpoint (2026-02-20)
- **`GET /api/config`**: New endpoint returning raw controller.yaml content (text/yaml). Used by Hub for live config diff and pull operations. Same auth as other config endpoints (Bearer token or session cookie).
### What was just completed (2026-02-20 session 64)
- **v0.21.0 — Hub Monitoring Takeover (Controller-side, Phases 5+6):**
Replaces external Healthchecks.io dependency with Hub-native event system. The controller now pushes structured events directly to the Hub's `/api/v1/event` endpoint, and the Hub handles dead man's switch detection, notification dispatch, and cooldown management.
**Phase 5 — Event Push System (`internal/notify/notifier.go`):**
- New core method `PushEvent(eventType, severity, message, details)` — non-blocking goroutine, 2 retries with 3s backoff, POSTs to Hub `/api/v1/event`
- 8 typed detail structs: `BackupDetails`, `DBDumpDetails`, `DiskDetails`, `HealthDetails`, `StorageDetails`, `UpdateDetails`, `AppDetails`, `CrossDriveDetails`
- Replaced all old `Notify*` methods with event-based equivalents:
- `NotifyBackupCompleted/Failed` → `backup_completed`/`backup_failed` events
- `NotifyDBDumpCompleted/Failed` → `db_dump_completed`/`db_dump_failed` events
- `NotifyIntegrityOK/Failed` → `backup_integrity_ok`/`backup_integrity_failed` events
- `NotifyHealthChange` → detects transitions, pushes `health_degraded`/`health_critical`/`health_recovered`
- `NotifyStorageDisconnected/Reconnected` → `storage_disconnected`/`storage_reconnected` events
- `NotifyControllerStarted` → `controller_started` event on startup
- `NotifyControllerUpdated` → `controller_updated` event (replaces `NotifyUpdateSuccess/Failed`)
- `NotifyAppDeployed/Removed` → `app_deployed`/`app_removed` events
- `NotifyCrossDriveCompleted/Failed` → `crossdrive_completed`/`crossdrive_failed` events
- `NotifyDRStarted/Completed` → `disaster_recovery_started`/`disaster_recovery_completed` events
- Removed old `/api/v1/notify` relay, `classifyWarning()`, and client-side cooldown logic (Hub handles cooldowns now)
- `SendTest()` now pushes `test` event type via `PushEvent`
- `SyncPreferences` updated to include `cooldownHours` parameter
**Phase 5 — Event Wiring:**
- `main.go`: Wired success events for backup, db-dump, integrity check; startup event with 5s delay; update event after `VerifyStartup()`
- `router.go`: Added `NotifyAppDeployed`/`NotifyAppRemoved` after successful deploy/remove via API
- `handler_restore.go`: Added `NotifyDRStarted`/`NotifyDRCompleted` in DR restore flow
- `server.go`: New `HubPushStatusData` struct and `SetHubPushStatus` callback for monitoring page
**Phase 5 — Hub Connection Monitoring:**
- `pusher.go`: Added `PushStatus` tracking (LastAttempt, LastSuccess, LastError, Consecutive failures) to report Pusher
- `handlers.go`: Monitoring page now shows Hub connection status (connected/unreachable, URL, customer ID, last success, last error) instead of Healthchecks ping UUIDs
- `monitoring.html`: Replaced "Távoli monitoring" section with "Hub kapcsolat" section
- `alerts.go`: Replaced "Missing ping UUIDs" alert with Hub connection alerts (`hub-disabled` warning, `hub-unreachable` error)
**Phase 5 — Expanded Notification Settings:**
- `settings.html`: Expanded from 4 checkboxes to 11 grouped toggles in two categories:
- "Hibák és figyelmeztetések": backup_failed, db_dump_failed, backup_integrity_failed, crossdrive_failed, disk alerts, storage_disconnected, node_down, health_critical, expected missed
- "Tájékoztató": storage_reconnected, health_recovered
- Compound toggles: "Lemez figyelmeztetés" maps to `disk_warning` + `disk_critical`; "Elvárt mentés elmaradt" maps to `expected_backup_missed` + `expected_dbdump_missed`
- `settings.go`: Updated `DefaultEnabledEvents` to new Hub event types
- `handlers.go`: Updated settings POST handler for expanded event names and compound toggles
**Phase 6 — Config Cleanup:**
- `main.go`: Deprecation log on startup when ping UUIDs are configured: `[INFO] Healthchecks ping UUIDs configured but no longer used — monitoring is now handled by the Hub`
- Pinger still runs for transitional backward compatibility
### What was just completed (2026-02-20 session 63)
- **v0.20.0 — Hub Config Management (Phase B):**
Two new features enabling the Hub to manage and compare controller configuration remotely.
**Feature A — Config Apply Endpoint:**
- `router.go`: Added `POST /api/config/apply` — accepts YAML body from Hub, validates it's parseable via `config.LoadFromBytes()`, writes atomically to controller.yaml (`.tmp` + `os.Rename`), returns success JSON. Restart required to apply.
- `router.go`: Added `GET /api/config/hash` — returns SHA256 hex digest of current controller.yaml
- `router.go`: Router struct gained `configPath string` field; `NewRouter()` signature updated
- `config.go`: Added `LoadFromBytes([]byte)` — parses YAML without file I/O (for validation)
- `config.go`: Added `FileHash(path)` — SHA256 hex digest helper
- `main.go`: Config endpoints use same dual auth middleware as self-update (session OR Hub API key Bearer token)
- `main.go`: Added `/api/config/` mux entry with `selfUpdateAuthMiddleware`
**Feature B — Config Hash in Reports:**
- `types.go`: Added `ConfigHash string` field to `Report` struct (JSON: `config_hash`)
- `builder.go`: `BuildReport()` now accepts `configPath string` parameter, computes SHA256 of controller.yaml and includes it in every report
- `main.go`: All 4 `BuildReport()` call sites updated to pass `*configPath`
- Hub uses this hash to compare against its generated YAML — shows "In sync" / "Config mismatch" / "Unknown" on the unified customer detail page
### What was just completed (2026-02-20 session 62)
- **docker-setup.sh — Hub Config Download:**
- Added `--hub-customer` and `--hub-password` CLI flags for downloading pre-configured controller.yaml from Felhom Hub
- Added `HUB_URL` global variable (default: `https://hub.felhom.eu`)
- Hub download logic at start of `run_config_wizard()`: downloads YAML via `curl` with `X-Retrieval-Password` header, validates response, extracts key variables (domain, CF tokens, email), sets global variables for subsequent setup steps
- Falls back to interactive wizard if download fails or credentials not provided
### What was just completed (2026-02-20 session 61)
- **v0.19.0 — Deployed App Removal + Missing Field Injection:**
Two new features: "Eltávolítás" (Remove) action for deployed stacks and automatic missing deploy field injection on template updates.
**Feature A — Deployed App Removal ("Eltávolítás"):**
- `delete.go`: Added `RemoveStack()` — removes deployed (non-orphaned) stack: `docker compose down --volumes`, optional HDD data cleanup, optional backup data cleanup (DB dumps + cross-drive rsync), removes `app.yaml` only (template files preserved for redeploy); stack reverts to "Nincs telepítve" state
- `delete.go`: Added `GetStackBackupData()` — returns backup path info (DB dump dir + cross-drive rsync dir) with sizes and existence status
- `delete.go`: Added `RemoveResponse`, `BackupDataResponse` structs, `buildPathInfo()` helper
- `router.go`: Added `POST /api/stacks/{name}/remove` endpoint — accepts `{remove_hdd_data, remove_backups}`, computes backup paths via `AppDBDumpPath()`/`AppSecondaryRsyncPath()`, cleans cross-drive config on success
- `router.go`: Added `GET /api/stacks/{name}/backup-data` endpoint — returns backup data paths with sizes
- `crossdrive.go`: Made `getAppDrivePath` → `GetAppDrivePath` (public) for use by router
- `stacks.html`: Added "Eltávolítás" button for stopped, deployed, non-orphaned, non-protected stacks
- `dashboard.html`: Same button in compact card layout
- `layout.html`: Added `removeStack()` modal — fetches HDD + backup data in parallel, 3-section layout (always removed / HDD data with checkbox / backup data with checkbox), reimport warning for preserved HDD data, restic retention note
- `layout.html`: Added `confirmRemoveStack()` — POST to `/remove`, shows result summary with removed/preserved paths
**Feature B — Missing Deploy Field Injection:**
- `deploy.go`: Added `InjectMissingFields(stackNames)` — iterates deployed stacks, compares `.felhom.yml` deploy_fields against `app.yaml` env vars, auto-generates values for missing `secret` (using generator spec) and `domain` fields, saves updated `app.yaml`
- `deploy.go`: Added `base64key` generator type — produces `base64:<N random bytes base64-encoded>` (for Laravel APP_KEY and similar)
- `deploy.go`: Added `containsStr()` helper
- `manager.go`: Added `DeployedStackNames()` — returns names of all deployed stacks
- `sync.go`: Added `postSyncHook func(updated []string)` field to `Syncer`; `New()` accepts optional hook; hook called in `doSync()` after rescan with names of updated stacks
- `main.go`: Wired injection on startup (all deployed stacks) and after sync (updated stacks only)
### v0.18.0 (2026-02-19 session 60)
- **v0.18.0 — Drive Migration & Tier 2 Restic Deprecation:**
Full drive replacement workflow with decommissioned state, enhanced per-app migration with backup awareness, and deprecation of restic as a Tier 2 cross-drive backup method (rsync only).
**Phase 1 — Restic Tier 2 Deprecation:**
- `settings.go`: Auto-migrate restic→rsync on startup via `migrateResticToRsync()` in `Load()`
- `crossdrive.go`: Removed `runResticBackup()`, `pruneResticRepo()`, `ensureResticRepo()`; `RunAppBackup()` calls rsync directly
- `backup.go`: Removed Tier 2 secondary restic scanning from `ListAllSnapshots()`
- `settings.go`: Removed cross-drive restic password methods (`GetOrCreateCrossDrivePassword`, etc.)
- `deploy.html`: Removed method dropdown (rsync/restic selector)
- `handlers.go`: Simplified `Tier2DriveGroup` (flat `Items` list), removed method handling from `settingsCrossBackupHandler()`
- `backups.html`: Removed method split in Tier 2 details section
- `router.go`: Always set method to "rsync" in cross-backup API
- `infra_backup.go`: Removed cross-drive password block from `CollectInfraBackup()`
- `main.go`: Removed `SetCrossDriveResticPassword` restore block
**Phase 2 — Enhanced Per-App Migration:**
- `backup.go`: Extracted `backupDrive()` from `runBackupInternal()` loop; added `TryRunDriveBackup()` with non-blocking lock
- `crossdrive.go`: Added `AnyRunning()` method
- `migrate.go`: Added `BackupTrigger` interface, `MigrateOrchestrator`, `RunEnhancedMigration()` with post-migration steps (DB dump copy, Tier 2 conflict clearing, auto-delete stale data, immediate Tier 1 backup)
- `storage_handlers.go`: Wired orchestrator into migration handler with `auto_delete_stale` support
- `migrate.html`: Added auto-delete checkbox, "cleaning" + "backing_up" progress steps
**Phase 3 — Full Drive Migration:**
- `settings.go`: Added `Decommissioned`/`DecommissionedAt`/`MigratedTo` fields to `StoragePath`; added `SetDecommissioned()`, `ClearDecommissioned()`, `IsDecommissioned()`, `GetDecommissionedPaths()`, `GetStorageLabel()`; `GetConnectedPaths()`/`GetSchedulableStoragePaths()` exclude decommissioned
- `migrate_drive.go` (NEW): `DriveMigrator` with `MigrateDrive()` 10-step flow (validate→stop→rsync→verify→configure→decommission→Tier2→start→backup→notify), `migrationTx` rollback pattern, excludes restic repos from rsync
- `settings.html`: Decommissioned card variant with "Kiváltva" badge, "Összes adat átköltöztetése" button on connected cards
- `migrate_drive.html` (NEW): Drive migration wizard (form + progress + done cards)
- `storage_handlers.go`: Added `/api/storage/migrate-drive`, `/api/storage/migrate-drive/status`, `/api/storage/decommission/remove` endpoints
- `server.go`: Added `/settings/storage/migrate-drive` route, `SetDriveMigrator()` setter
- `watchdog.go`: Skip decommissioned drives in `Check()`; block `SafeDisconnect()` for decommissioned
- `healthcheck.go`: Skip decommissioned paths in `checkStoragePaths()`
- `backup.go`: Skip decommissioned drives in `backupDrive()`/`runDBDumpsInternal()`; added `MigrationActiveCheck` callback to skip nightly backup during migration
- `crossdrive.go`: Reject decommissioned destinations in `ValidateDestination()`; skip decommissioned paths in `AutoEnableSmallApps()`
- `handlers.go`: Skip decommissioned drives in `buildStorageBars()`; made `SyncFileBrowserMounts()` public
- `main.go`: Added `driveMigrateStackAdapter`, wired `DriveMigrator` with all dependencies
**Phase 4 — Hub Changes:**
- `report/types.go`: Added `Decommissioned`/`MigratedTo` fields to `StorageReport`
- `report/builder.go`: Include decommissioned drives in report with flag
**Files modified:** 21 files modified + 2 new files (`migrate_drive.go`, `migrate_drive.html`).
### What was just completed (2026-02-19 session 59)
- **v0.16.1 + hub v0.1.8 — Hub Update Trigger + Controller URL Reporting:**
Controller now includes its external URL (`controller_url`) in periodic hub reports so the hub can trigger self-updates remotely. Hub tracks the URL in a new `controller_url` DB column, checks the Gitea registry for the latest controller image version (VersionChecker goroutine, `web/version.go`), and shows a "Controller Update" card on the customer detail page.
**Controller (v0.16.1):**
- `internal/report/types.go`: Added `ControllerURL string` field to Report struct.
- `internal/report/builder.go`: Sets `ControllerURL` from `cfg.Customer.Domain` → `https://felhom.<domain>`.
- `internal/api/router.go`: **Bug fix** — moved selfupdate routes to before `hasSuffix(path, "/update")` stack case (which was catching `/selfupdate/update` first).
**Hub (v0.1.8):**
- `cmd/hub/main.go`: Added `Registry` config section + defaults; creates `VersionChecker` goroutine if credentials configured; passes `apiKey` to `web.New()`.
- `internal/store/store.go`: Added `ControllerURL` to `CustomerSummary`; idempotent `ALTER TABLE reports ADD COLUMN controller_url TEXT` migration; updated `SaveReport`, `GetCustomers`, `GetCustomer`, `GetCustomerHistory` queries.
- `internal/web/version.go` (NEW): `VersionChecker` type — polls Gitea Docker Registry V2 API (`/v2/<owner>/<repo>/tags/list`) every 6h; parses semver tags; stores latest version thread-safely.
- `internal/web/server.go`: Added `apiKey`, `versionChecker` fields; updated `New()` signature; added `SetVersionChecker()`; added `handleTriggerUpdate` handler that proxies POST to controller's `/api/selfupdate/update`; added trigger-update route (before `/customers/` catch-all); updated `handleCustomerDetail` with `ControllerURL`, `LatestVersion`, `UpdateAvailable` template data; added `compareVersions` helper.
- `internal/web/templates/customer.html`: New "Controller Update" section between Health and Notifications — shows current/latest version with update indicator, controller URL link, and conditional "Trigger Update" button with JS.
- `internal/api/handler.go`: Added `ControllerURL` to `/api/v1/customers` JSON response.
- Hub config (`hub.yaml`): Added `registry:` section with Gitea admin credentials.
**Files modified/created:** controller: 3 files; hub: 5 modified + 1 created (version.go).
### What was just completed (2026-02-19 session 58)
- **v0.16.0 — Controller Self-Update:**
Watchtower-style self-update mechanism. New package `internal/selfupdate/` with 3 files: `version.go` (semver parsing/comparison), `state.go` (audit log state file I/O), `updater.go` (registry check via Gitea V2 API, update trigger, startup verification).
**Flow:** Gitea registry tag list → `docker pull` → atomic compose file rewrite → `docker compose up -d` → process replaced. State file (`update-state.json`) persists across restart as audit log; verified on next startup to detect success/failure.
**Config:** `SelfUpdateConfig` extended with `AutoUpdateTime` field + defaults for `Image` and `AutoUpdateTime`. Scheduler jobs: periodic check every `check_interval` (default 6h); optional daily auto-update at `auto_update_time` (default 04:30).
**API:** 3 new endpoints under `/api/selfupdate/` (`status`, `check`, `update`). Auth via session cookie OR `Authorization: Bearer <hub_api_key>` header (for external triggering from build scripts).
**UI:** Settings page "Verzió és frissítés" card shows current/latest version, check time, auto-update status, last update result. "Frissítés keresése" button queries registry; "Frissítés telepítése" button appears when update is available. `pollUntilBack()` JS polls `/api/health` after triggering update and reloads when container is back up.
**Notifications:** `NotifyUpdateSuccess()` and `NotifyUpdateFailed()` added to notifier for post-update startup verification results.
**Alert:** Dashboard shows "Új controller verzió elérhető" info alert when update is available.
**docker-compose.yml:** Added `/opt/docker/felhom-controller:/opt/docker/felhom-controller` directory bind mount (required for compose file access during self-update); named volume and read-only config override on top.
**Files modified/created (12):** `internal/selfupdate/version.go` (NEW), `internal/selfupdate/state.go` (NEW), `internal/selfupdate/updater.go` (NEW), `internal/config/config.go`, `internal/notify/notifier.go`, `internal/api/router.go`, `internal/web/server.go`, `internal/web/handlers.go`, `internal/web/alerts.go`, `internal/web/templates/settings.html`, `cmd/controller/main.go`, `docker-compose.yml`
### What was just completed (2026-02-19 session 57)
- **v0.15.7 — Fix backup page storage display & rename system drive label:**
Backup page ("Biztonsági mentés") now shows all registered storage paths instead of only a single "Külső HDD". Added `data["StorageBars"] = s.buildStorageBars()` to `backupsHandler` (was missing unlike dashboard/monitoring handlers). Updated `backups.html` storage bars section to use `StorageBars` loop (same pattern as monitoring page), replacing the old `{{if .HDDConfigured}}` single-HDD block.
Renamed system root partition label from "SSD (/)" to "Rendszer (/)" on all three pages (backup, monitoring, dashboard), as the root filesystem is not necessarily on an SSD.
**Files modified (4):** `internal/web/handlers.go`, `internal/web/templates/backups.html`, `internal/web/templates/monitoring.html`, `internal/web/templates/dashboard.html`
### What was just completed (2026-02-19 session 56)
- **v0.15.6 (controller) + hub v0.1.7 — Bug hunt fixes (BUGHUNT.md):**
**Controller — Restore race conditions (P0-P1):** All 4 restore handlers (`restorePageHandler`, `apiRestoreStatus`, `apiRestoreAll`, `apiRestoreSkip`) now hold `restoreMu.RLock()` across nil-check and field reads. `apiRestoreAll` uses new `TryStartRestore()` method for atomic check-and-set (eliminates double-restore race). `executeAllRestores()` snapshots plan under lock, uses `SetStatus("done")` instead of direct write. Removed dead no-op goroutine.
**Controller — restore_scan.go:** `dirIsEmpty()` now returns `false` on read errors (was silently treating unreadable dirs as empty, losing backup data). `Snapshot()` deep-copies Apps and Drives slices. Added `TryStartRestore()`, `SetStatus()`, `GetStatus()` helper methods.
**Controller — infra_backup.go (P0):** `controller.yaml` read failure now returns a real error (was silently creating empty backup). `settings.json` and restic password read failures now logged. Added `logger *log.Logger` parameter to `BuildInfraBackup`.
**Controller — main.go DR wiring:** Fixed ordering — `restoreSettingsFromHub` + settings reload now happens before `restorePasswordsFromHub` (prevents cross-drive password loss). Nil check after `ScanDrivesForBackups`. `os.MkdirAll` error now logged. `os.MkdirAll` added to `restoreSettingsFromHub` before write.
**Hub — store.go (P2):** 5 `json.Unmarshal` calls now log `[WARN]` on failure. `GetInfraBackupMeta` logs unmarshal error instead of silently returning wrong counts.
**docker-setup.sh (P0-P2):** DRY_RUN check moved to top of `run_config_wizard()` with dummy values (was prompting interactively even in dry-run). CF tunnel token quoted in docker-compose env. `htpasswd` uses `cut -d: -f2` + bcrypt format validation. `grep -qF` for literal path matching. Volume paths quoted in YAML output. Post-wizard validation rejects default `demo-felhom`/`homeserver.local` values.
**restore.html (P2-P3):** Error text uses `textContent` instead of `innerHTML`. Poll errors counted; after 10 failures shows "Kapcsolat megszakadt" message instead of polling silently forever.
**Files modified (controller, 6):** `internal/backup/restore_scan.go`, `internal/web/handler_restore.go`, `internal/report/infra_backup.go`, `cmd/controller/main.go`, `internal/web/templates/restore.html`, `scripts/docker-setup.sh`
**Files modified (hub, 1):** `hub/internal/store/store.go`
### What was just completed (2026-02-19 session 55)
- **v0.15.5 — Fix startup hub report silently failing:**
`Push()` now returns actual errors instead of always `nil`. Previously, push failures were logged internally but the caller could never detect them, leading to a misleading `[INFO] Startup hub report sent` log even when the push actually failed (e.g., hub returning HTTP 503 during simultaneous deployment). Removed the "Never returns error to caller" behavior: marshal error returns a wrapped error, and after 3 failed retries the error is returned to the caller (the internal `[WARN]` log before `return nil` is gone).
Startup hub push now retries 3 times with 15-second delays between outer attempts, giving the hub time to come up when both are deployed together. Each outer attempt uses `Push()`'s own internal 3-retry logic (5s backoff), so the hub gets up to ~40s total to become ready. If all 3 outer attempts fail, logs a clear warning with the next scheduled push interval.
**Files modified (2):** `internal/report/pusher.go`, `cmd/controller/main.go`
### What was just completed (2026-02-19 session 54)
- **v0.15.4 (controller) + hub v0.1.6 — Hub reporting improvements:**
**Controller:** When `hub.enabled: false` but URL+API key are configured, the controller now creates the `Pusher` and sends a one-time "disabled" notification on startup (`health.status = "disabled"`, `reporting_disabled: true`). This replaces the old behavior where a disabled controller was indistinguishable from a crashed node. Added `PushOnce()` method to `Pusher` (bypasses the `enabled` flag). Added `ReportingDisabled` field to the `Report` struct.
**Hub:** Added "disabled" status handling — when the latest report has `health_status = "disabled"`, the overall status is "disabled" (checked BEFORE the stale-time logic, so it stays "PAUSED" even after 30min+). Dashboard shows gray "PAUSED" badge. Customer detail shows "Reporting has been disabled on this node" with a hint to re-enable. Storage labels now shown (`label` field with fallback to `mount`). Report history timestamps now show date + time ("Feb 19 09:46" instead of "09:46:54"). New `.status-badge-disabled` CSS (neutral gray `#475569`).
**Files modified (controller):** `internal/report/types.go`, `internal/report/pusher.go`, `cmd/controller/main.go`
**Files modified (hub):** `hub/internal/web/server.go`, `hub/internal/web/templates/dashboard.html`, `hub/internal/web/templates/customer.html`, `hub/internal/web/templates/style.css`
### What was just completed (2026-02-19 session 53)
- **v0.15.3 — Show all storage paths on dashboard + fix hub report:**
Dashboard ("Vezérlőpult") and monitoring ("Rendszermonitor") pages now show usage bars for ALL registered storage paths instead of just one hardcoded "Külső HDD" bar. New `StorageBarInfo` type and `buildStorageBars()` helper build bars from `settings.GetStoragePaths()`. Each bar shows the storage label and live disk usage.
Hub storage report now correctly includes all registered storage paths with proper mount paths and labels. Previously it sent only root `/` plus one HDD entry using the deprecated (empty) `cfg.Paths.HDDPath`. Now uses `system.GetDiskUsage()` per storage path, same as the dashboard bars. Added `Label` field to `StorageReport` in `types.go`.
**Files modified (5):** `internal/web/handlers.go`, `internal/web/templates/dashboard.html`, `internal/web/templates/monitoring.html`, `internal/report/builder.go`, `internal/report/types.go`
### What was just completed (2026-02-19 session 52)
- **v0.15.2 — Fix data loss on container restart (2 bugs):**
**Bug 1:** Snapshot history delta stats (HOZZÁADOTT, ÚJ FÁJL, VÁLTOZOTT) showed 0 after container restart because restic doesn't store these stats — they were only in memory. Fixed by persisting the snapshot history ring buffer to `data/snapshot-history.json`. On startup, persisted stats are merged with restic repo snapshots. Added `saveSnapshotHistory()` (atomic write via tmp+rename), `loadSnapshotHistoryFromFile()`, updated `appendSnapshotRecord()` to save after each backup, and updated `LoadSnapshotHistory()` to merge persisted + restic data.
**Bug 2:** DB validation (ÉRVÉNYESÍTÉS column) showed "–" after restart because the synthesized `LastDBDump.Results` didn't copy `Validation` from `DumpFileInfo`. One-line fix: added `Validation: f.Validation` to the synthesized `DumpResult` in `GetFullStatus()`.
**Files modified:** `internal/backup/backup.go`
### What was just completed (2026-02-19 session 51)
- **v0.15.1 — Backup Page "Részletek" Overhaul:**
Replaced the "Tároló" section on the backup page with a new "Részletek" section containing 3 collapsible tier sections with per-drive breakdowns.
**Tier 1 (Helyi mentés):** Shows per-drive restic repo stats (size, snapshot count) with storage labels. Includes aggregated totals when multiple drives exist, plus DB dump summary, integrity check, and encryption key (all carried over).
**Tier 2 (Másodlagos másolat):** Groups cross-drive backup items by destination drive, separated into restic and rsync method sections with per-app sizes.
**Tier 3 (Távoli mentés):** Placeholder for future B2/S3/SFTP remote backup.
**Restore UI improvements:** Snapshot dropdown now groups by tier (optgroup), shows tier label + drive name per snapshot (e.g., "1. szint, hdd_1"), and marks Tier 1 as recommended. Also lists Tier 2 (secondary restic) snapshots for visibility.
**Backend:** New `DriveRepoInfo` struct, `perDriveRepoStats()` method, `ListAllSnapshots()` that includes secondary restic repos, and `Tier2DriveGroup` handler struct. `SnapshotInfo` now carries `Tier` and `DriveLabel` fields.
**Files modified (5):** `internal/backup/backup.go`, `internal/backup/restic.go`, `internal/web/handlers.go`, `internal/api/router.go`, `internal/web/templates/backups.html`, `internal/web/templates/style.css`
### What was just completed (2026-02-18 session 50)
- **v0.15.0 — Attach Existing Drive (bind mount wizard):**
New feature: Settings → "Meglévő meghajtó csatolása" wizard. Allows attaching a drive that already has a filesystem (ext4, etc.) without formatting. Solves the real-world scenario where a customer's drive contains existing data that must be preserved.
**How it works:** The partition is mounted read-only at a hidden staging path (`/mnt/.felhom-raw/<label>`). A directory browser lets the user navigate the drive's contents and create a new folder. The selected folder is bind-mounted at `/mnt/<hdd-name>`, keeping the controller's data isolated from existing files. Two fstab entries (raw + bind, both with `nofail`) ensure the mount survives reboots.
**Wizard flow:** Scan → Select partition (only shows partitions with existing FS) → Mount raw + Browse directories → Create folder if needed → Configure mount name + label → Finalize (bind mount + fstab + permissions + register). Cancel cleans up the temp mount.
**New files (4):** `internal/storage/attach.go`, `internal/storage/attach_linux.go`, `internal/storage/attach_other.go`, `internal/web/templates/storage_attach.html`
**Modified files (3):** `internal/web/storage_handlers.go` (6 new API handlers), `internal/web/server.go` (route + activeRawMount field), `internal/web/templates/settings.html` (button)
### What was just completed (2026-02-18 session 49)
- **v0.14.2 — Backup Bug Fixes (4 fixes from code review):**
**Bug 1 (HIGH):** rsync `--delete` was destroying `_db/` and `_config/` directories on every single-mount run. Fixed by adding `--exclude _*` to the rsync command in `runRsyncBackup()`. Controller-managed directories (underscore prefix) are now excluded from `--delete` cleanup. (`crossdrive.go`)
**Bug 2 (MEDIUM):** Scheduled backups (`RunBackup`, `RunDBDumps`) did not set `m.running`, so UI showed "not running" during nightly jobs and restore could overlap. Fixed by extracting `acquireRunning()` / `releaseRunning()` helpers and `runDBDumpsInternal()` / `runBackupInternal()` internal methods. All three public entry points now guard with the running flag; `RunFullBackup()` calls the internal methods directly to avoid deadlock. (`backup.go`)
**Bug 3 (MEDIUM):** `ValidateDestination` silently succeeded when `GetDiskUsage` returned nil (exotic filesystems, FUSE, NFS). Fixed by logging `[WARN]` and returning nil (backward-compatible). (`crossdrive.go`)
**Bug 4 (MEDIUM):** Empty `systemDataPath` produced relative dump paths. Fixed with: startup `[WARN]` in `NewManager()`, `[ERROR]` log in `GetAppDrivePath()`, and explicit guard in `DumpStackDB()` that returns an error when path is empty or non-absolute. (`backup.go`)
**Files modified (2):** `internal/backup/backup.go`, `internal/backup/crossdrive.go`
### What was just completed (2026-02-18 session 48)
- **v0.13.1 — UI Polish Fixes Round 2 (4 fixes):**
**Fix 1:** Deploy page "Biztonsági mentés" section now has proper card border. Root cause: `.deploy-cross-drive` used undefined CSS variables `--card-bg` and `--border` (only `--bg-secondary` and `--border-color` exist). Fixed by using correct vars (`style.css`).
**Fix 2:** Auto-generated env values section cleaned up (`deploy.html`, `style.css`). Badge moved inline with label. "Másolás" buttons removed (native select+copy sufficient). Secret fields keep show/hide toggle. Non-secret fields now plain readonly input without button wrapper. Removed `copyAutoField()` JS. CSS updated: `.form-group-auto` now block layout (was flex row), label uses `display: flex; gap: .5rem`, badge downsized to `0.75rem / normal weight`, readonly inputs get muted background.
**Fix 3:** Snapshot table n/a → 0 (`backups.html`). Replaced `<span class="col-na" title="...">n/a</span>` with plain `0` in all three stats columns. Removed `.col-na` CSS class (no longer used).
**Fix 4:** Disk warnings moved from top banner to inline under storage bars (`alerts.go`, `layout.html`, `handlers.go`, `dashboard.html`, `monitoring.html`, `style.css`). Added `Inline bool` field to `Alert` struct. Disk-related warnings set `Inline: true`. Layout banner skips inline alerts. New `GetInlineAlerts(page)` method on `AlertManager`. Dashboard and monitoring handlers pass `DiskWarnings`. Inline warning block rendered below storage bars. New `.inline-warning*` CSS classes (compact, subtle, colored).
**Files modified (8):** `alerts.go`, `handlers.go`, `templates/style.css`, `templates/dashboard.html`, `templates/backups.html`, `templates/deploy.html`, `templates/monitoring.html`, `templates/layout.html`
### What was just completed (2026-02-18 session 47)
- **v0.13.0 — UI Polish Fixes (8 independent fixes):**
**Fix 1:** backup-status-card border already correct (verified same styling as system-info-card).
**Fix 2:** Deploy page auto-generated fields now show actual values for deployed apps (`deploy.html`, `handlers.go`). Secrets show as password fields with show/hide toggle; domain/plain values show as readonly text with copy button. JS helpers `toggleAutoField()` / `copyAutoField()` added.
**Fix 3:** Temperature display made more prominent (`dashboard.html`, `style.css`). Dot enlarged to 11px; value wrapped in colored pill badge (`.temp-value-pill` / `.temp-pill-{green|yellow|red}`).
**Fix 4:** Dashboard backup card reworked (`dashboard.html`, `handlers.go`). Removed "Mentés most" button and `triggerBackup()` JS. Removed "Tároló méret" line. Added Tier 2 status line (configured/total apps) + warning row for failed cross-drive backups. Handler now computes `CrossDriveTotal`, `CrossDriveConfigured`, `CrossDriveFailed`.
**Fix 5:** HDD warning banner scoped to dashboard + monitoring pages only (`alerts.go`, `layout.html`, `funcmap.go`). Added `PageOnly []string` field to `Alert` struct. Disk-related warnings (keywords "meghajtón", "adattároló") get stable ID `"disk-not-separate"` + `PageOnly: ["dashboard", "monitoring"]`. `pageMatch()` template function added. Layout renders alerts conditionally.
**Fix 6:** Tárhely section moved up in Rendszermonitor — now appears right after "Rendszer áttekintés", before "Távoli monitoring" (`monitoring.html`).
**Fix 7:** Snapshot table improvements (`backups.html`, `style.css`). "MÉRET" renamed to "HOZZÁADOTT (új adat)". `–` for unavailable data replaced with `n/a` (with tooltip explaining restic limitations). New `.col-subtitle` and `.col-na` CSS classes.
**Fix 8:** Tároló section restructured into tiers (`backups.html`, `handlers.go`, `style.css`). Tier 1 (restic local), Tier 2 (cross-drive, only shown if configured), DB dump directory + total size. Removed "Távoli másolat: Nincs beállítva" placeholder. Handler passes `DBDumpDir`, `DBDumpTotalBytes`, `Tier2Dests` (deduplicated). New `.repo-tier` / `.repo-tier-title` CSS.
**Files modified (9):** `alerts.go`, `funcmap.go`, `handlers.go`, `templates/style.css`, `templates/dashboard.html`, `templates/backups.html`, `templates/deploy.html`, `templates/monitoring.html`, `templates/layout.html`
### What was just completed (2026-02-18 session 46)
- **v0.12.9 — Tier 2 for All Apps + Status Dot Update:**
**Fix 1: Tier 2 now configurable for ALL apps — not just HDD apps (`crossdrive.go`)**
- Removed `len(mounts) == 0` error gate from `RunAppBackup()` — empty mounts = config-only backup
- rsync: DB dump copy (`_db/`) + config rsync (`_config/`) still runs even with zero HDD mounts
- restic: config dir + DB dump dir still appended even without mount paths
- Non-HDD apps (Mealie, Gokapi, etc.) can now be protected against drive failure via Tier 2
**Fix 2: Status dot logic updated, HasHDDData gate removed (`handlers.go`)**
- `buildAppBackupRows()`: "auto" (gray) status removed — all apps start yellow ("Csak helyi mentés")
- Green requires Tier 2 configured + last status "ok" (not just "configured but never run")
- Tier2 section is now unconditional — no `if app.HasHDDData` gate
- Cross-drive summary loop: removed `if !app.HasHDDData { continue }` — all apps in summary
**Fix 3: Backup page template updates (`backups.html`)**
- Tier 2 row shown for all apps (removed `{{if .HasHDDData}}` gate)
- Meta badge: non-HDD apps show "Konfig" or "Konfig + DB" instead of "Auto"
- Tier 3 placeholder row added (grayed out "Hamarosan / távoli offsite")
- Button text: "Összes HDD mentés" → "Összes 2. mentés futtatása most"
**Fix 4: Deploy page cross-drive section visible for all deployed apps (`deploy.html`)**
- Removed `{{if .StorageInfo}}` double-gate — section now shows for all deployed apps
- Updated heading: "Másolat másik meghajtóra (felhasználói adatok)" → "2. mentés — másolat másik meghajtóra"
- Updated hint: "mint az alkalmazás adattárolója" → "a meghibásodás elleni védelem érdekében"
**Files modified (4):** `internal/backup/crossdrive.go`, `internal/web/handlers.go`, `internal/web/templates/backups.html`, `internal/web/templates/deploy.html`
### What was just completed (2026-02-18 session 45)
- **v0.12.8 — Complete Cross-Drive Backup + Per-Tier UI:**
**Fix 1: Cross-drive backup now includes DB dumps + app config (`crossdrive.go`, `main.go`)**
- `CrossDriveRunner` gets `dbDumpDir` field + `SetDBDumpDir(dir string)` setter
- `copyStackDBDumps()` helper copies `<stackName>_*.sql` files to `_db/` subfolder in rsync dest
- `runRsyncBackup()`: after HDD mount rsync loop, copies DB dumps to `_db/` and rsyncs config dir to `_config/` — both non-fatal on error
- `runResticBackup()`: appends config dir and full DB dump dir to restic paths (restic deduplicates)
- rsync destination layout: `backups/rsync/<app>/_db/` (dumps) + `_config/` (compose+yaml) + user data
- `main.go`: `crossDriveRunner.SetDBDumpDir(cfg.Paths.DBDumpDir)` wired after runner init
**Fix 2: UI restructured from per-layer to per-tier (`handlers.go`, `backups.html`, `style.css`)**
- `AppBackupRow` struct rebuilt: dropped old `DBLastRun/Status`, `VolumeLastRun/Status`, `HasUserData`, `UserDataConfigured/Method/Dest/Schedule/LastRun/LastStatus/LastError/StatusBadge` fields
- New fields: `BackupContents` (e.g., "DB + Konfig + Adatok"), `Tier1LastRun/LastStatus/DBStatus`, `Tier2Configured/Method/MethodLabel/Dest/Schedule/LastRun/LastStatus/LastError/StatusBadge/SizeHuman/Browsable`
- `buildAppBackupRows()` rewritten: destination health now via `s.crossDriveRunner.ValidateDestination()` instead of `system.CheckBackupDestination()`
- `backups.html`: two tier rows (1. mentés / 2. mentés) replace the old three layer rows (DB / Konfig / Userdata)
- `style.css`: added `.tier-label`, `.tier-location`, `.tier-contents`, `.tier-size`, `.tier-browsable` classes
**Fix 3: Cleanup (`router.go`)**
- `filterSnapshotsByPaths()` and `pathCovers()` deleted (were unused since v0.12.7a)
**Files modified (6):** `internal/backup/crossdrive.go`, `cmd/controller/main.go`, `internal/web/handlers.go`, `internal/web/templates/backups.html`, `internal/web/templates/style.css`, `internal/api/router.go`
### What was just completed (2026-02-18 session 44)
- **v0.12.7a — Post-deploy fixes:**
**Fix A: Restore now shows snapshots for all apps (`internal/api/router.go`)**
- Root cause: `filterSnapshotsByPaths` filtered older snapshots (pre-v0.12.7) by HDD paths. Older snapshots don't contain HDD paths (backup wasn't mandatory yet), so Immich got zero snapshots.
- Fix: removed HDD path filtering entirely from `backupSnapshots`. All snapshots contain config + DB dumps and are useful for any app. `RestoreApp` extracts whatever paths are available from the chosen snapshot.
- `filterSnapshotsByPaths` and `pathCovers` functions kept (unused, no compile error).
**Fix B: Clarified "no cross-drive" warning (`internal/web/handlers.go`, `backups.html`, `style.css`)**
- Root cause: "Nincs beállítva" / red dot implied no backup at all — misleading since nightly restic now always covers HDD data.
- `handlers.go`: status `"red"` → `"yellow"`, StatusText → `"Nincs második másolat (csak helyi mentés)"`
- `backups.html`: added `✓ Helyi mentés auto` badge before the `⚠ Nincs 2. másolat` warning
- `style.css`: `.layer-auto-ok` class added (green text for the auto badge)
**Files modified (3):** `internal/api/router.go`, `internal/web/handlers.go`, `internal/web/templates/backups.html`, `internal/web/templates/style.css`
### What was just completed (2026-02-18 session 43)
- **v0.12.7 — Backup Architecture Overhaul (mandatory HDD backup, pre-dump, restore for all apps):**
**Fix 1: HDD data backup now mandatory (`backup.go`, `appdata.go`, `settings.go`)**
- `resolveAppBackupPaths()` rewrote to iterate ALL deployed stacks via `ListDeployedStacks()` — no longer reads `GetAppBackupMap()` or checks `Enabled` flag
- `DiscoverAppData()` signature simplified: dropped `backupPrefs map[string]bool` parameter; `BackupEnabled` is now derived from `HasHDDData` (if app has HDD data, it's always backed up)
- `RefreshCache()` updated to call new `DiscoverAppData(m.stackProvider, status.DiscoveredDBs)` signature
- 5 dead settings methods deleted: `IsAppBackupEnabled`, `SetAppBackup`, `GetAppBackupMap`, `SetAppBackupBulk`, `GetAppBackupPrefs` — `AppBackupPrefs.Enabled` field kept in struct for backward-compat JSON loading
**Fix 2: Cross-drive backup triggers fresh DB dump first (`crossdrive.go`, `backup.go`, `main.go`)**
- New `DBDumper` interface with `DumpStackDB(ctx, stackName)` in `crossdrive.go`
- `CrossDriveRunner` gets `dbDumper` field + `SetDBDumper(d DBDumper)` setter
- `Manager.DumpStackDB()` discovers containers for that stack via `DiscoverDatabases()`, runs `DumpAll()`, persists validation cache — same logic as nightly dump but scoped to one stack
- `RunAppBackup()` calls `DumpStackDB()` before `ValidateDestination()` — non-fatal on failure (logs warn, proceeds with user data)
- `main.go` wires `crossDriveRunner.SetDBDumper(backupMgr)` after both are initialized
**Fix 3: Restore dropdown shows ALL deployed apps (`backups.html`, `restore.go`, `router.go`)**
- `restore.go` rewritten: no `IsAppBackupEnabled()` check; resolves `GetStackComposePath` + `DBDumpDir` + HDD mounts; always restores config+DB, adds user data if `hasHDD`; logs restore type (`config+DB` vs `full (config+DB+userdata)`)
- Restore dropdown template: removed `{{if and .HasHDDData .BackupEnabled}}` filter; every app gets an `<option>` with `data-has-hdd` and `data-has-db` attributes
- New `#restore-type-info` div added between snapshot selector and warnings
- `onRestoreAppChange()` JS updated: reads `data-has-hdd`/`data-has-db` from selected option, shows Hungarian restore type banner (full / config+DB / config only) with color-coded styling
- `router.go` `backupSnapshots`: added clarifying comment for non-HDD apps (no filter = all snapshots returned)
**Fix 4: Honest UI label (`backups.html`)**
- "Docker kötetek" renamed to "Konfiguráció" — Docker named volumes at `/var/lib/docker/volumes/` are NOT in the restic backup paths; what's actually backed up is compose files + app.yaml + .felhom.yml
**CSS: `.restore-info` and `.restore-info-partial` classes added to `style.css`**
**Files modified (9):** `internal/backup/backup.go`, `internal/backup/appdata.go`, `internal/settings/settings.go`, `internal/backup/crossdrive.go`, `internal/backup/restore.go`, `cmd/controller/main.go`, `internal/web/templates/backups.html`, `internal/web/templates/style.css`, `internal/api/router.go`
### What was just completed (2026-02-18 session 42)
- **v0.12.6 — Cross-Drive Backup Rsync Fixes:**
**Context:** After fixing mount-point validation and system-drive thresholds (v0.12.5), testing revealed two more rsync issues for Immich.
**Fix 3: Simplified rsync destination path structure (`internal/backup/crossdrive.go` `runRsyncBackup`)**
- Old logic stripped only the first 2 path segments and kept the rest as a subpath, producing redundant nesting: `backups/rsync/immich/storage/immich/<data>` instead of `backups/rsync/immich/<data>`
- New logic: if app has a single mount, rsync directly into the stack folder (`backups/rsync/immich/`); if multiple mounts, use each mount's leaf directory name as subfolder
- Duplicate leaf names disambiguated by appending `_N` index suffix
- Loop variable changed from `_, srcMount` to `i, srcMount` to support the index-based disambiguation
- Old nested `storage/immich/` folder will remain orphaned after first run (no data loss; `--delete` only affects the target subtree)
**Fix 4: Exclude app-internal DB dump files from rsync (`internal/backup/crossdrive.go` `runRsyncBackup`)**
- Apps like Immich store their own periodic DB dumps in `<data>/backups/*.sql.gz` (~16 MB/day)
- The controller already handles DB backups via `pg_dump` separately — copying these again via rsync is redundant and wastes space
- Added `--exclude backups/*.sql.gz`, `--exclude backups/*.sql`, `--exclude backups/*.dump` to rsync command
- The `backups/` directory itself and non-dump files within it are preserved
**Files modified (1):** `internal/backup/crossdrive.go`
### What was just completed (2026-02-18 session 41)
- **v0.12.5 — Cross-Drive Backup Validation Fix:**
**Root cause:** Immich cross-drive backup failed with `destination /mnt/hdd_placeholder is not a mount point` because `ValidateDestination()` hard-blocked non-mount-point destinations. The `/mnt/hdd_placeholder` folder is on the internal SSD (not a separate mount), so the device-ID check returned false.
**Fix 1: Drive-type-aware space checks in `ValidateDestination` (`internal/backup/crossdrive.go`)**
- `onSystemDrive` flag replaces the previous boolean-only mount-point check
- System-drive destinations: require **≥10 GB free** and **<90% usage** to protect OS stability
- External-drive destinations: require **≥100 MB free** (original threshold)
- Updated function comment to reflect the new tiered logic
**Fix 2: Aligned `CheckBackupDestination` UI thresholds for system drives (`internal/system/mounts_linux.go`)**
- Tier 4 disk checks now branch on `h.SystemDrive` flag (set in Tier 3)
- System drive: block at <10 GB free OR ≥90% used (matches runner enforcement); Hungarian warning messages
- External drive: warn at ≥90% used, block at ≥95% used (unchanged)
- Removed the `&& h.Severity == "ok"` guard that prevented system-drive warnings from being overridden properly
**Files modified (2):** `internal/backup/crossdrive.go`, `internal/system/mounts_linux.go`
### What was just completed (2026-02-18 session 40)
- **v0.12.4 — Correctness & Robustness Bug Fixes (TASK.md — 15 bugs fixed):**
**CRITICAL fixes (data loss, panics):**
- **C1: `SetAppBackupBulk` data loss + nil map panic** — Fixed: now updates map IN PLACE instead of replacing it, so stacks absent from the input are preserved. Added nil guard for `s.AppBackup`. (`internal/settings/settings.go`)
- **C2: `UpdateStackConfig` nil Env map panic** — Added nil check `if appCfg.Env == nil { appCfg.Env = make(...) }` before the field assignment loop. (`internal/stacks/deploy.go`)
- **C3: `ValidateDump` missing scanner.Err() check** — Added `if err := scanner.Err()` check after the scan loop so I/O errors don't silently mark a partial dump as valid. (`internal/backup/dbdump.go`)
**HIGH fixes (logic errors, resource leaks):**
- **H1: `nextDailyRun` DST bug** — Replaced `next.Add(24 * time.Hour)` with `time.Date(day+1, ...)` for correct scheduling across Europe/Budapest DST transitions. (`internal/scheduler/scheduler.go`)
- **H2: `nextDailyRun` repeated `LoadLocation`** — Cached timezone in package-level `sync.Once` variable; `getBudapestLocation()` now loaded only once. (`internal/scheduler/scheduler.go`)
- **H3: `settings.save()` .tmp file leak** — Added `os.Remove(tmpPath)` cleanup on `WriteFile` failure path. (`internal/settings/settings.go`)
- **H4: `SetNotificationPrefs` nil pointer panic** — Added nil guard at start of function, returns error instead of panicking. (`internal/settings/settings.go`)
- **H5: `appDirSize` ignores `Sscanf` return value** — Now checks `n != 1` and returns `(0, "?")` on parse failure. Same fix applied to `getDirSizeBytes` in `stacks/delete.go`. (`internal/backup/appdata.go`, `internal/stacks/delete.go`)
- **H6: `getDirSizeBytes` no timeout** — Added `exec.CommandContext` with 30s timeout. Added `"context"` import. (`internal/stacks/delete.go`)
- **H7: `dbdump.go` tmpFile not using `defer Close`** — Replaced explicit `tmpFile.Close()` call with `defer tmpFile.Close()` so the file handle is released even on panic. (`internal/backup/dbdump.go`)
- **H8: `UpdateCrossDriveStatus` misleading comment** — Updated comment to accurately describe the "does nothing if nil" behavior instead of claiming it "creates one if nil". (`internal/settings/settings.go`)
**MEDIUM fixes (code quality, edge cases):**
- **M1: Custom `contains`/`containsBytes` replaced** — Removed bespoke `containsBytes` and simplified `contains` to delegate to `strings.Contains`. Added `"strings"` import. (`internal/notify/notifier.go`)
- **M2: `scheduler.Every()` doesn't validate interval** — Added early return with error log if `interval <= 0` to prevent panic in `time.NewTicker`. (`internal/scheduler/scheduler.go`)
- **M3: `executeJob` panic recovery missing `LastRun`** — Panic recovery defer now also sets `job.LastRun = time.Now()` so the job status shows a timestamp after a panic. (`internal/scheduler/scheduler.go`)
- **M4: `logPostStartStatus` goroutine captures env by reference** — Copies the env slice before launching the goroutine (`envCopy`). (`internal/stacks/manager.go`)
- **M5: Multiple `time.LoadLocation` calls in web package** — Added package-level `getTimezone()` with `sync.Once` in `funcmap.go`. Replaced all `time.LoadLocation("Europe/Budapest")` calls in the web package with `getTimezone()`. (`internal/web/funcmap.go`, `internal/web/handlers.go`)
**Files modified (8):** `internal/settings/settings.go`, `internal/stacks/deploy.go`, `internal/backup/dbdump.go`, `internal/scheduler/scheduler.go`, `internal/backup/appdata.go`, `internal/stacks/delete.go`, `internal/stacks/manager.go`, `internal/notify/notifier.go`, `internal/web/funcmap.go`, `internal/web/handlers.go`
### What was just completed (2026-02-17 session 39)
- **v0.12.3 — Security & Correctness Bug Fixes (TASK.md — 33 bugs fixed):**
**CRITICAL fixes (data races, security vulnerabilities):**
- **C1: Data race in RefreshCache** — Moved `m.lastDBDump.Results` mutation inside `m.mu.Lock()`. Was previously mutating shared state without the lock, causing potential torn writes visible to `GetFullStatus()` goroutines. (`internal/backup/backup.go`)
- **C2: SnapshotHistory reversed after unlock** — Moved snapshot reversal loop before `m.cachedStatus = status` (inside the lock). Previously reversed after `Unlock()`, so `m.cachedStatus.SnapshotHistory` was reversed without protection. (`internal/backup/backup.go`)
- **C3: SetStackProvider write without lock** — `m.stackProvider = provider` now wrapped in `m.mu.Lock()`. Read by `resolveAppBackupPaths()` concurrently. (`internal/backup/backup.go`)
- **C4: GetFullStatus shallow-copies mutable pointers** — `LastDBDump` and `LastBackup` are now deep-copied (struct + Results slice) so callers cannot mutate shared manager state. (`internal/backup/backup.go`)
- **C5: IsSystemDisk 8-bit major mask** — Replaced `>> 8 & 0xff` with `unix.Major()`/`unix.Minor()` (12-bit extraction). Also compares disk-portion of minor (groups of 16) to correctly distinguish physical disks of the same type. Adds `golang.org/x/sys/unix` import. (`internal/storage/safety_linux.go`)
- **C6: No /dev/ prefix validation on DevicePath** — `FormatAndMount` now validates `DevicePath` starts with `/dev/` and does not contain `..` before any disk operations. (`internal/storage/format_linux.go`)
- **C7: Path traversal in extractName** — `extractName()` now rejects empty string, `.`, `..`, and names containing `/` or `\`. (`internal/api/router.go`)
- **C8: Path traversal in TargetPath** — Migration API validates `TargetPath` against registered storage paths from settings before starting migration job. (`internal/web/storage_handlers.go`)
- **C9: Path traversal in DestinationPath** — Cross-drive backup config API validates `DestinationPath` against registered storage paths when `enabled=true`. (`internal/api/router.go`)
- **C10: Path traversal in ParseComposeHDDMounts** — `filepath.Clean()` applied before prefix check; uses separator-aware check `cleanHDD + string(filepath.Separator)` to prevent `${HDD_PATH}/../../etc/passwd` escaping. (`internal/stacks/delete.go`)
**HIGH fixes (logic errors, resource leaks):**
- **H1: ValidateDump reads entire file into memory** — Replaced `os.ReadFile` with `bufio.Scanner` reading line-by-line. 256KB per-line buffer prevents OOM on large (500MB+) SQL dumps during 5-min cache refresh. (`internal/backup/dbdump.go`)
- **H2/H3: Double du invocation per mount + no timeout** — Replaced `appDirSizeHuman()`+`appDirSizeBytes()` with single `appDirSize()` function using `exec.CommandContext` with 30s timeout. Halves subprocess calls per mount point. (`internal/backup/appdata.go`)
- **H4: Snapshot validation only checks first 100** — Replaced `ListSnapshots(100)` existence check with regex validation (`^[0-9a-f]{8,64}$`). Allows restoring any snapshot; `restic restore` returns a clear error for non-existent IDs. (`internal/backup/restore.go`)
- **H5: No pruning for cross-drive restic repos** — Added `pruneResticRepo()` called after each successful cross-drive restic backup (`forget --keep-daily 7 --keep-weekly 4 --prune`). Non-fatal — logs warning on failure. (`internal/backup/crossdrive.go`)
- **H6: Temp password file management** — Reorganized temp file lifecycle: close before deferred remove, remove-on-write-error cleanup. (`internal/backup/crossdrive.go`)
- **H7: dirSizeBytes swallows walk errors** — `filepath.Walk` callback now returns errors instead of `nil`, propagating permission/IO issues. (`internal/backup/crossdrive.go`)
- **H8: Non-atomic fstab write** — `AppendFstabEntry` now reads existing fstab, writes to `.tmp`, then atomically renames. Crash-safe. (`internal/storage/safety_linux.go`)
- **H9: IsDeviceMounted naive prefix matching** — After prefix check, next character must be digit (`0-9`) or `p` (partition marker). Prevents `/dev/sdb` matching `/dev/sdba`. (`internal/storage/safety_linux.go`)
- **H10: eMMC device mapping bug** — `partitionToParentDisk` now handles `mmcblk0p1 → mmcblk0` and `nvme0n1p1 → nvme0n1` patterns. Uses `LastIndex("p")` with digit-suffix check before falling back to `TrimRight("0-9")`. (`internal/storage/scan_linux.go`)
- **H11: Data race on bytesCopied in rsync error path** — Error return path in `runRsync` now reads `bytesCopied` under mutex lock. (`internal/storage/migrate.go`)
- **H13: Path prefix match without separator** — Migration source path check now uses `srcPath == req.CurrentHDDPath || strings.HasPrefix(srcPath, req.CurrentHDDPath+"/")`. Prevents `/mnt/hdd` matching `/mnt/hdd_backup/data`. (`internal/storage/migrate.go`)
- **H14: DeleteStack continues after failed compose down** — `docker compose down` failure now returns an error immediately, preventing deletion of files while containers are still running. (`internal/stacks/delete.go`)
- **H16: exec.Command("docker") without timeout** — `syncFileBrowserMounts()` now uses `exec.CommandContext` with 60s timeout. (`internal/web/handlers.go`)
- **H17: SetNotificationPrefs stores caller's pointer** — Deep-copies `NotificationPrefs` struct and `EnabledEvents` slice before storing. (`internal/settings/settings.go`)
- **H18: wipefs error silently discarded** — wipefs failure logged as warning via progress channel; continues (wipefs may not be installed). (`internal/storage/format_linux.go`)
- **H19: Orphaned fstab entry on mount failure** — New `RemoveFstabEntry()` function atomically removes UUID entry. Called as rollback on `mount` failure and `findmnt` verify failure. (`internal/storage/safety_linux.go`, `format_linux.go`)
**MEDIUM fixes (edge cases, code quality):**
- **M1: formatBytes duplicate in dbdump.go** — Removed `formatBytes()` from `dbdump.go`; all callers (backup.go, restic.go, dbdump.go) now use `humanizeBytes()` from appdata.go. (`internal/backup/dbdump.go`, `backup.go`, `restic.go`)
- **M2: Dead code .tmp suffix check** — Reordered filter in `ListDumpFiles`: `.tmp` check now comes before `.sql` check to correctly skip `.sql.tmp` temp files (was unreachable before). (`internal/backup/dbdump.go`)
- **M3: sizeBytes() returns 0 for string types** — Added `case string:` to `sizeBytes()` using `strconv.ParseUint`. (`internal/storage/scan_linux.go`)
- **M6: Dead elapsed variable** — Removed `_ = elapsed`; elapsed time now shown inline in the "done" progress message. (`internal/storage/migrate.go`)
- **M7: time.LoadLocation error silently discarded** — Two locations in handlers.go now handle `LoadLocation` error, falling back to `time.UTC`. (`internal/web/handlers.go`)
- **M10: filterSnapshotsByPaths imprecise prefix** — Added `pathCovers()` helper using separator-aware prefix check. Prevents `/mnt/hdd_1` matching `/mnt/hdd_10/data`. (`internal/api/router.go`)
- **M11: XSS in editStorageLabel innerHTML** — `cancelEditLabel()` in settings.html now uses DOM manipulation (`document.createElement`, `.textContent`) instead of `innerHTML` for the label text. (`internal/web/templates/settings.html`)
**Files modified (15):** `internal/backup/backup.go`, `internal/backup/appdata.go`, `internal/backup/dbdump.go`, `internal/backup/restore.go`, `internal/backup/crossdrive.go`, `internal/backup/restic.go`, `internal/storage/safety_linux.go`, `internal/storage/format_linux.go`, `internal/storage/scan_linux.go`, `internal/storage/migrate.go`, `internal/stacks/delete.go`, `internal/api/router.go`, `internal/web/handlers.go`, `internal/web/storage_handlers.go`, `internal/settings/settings.go`, `internal/web/templates/settings.html`
### What was just completed (2026-02-17 session 38)
- **v0.12.2 — Restore Section Simplification (Bug 4 from v0.12.1 TASK.md):**
- **Feature: Snapshot filtering by app** — `GET /api/backup/snapshots?stack={name}` now filters snapshots to those whose `Paths` overlap with the app's HDD mount paths. Uses prefix matching (snapshot path is prefix of required, or vice versa). New `filterSnapshotsByPaths()` helper in `internal/api/router.go`. Manager gains `GetStackHDDMounts()` method to expose stackProvider's mount resolution.
- **Feature: Auto-stop/restart on restore** — `RestoreApp()` now stops the app's containers before running `restic restore` and restarts them after (even on failure). Avoids data corruption from live writes during restore. Eliminates the "Javasoljuk az alkalmazás leállítását" advisory from the UI.
- **Interface extension: StackDataProvider** — Added `StopStack(name string) error` and `StartStack(name string) error` to the `backup.StackDataProvider` interface in `internal/backup/appdata.go`. `stackAdapter` in `cmd/controller/main.go` wires these through to `stacks.Manager`.
- **UI simplification: Restore section** — Removed confusing "Visszaállítandó útvonalak" path list (technical detail not needed by customer). Snapshot dropdown now populated per-app (filtered) with human-friendly format: `2026-02-17 hétfő 03:00 (a3f2b1)`. Single calm warning replacing the triple-exclamation block. Empty filtered result shows inline message instead of empty dropdown. `data-paths` attribute removed from app dropdown options.
- **Files modified (6):** `internal/backup/appdata.go`, `internal/backup/backup.go`, `internal/backup/restore.go`, `internal/api/router.go`, `internal/web/templates/backups.html`, `cmd/controller/main.go`
### What was just completed (2026-02-17 session 37)
- **v0.12.0 — Backup Page Overhaul — Unified App Backup Status & Bug Fixes:**
- **Bug Fix 1: Duplicate unconfigured apps** — `GetFullStatus()` now returns a deep copy of the cached status. `CrossDriveSummary`, `UnconfiguredApps`, and `CrossDriveWarnings` slices are always nil in the returned copy so the handler builds them fresh on every page load. Previously the handler appended to the cached slices, causing 3× duplication on 3 page loads.
- **Bug Fix 2: Misleading "drive disconnected" error** — Replaced the binary `IsMountPoint || !IsWritable` check with tiered `CheckBackupDestination()` validation (new in `internal/system/mounts_linux.go` and stub in `mounts_other.go`). Tiers: path doesn't exist (critical/blocked), not writable (critical/blocked), same block device as `/` (warning/allowed with note about system drive), disk >95% full (critical/blocked), disk >90% (warning/allowed). `isSameBlockDevice()` replaces `IsMountPoint()` for source/dest same-device detection. Used in both `deployHandler()` and `backupsHandler()` for display, and in `crossdrive.go` logic via `CheckBackupDestination()`.
- **Bug Fix 3: Dead BackupEnabled toggle** — Removed `settingsAppBackupHandler()` from handlers.go and its `POST /settings/app-backup` route from server.go. The toggle wrote to settings.json but nothing read it to skip apps. UI nightly backup section in deploy.html now shows an informational note instead of the toggle.
- **Architecture: Unified per-app backup rows** — New `AppBackupRow` struct and `buildAppBackupRows()` in handlers.go. Replaces old "Alkalmazás adatok" + "Másolatok másik meghajtóra" sections with a single expandable row per app showing all 3 backup layers (DB, Docker volumes, user data). Status dot: green=fully covered, yellow=warning (failed run, system drive, disk full), red=HDD data without cross-drive configured, auto=no user data. Expandable JS toggle with ▶/▼ icon.
- **Architecture: Sequential backup chaining** — Removed independent `cross-drive-daily` (03:30) and `cross-drive-weekly` (04:30) scheduler jobs. Cross-drive backups now run immediately after the restic backup completes (daily jobs every night; weekly jobs on Sunday). This ensures DB dump → restic → cross-drive happen in the same window for file/DB consistency on restore.
- **Architecture: Deploy page schedule dropdown** — Removed "Csak kézi indítás" option (schedule="manual"). Two options remain: "Naponta (az éjszakai mentés után)" and "Hetente, vasárnap (az éjszakai mentés után)". Weekly option shows informational note about DB consistency implications. Existing "manual" configs treated as "weekly" in the dropdown.
- **CSS added:** `.app-backup-row`, `.app-backup-row-header`, `.app-backup-row-name`, `.app-backup-row-meta`, `.app-backup-row-detail`, `.status-dot` (green/yellow/red/auto), `.backup-layers`, `.backup-layer-row`, `.layer-label`, `.layer-badge`, `.layer-na`, `.layer-method`, `.layer-dest`, `.layer-schedule`, `.layer-last`, `.layer-unconfigured`, `.layer-actions`, `.layer-warnings`, `.backup-layer-warning`, `.btn-xs`, `.text-ok`, `.text-error`.
- **Files modified (9):** `internal/backup/backup.go`, `internal/system/mounts_linux.go`, `internal/system/mounts_other.go`, `internal/web/handlers.go`, `internal/web/server.go`, `internal/web/templates/backups.html`, `internal/web/templates/deploy.html`, `internal/web/templates/style.css`, `cmd/controller/main.go`
### What was just completed (2026-02-17 session 36)
- **v0.11.9 — UI Polish Fixes for deploy/settings backup section:**
- **Fix 1: Spacing** — `.deploy-cross-drive` `margin-bottom` increased from `1rem` to `1.5rem` for consistent spacing before deploy form.
- **Fix 2: Tooltip on "Módszer"** — Renamed "Verziózott mentés (restic)" to "Titkosított mentés (restic)". Added info `(i)` tooltip explaining rsync vs restic tradeoffs.
- **Fix 3: Nightly backup indicator** — Replaced disabled checkbox (with confusing pointer cursor) with a non-interactive green/gray dot indicator.
- **Fix 4: Progressive disclosure** — Dest/method/schedule selects are disabled until "Engedélyezve" is checked. JS `toggleCrossDriveFields()` enables/disables them. Backend handler updated to preserve existing config when disabling (disabled fields not submitted).
- **Fix 5: Emoji cleanup** — Removed all emoji from `deploy.html` backup section (h4, warning, status, hint, stale data) and `backups.html` cross-drive summary (status badges, schedule badge, unconfigured warning). JS callbacks also cleaned up.
- **CSS added:** `.info-tooltip`, `.info-icon`, `.info-tooltip-text`, `.cross-drive-nightly-status`, `.nightly-status-indicator`, `.nightly-enabled`, `.nightly-disabled`, `.meta-badge-fail`.
- **Files modified (4):** `web/templates/deploy.html`, `web/templates/backups.html`, `web/templates/style.css`, `web/handlers.go`
### What was just completed (2026-02-17 session 35)
- **v0.11.8 — Per-App Cross-Drive Backup (3-2-1 rule, second copy on different media):**
- **Feature: CrossDriveBackup data model** — `AppBackupPrefs` extended with `CrossDrive *CrossDriveBackup` field in `settings.go`. New methods: `GetCrossDriveConfig`, `SetCrossDriveConfig`, `UpdateCrossDriveStatus`, `GetAllCrossDriveConfigs`, `GetOrCreateCrossDrivePassword`. Existing `SetAppBackup`/`SetAppBackupBulk` now preserve cross-drive config. Auto-generated restic password stored in `settings.json`.
- **Feature: CrossDriveRunner** — New `internal/backup/crossdrive.go`. Supports rsync (simple mirror with `--delete`) and restic (versioned, deduplicated, shared repo). Safety guards: destination ≠ source, mount point check, writable check, per-app concurrency lock. `RunAllScheduled(ctx, schedule)` iterates all apps matching the given schedule. Status (last_run, last_status, last_error, last_duration, last_size_human) persisted to settings.json after each run.
- **Feature: Scheduler jobs** — Two new daily jobs: `cross-drive-daily` at 03:30 (for apps with `schedule: daily`), `cross-drive-weekly` at 04:30 Sundays only (for `schedule: weekly`).
- **Feature: API endpoints** — 4 new routes: `POST /api/stacks/{name}/cross-backup`, `POST /api/stacks/{name}/cross-backup/run`, `GET /api/stacks/{name}/cross-backup/status`, `POST /api/backup/cross-drive/run-all`.
- **Feature: Deploy/Settings page UI** — New "Biztonsági mentés" card on the deploy page for apps with HDD data. Shows nightly backup toggle (read-only link), cross-drive dropdowns (destination, method, schedule), last run status, manual trigger button. States: no other storage (info message), configured, destination unreachable (warning). Flash messages on save redirect.
- **Feature: Backup page summary** — New "Másolatok másik meghajtóra" section showing all configured apps with method, destination, last status, size. Warns about unconfigured apps with HDD data. Destination health warnings. "Összes futtatása most" button.
- **CSS:** `margin-bottom: 1.5rem` added to `.deploy-stale-data`. New styles: `.deploy-cross-drive`, `.cross-drive-list`, `.cross-drive-item`, `.cross-drive-header`, `.cross-drive-meta`, `.cross-drive-actions`.
- **Files modified (10):** `settings/settings.go`, `backup/crossdrive.go` (new), `backup/backup.go`, `api/router.go`, `web/handlers.go`, `web/server.go`, `web/templates/deploy.html`, `web/templates/backups.html`, `web/templates/style.css`, `cmd/controller/main.go`
### What was just completed (2026-02-17 session 34)
- **v0.11.7 — Stale Data Cleanup + FileBrowser Sync + UI Title Fix:**
- **Feature: Stale data cleanup** — After app data migration, the deploy/settings page now shows leftover data on previous storage paths with size info and a delete button. Two-step confirmation required before deletion. Protected paths (storage root, media, Dokumentumok, appdata) cannot be deleted. Also available immediately after migration on the migration-done page.
- **Fix: FileBrowser sync after migration** — `syncFileBrowserMounts()` now called after successful data migration, ensuring FileBrowser mounts reflect the current storage layout.
- **Fix: Deploy page title** — Already-deployed apps now show "Beállítások" (Settings) instead of "Telepítés" (Deploy) in both the browser page title and the `<h2>` heading.
- **Internal: Exported `ProtectedHDDPaths()`** from stacks package for reuse in web handlers.
- **Files modified (7):** `internal/stacks/delete.go`, `internal/web/handlers.go`, `internal/web/storage_handlers.go`, `internal/web/templates/deploy.html`, `internal/web/templates/migrate.html`, `internal/web/templates/style.css`
### What was just completed (2026-02-17 session 33)
- **v0.11.6 — FileBrowser Auto-Mount Sync + UI Polish (3 fixes):**
- **Feature: FileBrowser auto-mount sync** — Added `syncFileBrowserMounts()` and `generateFileBrowserCompose()` to `handlers.go`. After a storage path is added (via storage init wizard) or removed, the controller regenerates `/opt/docker/stacks/filebrowser/docker-compose.yml` with volume mounts for all registered paths (`/mnt/hdd_1:/srv/hdd_1` etc.), then recreates the FileBrowser container. Domain is read from FileBrowser's `.env`. If FileBrowser isn't deployed, the function silently returns. The generated compose is self-contained (no env vars).
- **UI Fix 1: Badge color fix** — `settings.html`: changed "Nincs csatolva!" (red `state-red`) badge to "Rendszermeghajtón" (yellow `badge-warn`). The path is on the system SSD, which isn't an error — just informational. Added `.badge-warn { background: rgba(250, 204, 21, 0.15); color: #facc15; }` to `style.css`.
- **UI Fix 2: Progress bar fix** — `storage_init.html`: replaced the disk-usage gradient progress bar (green→yellow→red zones, alarming at 30%) with a clean single-color `progress-bar-task` bar. Added `.progress-bar-task` and `.progress-bar-task .progress-fill` CSS classes to `style.css`.
- **UI Fix 3: Button text fix** — `settings.html`: "Alapértelmezett" button (reads as status, confusing) → "Legyen alapértelmezett" (clear action verb).
- **Files modified (5):** `web/handlers.go`, `web/storage_handlers.go`, `web/templates/settings.html`, `web/templates/storage_init.html`, `web/templates/style.css`
### What was just completed (2026-02-17 session 32)
- **v0.11.4 — Bugfix: Storage Initialization (FormatAndMount) — 3 bugs + 4 safety improvements:**
- **Bug 1 (sfdisk):** Added `wipefs -a` before sfdisk; changed sfdisk input from `,,,L` (unsupported GPT type shorthand) to `,,` (default Linux GUID); added `--force --wipe always` flags. Previous table confusing sfdisk and `L` type not accepted for GPT.
- **Bug 2 (mount):** Replaced `mount mountPath` (fstab lookup — uses container's /etc/fstab, not host's) with explicit `mount -t ext4 -o defaults,noatime /host-dev/sdb1 /mnt/hdd_1`. fstab entry still written to `/host-fstab` for host reboot persistence.
- **Bug 3 (mount propagation):** Changed `/mnt` volume in compose to long-form bind with `propagation: rshared`. Also ran `mount --bind /mnt /mnt && mount --make-rshared /mnt` on demo host. Confirmed `Propagation=rshared` in `docker inspect`. Mounts created inside container now propagate to host.
- **Safety 1 (post-mount verification):** Added `findmnt` check after mount — fails with clear error if mount isn't actually visible.
- **Safety 2 (ASCII label):** Use `req.MountName` (always ASCII) for ext4 `-L` label (16-byte limit). Display label (`req.Label`, may contain UTF-8 Hungarian chars) stays only in settings.json.
- **Safety 3 (smart partition):** In `storageInitAPIHandler`, if disk has exactly 1 empty partition (no filesystem), skip wipefs+sfdisk entirely and format existing partition directly. Handles demo sdb case (sdb1 exists, no FS).
- **Safety 4 (progress messages):** Updated `send()` calls to include command details (device paths, flags) for remote debugging via UI progress panel.
- **Files modified (3):** `storage/format_linux.go`, `docker-compose.yml`, `web/storage_handlers.go`
### What was just completed (2026-02-17 session 31)
- **v0.11.3 — Bugfix: Missing sfdisk in container (fdisk package):**
- `sfdisk` is in the `fdisk` package on Debian bookworm, not `util-linux`. Dockerfile had `util-linux` but not `fdisk`, so `sfdisk` was missing and partitioning failed.
- Added `fdisk` to Dockerfile's `apt-get install` list. Updated comment to clarify which package provides what.
- Verified: all six disk tools now present in container (`sfdisk`, `mkfs.ext4`, `blkid`, `mount`, `lsblk`, `partprobe`).
- **Files modified (1):** `Dockerfile`
### What was just completed (2026-02-17 session 30)
- **v0.11.2 — Bugfix: /dev/sdb not accessible inside container:**
- **Root cause:** Docker always creates a fresh tmpfs at `/dev` inside containers. Even with `privileged: true`, the bind mount `- /dev:/dev` is silently dropped. Block device nodes like `/dev/sdb` don't exist inside the container.
- **Fix:** Mount host `/dev` at `/host-dev` instead. With `privileged: true`, the kernel allows I/O to the device nodes regardless of path inside the container.
- **docker-compose.yml:** Changed `- /dev:/dev` → `- /dev:/host-dev:rw`. Also applied missing `privileged: true`, `/etc/fstab:/host-fstab`, and `/run/udev:/run/udev:ro` to demo node's live compose (never applied after v0.11.0).
- **safety.go:** Added `HostDevPath = "/host-dev"` constant and `HostDevicePath(devPath) string` helper (`/dev/sdb` → `/host-dev/sdb`).
- **format_linux.go:** All device operations (os.Stat, sfdisk, partprobe, mkfs.ext4, blkid UUID) use `HostDevicePath()`.
- **safety_linux.go:** `IsSystemDisk()` stats device via `HostDevicePath()`.
- **scan_linux.go:** `enrichWithBlkid()` probes each partition individually (`blkid -o value -s TYPE/UUID/LABEL /host-dev/sdXN`) instead of batch `blkid -o export` (which fails when `/dev` is Docker's minimal tmpfs).
- **Verified:** `/host-dev/sda`, `/host-dev/sdb`, partitions visible; `blkid /host-dev/sdb1` returns correct UUID/fstype/label.
- **Files modified (5):** `storage/safety.go`, `storage/safety_linux.go`, `storage/format_linux.go`, `storage/scan_linux.go`, `docker-compose.yml`
### What was just completed (2026-02-17 session 29)
- **v0.11.1 — Bugfix: Storage Scan — System Disk Detection & FSType in Container:**
- **Bug 1 fix: System disk detection** — Replaced mount-point string comparison (`== "/"`, `"/boot"`, `"/boot/efi"`) with host fstab parsing. Inside the container, `lsblk` reports container mount points (e.g. `/opt/docker/felhom-controller/data`), not host mount points. New `getSystemDiskNames()` reads `/host-fstab` (fallback: `/etc/fstab`), finds system entries (`/`, `/boot`, `/boot/efi`, `swap`), resolves `UUID=` entries to device paths via `blkid -U`, and marks parent disks as system. `partitionToParentDisk()` handles both standard (`sda2→sda`) and NVMe (`nvme0n1p2→nvme0n1`) naming.
- **Bug 2 fix: FSType enrichment** — `lsblk` returns null fstype in containers (udev/blkid cache incomplete). New `enrichWithBlkid()` runs `blkid -o export` after lsblk scan and fills in missing `FSType`, `UUID`, `Label` per partition from direct device probing. Runs on both `AvailableDisks` and `SystemDisks`.
- **Result:** sda (system SSD) now correctly appears in SystemDisks; sdb (USB HDD) appears in AvailableDisks; partition fstypes (vfat/ext4/swap) correctly shown; sdb1 genuinely shows "(nincs fájlrendszer)".
- **Files modified (1):** `storage/scan_linux.go`
### What was just completed (2026-02-17 session 28)
- **v0.11.0 — Phase C: Storage Init, Data Migration & Startup Fixes:**
- **Step 0: Startup ping + hub report** — Controller now fires heartbeat ping, system_health ping, and hub report immediately on startup (5s delay) instead of waiting for first scheduler tick (5-15 min). `hubPusher` instance created once and reused for both startup and periodic reports. Prevents Healthchecks showing stale "Last Ping: X ago" after restarts.
- **Step 1-3: Storage initialization wizard** — New `internal/storage/` package (`scan.go`, `format.go`, `safety.go`, `format_linux.go`, `safety_linux.go`, `scan_linux.go` + non-linux stubs). `ScanDisks()` via `lsblk -J`. `FormatAndMount()` with progress channel (partition via sfdisk → mkfs.ext4 → blkid UUID → fstab backup + UUID-based entry → mount → chown + subdirs). Safety guards: system disk detection via major device numbers, mount path conflict, confirmation "FORMÁZÁS" required. New wizard page at `/settings/storage/init`. JSON API endpoints at `/api/storage/scan`, `/api/storage/init`, `/api/storage/init/status`. Auto-registers storage path in settings.json after success.
- **Step 4-5: Data migration** — New `MigrateAppData()` in `internal/storage/migrate.go`. Per-app "Mozgatás" button on deploy page (for deployed apps with HDD data) and settings page storage app list. Migration flow: stop app → rsync with `--info=progress2` progress parsing → update `app.yaml` HDD_PATH → start app. Rollback on failure (revert config + restart with original path). Old data preserved. New migration page at `/stacks/{name}/migrate`. JSON API at `/api/storage/migrate`, `/api/storage/migrate/status`.
- **Step 6: Per-app storage display** — Deploy page (read-only mode) now shows "Adattárolás" section for deployed apps: current path + label, data size, free space. "Mozgatás" link shown when other storage paths exist.
- **Step 7: Container setup** — Added `privileged: true` to `docker-compose.yml`. New volume mounts: `/dev:/dev`, `/etc/fstab:/host-fstab`, `/run/udev:/run/udev:ro`. Docker socket changed from `:ro` to writable. `Dockerfile` adds: `util-linux`, `e2fsprogs`, `rsync`, `parted`.
- **Storage API routing** — New `/api/storage/` prefix registered in `main.go` before `/api/` catch-all (longer prefix takes priority in Go ServeMux). `ServeStorageAPI` method on web.Server handles all storage JSON endpoints.
- **CSS additions** — `.disk-step`, `.disk-step-active`, `.disk-step-done`, `.disk-progress-steps`, `.disk-progress-bar-wrap`, `.deploy-storage-info` styles.
- **Files created (13):** `storage/scan.go`, `storage/scan_linux.go`, `storage/scan_other.go`, `storage/safety.go`, `storage/safety_linux.go`, `storage/safety_other.go`, `storage/format.go`, `storage/format_linux.go`, `storage/format_other.go`, `storage/migrate.go`, `web/storage_handlers.go`, `templates/storage_init.html`, `templates/migrate.html`
- **Files modified (8):** `main.go`, `web/server.go`, `web/handlers.go`, `templates/settings.html`, `templates/deploy.html`, `templates/style.css`, `docker-compose.yml`, `Dockerfile`
### What was just completed (2026-02-17 session 27)
- **v0.10.0 — Phase B: Storage Management UI Polish & Health Severity Fix:**
- **Step 0: Health severity fix** — `checkStoragePaths()` mount-point check reclassified from **issue** (FAIL) to **warning** (WARN). All storage health messages translated to Hungarian. Added `.monitoring-banner-warn` CSS class for yellow warning banners. Prevents false FAIL status on demo/test environments where storage is intentionally on SSD.
- **Step 1: Success flash messages** — All 4 storage handlers (add/remove/set-default/toggle-schedulable) now redirect with `?storage_msg=success&storage_detail=...` query params. Settings page displays green "alert-info" flash on success. Consistent with backup page flash pattern.
- **Step 2: Edit storage path labels** — New `SetStorageLabel()` method in `settings.go`. New `POST /settings/storage/label` route + handler. Inline edit UI with ✏️ button, text input, OK/Cancel. Added `.btn-ghost` CSS class.
- **Step 3: App details per storage path** — Settings page now shows expandable `<details>` list per storage path with app names, sizes, and links to deploy page. New `StorageAppDetail` struct + `appDetailsForPath()` helper. Added CSS for `.storage-app-details`, `.storage-app-list`, `.storage-app-row`.
- **Step 4: Storage badge on stacks page** — Deployed app cards show "💾 Label" badge indicating which registered storage path the app uses. `StorageLabels` map built from deployed apps' HDD_PATH → registered storage path label lookup. Added `.meta-badge-storage` CSS.
- **Step 5: Deploy dropdown enhancements** — Storage path dropdown now shows free space ("234 GB szabad"). `DeployStoragePath` struct wraps `StoragePath` with `FreeHuman`/`FreePercent` from `GetDiskUsage()`. JS `checkStorageSpace()` shows yellow warning when selected storage has <20% free.
- **Step 6: Filesystem & disk info** — New `FSInfo` struct + `GetFSInfo()` in `mounts_linux.go` using `findmnt` command + `/sys/block/` sysfs reads for disk model. Settings page shows "ext4 · /dev/sdb1 · WD Elements" below disk usage bar. Non-Linux stub returns nil.
- **Step 7: Backup page storage context** — Added `StorageLabel` field to `AppBackupInfo`. Backup page shows storage label badge per app by matching HDD path prefixes against registered storage paths. Uses existing `.meta-badge-storage` CSS.
- **Files modified (12):** `healthcheck.go`, `settings.go`, `mounts_linux.go`, `mounts_other.go`, `appdata.go`, `handlers.go`, `server.go`, `settings.html`, `stacks.html`, `deploy.html`, `backups.html`, `style.css`
### What was previously completed (2026-02-17 session 26)
- **v0.9.0 — Phase A: Storage Paths Foundation & Backup Toggle Fix:**
- **Root cause:** Per-app backup toggles (v0.8.0) didn't appear because `controller.yaml` had no `paths.hdd_path` set → `ParseComposeHDDMounts` returned nil. Even with global hdd_path, apps with different HDD_PATH values wouldn't match.
- **Core fix: Per-app HDD_PATH resolution** — `stackAdapter.GetStackHDDMounts()` now reads each app's own `HDD_PATH` from its `app.yaml` env section (Priority 1), falling back to all registered storage paths (Priority 2). Removed dependency on global `cfg.Paths.HDDPath`.
- **Storage paths registry** (`settings.json`) — new `StoragePath` struct with Path, Label, IsDefault, Schedulable, AddedAt. Thread-safe CRUD methods in `settings.go` (Get/Add/Remove/SetDefault/SetSchedulable). Multiple external storage paths supported.
- **Auto-discovery** — On startup, `discoverHDDPaths()` scans deployed apps' `app.yaml` for `HDD_PATH` values. `AutoDiscoverStoragePaths()` registers discovered paths with inferred labels. Legacy `cfg.Paths.HDDPath` used as fallback.
- **Mount-point validation** — New `mounts_linux.go` (build-tagged): `IsMountPoint()` via `syscall.Stat_t.Dev` comparison, `IsWritable()`, `PathsOverlap()`, `GetDiskUsage()` via `syscall.Statfs`. Non-Linux stubs in `mounts_other.go`.
- **Settings page "Adattárolók" section** — Lists registered paths with label, path, disk usage bar, app count, badges (default/active/unmounted). Actions: set default, toggle schedulable, remove (with guards). Expandable "Új adattároló hozzáadása" form with 5-step validation (exists, mount point, writable, no overlap, no duplicate).
- **Deploy page storage dropdown** — `path` field type renders as `<select>` dropdown of schedulable storage paths. Falls back to text input with warning if no paths registered.
- **Health check storage monitoring** — `RunHealthCheck()` now accepts `storagePaths` parameter. Checks: path accessible (warning), not a mount point (issue — data writes to SSD!), disk usage ≥95% (issue) / ≥90% (warning).
- **Controller docker-compose.yml** — Changed HDD mount from `${HDD_PATH:-/mnt/hdd_placeholder}:...:ro` to `/mnt:/mnt:rw` for multi-storage support + restore capability.
- **Removed unused `hddPath` param** from `DiscoverAppData()` signature in backup/appdata.go.
- **Files created (2):** `system/mounts_linux.go`, `system/mounts_other.go`
- **Files modified (11):** `settings.go`, `main.go`, `appdata.go`, `backup.go`, `handlers.go`, `server.go`, `settings.html`, `deploy.html`, `style.css`, `healthcheck.go`, `docker-compose.yml`, `report/builder.go`
### What was previously completed (2026-02-16 session 25)
- **v0.8.0 — Phase 7: Storage Overview, Per-App Backup Toggles & Limited Restore:**
- **Storage overview on backup page** — new "Tárhely áttekintés" section as first section on backup page showing SSD/HDD progress bars + backup repo stats (repo size, dump file count, snapshot count). Reuses existing `system.GetInfo()` and `RepoStats`.
- **Restic password visibility** — new "Titkosítási kulcs" section inside the repository card. Masked password field with show/copy buttons (JS toggle). Password synced to hub via periodic report for disaster recovery (`ResticPassword` field added to `BackupReport`).
- **App data discovery** — new `internal/backup/appdata.go`:
- `StackDataProvider` interface to avoid circular imports between backup and stacks packages
- `AppBackupInfo`, `AppDataPath`, `AppDockerVolume` structs
- `DiscoverAppData()` iterates deployed stacks, discovers HDD bind mounts (via adapter calling `ParseComposeHDDMounts`), Docker named volumes (via `parseComposeNamedVolumes` using YAML parser), and DB dump status
- Stack adapter in `main.go` implements `StackDataProvider` using `stacks.Manager`
- **Per-app backup toggles** — new "Alkalmazás adatok" section on backup page:
- Toggle checkbox per app (only for apps with HDD data)
- Shows HDD paths with sizes, Docker volume info, DB dump notes
- `POST /settings/app-backup` handler saves preferences to `settings.json`
- `AppBackupPrefs` struct + bulk getter/setter in `settings.go`
- `RefreshCache()` populates `AppDataInfo` via `DiscoverAppData()`
- **Dynamic backup paths** — `RunBackup()` now includes enabled app HDD data paths:
- `resolveAppBackupPaths()` reads enabled apps from settings, resolves HDD paths via provider
- Paths logged at INFO level, included in restic snapshot
- `BackupPaths` display on backup page includes app data paths
- **Limited app restore** — new restore section on backup page:
- `RestoreApp()` in `restore.go`: validates enabled, resolves HDD paths, validates snapshot exists, uses running mutex
- `RestoreAppData()` on `ResticManager`: runs `restic restore` with `--include` flags for specific paths
- `POST /backup/restore` web handler with confirmation flow
- `GET /api/backup/snapshots` JSON endpoint for restore dropdown
- UI: app/snapshot dropdowns, warning box, confirmation checkbox, JS-driven form submission
- **Exported `ParseComposeHDDMounts`** from stacks package (was unexported `parseComposeHDDMounts`)
- **Flash messages** on backup page via query params (success/error redirects from handlers)
- **CSS**: New styles for storage overview grid, app backup toggles, encryption key field, restore section, flash messages
- **Files created**: `appdata.go`, `restore.go`
- **Files modified**: `backup.go`, `restic.go`, `handlers.go`, `server.go`, `backups.html`, `style.css`, `settings.go`, `delete.go`, `router.go`, `types.go`, `builder.go`, `main.go`
### What was previously completed (2026-02-16 session 24)
- **v0.7.2 — Fix Notification Preferences Sync (Controller → Hub):**
- **Two repos changed** (deploy-felhom-compose + felhom.eu):
- **Hub: `POST /api/v1/preferences` endpoint** (`hub/internal/api/handler.go`):
- New route in API handler: same Bearer token auth as /report and /notify
- Accepts JSON payload: `{customer_id, email, enabled_events}`
- Calls existing `store.SaveNotificationPrefs()` — no store changes needed
- Logs preference updates at INFO level
- **Hub: Notification section on customer detail page** (`hub/internal/web/`, `hub/internal/store/store.go`):
- New `GetRecentNotifications()` store method returns last N notification_log entries
- `handleCustomerDetail()` loads NotifPrefs + RecentNotifications
- `joinStrings` template function added for event list display
- `customer.html` template: new "Notifications" section showing email, events, and last 10 notification log entries (time, event, status, message)
- **Controller: `SyncPreferences` method** (`internal/notify/notifier.go`):
- New `preferencesRequest` struct for JSON payload
- `SyncPreferences(email, enabledEvents)` — synchronous POST to hub `/api/v1/preferences`
- `IsEnabled()` getter for checking hub connectivity
- Hungarian error messages for user-facing feedback
- **Controller: Sync on settings save** (`internal/web/handlers.go`):
- `settingsNotificationsHandler` now calls `SyncPreferences` after saving to `settings.json`
- Three flash message variants: success (synced), warning (local save OK, sync failed), error (save failed)
- Local save always succeeds even if hub sync fails
- **Controller: Sync on startup** (`cmd/controller/main.go`):
- Non-blocking goroutine syncs preferences to hub when controller starts
- Only runs if hub is enabled and email is configured
- Handles hub DB rebuild recovery (re-populates preferences after hub redeployment)
- **Files changed**: hub (3 files: handler.go, store.go, server.go, customer.html), controller (3 files: notifier.go, handlers.go, main.go)
- **Documentation**: README.md updated (version, notify module, phase checklist), CONTEXT.md updated
### What was previously completed (2026-02-16 session 23)
- **v0.7.1 — Phase 2: Monitoring Warnings, Dashboard Alerts & Notification System:**
- **Three workstreams across two repos** (deploy-felhom-compose + felhom.eu):
- **Monitoring page "Távoli monitoring" section** (`monitoring.html`, `handlers.go`):
- New section between System Overview and System Metrics showing healthcheck ping UUID status
- 5 rows: Heartbeat, System Health, DB Dump, Backup, Backup Integrity — each shows ✅ configured or ⚠️ missing
- Banner: green (all configured), yellow (some missing), red (monitoring disabled)
- `isPingConfigured()` helper checks non-empty AND not "CHANGEME" prefix
- **Dashboard alert banners** (new `alerts.go`, `layout.html`):
- `AlertManager` struct with `Refresh()` + `GetAlerts()` — generates alerts from health report, missing pings, backup disabled
- Alert types: `Alert{ID, Level, Message, Link, LinkText}` — levels: error/warning/info
- Renders colored banners (red/yellow/blue) after `<main class="content">` on all pages
- Caps at 5 alerts with "+N more" overflow; monitoring page excludes "pings-missing" (shown in table instead)
- Refreshed every 5 min via system-health scheduler task + once at startup
- **Hub notification relay** (felhom.eu repo — `hub/internal/api/handler.go`, `hub/internal/store/store.go`):
- `POST /api/v1/notify` endpoint: Bearer auth, JSON payload (customer_id, event_type, severity, message, details)
- New `customer_notifications` table (email, enabled_events JSON) + `notification_log` audit table
- Resend email integration: direct HTTP POST to `https://api.resend.com/emails`
- Hungarian email template with event details, timestamp, severity
- `hub.yaml.example` updated with notifications config section
- **Controller-side notifier** (new `internal/notify/notifier.go`):
- `Notifier` struct: fires HTTP POST to hub `/api/v1/notify`, non-blocking (goroutine)
- Cooldown tracking per event type (default 6h, configurable via UI)
- Checks notification preferences (email configured + event enabled) before sending
- `NotifyHealthChange()`: only notifies on status degradation (ok→warn, ok→fail, warn→fail)
- `NotifyBackupFailed/NotifyDBDumpFailed/NotifyIntegrityFailed` convenience methods
- `SendTest()` for test email flow
- Wired into scheduler: system-health task calls `NotifyHealthChange()`, backup tasks call failure notifiers
- **Notification preferences UI** (`settings.html`, `handlers.go`):
- New "Értesítések" Section C on Settings page (only shown when hub enabled)
- Email input, 4 event checkboxes (disk_warning, backup_failed, update_available, security_update)
- Cooldown hours input (default 6)
- "Mentés" + "Teszt email küldése" buttons
- Saved to `settings.json` via `NotificationPrefs` struct (Email, EnabledEvents, CooldownHours)
- **Settings persistence expanded** (`settings.go`):
- `NotificationPrefs` struct with Email, EnabledEvents, CooldownHours
- `DefaultEnabledEvents`: disk_warning, backup_failed, update_available
- `GetNotificationPrefs()` returns defaults if nil, `SetNotificationPrefs()` saves atomically
- **Files changed**: 3 new (alerts.go, notifier.go, notify package), ~12 modified across both repos
- **Deployed:** Controller v0.7.1 to demo-felhom.eu, verified healthy (0 alerts on clean system)
### What was previously completed (2026-02-16 session 22)
- **v0.7.0 — Phase 1: Authentication, Persistence & Settings Page:**
- **New `internal/settings/settings.go`:** Shared persistence layer via `settings.json` in the data directory. Atomic writes (tmp + rename), thread-safe with `sync.RWMutex`. Stores password hash overrides and DB validation cache. Graceful handling if file doesn't exist.
- **Auth improvements:**
- Password resolution priority: `settings.json` → `controller.yaml` → none (open dashboard)
- Startup logs which source is active: `Auth: using password from settings.json/controller.yaml/no password configured`
- Session duration extended to 7 days (was 24h)
- `?next=` redirect after session expiry — returns user to the page they were on
- Flash messages on login page (green info box, used after password change)
- Conditional logout link — hidden when auth is disabled (no password configured)
- `invalidateAllSessions()` method for password change flow
- **New Settings page (`/settings`):**
- "Rendszer konfiguráció" section: read-only display of controller.yaml values (customer ID/name/domain, git repo/sync interval, backup enabled/schedule, monitoring, healthchecks URL, hub status, controller version)
- "Jelszó módosítás" section: form with current password, new password, confirm — validates min 8 chars, match check, bcrypt comparison
- Password saved to `settings.json`, all sessions invalidated, redirect to login with flash message
- Only shown if auth is enabled; otherwise shows info message to contact operator
- **Sidebar update:**
- "Beállítások" menu item with ⚙ icon pinned to bottom (above version/logout)
- Version and logout link separated from nav links
- Logout link conditionally shown only when auth is enabled
- **DB validation persistence:**
- After each successful dump, validation results saved to `settings.json` (`db_validations` map keyed by filename)
- Cached data survives container restarts
- `DBValidationCache` struct with `validated_at`, `table_count`, `has_header`, `error`
- **10 files changed** (3 new: settings.go, settings.html; 7 modified: main.go, backup.go, auth.go, handlers.go, server.go, layout.html, login.html, style.css)
- **Deployed:** Controller v0.7.0 to demo-felhom.eu, verified healthy
### What was previously completed (2026-02-16 session 21)
- **v0.6.3 — Bug fixes from v0.6.2 code scan (4 minor fixes):**
- **Bug 1:** `--hdd-path` in `docker-setup.sh` now uses `require_arg` validation like all other flags. Previously, `--hdd-path` as the last argument without a value would crash with a cryptic bash error under `set -u` instead of a friendly message.
- **Bug 2:** `stackAction()` in `layout.html` now receives `event` as an explicit parameter instead of relying on the deprecated implicit `window.event`. All 10 onclick call sites in `dashboard.html` and `stacks.html` updated to pass `event` as first argument.
- **Bug 3:** Page `<title>` now has an em dash separator: `"Vezérlőpult — Felhom.eu"` instead of `"VezérlőpultFelhom.eu"`.
- **Bug 4:** `nextPruneLabel()` in `funcmap.go` now returns `"ma"` (Hungarian for "today") on Sunday before 4am, consistent with the `nextRunLabel` function. Previously returned the date in `"2006-01-02"` format.
- **Deployed:** Controller v0.6.3 to demo-felhom.eu, verified healthy
### What was previously completed (2026-02-16 session 20)
- **Hub Dashboard Bugs + Backup Validation Fix (3 bugs):**
- **Bug 1&2 (Hub repo, felhom-hub v0.1.2):** Hub timestamp parsing failure — `time.Parse` with single hardcoded format silently failed for formats returned by `modernc.org/sqlite`. Added `parseSQLiteTime()` that tries 6 common formats. Fixed: hub main page showing DOWN despite OK status, and report history timestamps showing 00:00:00.
- **Bug 3 (Controller repo, v0.6.2):** Backup page showing "Hiba" for all DB validations — zero-value `DumpValidation{}` (never assigned) hit the `{{else}}` branch in template. Three fixes:
- Template: 4-branch guard (Valid → OK / Error → Hiba / zero-value → "–" with tooltip)
- Debug logging: Added `[DEBUG]` and `[WARN]` log lines to all `ValidateDump()` code paths
- Re-validation: `RefreshCache()` now cross-checks `lastDBDump` results against fresh `ListDumpFiles()` validation, healing stale in-memory state
- **Deployed:** Hub v0.1.2 to k3s, Controller v0.6.2 to demo-felhom
- **Verified:** Controller logs show `ValidateDump OK` for all 3 databases (immich: 60 tables, paperless: 67 tables, romm: 14 tables)
### What was previously completed (2026-02-16 session 19)
- **v0.6.1 — Code Review Bugfixes (7 fixes):**
- **Fix 1:** `http.NotFound(w, nil)` → pass actual `*http.Request` in `deployHandler` and `appDetailHandler`
- **Fix 2:** Dashboard running/stopped counts now computed from the filtered `deployedStacks` set (was counting ALL stacks including non-deployed)
- **Fix 3:** Session cookie `Secure` flag now dynamic based on `r.TLS != nil || X-Forwarded-Proto == "https"`. `SameSite` changed from `Strict` to `Lax` (Strict breaks Cloudflare Tunnel redirects)
- **Fix 4:** Removed misleading `subtle.ConstantTimeCompare` from `isValidSession()` (map lookup already leaks timing; comparing token to itself is meaningless). Removed unused `token` field from `session` struct. Removed `crypto/subtle` import.
- **Fix 5:** Replaced `time.Tick()` (goroutine leak) with proper `time.NewTicker` + `done` channel in `cleanupSessions()`. Added `Close()` method to Server. Added `done chan struct{}` to Server struct.
- **Fix 6:** Added `http.MaxBytesReader(w, req.Body, 1<<20)` (1MB limit) to `deployStack`, `updateOptionalConfig`, `deleteStack` API handlers via `limitBody()` helper.
- **Fix 7:** Cached `time.LoadLocation("Europe/Budapest")` once at top of `templateFuncMap()`, removed 5 per-function `LoadLocation` calls (timeAgo, fmtTime, fmtTimeShort, nextRunLabel, nextPruneLabel).
- **Post-fix verification:** All 4 grep checks pass (0 results for NotFound(w,nil), ConstantTimeCompare, time.Tick(, Secure:.*true). `go vet ./...` clean.
- **Controller version:** v0.6.1 — deployed and verified on demo-felhom.eu
### What was previously completed (2026-02-16 session 18)
- **v0.6.0 — Healthcheck Implementation + Central Push + Hub Dashboard:**
- **Part 1 — Healthcheck enhancements (controller-side):**
- Added `heartbeat` ping — lightweight "I'm alive" signal every 5 min (no logic, just ping)
- Added `backup_integrity` ping — weekly `restic check` on Sunday 04:00, pings healthchecks with result
- Added `Heartbeat` and `BackupIntegrity` fields to `PingUUIDsConfig`
- Added `RunIntegrityCheck()` to backup Manager (calls restic Check(), updates lastCheckTime/lastCheckOK, pings)
- Updated `controller.yaml.example` with new monitoring ping_uuids
- Created `monitoring/DEPRECATED.md` for legacy bash monitoring scripts
- **Part 2 — Central hub reporting (controller-side):**
- New `internal/report/` package: types.go (Report struct), builder.go (BuildReport), pusher.go (HTTP push)
- Report builder gathers data from all subsystems: system info (via metrics.GetStaticInfo + system.GetInfo), container stats (via metricsStore.QueryContainerSummary), backup status (via backupMgr.GetFullStatus), health (via monitor.RunHealthCheck), stacks (via stackMgr.GetStacks)
- Report pusher: POST JSON to hub with Bearer token auth, 3 retries with 5s backoff, never fails caller
- Added `HubConfig` to config.go (enabled, url, api_key, push_interval)
- Wired hub reporting into scheduler (configurable interval, default 15m)
- Hub reporting disabled by default (hub.enabled: false)
- **Part 3 — Hub service (felhom.eu repo, new `hub/` subfolder):**
- Full Go service: `cmd/hub/main.go`, `internal/api/handler.go`, `internal/store/store.go`, `internal/web/server.go`
- SQLite store with WAL mode, auto-migration, denormalized fields for fast queries
- REST API: POST /api/v1/report (Bearer token auth), GET /api/v1/customers, GET /api/v1/customers/{id}, GET /api/v1/customers/{id}/history
- Dark theme dashboard (English): multi-customer overview table with status indicators, customer detail page with system/storage/containers/backup/health sections
- Color coding: green (OK, <30min), yellow (warn or 30-60min), red (fail or >60min)
- K8s manifest: Deployment + Service + Ingress for hub.felhom.eu in felhom-system namespace
- Dockerfile, Makefile, hub.yaml.example config
- 90-day report retention with daily auto-prune
- **Controller version:** v0.6.0 — deployed and verified on demo-felhom.eu (9 scheduler jobs, all new jobs registered)
- **Manual steps remaining for Viktor (Part 4 of TASK.md):**
- Create 5 healthcheck checks on status.felhom.eu (heartbeat, system-health, db-dump, backup, backup-integrity)
- Update controller.yaml on demo-felhom with real UUIDs
- Build and deploy felhom-hub to k3s cluster
- Configure hub.felhom.eu DNS in Cloudflare
- Enable hub reporting on demo-felhom controller.yaml
### What was previously completed (2026-02-16 session 17)
- **v0.5.4 — Monitoring Page Frontend Fixes (4 bugs, frontend-only):**
- **Bug 1: Tooltip "Invalid Date"** — `items[0].parsed.x` unreliable across Chart.js versions. Fixed tooltip callback to use `items[0].raw.x` (direct {x,y} data access) with `parsed.x` as fallback.
- **Bug 2: Charts fill full width regardless of data density** — `setChartXBounds()` setting `min/max` at runtime was ignored because the scale was created without them. Fixed by including `min: now - defaultRangeMs, max: now` in the initial `chartOpts()` options. Now "7 nap" shows full 7-day x-axis with data clustered on the right.
- **Bug 3: Sysinfo values not consistently right-aligned** — `.sysinfo-grid` used `auto-fill` creating variable-width cells. Fixed to `1fr 1fr` (fixed 2-column). Added `align-items: baseline`, `gap: 1rem`, `white-space: nowrap` on labels, `font-weight: 600` + `word-break: break-word` on values. Removed redundant `<style>` block from monitoring.html (styles now in style.css).
- **Bug 4: Charts overflow on mobile** — Added `min-width: 0` on `.chart-box` (critical CSS grid fix), `overflow: hidden` + `max-width: 100%` on `.chart-wrap` and `.chart-wrap-bar`, `max-width: 100%` on canvas.
- **Controller version:** v0.5.4 — deployed and verified on demo-felhom.eu
### What was previously completed (2026-02-16 session 16)
- **v0.5.1 — Monitoring Page Bugfixes:**
- **Bug 1: Hostname** — `os.Hostname()` returns the container ID inside Docker. Fixed by mounting `/etc/hostname:/host/etc/hostname:ro` and reading it first in `sysinfo.go`. Now shows `demo-felhom`.
- **Bug 2: Tooltip timestamps** — Chart.js tooltip callback used `items[0].parsed.x` (category index 0,1,2...) instead of `items[0].label` (actual timestamp). Index 0 worked by accident (`0 || label` falls through), but all other points showed 1970-01-01.
- **Bug 3+4: Default range + empty charts** — Default range was `24h` but new system had only minutes of data. Changed to `1h` default for both system and container detail charts. Moved `active` class to "1 óra" button.
- **Controller version:** v0.5.1 — deployed and verified on demo-felhom.eu
### What was previously completed (2026-02-16 session 15)
- **v0.5.0 — Backup Bugfixes + Monitoring Page with Metrics Store:**
- **Task 1: Fixed "Helyi mentés" showing "–" after restart** — `GetFullStatus()` now synthesizes `LastBackup` from `SnapshotHistory` and `LastDBDump` from `DumpFiles` on disk when the in-memory values are nil (e.g., after controller restart). Dashboard handler also updated to use `GetFullStatus()` instead of `GetStatus()` for consistent behavior.
- **Task 2: Verified backup page caching** — Already implemented in v0.4.7 (`RefreshCache`, scheduler job, `AfterBackup` callback). No changes needed.
- **Task 3: New Monitoring Page ("Rendszermonitor")** — Full system monitoring subsystem:
- **SQLite metrics store** (`internal/metrics/store.go`, `types.go`): WAL-mode SQLite via `modernc.org/sqlite` (pure Go, no CGO). Stores system metrics (CPU%, memory, temperature, load) and container metrics (CPU%, memory, net/block I/O) with timestamp. Downsampled queries via bucket-based `GROUP BY` for Chart.js. 30-day auto-prune via daily scheduler job at 04:00.
- **Metrics collector** (`internal/metrics/collector.go`): Background goroutine collects system + container metrics every 60 seconds. System data from `system.GetInfo()`, container data from `docker stats --no-stream` with tab-separated format parsing.
- **System info provider** (`internal/metrics/sysinfo.go`, `sysinfo_other.go`): Reads hostname, OS, kernel, CPU model/cores, uptime from `/proc` filesystem. Linux-specific with build-tag fallback for cross-compilation.
- **REST API endpoints** (4 new routes in `router.go`): `GET /api/metrics/system` (time-series with range presets), `GET /api/metrics/containers/summary` (current stats), `GET /api/metrics/containers/{name}` (per-container time-series), `GET /api/metrics/sysinfo` (static system info).
- **Monitoring page template** (`monitoring.html`): 5 sections — System Overview (sysinfo via API), System Metrics Charts (4 line charts: CPU, Memory, Temperature, Load in 2×2 grid), Container Resources (2 horizontal bar charts: CPU% and Memory), Per-container Detail (click to expand with historical charts), Storage (server-rendered progress bars). Time range selectors (1h/6h/24h/7d/30d). Auto-refresh every 60s.
- **Chart.js 4.4.7** embedded locally (offline environments, ~200KB UMD), dark theme configuration matching site design.
- **CSS**: ~100 lines added for monitoring page (`.monitor-card`, `.charts-grid`, `.chart-box`, `.container-charts-row`, `.storage-bars`, responsive rules).
- **Wiring**: 4th sidebar nav item "Rendszermonitor", metrics DB path in named volume (`data/metrics.db`), `/etc/os-release:/host/etc/os-release:ro` volume mount in docker-compose.yml, Dockerfile updated to `golang:1.24-bookworm` (required by `modernc.org/sqlite`), `go.mod` upgraded to `go 1.24.0`.
- **Controller version:** v0.5.0 — deployed and verified on demo-felhom.eu (metrics collecting, 16 containers reporting, sysinfo showing Intel N100 correctly)
### What was previously completed (2026-02-16 session 14)
- **v0.4.7 — Protected Stack Detail Pages + Backup Page Caching:**
- **Protected stacks clickable** — `data-href` gating changed from `{{if not .Protected}}` to `{{if .Meta.Slug}}` on both `stacks.html` and `dashboard.html`. Protected stacks with `.felhom.yml` (i.e. a slug) are now clickable, linking to `/apps/{slug}`. Stacks without `.felhom.yml` remain non-clickable.
- **"Részletek" button for protected stacks** — Protected stack action section in `stacks.html` now shows a "Részletek" link when the stack has a slug, next to the restart button.
- **FileBrowser `.felhom.yml` resources** — Added `resources` section (mem_request: 128M, mem_limit: 256M, pi_compatible: true, needs_hdd: true) to both `install_filebrowser()` in `docker-setup.sh` and manually on the demo node. FileBrowser detail page now shows memory/Pi/HDD badges.
- **Backup page caching** — `GetFullStatus()` no longer runs expensive subprocess calls (restic stats, docker inspect, disk listing) on every page load. Instead, a new `RefreshCache()` method runs these in the background:
- Every 5 minutes via `backup-cache` scheduler job
- After each successful backup via `AfterBackup` callback
- On startup via a goroutine (non-blocking)
- `GetFullStatus()` returns the cached `FullBackupStatus` instantly, updating only dynamic fields (running flag, next run times, snapshot history). Falls back to a minimal status if cache hasn't populated yet.
- **Controller version:** v0.4.7 — deployed and verified on demo-felhom.eu
### What was previously completed (2026-02-16 session 13)
- **v0.4.6 — MariaDB Validation Fix + Dashboard & Protected Stack UX:**
- **Bugfix: MariaDB dump validation false positive** — MariaDB 11.4+ prepends `/*M!999999\- enable the sandbox mode */` before the dump header comment. `ValidateDump()` now scans the first 10 lines for the expected header pattern instead of just checking line 1. Accepts `-- MariaDB dump`, `-- MySQL dump`, `-- mysqldump` for MariaDB and `-- PostgreSQL database dump` for PostgreSQL.
- **Dashboard shows deployed apps only** — `dashboardHandler()` filters to deployed + protected stacks only. Non-deployed apps remain on the Alkalmazások page. Section heading changed to "Telepített alkalmazások". `TotalCount` stat card still shows all 52 apps.
- **Protected stack restart button** — Protected stacks (traefik, cloudflared, felhom-controller, filebrowser) now show an "Újraindítás" restart button when operational, on both dashboard (compact ↻) and Alkalmazások page (full button). "Védett" / "Védett rendszerkomponens" badge still shown.
- **API protection guard** — Centralized guard in `actionStack()` blocks all actions except `restart` on protected stacks (HTTP 403). Defense-in-depth: `StopStack()` and `DeleteStack()` retain their own guards.
- **FileBrowser `.felhom.yml`** — `install_filebrowser()` in `docker-setup.sh` now creates `.felhom.yml` with `subdomain: files` metadata, so the controller shows the `files.DOMAIN ↗` URL link. Manually created on demo node.
- **Controller version:** v0.4.6 — deployed and verified on demo-felhom.eu
### What was previously completed (2026-02-16 session 12)
- **v0.4.5 — Dedicated Backup Page ("Biztonsági mentés"):**
- **New `/backups` page** with full backup system visibility — 5 sections:
1. **Status overview cards**: Local backup status (green/gray), remote placeholder (gray), DB count, repo size
2. **Schedule section**: DB dump/restic/prune schedule with next-run times, last backup time + duration, retention policy, "Mentés most" button
3. **Database table**: Lists all discovered DBs with type badge (PostgreSQL/MariaDB), dump file size, last dump time, validation (table count), status
4. **Snapshot history table**: Last 20 snapshots with ID, time, data added, files new/changed
5. **Repository info card**: Path, size, snapshot count, integrity check status, backed-up paths list, remote copy placeholder
- **Backend extensions:**
- `SnapshotRecord` type + ring buffer (20 entries) in Manager for per-snapshot stats
- `DumpValidation` — scans dump files for CREATE TABLE statements, validates header and file size
- `ValidateDump()` runs after each successful dump in `DumpOne()`
- `ListDumpFiles()` scans dump directory for existing `.sql` files (fallback when in-memory results empty)
- `ListSnapshots()` on ResticManager — returns all snapshots from restic (newest first)
- `GetFullStatus()` on Manager — single call returns everything the page needs
- `LoadSnapshotHistory()` populates history from restic on startup (without delta stats)
- Restic check result tracking (`lastCheckTime`, `lastCheckOK`)
- `NextDailyRun()` exported from scheduler for next-run time calculation
- **Server wiring:**
- `Server` struct now holds `*scheduler.Scheduler`
- `NewServer()` accepts scheduler parameter
- `/backups` route + `backupsHandler()` in handlers.go
- **New template functions** (`funcmap.go`): `timeAgo`, `fmtTime`, `fmtTimeShort`, `dbTypeLabel`, `nextRunLabel`, `pruneLabel`, `nextPruneLabel`, `fmtDuration`, `fmtBytes`, `shortID`
- **Navigation**: Sidebar now has 3 items (Vezérlőpult, Alkalmazások, Biztonsági mentés)
- **Dashboard**: Backup card title is now a clickable link to `/backups`
- **Auto-refresh**: Page polls `/api/backup/status` every 3s during backup-in-progress, reloads when complete
- **CSS**: Full dark-theme styles for schedule card, database table, snapshot table, repository card, validation badges, DB type badges, empty state
- **Controller version:** v0.4.5 — deployed and verified on demo-felhom.eu (2 historical snapshots loaded)
### What was previously completed (2026-02-15 session 11)
- **v0.4.1 — App Filtering + Bugfixes:**
- **Filter bar on Alkalmazások page**: Four pill-shaped filter buttons (Mind/Futó/Leállítva/Telepíthető) with live count badges computed from DOM. Filters stack cards via `display: none`, updates URL with `?filter=running` via `history.replaceState`. Reads filter from URL on page load for deep-linking support.
- **New `filterCategory` template function** (`funcmap.go`): Maps container state + deployed flag to filter categories (running/stopped/available). Each stack card gets a `data-filter-state` attribute for client-side filtering.
- **Clickable dashboard stat cards**: Stat cards (Futó/Leállítva/Összes) changed from `<div>` to `<a>` with `href` linking to `/stacks?filter=running`, `/stacks?filter=stopped`, `/stacks` respectively. Hover effect with translateY + box-shadow.
- **docker-compose.yml synced to demo node**: Fixed the stale compose file that still had `dashboard.${DOMAIN}` Traefik label (from pre-v0.3.0). Now uses correct `felhom.${DOMAIN}` label + `/sys:/host/sys:ro` mount.
- **Controller version:** v0.4.1 — deployed and verified on demo-felhom.eu
- **Remaining manual tasks for Viktor (Task 2 & 3 from TASK.md):**
- Verify `felhom.demo-felhom.eu` resolves correctly (Cloudflare Tunnel public hostname may need updating from `dashboard.*` to `felhom.*`)
- Update Pi-hole local DNS if applicable
- Enable backup in `controller.yaml` on demo node (`backup.enabled: true`)
- Create `/srv/backups` directories on demo node
### What was previously completed (2026-02-15 session 10)
- **v0.4.0 — Monitoring & Health + Backups (Phase 2 & 3):**
- **Central job scheduler** (`internal/scheduler/scheduler.go`):
- Replaces ad-hoc goroutines in main.go with a unified scheduler
- `Every(name, interval, fn)` for periodic jobs, `Daily(name, timeStr, fn)` for scheduled tasks
- Panic recovery, skip-if-running, quiet mode for high-frequency jobs (≤30s)
- Daily jobs use `Europe/Budapest` timezone with `time.Timer` for DST correctness
- Graceful shutdown with 30s timeout for running jobs
- **CPU usage collector** (`internal/system/cpu_linux.go`):
- Background goroutine samples `/proc/stat` every 5s, computes delta-based CPU %
- Platform stubs for non-Linux in `cpu_other.go`
- **Temperature & load metrics** (`internal/system/info_linux.go`):
- Reads `/proc/loadavg` for 1/5/15 min load averages
- Reads thermal zones from `/host/sys/class/thermal/` (Docker mount) with `/sys/` fallback
- Handles millidegree values, picks highest zone, with hwmon fallback
- **Healthchecks.io pinger** (`internal/monitor/pinger.go`):
- HTTP ping client for Healthchecks.io-compatible endpoints
- POST to `/ping/{uuid}` (success), `/fail` (failure), `/start` (started)
- 10s timeout, 3 retries with 2s backoff, skips CHANGEME UUIDs
- **System health checks** (`internal/monitor/healthcheck.go`):
- Checks disk, memory, CPU, temperature, Docker reachability, protected containers
- Returns HealthReport with status "ok"/"warn"/"fail" + formatted message for pings
- **Database dump engine** (`internal/backup/dbdump.go`):
- Auto-discovers PostgreSQL/MariaDB containers via `docker ps` + `docker inspect`
- Dumps via `docker exec pg_dump`/`mariadb-dump` with 5min timeout
- Atomic writes (`.tmp` → `.sql`), empty file detection, stale temp cleanup
- **Restic integration** (`internal/backup/restic.go`):
- Auto-generates repository password (32 random bytes, base64url)
- Init, snapshot (JSON output), prune, check, stats, latest snapshot
- Stale lock detection with automatic unlock + retry
- **Backup orchestrator** (`internal/backup/backup.go`):
- DB dumps + restic snapshots, weekly prune on Sundays
- Thread-safe running flag, Healthchecks.io pings with results
- `RunFullBackup()` for manual trigger (sequential: dumps → snapshot)
- **Wiring updates:**
- `main.go`: scheduler-based job registration, cpuCollector lifecycle, pinger + backupMgr init
- `api/router.go`: `GET /api/backup/status`, `POST /api/backup/run`
- `web/server.go` + `handlers.go`: pass cpuCollector to GetInfo(), backup status on dashboard
- `funcmap.go`: `tempColor`, `fmtTemp`, `fmtLoad` template functions
- **Dashboard UI enhancements:**
- CPU usage bar with load average display below
- Temperature with colored indicator dot (green/yellow/red at 60°/75°C)
- Backup status card: last run time, DB count, repo size/snapshots
- "Mentés most" button triggers manual backup via API
- **Config updates:**
- `controller.yaml.example`: added `system_health_interval`, `hdd_path`, `system.reserved_memory_mb`
- `docker-compose.yml`: added `/sys:/host/sys:ro` mount for temperature reading
- `restic_password_file` default changed to `data/` subdir (auto-generated in named volume)
- **Controller version:** v0.4.0 — deployed and verified on demo-felhom.eu
### What was previously completed (2026-02-15 session 9)
- **v0.3.0 — Structural refactoring (templates + server split + domain rename):**
- **Templates: go:embed migration** — moved all 7 HTML templates + CSS from Go string constants to individual files in `internal/web/templates/`. Created `embed.go` with `//go:embed` directive. Template loading now uses `ParseFS()` instead of `Parse()`. CSS served from embed.FS via `ReadFile()`. Zero runtime file dependencies — still compiled into the binary.
- **Server decomposition** — split monolithic `server.go` (540 lines) into focused files:
- `auth.go`: session struct, auth middleware, login/logout handlers, session management
- `handlers.go`: page handlers (dashboard, stacks, logs, deploy, app detail)
- `funcmap.go`: template FuncMap with 14 custom functions
- `server.go`: Server struct, NewServer, loadTemplates (3-liner), ServeHTTP routing, render helper, static file serving
- **Domain rename** — controller subdomain changed from `dashboard.*` to `felhom.*` in Traefik labels and setup script
- **Documentation updated** — CLAUDE.md, README.md, CONTEXT.md all reflect new file structure
- **Reminder for Viktor:** Update Cloudflare Tunnel public hostname (`dashboard.demo-felhom.eu` → `felhom.demo-felhom.eu`) and Pi-hole DNS if needed
- **Controller version:** v0.3.0
### What was previously completed (2026-02-15 session 8)
- **FileBrowser as infrastructure service:**
- Created `scripts/hdd-setup.sh` (adapted from deploy-portainer) — sets up HDD folder structure with `Dokumentumok` user dir
- Created `scripts/docker-setup.sh` (adapted from deploy-portainer) — installs Docker, Traefik, FileBrowser as infra services
- Added `filebrowser` to protected stacks in `controller.yaml.example`
- Removed `templates/filebrowser/` from app-catalog-felhom.eu (no longer a catalog app)
- **Orphan stack detection and deletion:**
- Added `Orphaned` field to Stack struct + `getCatalogTemplateSlugs()` helper
- Orphan detection in `ScanStacks()` — deployed stacks with no matching catalog template marked as orphaned
- New `delete.go`: `DeleteStack()` (compose down + HDD cleanup + dir removal), `GetStackHDDData()`, `parseComposeHDDMounts()`
- Safety: protected HDD paths (root, media, storage, Dokumentumok, appdata) can never be deleted
- New API endpoints: `DELETE /api/stacks/{name}` and `GET /api/stacks/{name}/hdd-data`
- UI: orange "Elavult" badge on orphaned stacks, "Törlés" button, delete confirmation modal
- Modal shows HDD data paths/sizes, checkbox for "Felhasználói adatok törlése a merevlemezről"
- Hides "Frissítés" and "Részletek" buttons for orphaned stacks
- **Verified:** 1 orphaned stack detected on startup (filebrowser — now infra, removed from catalog)
- **Controller version:** v0.2.15
### Previously completed (2026-02-14 session 7)
- **Fixed YAML parse error in romm `.felhom.yml`** (app-catalog repo):
- Root cause: Hungarian opening quote `„` (U+201E) paired with ASCII `"` (0x22) inside YAML double-quoted strings terminated the string prematurely
- Affected lines: `help_text` for IGDB Client Secret and SteamGridDB API Key fields
- Fix: escaped inner ASCII double quotes with `\"` in the YAML strings
- This caused `LoadMetadata()` to silently fail and return empty defaults for ALL romm metadata (tagline, resources, category — everything)
- **Added error logging to `LoadMetadata()`** in `metadata.go`:
- `[ERROR]` log on YAML parse failure (was silently swallowed — critical bug)
- Temporary `[DEBUG]` log used for diagnosis, then removed
- **Fixed deploy command in CLAUDE.md**:
- `sed` pattern now targets only `image:` lines (was matching service name too, breaking YAML)
- Added `sudo` for both sed and docker compose (directory is root-owned)
- **Controller version:** v0.2.14
### Previously completed (2026-02-14 session 6)
- **Bug fix: App info logo SVG rendering** — `.app-info-logo` CSS in `templates.go`:
- Added `min-width`, `min-height`, `max-width`, `max-height: 80px` and `overflow: hidden`
- Prevents SVG images with explicit dimensions or no viewBox from overflowing container
- Logo now reliably renders at 80x80 regardless of SVG intrinsic size
- **Controller version:** v0.2.12
### Previously completed (2026-02-14 session 5)
- **App detail/info pages** — new feature:
- New route: `GET /apps/{slug}` renders a full info page (was redirect to deploy page)
- Hero section with logo, tagline, resource badges
- Screenshots section (graceful — hidden via `onerror` if assets don't exist)
- Info cards: use cases, first steps, prerequisites, default credentials, docs link
- Optional config form with AJAX save (POST `/api/stacks/{name}/optional-config`)
- New `.felhom.yml` fields: `app_info` (tagline, use_cases, first_steps, prerequisites, default_creds, docs_url) and `optional_config` (groups of env var fields)
- New structs in `metadata.go`: `AppInfo`, `OptionalConfigGroup`, `OptionalConfigField`
- `UpdateOptionalConfig` in `deploy.go`: saves optional env vars to `app.yaml`, restarts deployed stacks with `docker compose up -d` to pick up new env vars
- Navigation updated: stack cards on dashboard/stacks pages now link to `/apps/{slug}`, deploy page has "Részletek" link back to info page
- **RoMM metadata updated** (app-catalog repo):
- Full `app_info` section: tagline, 5 use cases, 6 first steps, 3 prerequisites, default creds, docs URL
- 6 optional config fields for metadata providers: IGDB (client_id + secret), SteamGridDB, ScreenScraper (user + password), MobyGames
- docker-compose.yml updated with SCREENSCRAPER_USER, SCREENSCRAPER_PASSWORD, MOBYGAMES_API_KEY env vars
- Display name fixed: "ROMM" → "RomM"
- **Controller version:** v0.2.11
### Previously completed (2026-02-14 session 4)
- **Fixed deploy race condition** in `internal/stacks/deploy.go`:
- In-memory `Deployed` flag now set BEFORE `docker compose up -d` (compose up can take 30-60s for image pulls)
- On failure: both in-memory state and disk (app.yaml) are reverted
- Eliminates stale "Telepítés" button during long compose operations
- **Added `checkBeforeDeploy()` JS guard** in `internal/web/templates.go`:
- Telepítés buttons on Vezérlőpult and Alkalmazások pages now fetch live state from `/api/stacks/{name}` before navigating
- If app is already deployed (e.g., another tab deployed it), shows alert and reloads page instead of navigating to deploy form
- Catches stale UI state gracefully
### Previously completed (2026-02-14 session 3)
- **Enhanced debug logging** across all stack operations in `internal/stacks/`:
- **Operation timing**: All stack ops (start, stop, restart, update, deploy) now log elapsed time
- **Post-start container state check**: Async goroutine after start/restart/update/deploy
- **Image pull detection**: Checks local images before deploy/update (debug level)
- **GetLogs/ScanStacks improvements**: Byte count logging, deployed/available counts
- All verbose checks gated on `cfg.Logging.Level == "debug"`; timing always at INFO
- **UI improvements** in `internal/web/templates.go` and `server.go`:
- **Memory bar fix on deploy page**: Bar segments now always visible (min-width: 3px), new app segment uses translucent green with distinct border for clear visual separation from committed memory
- **Clickable app cards**: Cards on Vezérlőpult and Alkalmazások pages are now clickable (navigates to deploy/detail page). Uses `data-href` attribute + delegated click handler. Protected stacks excluded. Actions area (buttons, state labels) excluded from click-to-navigate
- **Live-scrolling logs**: Logs page now auto-refreshes every 3s via AJAX polling (`?raw=1` returns plain text). Fixed-height container (70vh) with auto-scroll to bottom. Pulsing green "Élő" indicator. Pause/resume toggle ("Szüneteltetés"/"Folytatás"). User scroll position preserved when scrolled up to read history
- **Deployment progress UI**: Deploy button no longer shows alert+redirect immediately. Instead shows 3-step progress panel: config saved → containers starting → app initializing. Polls `GET /api/stacks/{name}` every 3s to track actual container health state. Handles running (auto-redirect), starting (keep polling), unhealthy (warning), exited (error), and 120s timeout. Shows elapsed time counter
- **Mealie healthcheck fix** (app-catalog-felhom.eu):
- `wget --spider` replaced with Python TCP socket check — mealie image doesn't include wget
- `start_period` increased to 60s (DB migrations take ~40s on first start)
- **Healthcheck audit**: filebrowser (Alpine, has BusyBox wget — OK), stirling-pdf (Ubuntu, has wget — OK)
### Previously completed (2026-02-15 session 2)
- **Phase 4: Git Sync + App Catalog Audit** — major milestone
- **Git sync module** (`internal/sync/sync.go`):
- Clones/pulls app-catalog-felhom.eu repo to local cache on startup
- Periodic sync based on `git.sync_interval` (default 15m)
- Copies `docker-compose.yml` + `.felhom.yml` to stacks dir (never overwrites `app.yaml`/`.env`)
- SHA-256 content comparison — only writes changed files
- Triggers `ScanStacks()` after sync so dashboard updates immediately
- Uses `os/exec` git CLI — no Go git library dependency
- **Manual sync button** ("Sablonok frissítése") on Alkalmazások page:
- `POST /api/sync` endpoint with 30s debounce
- Toast notification shows result (success/failure/what changed)
- Auto-reloads page if new apps or updates detected
- **Sync status** added to `/api/system/info` (last_sync, last_status, syncing flag)
- **.felhom.yml files created for all 10 apps** (paperless-ngx already had one):
- actualbudget, docmost, filebrowser, homebox, immich, mealie, romm, stirling-pdf, vaultwarden
- All follow the same format: display_name, description, category, subdomain, resources, deploy_fields
- **Docker Compose templates audited and fixed** for all 10 apps:
- Fixed `{{DOMAIN}}` → `${DOMAIN}` syntax in homebox, mealie, romm, stirling-pdf
- Fixed `{{HDD_PATH}}` → `${HDD_PATH}` in romm
- Added `deploy.resources.limits.memory` to all services across all templates
- Added `TZ=Europe/Budapest` to all sidecar services (postgres, redis, mariadb)
- Added healthcheck to romm main service
- Added `romm-redis` `condition: service_healthy` (was `service_started`)
- Standardized header comment blocks across all templates
- **Documentation updated**: app-catalog README, CLAUDE.md, CONTEXT.md
### Previously completed (2026-02-15 session 1)
- **Memory validation during deployment**:
- Pre-deploy memory check: compares `mem_request` sum against usable system RAM
- Hard block if requests exceed usable memory (total - 384MB reserved)
- Soft warning if `mem_limit` sum exceeds total RAM (overcommit OK for limits)
- `ParseMemoryMB()` supports "500M", "1G", "1.5G", "1024" formats
- `CommittedMemory()` sums requests/limits across all deployed stacks
- Memory summary bar shown on deploy page before user clicks deploy
- `system.reserved_memory_mb` configurable in controller.yaml (default: 384)
- **Display: `~` prefix on mem_request** in UI badges (display-only, exact value stored)
- **Felhom.eu logo** replaced text logos in sidebar and login page with actual SVG logo
- Logo SVG embedded as Go string constant, served at `/static/felhom-logo.svg`
### Previously completed (2026-02-14)
- **System info bar on Vezérlőpult dashboard**: RAM, SSD, and optional HDD usage
- Progress bars with color coding (green < 70%, yellow 70-85%, red > 85%)
- New `internal/system` package reads `/proc/meminfo` + `syscall.Statfs`
- Platform-specific: Linux impl + non-Linux stub (build tags)
- Hungarian labels: "Memória", "SSD tárhely", "Külső HDD"
- **Docker Compose memory limits** on paperless-ngx template:
- paperless-webserver: 768M, postgres: 256M, redis: 128M
- Added `mem_limit` field to `.felhom.yml` ResourceHints (total: 1152M)
- **`/api/system/info` endpoint** now returns live system metrics (was customer info)
- **Config**: Added `paths.hdd_path` for external HDD monitoring
- Controller image builds via build.sh, pushes to Gitea container registry
### Previously completed (2026-02-13)
- Built the entire felhom-controller from scratch (Go, no frameworks)
- Debugged and fixed 7 issues during first real deployment:
1. Password validation (empty passwords accepted)
2. In-memory Deployed flag not updating after deploy
3. Health-aware state parsing (starting/unhealthy detection)
4. Random card ordering (Go map iteration)
5. "Részletek" button redirect for deployed apps
6. Paperless OCR language installation (LANGUAGES vs LANGUAGE env var)
7. Documentation: restart vs up -d for image updates
### What's next (priorities)
1. **Test per-app backup** — enable backup for Paperless-ngx HDD data, trigger manual backup, verify restic snapshot includes HDD paths
2. **Test restore** — restore app data from snapshot, verify file recovery (now possible with /mnt:rw mount)
3. **Deploy Immich** — tests HDD path + secrets + multi-storage (biggest real-world test)
4. Add `app_info` + `optional_config` to more apps (Immich, Mealie, Vaultwarden)
5. Test on Raspberry Pi (pi-customer-1)
6. Self-update mechanism
7. Hub alerting (webhook to Healthchecks for stale customers)
8. Docker volume backup (mount `/var/lib/docker/volumes:ro` into controller)
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->