- internal/family: the family list (bcrypt, generated 4x4 passwords shown once) + 30-day sessions in family.json
(0600, atomic); a reset (generation), a removal or a logout ends sessions at the next request.
- internal/stacks/family_gate.go: family_gate / family_gate_except / min_controller in .felhom.yml; the door is written
BEFORE the first start (install and a removed app's restore), a life record in app.yaml, reconciled by the gate loop;
priority below the install hold, setup gate and sign-up block; every exception anchored ^/prefix(/|$) (finding F1).
- internal/web/family_gate.go: forwardAuth /__felhom_gate/family (app cookie felhom_famgate, host-only, names a store
session); /__family/start|login|logout on the dashboard host (session cookie felhom_family, Path=/__family);
sign-in counted per visitor (clientIP) AND per name, short windows; the household's dashboard session vouches.
RequireAuth never reads a family cookie. The "Család" card on the security page: add / new password / remove.
Red-proofs RP-F1..RP-F7 (felhom.eu audits/family-gate-2026-10-02/A/).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- felhom-tunnel network 172.16.253.0/29 (ip-range .4/30): cloudflared alone at .2, traefik at .3; traefik's websecure
trusts forwarded headers from 172.16.253.2/32 only, and every request passes felhom-forwarded@file, which removes
the client-writable host/path/address headers (X-Forwarded-Host/-Uri/-Method/-Prefix, Forwarded, True-Client-Ip, …)
and fixes X-Forwarded-Port to 443 (measured: Cloudflare passes a client's X-Forwarded-Host/-Port).
- EnsureBaseStack reconciles a RUNNING traefik/cloudflared whose rendered files changed (recreate), refuses a rewrite
that would drop a certificate resolver, and moves cloudflared only once traefik is on the tunnel network.
- clientIP: believed only when the TCP peer is traefik; the rightmost X-Forwarded-For entry (the hop traefik saw);
the tunnel hop → CF-Connecting-IP (the edge refuses a client-sent one, measured 403). rateKey: IPv6 per /64.
Dashboard login, claim, share and escrow counters key on it; the setup gate logs it.
- Dashboard login messages: keys, informal voice, both languages.
Red-proofs: RP-A1 (leftmost hop), RP-A2 (shared tunnel key), RP-A3 (no reconcile) — felhom.eu audits/visitors-2026-10-01/A.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: the product pins tag@digest, and such images are stored untagged (repo:<none>);
`docker image ls` without -a did not list them, so the retention saw almost no app image. Now `image ls -a`;
an anonymous <none>:<none> entry is never a candidate; the one-time marker is v2 so the corrected sweep runs
once everywhere. Test TestImageRetention_SeesUntaggedDigestPulledImages, red-proofed. 0.284.0/0.284.1 were
never floored.
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: v0.284.0 wired the remove half into DeleteStack only; the app page's Remove runs RemoveStack.
Now RemoveStack reads the app's image repositories before its compose down and runs the retention after. Every
retention pass logs one line (images seen, candidates, deleted), so a pass that kept everything is visible.
Tests TestImageRetention_TheRemoveButtonRunsIt / ADoneUpdateRunsItWithThePrevious, red-proofed. v0.284.0 was never
floored (scratch 9202 only).
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Image retention: after a done/undone guarded Update and at remove, an app's images older than its running
and previous one are deleted — never an image any container, installed compose or installed/previous record
names (box-wide keep set read at delete time); exact id, never forced or pruned; paused while any update runs;
a one-time sweep of catalog app images at the first start. Install hold: an after_install app is installed
behind the setup gate's door and opens when after_install succeeds or the household says it changed the login.
Tests TestImageRetention_* and TestInstallHold_* with red-proofs; parity fixture for the held card.
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A drive move persisted through the restore's fresh app.yaml write and dropped the pin: the syncer
then copied the catalog verbatim and the next start jumped the app past its ladder (R-700).
persistDriveFlip now changes HDD_PATH and nothing else. The restore's write carries the life
records (conversion copies, desired_state, update history) from the app.yaml it replaces, and a
second conversion no longer overwrites the first kept copy's record (R-697).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
An install over an app's kept drive folder (appdata/<app> non-empty) asks the household:
"use my kept data" (a load from the newest copy of THIS drive's install, own unit or
second-drive mirror, then the template's after_load) or "start fresh" (the folder is
renamed into <drive>/kept/<app>/<date>/ with the removed app's unit; nothing deleted).
The install API answers 409 kept_data_choice until one is chosen; DeployStack refuses
too. New page Megorzott adatok / Kept data (/kept-data): Load / Look / Delete (typed
confirmation, the only deletion of kept data). FileBrowser gets a read-only source.
The drive-full warning names the kept folders. <drive>/kept is protected and outside
every backup leg.
R-690: the removed-app restore (R-487) never found a unit on a DATA drive — it asked
GetStackComposePath (true for every catalog app) and restored nextcloud with no env.
Now isStackDeployed; pinned with a production-shaped provider.
Red-proofs: audits/night-2026-09-26/E/redproofs/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A step whose ladder entry carries engine_conversion {service, engine, from, to}
converts the database: the old engine alone, the check (owners, roles,
extensions, per-table row counts), pg_dumpall validated by its completion line,
the volume emptied only after the undo copy's marker is validated again, the new
engine alone, the load with ON_ERROR_STOP, the check again + PG_VERSION. Any
failure goes to the existing undo; a restart during converting is undone.
A PostgreSQL major move without the mark is refused before anything moves.
The old datadir's copy is kept until a backup is proven after the conversion.
17 tests, 9 red-proofs (audits/night-2026-09-26/B/).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (night 2026-09-24 Part B): the sync rendered the ladder's newest
tested digest into a RUNNING app's compose, so the next restart would pull a new image
with no backup and no undo. stacks.CarryDigests keeps the running digest for an
installed app; a fresh install still takes the tested digest. Red-proofed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (romm): .felhom.yml flows into the stack dir on every
catalog sync, so "the old .felhom.yml" saved at update time was already the
new one, and the serving old version was judged with the new probe.
New record applied-meta/.felhom.yml, written whenever a version is pinned
(deploy, adoption, pin advance) and put back by the undo, like
applied-compose.yml. The fixture now places the new file at sync time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: the periodic probe (current .felhom.yml, new port) flips
the app to unhealthy, and the update's health wait probed only 'running'
apps - so the undo's old probe was never asked and a serving old version was
judged "did not start". With the undo's override, an unhealthy app is probed
and the old check decides; never settled on container state.
New seam probeRunFn; the test drives the real wait loop and reproduces the
live message when the fix is switched off.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.
- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
(pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.
MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Caught by the live proof on 9202, not by a test. The v0.262.0 guard fired exactly right and
answered HTTP 500: router.go maps remove errors by grepping the error TEXT for "not deployed" /
"still running" / "not found" / "protected", and the busy sentence contains none of them.
A 500 tells the UI something broke; this is "wait a moment". Now a typed *stacks.RemoveBusyError
matched with errors.As and answered 409, carrying both the Hungarian bytes and the bundle key.
Its test asserts the sentence contains none of the words the text mapping greps for, so the type is
load-bearing rather than decorative. Red-proof seen failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path
sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a
working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404
after. It now falls through to the same settle path with a WARN naming the candidates.
The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve
by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates
returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence.
R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose
file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is
still not diagnosed and R-634 stays open for it.
R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own
sentence - the product already refused this clash for update and for restore. And because `down`
returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying
its label is removed by name with its labels logged, and the answer carries `verified`.
R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down
that destroys them. Two existing tests pin the compose sequence and correctly caught the new step;
their expectations are updated with the reason that the ORDER is the assertion.
R-614: RemoveStack calls ClearUpdateState.
NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done.
Three new sentences, each born as a key in both bundles. Four red-proofs seen failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The controller self-updates daily at 04:30 by default, and after any hub report
once a floor sits above the box. That swap restarts the controller container.
The window proposed for automatic app updates is 02:30-05:00. It contains 04:30.
R-608 — a two-way lock, wired in main.go (stacks never imports selfupdate):
- stacks.Manager.AnyUpdating() -> Updater.SetAppUpdatingCheck, consulted in the
same three places as the existing backupRunning gate.
- Updater.IsUpdateRunning -> Manager.SetSelfUpdatingCheck; UpdatePreflight
refuses `self_updating`.
- MEASURED: the gap was narrower than assumed. The update's `backing-up` phase
already takes the backup single-flight, so that one phase was covered. The
other six were not, and `starting`/`verifying` are where data may have moved.
- The lock must NOT latch: a held app does not block the controller's own
updates, including the release that might fix the hold.
R-609 — the 409 carries `data.reason`, additively. transient (busy, updating,
deploying, migrating, self_updating) vs terminal (held, downgrade). Found while
writing the test: the router refuses a HELD app on its own line before the
preflight, so `held` would have been the one reason missing.
Five red-proofs, each seen to fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Recounted at catalog 18a6d2d8: 66 unique pins — 48 full X.Y.Z, 6 two-part
lines, 4 major lines (10 float), 8 exact versions wearing a variant suffix.
The '23' carried since v0.233.0 matches no definition the catalog supports.
Definition written down beside the number so it can be rechecked.
Also: CONTEXT said the fleet floor was 0.257.0; the hub says 0.259.0.
Comment and doc only — no behaviour change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
MEASURED 2026-09-15 (BIGNIGHT Phase 6): privatebin updated 2.0.5 -> 2.0.6, catalog
reverted to 2.0.5, and the box read „Frissítés elérhető — ma" over an Update that
would have moved the pin BACKWARDS onto a possibly-migrated datadir.
- stacks.CatalogOrder: the comparison gains a fourth answer (Ahead) and moves out of
web, so the badge and UpdatePreflight cannot drift apart.
- The badge: ahead reads „Naprakész"/"Up to date", tag-ok, with a title saying why.
- The refusal: UpdatePreflight returns `downgrade` (409), born as a bundle key; the
API now renders update refusals through errText so it reaches English households.
- Ahead is narrow: every differing service must be orderable AND newer, else Behind.
- Ordering is util.Version.Compare behind a tag normaliser — no second comparator.
- Three red-proofs, each seen to fail.
R-589 was already fixed in v0.258.0; only its register row was stale.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The READ PATH for a second language in `.felhom.yml`. An `i18n: {en: …}` sibling
block inside the same file; `Metadata.For(lang)` merges it FIELD BY FIELD over the
Hungarian, so a missing or blank English field shows the Hungarian one and a
half-translated app is a legal, shippable state.
`For("hu")` is the parsed struct with `I18n` cleared and nothing else — measured
against all 53 real catalog files, copied into `internal/stacks/testdata/catalog/`.
Lists replace whole; every other list is matched by its own key, never by position.
`For` never writes through the receiver: the metadata is the stack manager's, shared
by concurrent requests, and an in-place merge would leak one household's language
into another household's page.
Pages reach catalog copy only through `LocalizeStacks`/`LocalizeStackPtr`/`MetaFor`,
and `TestNoDirectMetaCopyReadOnPages` keeps a named, reasoned allow-list of every
direct `.Meta.<copy>` read in `internal/web` so the NEXT page to read one fails the
suite instead of quietly rendering Hungarian to an English household.
Eight red-proofs. Two of them convicted a hollow TEST rather than the code: a struct
copy shares its slices' backing arrays, so the obvious DeepEqual mutation check
passed a deliberately broken merge; and a one-entry fixture cannot tell key matching
from position matching. Both rewritten, both then seen to fail.
MinAgent: 0.131.0 (unchanged). Older controllers are unaffected — `LoadMetadata`
uses non-strict `yaml.Unmarshal`, so a pre-0.257.0 box drops the whole block.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
179 Hungarian sentences were built deep inside a package with fmt.Errorf and printed by
whoever caught them: too late to translate where they are shown, too early where they are
made. Every one now carries its key across that gap. ZERO Hungarian error literals remain.
util.MsgError does three things at once, each earned:
- Error() is the Hungarian, byte for byte, so every un-converted printer is unchanged;
- errors.Is answers for the kind AND for a wrapped cause (KindErrorf dropped the cause);
- an error ARGUMENT renders recursively, so "formázás sikertelen: %w" translates whole.
A foreign error — restic, docker, ssh, the stdlib — prints verbatim. It is not ours.
76 display sites go through errText, and TestNoErrErrorInPageOutput convicts any that do
not. memoryVerdict returns an error rather than a sentence, so the deploy's 409 and the
household's language come from one value; UpdateRefusal gained a Cause to carry it.
Plurals, one rule, stated once: a key with .one/.other takes its COUNT first. Not a
per-call-site flag — the producer somebody forgot would read "3 app is not running". The
guard caught a real key collision (alert.deadapp.one) the day the rule landed.
TWO DEFECTS FOUND IN MY OWN TOOLING, recorded rather than quietly fixed. The bulk converter
silently dropped multi-line concatenations, damaging 7 producers — and the parity gate could
not see it, because every surviving fragment WAS a real base literal while the CALL had lost
text; two behaviour tests caught it. And the counting script was case-sensitive, so it said
"0 left" while five remained.
MinAgent: 0.131.0 (unchanged). No hub release needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Every Hungarian sentence is byte-identical; each decision now reads a signal set where the message is
made. util.KindErrorf builds the same bytes fmt.Errorf did while carrying a sentinel for errors.Is.
- Deploy status (api/router.go): deployStatusFor() by kind — stacks.ErrAlreadyDeployed (409),
ErrRequiredField / ErrPathMissing / ErrNotEnoughMemory (400). The „kötelező" / „memória" /
"does not exist" / "already deployed" text chain is gone.
- Off-site failure class (backup/offbox.go): ErrOffsiteQuota replaces the „tárhelykeretet" match. The
restic/ssh signatures stay text matches on purpose — that output is not ours and is not translated.
- Alert placement (web/alerts.go): monitor.HealthReport carries WarningKinds parallel to Warnings;
the "not on a separate drive" warning is inline by KIND. The hub report is untouched (builder.go
copies Status/Issues/Warnings only) — pinned by a wire test.
- Stale off-site note (web/handlers.go): settings LastWarningKind + backup.OffboxWarnNoAppsSelected.
The text test survives ONLY for kind == "" (a box whose last run predates 0.251.0) and is removed
when R-570 closes; slice 2 must not translate that producer before then.
Tests (all red-proofed by restoring the pre-fix predicate — see the audit's redproofs.txt):
TestR553_Deploy_DecisionSurvivesWordingChange, TestR553_DeployHandlerUsesTheKind,
TestR553_DeployProducersCarryKindAndKeepTheirWords (through the real DeployStack),
TestR553_OffsiteQuota_{Decision,HeadLine}SurvivesWordingChange, TestR553_OffboxRunRecordsTheKind,
TestR553_StorageWarningsCarryKindsAndKeepTheirWords, TestR553_DiskWarningPlacementSurvivesWordingChange,
TestR553_HubReportWarningsAreUnchangedOnTheWire, TestR553_StaleNote*, TestR553_WarningKindIsPersistedAndCopied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.
R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.
R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.
Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
R-486 (P1): removing an app with its backups KEPT keeps its Tier-2 record,
so the second-drive restore is no longer refused over an intact mirror.
R-484: postgis/pgvector/timescaledb images are Postgres (logical dumps).
R-485: the backup card sizes the recovery unit and the mirror(s).
R-480: a held update's sentence leaves the card once the hold is lifted.
R-477: the update's off-site lookup is one snapshots call, no stats.
R-478: a copy older than this install's deploy does not count.
R-474: "delete backups" deletes the unit, the mirror(s) and the prefs.
Tests and red-proofs per row; evidence in felhom.eu
documentation/audits/v0240-2026-09-13/ and nightly-2026-09-13-adventurelog/.
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.
Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS