Image retention: after a done/undone guarded Update and at remove, an app's images older than its running
and previous one are deleted — never an image any container, installed compose or installed/previous record
names (box-wide keep set read at delete time); exact id, never forced or pruned; paused while any update runs;
a one-time sweep of catalog app images at the first start. Install hold: an after_install app is installed
behind the setup gate's door and opens when after_install succeeds or the household says it changed the login.
Tests TestImageRetention_* and TestInstallHold_* with red-proofs; parity fixture for the held card.
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A drive move persisted through the restore's fresh app.yaml write and dropped the pin: the syncer
then copied the catalog verbatim and the next start jumped the app past its ladder (R-700).
persistDriveFlip now changes HDD_PATH and nothing else. The restore's write carries the life
records (conversion copies, desired_state, update history) from the app.yaml it replaces, and a
second conversion no longer overwrites the first kept copy's record (R-697).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
An install over an app's kept drive folder (appdata/<app> non-empty) asks the household:
"use my kept data" (a load from the newest copy of THIS drive's install, own unit or
second-drive mirror, then the template's after_load) or "start fresh" (the folder is
renamed into <drive>/kept/<app>/<date>/ with the removed app's unit; nothing deleted).
The install API answers 409 kept_data_choice until one is chosen; DeployStack refuses
too. New page Megorzott adatok / Kept data (/kept-data): Load / Look / Delete (typed
confirmation, the only deletion of kept data). FileBrowser gets a read-only source.
The drive-full warning names the kept folders. <drive>/kept is protected and outside
every backup leg.
R-690: the removed-app restore (R-487) never found a unit on a DATA drive — it asked
GetStackComposePath (true for every catalog app) and restored nextcloud with no env.
Now isStackDeployed; pinned with a production-shaped provider.
Red-proofs: audits/night-2026-09-26/E/redproofs/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A step whose ladder entry carries engine_conversion {service, engine, from, to}
converts the database: the old engine alone, the check (owners, roles,
extensions, per-table row counts), pg_dumpall validated by its completion line,
the volume emptied only after the undo copy's marker is validated again, the new
engine alone, the load with ON_ERROR_STOP, the check again + PG_VERSION. Any
failure goes to the existing undo; a restart during converting is undone.
A PostgreSQL major move without the mark is refused before anything moves.
The old datadir's copy is kept until a backup is proven after the conversion.
17 tests, 9 red-proofs (audits/night-2026-09-26/B/).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (night 2026-09-24 Part B): the sync rendered the ladder's newest
tested digest into a RUNNING app's compose, so the next restart would pull a new image
with no backup and no undo. stacks.CarryDigests keeps the running digest for an
installed app; a fresh install still takes the tested digest. Red-proofed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (romm): .felhom.yml flows into the stack dir on every
catalog sync, so "the old .felhom.yml" saved at update time was already the
new one, and the serving old version was judged with the new probe.
New record applied-meta/.felhom.yml, written whenever a version is pinned
(deploy, adoption, pin advance) and put back by the undo, like
applied-compose.yml. The fixture now places the new file at sync time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: the periodic probe (current .felhom.yml, new port) flips
the app to unhealthy, and the update's health wait probed only 'running'
apps - so the undo's old probe was never asked and a serving old version was
judged "did not start". With the undo's override, an unhealthy app is probed
and the old check decides; never settled on container state.
New seam probeRunFn; the test drives the real wait loop and reproduces the
live message when the fix is switched off.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.
- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
(pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.
MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Caught by the live proof on 9202, not by a test. The v0.262.0 guard fired exactly right and
answered HTTP 500: router.go maps remove errors by grepping the error TEXT for "not deployed" /
"still running" / "not found" / "protected", and the busy sentence contains none of them.
A 500 tells the UI something broke; this is "wait a moment". Now a typed *stacks.RemoveBusyError
matched with errors.As and answered 409, carrying both the Hungarian bytes and the bundle key.
Its test asserts the sentence contains none of the words the text mapping greps for, so the type is
load-bearing rather than decorative. Red-proof seen failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path
sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a
working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404
after. It now falls through to the same settle path with a WARN naming the candidates.
The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve
by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates
returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence.
R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose
file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is
still not diagnosed and R-634 stays open for it.
R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own
sentence - the product already refused this clash for update and for restore. And because `down`
returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying
its label is removed by name with its labels logged, and the answer carries `verified`.
R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down
that destroys them. Two existing tests pin the compose sequence and correctly caught the new step;
their expectations are updated with the reason that the ORDER is the assertion.
R-614: RemoveStack calls ClearUpdateState.
NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done.
Three new sentences, each born as a key in both bundles. Four red-proofs seen failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The controller self-updates daily at 04:30 by default, and after any hub report
once a floor sits above the box. That swap restarts the controller container.
The window proposed for automatic app updates is 02:30-05:00. It contains 04:30.
R-608 — a two-way lock, wired in main.go (stacks never imports selfupdate):
- stacks.Manager.AnyUpdating() -> Updater.SetAppUpdatingCheck, consulted in the
same three places as the existing backupRunning gate.
- Updater.IsUpdateRunning -> Manager.SetSelfUpdatingCheck; UpdatePreflight
refuses `self_updating`.
- MEASURED: the gap was narrower than assumed. The update's `backing-up` phase
already takes the backup single-flight, so that one phase was covered. The
other six were not, and `starting`/`verifying` are where data may have moved.
- The lock must NOT latch: a held app does not block the controller's own
updates, including the release that might fix the hold.
R-609 — the 409 carries `data.reason`, additively. transient (busy, updating,
deploying, migrating, self_updating) vs terminal (held, downgrade). Found while
writing the test: the router refuses a HELD app on its own line before the
preflight, so `held` would have been the one reason missing.
Five red-proofs, each seen to fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Recounted at catalog 18a6d2d8: 66 unique pins — 48 full X.Y.Z, 6 two-part
lines, 4 major lines (10 float), 8 exact versions wearing a variant suffix.
The '23' carried since v0.233.0 matches no definition the catalog supports.
Definition written down beside the number so it can be rechecked.
Also: CONTEXT said the fleet floor was 0.257.0; the hub says 0.259.0.
Comment and doc only — no behaviour change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
MEASURED 2026-09-15 (BIGNIGHT Phase 6): privatebin updated 2.0.5 -> 2.0.6, catalog
reverted to 2.0.5, and the box read „Frissítés elérhető — ma" over an Update that
would have moved the pin BACKWARDS onto a possibly-migrated datadir.
- stacks.CatalogOrder: the comparison gains a fourth answer (Ahead) and moves out of
web, so the badge and UpdatePreflight cannot drift apart.
- The badge: ahead reads „Naprakész"/"Up to date", tag-ok, with a title saying why.
- The refusal: UpdatePreflight returns `downgrade` (409), born as a bundle key; the
API now renders update refusals through errText so it reaches English households.
- Ahead is narrow: every differing service must be orderable AND newer, else Behind.
- Ordering is util.Version.Compare behind a tag normaliser — no second comparator.
- Three red-proofs, each seen to fail.
R-589 was already fixed in v0.258.0; only its register row was stale.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The READ PATH for a second language in `.felhom.yml`. An `i18n: {en: …}` sibling
block inside the same file; `Metadata.For(lang)` merges it FIELD BY FIELD over the
Hungarian, so a missing or blank English field shows the Hungarian one and a
half-translated app is a legal, shippable state.
`For("hu")` is the parsed struct with `I18n` cleared and nothing else — measured
against all 53 real catalog files, copied into `internal/stacks/testdata/catalog/`.
Lists replace whole; every other list is matched by its own key, never by position.
`For` never writes through the receiver: the metadata is the stack manager's, shared
by concurrent requests, and an in-place merge would leak one household's language
into another household's page.
Pages reach catalog copy only through `LocalizeStacks`/`LocalizeStackPtr`/`MetaFor`,
and `TestNoDirectMetaCopyReadOnPages` keeps a named, reasoned allow-list of every
direct `.Meta.<copy>` read in `internal/web` so the NEXT page to read one fails the
suite instead of quietly rendering Hungarian to an English household.
Eight red-proofs. Two of them convicted a hollow TEST rather than the code: a struct
copy shares its slices' backing arrays, so the obvious DeepEqual mutation check
passed a deliberately broken merge; and a one-entry fixture cannot tell key matching
from position matching. Both rewritten, both then seen to fail.
MinAgent: 0.131.0 (unchanged). Older controllers are unaffected — `LoadMetadata`
uses non-strict `yaml.Unmarshal`, so a pre-0.257.0 box drops the whole block.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
179 Hungarian sentences were built deep inside a package with fmt.Errorf and printed by
whoever caught them: too late to translate where they are shown, too early where they are
made. Every one now carries its key across that gap. ZERO Hungarian error literals remain.
util.MsgError does three things at once, each earned:
- Error() is the Hungarian, byte for byte, so every un-converted printer is unchanged;
- errors.Is answers for the kind AND for a wrapped cause (KindErrorf dropped the cause);
- an error ARGUMENT renders recursively, so "formázás sikertelen: %w" translates whole.
A foreign error — restic, docker, ssh, the stdlib — prints verbatim. It is not ours.
76 display sites go through errText, and TestNoErrErrorInPageOutput convicts any that do
not. memoryVerdict returns an error rather than a sentence, so the deploy's 409 and the
household's language come from one value; UpdateRefusal gained a Cause to carry it.
Plurals, one rule, stated once: a key with .one/.other takes its COUNT first. Not a
per-call-site flag — the producer somebody forgot would read "3 app is not running". The
guard caught a real key collision (alert.deadapp.one) the day the rule landed.
TWO DEFECTS FOUND IN MY OWN TOOLING, recorded rather than quietly fixed. The bulk converter
silently dropped multi-line concatenations, damaging 7 producers — and the parity gate could
not see it, because every surviving fragment WAS a real base literal while the CALL had lost
text; two behaviour tests caught it. And the counting script was case-sensitive, so it said
"0 left" while five remained.
MinAgent: 0.131.0 (unchanged). No hub release needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Every Hungarian sentence is byte-identical; each decision now reads a signal set where the message is
made. util.KindErrorf builds the same bytes fmt.Errorf did while carrying a sentinel for errors.Is.
- Deploy status (api/router.go): deployStatusFor() by kind — stacks.ErrAlreadyDeployed (409),
ErrRequiredField / ErrPathMissing / ErrNotEnoughMemory (400). The „kötelező" / „memória" /
"does not exist" / "already deployed" text chain is gone.
- Off-site failure class (backup/offbox.go): ErrOffsiteQuota replaces the „tárhelykeretet" match. The
restic/ssh signatures stay text matches on purpose — that output is not ours and is not translated.
- Alert placement (web/alerts.go): monitor.HealthReport carries WarningKinds parallel to Warnings;
the "not on a separate drive" warning is inline by KIND. The hub report is untouched (builder.go
copies Status/Issues/Warnings only) — pinned by a wire test.
- Stale off-site note (web/handlers.go): settings LastWarningKind + backup.OffboxWarnNoAppsSelected.
The text test survives ONLY for kind == "" (a box whose last run predates 0.251.0) and is removed
when R-570 closes; slice 2 must not translate that producer before then.
Tests (all red-proofed by restoring the pre-fix predicate — see the audit's redproofs.txt):
TestR553_Deploy_DecisionSurvivesWordingChange, TestR553_DeployHandlerUsesTheKind,
TestR553_DeployProducersCarryKindAndKeepTheirWords (through the real DeployStack),
TestR553_OffsiteQuota_{Decision,HeadLine}SurvivesWordingChange, TestR553_OffboxRunRecordsTheKind,
TestR553_StorageWarningsCarryKindsAndKeepTheirWords, TestR553_DiskWarningPlacementSurvivesWordingChange,
TestR553_HubReportWarningsAreUnchangedOnTheWire, TestR553_StaleNote*, TestR553_WarningKindIsPersistedAndCopied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.
R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.
R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.
Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
R-486 (P1): removing an app with its backups KEPT keeps its Tier-2 record,
so the second-drive restore is no longer refused over an intact mirror.
R-484: postgis/pgvector/timescaledb images are Postgres (logical dumps).
R-485: the backup card sizes the recovery unit and the mirror(s).
R-480: a held update's sentence leaves the card once the hold is lifted.
R-477: the update's off-site lookup is one snapshots call, no stats.
R-478: a copy older than this install's deploy does not count.
R-474: "delete backups" deletes the unit, the mirror(s) and the prefs.
Tests and red-proofs per row; evidence in felhom.eu
documentation/audits/v0240-2026-09-13/ and nightly-2026-09-13-adventurelog/.
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.
Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Removal resolves the drive from the app's own app.yaml HDD_PATH (the 07 ~L437
rule), never the global cfg.Paths.HDDPath which no box sets. A data removal
that cannot be resolved, or whose drive is absent, is refused with a typed
RemoveRefusedError -> 409 + exact Hungarian sentence, before compose down, and
the app is kept. SSD app -> hdd_paths_removed: [] never null; missing folders
stated; backup-path refusals reach the response.
15 tests, two red-proofs run (pre-fix fallback -> C fails with err=nil and the
handler 200s; "no drive refuses" -> D fails).
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found by the LIVE validation on demo-hp, not by review. Scenario A passed - a
non-image catalog change reached the pinned app on the real 15-minute cycle - and
that is exactly what exposed the gap: the stored applied-compose.yml is written
when the PIN is written, so the fix landed in the live compose file and not in the
store. The first time the catalog then moved a version, the freeze would have
rendered the pre-fix definition and reverted every fix delivered since - silently
undoing the half of the operator's ruling that says fixes keep flowing.
The equal-images branch now refreshes the store as it delivers. The images cannot
move in that branch by construction, so no version moves and no intent is
rewritten. RenderPlan gains StackDir so the syncer can write it.
TestFixRefreshesTheStoredDefinition asserts both halves: the fix reaches the
store, and it survives the freeze that follows.
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.
The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.
NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.
app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.
The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.
AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.
THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.
+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.
BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.
And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.
Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.
+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no
auto-update, the three lifecycle buttons byte-identical.
Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each
entry carrying ref + repo digest + first-seen timestamp. Written by
Manager.recordInstalledImages after a successful compose up from StartStack,
RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never
from docker-compose.yml: the syncer overwrites a deployed app's compose on a
15-minute cycle and the two disagreed for 25 minutes in the spike's own
measurement. A failed write NEVER refuses the action - the deliberate opposite
of SetDesiredState, because this is an observation and that is an intent. Not
called from StartStackServices (the R-47 DB-only window). Its own docker seam
with a context and a 30s timeout, which neither existing exec helper has.
Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the
recorded ref per service against what the current template pins and returns a
*MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS.
NO RECORD RENDERS NOTHING: absent means unknown and never means current. No
version number reaches the customer and no registry is queried.
Known limitation, filed not hidden: 23 catalog pins float, so those apps can read
Naprakesz when the image behind the tag has moved.
+17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real
RestartStack plus an AST walk of the four call sites. Three companion red-proofs
run and reverted.
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the
R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A
two-container app whose database exits goes unhealthy BECAUSE it cannot reach
that database - so the symptom the dead database causes was what suppressed the
alarm for it. unhealthy is not a down state, so classifyRunStates never marked
the app down and app_start_failed never fired.
Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the
F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich
failure, back through a different door.
Two things moved, and either alone leaves the defect standing: the supervised
test is hoisted above the unhealthy/starting/restarting returns, and "some
members are up" now counts ANY member not in the down bucket. The old guard was
running > 0, which made the R-51 block unreachable in exactly the case it was
written for.
IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy
container is running and folding it in reintroduces the flapping that exclusion
exists to stop. No new state was minted. Only the ORDER changed. The priority
comment was rewritten because it asserted an ordering the code no longer has.
Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they
asserted an unhealthy/starting/restarting member beat an exited peer on
unless-stopped, which pinned the defect as settled behaviour. They keep their
intent with the down member given a benign policy.
R-383. The double-failure message said the previous state's backup EXISTS,
built from the returned path without asking the filesystem - and a missing file
is one of the two ways that rollback fails. undoCopyPhrase now describes the
copy from disk: present, partial, missing (still naming where it should be), or
never written. Zero-length counts as missing.
Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two
halves of R-384 convict independently.
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.
PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.
PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.
PART 4 - measured before theorising, on the live off-site target:
snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
=> 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.
TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.
RED-PROOFS, mutation asserted applied then reverted to 0:
A three template guards dropped (count asserted 3) -> the blank form returned
P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential
DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.
NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
ExportDataMounts lives in delete.go, which reads as a destructive path. IT IS NOT: its single
production caller is the .fab export adapter, and nothing deletes based on its result. The
delete path's own guard, ProtectedHDDPaths, is layout-agnostic by construction -- it protects
BOTH <hdd>/... and <hdd>/felhom-data/... -- so deletion was never affected by the
namespace-root defect. That scope note is now in the function's doc comment, because the file
placement will mislead the next reader exactly as it misled the spec for this change.
Separated into its own commit anyway, so a change to a function whose filename says "delete"
is reviewable on its own.
An empty nsRoot falls back to hddPath -- the pre-R-203 shape -- so any caller not yet updated
keeps working on enrolled drives.
Tests cover both drive kinds and assert the NEGATIVE: no emitted path lies outside the app's
own data roots. Red-proof: leaving the site bare fails the system-drive row, emitting
/mnt/sys_drive/userdata where the canonical root is /mnt/sys_drive/felhom-data/userdata.
appbackup's path helpers take a NAMESPACE ROOT. Five call sites passed a bare DRIVE path.
On an enrolled drive the two coincide, so nothing showed; on the system-data fallback they
differ by exactly the felhom-data segment, and the app then bound a directory the off-site
capture set never looked at -- while the run reported ok. Measured live on demo-hp: the app
wrote to /mnt/sys_drive/userdata/media/books, the capture set looked for
/mnt/sys_drive/felhom-data/userdata/media/books.
THE RULE NOW HAS ONE EXPRESSION. appbackup.NamespaceRootFor / IsEnrolledDrive encode the
drive-kind comparison; backup.Manager.namespaceRoot and stacks.Manager.inGuest delegate to
it. There were already TWO copies and they differed -- the backup package's compared without
filepath.Clean, the stacks package's with it, so a trailing slash from config would have
flipped the mode in one and not the other.
Sites routed through it:
- stacks/deploy.go withPathVars -> ${USERDATA_PATH} (the live defect)
- appexport/fabplan.go + export.go (via a new provider method)
- web/handlers.go FileBrowser mounts (latent: the system drive is
deliberately never a registered StoragePath, so this is the identity today)
ComputeFabBuckets now receives the namespace root, which is what ComputeCaptureSet has always
received -- so the export's classified paths and the backup's capture set describe the same
directories by construction instead of by coincidence.
Tests are table-driven over BOTH drive kinds, because this survived by being invisible on the
kind that already worked. Red-proofs observed: restoring the bare-path call fails the
system-drive row with the two paths differing by /felhom-data; inverting the drive-kind
comparison fails every enrolled row.
R-171 (a regression v0.189.0 introduced, CONFIRMED on hardware before any fix
was written). Replacing isBootOrphan's container-count term with recorded intent
made a drive-gate-stopped app read as a boot orphan: the gate stops apps with
`compose down` (zero containers) and never touches desired_state, because it is
not the customer. Observed on 9201 with the drive held unmounted — the sweep
found and started it, burned both attempts, and handed it to the dead-app alarm.
The write hazard did not materialise (the unbound mountpoint is host-root-owned
and the guest is unprivileged) but that protection is accidental and untested.
New consumer-side seam bootrecon.StartGate, fail-safe (cannot determine ⇒ do not
start), wired in main.go. The rule is not new: the API's startGatedByMissingDrive
already refuses this; the sweep bypassed it.
R-157 mechanism A. The sweep looked once at T+5s, deriving candidates from a
fleet docker was still restoring — three of six hard resets. Now a settle-then-
sweep window: sample every 5s, settled after 3 identical samples, sweep ONCE at
the end; ends on settled or a 50s budget, and the log says which. The budget is
50s because settle+budget+one retry must stay under the 90s dead-app grace — a
test rejected 60s at 95s. A window that overruns emits a LATE RECOVERY warn
rather than the grace being widened to hide it.
Widening the window made two more holders reachable, so the one gate covers all
three: an absent drive, a quiesce, and an in-flight app-data operation — reusing
quiesce.SuppressedStacks() and a new read-only AppStopGuard.HeldStacks().
R-170. shouldRecreateOnBoot now reads desired_state with the identical three-way
table; absent keeps the old hasContainers behaviour exactly. Its comment argued
for the container count and was rewritten. presentStable is untouched. The two
gates' agreement is pinned from both sides against one fixture table.
27/27 packages green; 6 red-proofs observed FAIL then restored.