Use my kept data / Load consider the off-site snapshot when it is newer than every local copy or
the only one; the unit is downloaded alone, judged (drive, data, recorded data version) and only
then restored. The page names the copy and its date.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A drive move persisted through the restore's fresh app.yaml write and dropped the pin: the syncer
then copied the catalog verbatim and the next start jumped the app past its ladder (R-700).
persistDriveFlip now changes HDD_PATH and nothing else. The restore's write carries the life
records (conversion copies, desired_state, update history) from the app.yaml it replaces, and a
second conversion no longer overwrites the first kept copy's record (R-697).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A just-installed app's unit, captured by the status refresh before any backup, satisfied
the precondition on its manifest time; tandoor's PostgreSQL was converted with no backup
of its database. Listed still; never a copy on Tier 1 or Tier 2. Red-proof RP6.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
An install over an app's kept drive folder (appdata/<app> non-empty) asks the household:
"use my kept data" (a load from the newest copy of THIS drive's install, own unit or
second-drive mirror, then the template's after_load) or "start fresh" (the folder is
renamed into <drive>/kept/<app>/<date>/ with the removed app's unit; nothing deleted).
The install API answers 409 kept_data_choice until one is chosen; DeployStack refuses
too. New page Megorzott adatok / Kept data (/kept-data): Load / Look / Delete (typed
confirmation, the only deletion of kept data). FileBrowser gets a read-only source.
The drive-full warning names the kept folders. <drive>/kept is protected and outside
every backup leg.
R-690: the removed-app restore (R-487) never found a unit on a DATA drive — it asked
GetStackComposePath (true for every catalog app) and restored nextcloud with no env.
Now isStackDeployed; pinned with a production-shaped provider.
Red-proofs: audits/night-2026-09-26/E/redproofs/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A step whose ladder entry carries engine_conversion {service, engine, from, to}
converts the database: the old engine alone, the check (owners, roles,
extensions, per-table row counts), pg_dumpall validated by its completion line,
the volume emptied only after the undo copy's marker is validated again, the new
engine alone, the load with ON_ERROR_STOP, the check again + PG_VERSION. Any
failure goes to the existing undo; a restart during converting is undone.
A PostgreSQL major move without the mark is refused before anything moves.
The old datadir's copy is kept until a backup is proven after the conversion.
17 tests, 9 red-proofs (audits/night-2026-09-26/B/).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (night 2026-09-24 Part B): the sync rendered the ladder's newest
tested digest into a RUNNING app's compose, so the next restart would pull a new image
with no backup and no undo. stacks.CarryDigests keeps the running digest for an
installed app; a fresh install still takes the tested digest. Red-proofed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (romm): .felhom.yml flows into the stack dir on every
catalog sync, so "the old .felhom.yml" saved at update time was already the
new one, and the serving old version was judged with the new probe.
New record applied-meta/.felhom.yml, written whenever a version is pinned
(deploy, adoption, pin advance) and put back by the undo, like
applied-compose.yml. The fixture now places the new file at sync time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: the periodic probe (current .felhom.yml, new port) flips
the app to unhealthy, and the update's health wait probed only 'running'
apps - so the undo's old probe was never asked and a serving old version was
judged "did not start". With the undo's override, an unhealthy app is probed
and the old check decides; never settled on container state.
New seam probeRunFn; the test drives the real wait loop and reproduces the
live message when the fix is switched off.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.
- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
(pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.
MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Caught by the live proof on 9202, not by a test. The v0.262.0 guard fired exactly right and
answered HTTP 500: router.go maps remove errors by grepping the error TEXT for "not deployed" /
"still running" / "not found" / "protected", and the busy sentence contains none of them.
A 500 tells the UI something broke; this is "wait a moment". Now a typed *stacks.RemoveBusyError
matched with errors.As and answered 409, carrying both the Hungarian bytes and the bundle key.
Its test asserts the sentence contains none of the words the text mapping greps for, so the type is
load-bearing rather than decorative. Red-proof seen failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path
sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a
working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404
after. It now falls through to the same settle path with a WARN naming the candidates.
The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve
by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates
returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence.
R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose
file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is
still not diagnosed and R-634 stays open for it.
R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own
sentence - the product already refused this clash for update and for restore. And because `down`
returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying
its label is removed by name with its labels logged, and the answer carries `verified`.
R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down
that destroys them. Two existing tests pin the compose sequence and correctly caught the new step;
their expectations are updated with the reason that the ORDER is the assertion.
R-614: RemoveStack calls ClearUpdateState.
NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done.
Three new sentences, each born as a key in both bundles. Four red-proofs seen failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The controller self-updates daily at 04:30 by default, and after any hub report
once a floor sits above the box. That swap restarts the controller container.
The window proposed for automatic app updates is 02:30-05:00. It contains 04:30.
R-608 — a two-way lock, wired in main.go (stacks never imports selfupdate):
- stacks.Manager.AnyUpdating() -> Updater.SetAppUpdatingCheck, consulted in the
same three places as the existing backupRunning gate.
- Updater.IsUpdateRunning -> Manager.SetSelfUpdatingCheck; UpdatePreflight
refuses `self_updating`.
- MEASURED: the gap was narrower than assumed. The update's `backing-up` phase
already takes the backup single-flight, so that one phase was covered. The
other six were not, and `starting`/`verifying` are where data may have moved.
- The lock must NOT latch: a held app does not block the controller's own
updates, including the release that might fix the hold.
R-609 — the 409 carries `data.reason`, additively. transient (busy, updating,
deploying, migrating, self_updating) vs terminal (held, downgrade). Found while
writing the test: the router refuses a HELD app on its own line before the
preflight, so `held` would have been the one reason missing.
Five red-proofs, each seen to fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Recounted at catalog 18a6d2d8: 66 unique pins — 48 full X.Y.Z, 6 two-part
lines, 4 major lines (10 float), 8 exact versions wearing a variant suffix.
The '23' carried since v0.233.0 matches no definition the catalog supports.
Definition written down beside the number so it can be rechecked.
Also: CONTEXT said the fleet floor was 0.257.0; the hub says 0.259.0.
Comment and doc only — no behaviour change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
MEASURED 2026-09-15 (BIGNIGHT Phase 6): privatebin updated 2.0.5 -> 2.0.6, catalog
reverted to 2.0.5, and the box read „Frissítés elérhető — ma" over an Update that
would have moved the pin BACKWARDS onto a possibly-migrated datadir.
- stacks.CatalogOrder: the comparison gains a fourth answer (Ahead) and moves out of
web, so the badge and UpdatePreflight cannot drift apart.
- The badge: ahead reads „Naprakész"/"Up to date", tag-ok, with a title saying why.
- The refusal: UpdatePreflight returns `downgrade` (409), born as a bundle key; the
API now renders update refusals through errText so it reaches English households.
- Ahead is narrow: every differing service must be orderable AND newer, else Behind.
- Ordering is util.Version.Compare behind a tag normaliser — no second comparator.
- Three red-proofs, each seen to fail.
R-589 was already fixed in v0.258.0; only its register row was stale.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS