internal/sockheal: 60 s of refusals (never a timeout, only after Docker answered once) → exit 75 so
Docker's restart policy brings the controller back on the current socket; every 5 min it restarts any
other socket user (traefik) holding an older inode. Measured on 9202: only a docker.socket restart
re-creates the file; dockerd crash / docker.service restart keep it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: the product pins tag@digest, and such images are stored untagged (repo:<none>);
`docker image ls` without -a did not list them, so the retention saw almost no app image. Now `image ls -a`;
an anonymous <none>:<none> entry is never a candidate; the one-time marker is v2 so the corrected sweep runs
once everywhere. Test TestImageRetention_SeesUntaggedDigestPulledImages, red-proofed. 0.284.0/0.284.1 were
never floored.
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: v0.284.0 wired the remove half into DeleteStack only; the app page's Remove runs RemoveStack.
Now RemoveStack reads the app's image repositories before its compose down and runs the retention after. Every
retention pass logs one line (images seen, candidates, deleted), so a pass that kept everything is visible.
Tests TestImageRetention_TheRemoveButtonRunsIt / ADoneUpdateRunsItWithThePrevious, red-proofed. v0.284.0 was never
floored (scratch 9202 only).
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Image retention: after a done/undone guarded Update and at remove, an app's images older than its running
and previous one are deleted — never an image any container, installed compose or installed/previous record
names (box-wide keep set read at delete time); exact id, never forced or pruned; paused while any update runs;
a one-time sweep of catalog app images at the first start. Install hold: an after_install app is installed
behind the setup gate's door and opens when after_install succeeds or the household says it changed the login.
Tests TestImageRetention_* and TestInstallHold_* with red-proofs; parity fixture for the held card.
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Use my kept data / Load consider the off-site snapshot when it is newer than every local copy or
the only one; the unit is downloaded alone, judged (drive, data, recorded data version) and only
then restored. The page names the copy and its date.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A drive move persisted through the restore's fresh app.yaml write and dropped the pin: the syncer
then copied the catalog verbatim and the next start jumped the app past its ladder (R-700).
persistDriveFlip now changes HDD_PATH and nothing else. The restore's write carries the life
records (conversion copies, desired_state, update history) from the app.yaml it replaces, and a
second conversion no longer overwrites the first kept copy's record (R-697).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A just-installed app's unit, captured by the status refresh before any backup, satisfied
the precondition on its manifest time; tandoor's PostgreSQL was converted with no backup
of its database. Listed still; never a copy on Tier 1 or Tier 2. Red-proof RP6.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (night 2026-09-24 Part B): the sync rendered the ladder's newest
tested digest into a RUNNING app's compose, so the next restart would pull a new image
with no backup and no undo. stacks.CarryDigests keeps the running digest for an
installed app; a fresh install still takes the tested digest. Red-proofed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202 (romm): .felhom.yml flows into the stack dir on every
catalog sync, so "the old .felhom.yml" saved at update time was already the
new one, and the serving old version was judged with the new probe.
New record applied-meta/.felhom.yml, written whenever a version is pinned
(deploy, adoption, pin advance) and put back by the undo, like
applied-compose.yml. The fixture now places the new file at sync time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found live on 9202: the periodic probe (current .felhom.yml, new port) flips
the app to unhealthy, and the update's health wait probed only 'running'
apps - so the undo's old probe was never asked and a serving old version was
judged "did not start". With the undo's override, an unhealthy app is probed
and the old check decides; never settled on container state.
New seam probeRunFn; the test drives the real wait loop and reproduces the
live message when the fix is switched off.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.
- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
(pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.
MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS