Commit Graph

135 Commits

Author SHA1 Message Date
admin 6be6c53e29 v0.283.0: apps go off-site by themselves (decision 50) with a size warning; a Stop holds during a backup (R-721); page slips (R-724/R-725)
gates / gates (push) Successful in 25s
Red-proofs RP31-RP38. MinAgent 0.131.0 (unchanged).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:31:55 +02:00
admin 4110da50e9 v0.282.0: after_setup (the app's own sign-up switch), close sign-up now (decision 49), probes read lists + done_status (R-715)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-29 16:19:12 +02:00
admin 149467c795 v0.281.0: 'Done' asks the probe first; sign-up closed after the first admin (decision 47); R-713 code-bound values refused, ${NAME|base64}
gates / gates (push) Successful in 23s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-29 09:57:58 +02:00
admin 7fa8768cfd v0.280.0: the setup gate (decision 46); R-710 'I changed it' + absent-record window; R-709 password fields off the page; password:N:special generator
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-29 08:51:46 +02:00
admin 0c702f834a v0.279.0: after_install (decision 45), known default logins on the page, Part D empty-backup alarm, night chain (R-705), R-706
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-28 18:38:30 +02:00
admin 820e8efde1 v0.276.0: a restore and a drive move keep the app's records (R-697, R-700)
gates / gates (push) Successful in 26s
A drive move persisted through the restore's fresh app.yaml write and dropped the pin: the syncer
then copied the catalog verbatim and the next start jumped the app past its ladder (R-700).
persistDriveFlip now changes HDD_PATH and nothing else. The restore's write carries the life
records (conversion copies, desired_state, update history) from the app.yaml it replaces, and a
second conversion no longer overwrites the first kept copy's record (R-697).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 17:59:04 +02:00
admin b6810f14ff v0.275.0: a backup's data and its version travel together (R-696, 07 §6.6, D4 option A); R-695, R-691, R-694
gates / gates (push) Successful in 23s
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-26 10:35:22 +02:00
admin 064f23896a kept data: the list names a leftover by the app that binds it through HDD_PATH, never the file browser (R-692, found live on 9202)
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 13:51:26 +02:00
admin 43e99d160c kept data: the choice at reinstall, the list, the read-only view, the load (09 decision 36); R-690 fixed
gates / gates (push) Successful in 26s
An install over an app's kept drive folder (appdata/<app> non-empty) asks the household:
"use my kept data" (a load from the newest copy of THIS drive's install, own unit or
second-drive mirror, then the template's after_load) or "start fresh" (the folder is
renamed into <drive>/kept/<app>/<date>/ with the removed app's unit; nothing deleted).
The install API answers 409 kept_data_choice until one is chosen; DeployStack refuses
too. New page Megorzott adatok / Kept data (/kept-data): Load / Look / Delete (typed
confirmation, the only deletion of kept data). FileBrowser gets a read-only source.
The drive-full warning names the kept folders. <drive>/kept is protected and outside
every backup leg.

R-690: the removed-app restore (R-487) never found a unit on a DATA drive — it asked
GetStackComposePath (true for every catalog app) and restored nextcloud with no env.
Now isStackDeployed; pinned with a production-shaped provider.

Red-proofs: audits/night-2026-09-26/E/redproofs/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 13:32:30 +02:00
admin 1c5dd07604 stacks: R-687 — an empty leg reports steps [], and a taken files_may_change step names its whole copy
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 13:24:19 +02:00
admin b6399fdf0b stacks: the recovery line names CONVERTING, not UNDOING, for a restart during the conversion (found live on 9202)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 13:18:56 +02:00
admin 2caae38a71 stacks: the box converts a PostgreSQL major as a guarded-update step (09 6.4 part 10, decisions 35/37/38)
gates / gates (push) Successful in 25s
A step whose ladder entry carries engine_conversion {service, engine, from, to}
converts the database: the old engine alone, the check (owners, roles,
extensions, per-table row counts), pg_dumpall validated by its completion line,
the volume emptied only after the undo copy's marker is validated again, the new
engine alone, the load with ON_ERROR_STOP, the check again + PG_VERSION. Any
failure goes to the existing undo; a restart during converting is undone.
A PostgreSQL major move without the mark is refused before anything moves.
The old datadir's copy is kept until a backup is proven after the conversion.
17 tests, 9 red-proofs (audits/night-2026-09-26/B/).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 12:57:03 +02:00
admin 44ae4dea70 controller v0.272.0: the backup page says when a whole-box backup does not fit (R-685); R-671, R-670, R-677
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 11:19:27 +02:00
admin 9cf13a3add controller v0.271.0: automatic app updates — the update leg after the off-site copy, the backup gate waits, the switch (09 6.4 part 7; R-680, R-678, R-643)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:07:24 +02:00
admin 504eae018b v0.270.0: no update for a current app (R-679); an interrupted install is reported (R-681); a restore brings back the pinned version's health check (R-669); R-674
gates / gates (push) Successful in 24s
Five red-proofs. MinAgent 0.131.0 unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:49:43 +02:00
admin 5b1b191ffc v0.269.1: an installed app keeps the image digest it runs until a guarded Update moves it
gates / gates (push) Successful in 26s
Found live on 9202 (night 2026-09-24 Part B): the sync rendered the ladder's newest
tested digest into a RUNNING app's compose, so the next restart would pull a new image
with no backup and no undo. stacks.CarryDigests keeps the running digest for an
installed app; a fresh install still takes the tested digest. Red-proofed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 13:01:31 +02:00
admin 3c6b49b31c controller v0.269.0: whole restore from the second drive; crash loops stopped; exact image digests; steps judged by their own .felhom.yml (decisions 26-28, R-661 R-666 R-667 R-668 R-664 R-665 R-662, 09 6.4 part 6)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 12:18:39 +02:00
admin 206b0357d1 controller v0.268.0: the undo finds volumes by definition; a held app names only a whole copy; one press = one tested step (R-658, R-659, R-660, R-651; 09 §6.4 part 5)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 08:15:27 +02:00
admin 80e6ad8c47 controller v0.267.0: tests off DooPlex's Docker, cut-off copies refused, two pages true
gates / gates (push) Successful in 26s
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 20:25:28 +02:00
admin 964ae7538f controller v0.266.0: a failed install removes what it started (R-649, operator ruling)
gates / gates (push) Successful in 27s
compose down (volumes kept) before the record reads not deployed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 18:06:29 +02:00
admin 0054d4bd69 controller v0.265.0: R-634 cause fixed, held apps say so, OOM storm alarm, R-647 leftovers
gates / gates (push) Successful in 27s
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 17:16:57 +02:00
admin bc278944a3 controller v0.264.0: the household is told when an update is undone or held, in its language
gates / gates (push) Successful in 25s
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 13:51:21 +02:00
admin 2cd66663f3 v0.263.2: the undo keeps the probe of the pinned version (R-637)
gates / gates (push) Successful in 25s
Found live on 9202 (romm): .felhom.yml flows into the stack dir on every
catalog sync, so "the old .felhom.yml" saved at update time was already the
new one, and the serving old version was judged with the new probe.

New record applied-meta/.felhom.yml, written whenever a version is pinned
(deploy, adoption, pin advance) and put back by the undo, like
applied-compose.yml. The fixture now places the new file at sync time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 11:53:53 +02:00
admin 5d38573a1a v0.263.1: the undo asks the old probe even on an app marked unhealthy (R-637)
gates / gates (push) Successful in 27s
Found live on 9202: the periodic probe (current .felhom.yml, new port) flips
the app to unhealthy, and the update's health wait probed only 'running'
apps - so the undo's old probe was never asked and a serving old version was
judged "did not start". With the undo's override, an unhealthy app is probed
and the old check decides; never settled on container state.

New seam probeRunFn; the test drives the real wait loop and reproduces the
live message when the fix is switched off.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 11:35:25 +02:00
admin 8fc2b4a1a9 v0.263.0: a failed update puts the app back by itself (09 decision 15, R-637)
gates / gates (push) Successful in 26s
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.

- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
  (pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.

MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 11:12:49 +02:00
admin 3d41758be0 v0.262.1: the busy refusal is a CONFLICT, not a server error (R-633)
gates / gates (push) Successful in 24s
Caught by the live proof on 9202, not by a test. The v0.262.0 guard fired exactly right and
answered HTTP 500: router.go maps remove errors by grepping the error TEXT for "not deployed" /
"still running" / "not found" / "protected", and the busy sentence contains none of them.

A 500 tells the UI something broke; this is "wait a moment". Now a typed *stacks.RemoveBusyError
matched with errors.As and answered 409, carrying both the Hungarian bytes and the bundle key.

Its test asserts the sentence contains none of the words the text mapping greps for, so the type is
load-bearing rather than decorative. Red-proof seen failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 21:48:01 +02:00
admin b793638484 v0.262.0: six defects two drill nights found in the update, remove and hold paths
gates / gates (push) Successful in 26s
R-630 (P1): waitUpdateHealthy kept the probe inside `if hc != nil && len(hc.Checks) > 0`, and when
findProbeContainer returned "" its else set last="no probe container" and LOOPED - the settle path
sat in the outer else, unreachable. So verifying could only time out and failAndHold then stopped a
working app. Measured on paperless-ngx: three containers healthy, failed at +313.0s, front door 404
after. It now falls through to the same settle path with a WARN naming the candidates.

The probe target is decidable now: HealthCheckConfig.Container plus findProbeContainerMeta resolve
by exact stack name -> explicit container -> a UNIQUE prefix -> nothing with the candidates
returned. The old rule took the FIRST prefix match. A skipped stack records why instead of silence.

R-634 (half): RemoveStack refused on the !Deployed FLAG while the machine had containers, a compose
file and an app.yaml. It now asks whether anything EXISTS. The mechanism producing the bad record is
still not diagnosed and R-634 stays open for it.

R-633/R-626: RemoveStack consults UpdateGuards.Busy and IsUpdating and refuses with the app's own
sentence - the product already refused this clash for update and for restore. And because `down`
returning 0 is a request not a result, the project is watched for 25s afterwards, anything carrying
its label is removed by name with its labels logged, and the answer carries `verified`.

R-621: failAndHold writes compose logs --tail 400 into <stackdir>/hold-logs/<ts>/ BEFORE the down
that destroys them. Two existing tests pin the compose sequence and correctly caught the new step;
their expectations are updated with the reason that the ORDER is the assertion.

R-614: RemoveStack calls ClearUpdateState.

NOT in this release: R-625 (a held app still renders an Update button). Named, not half-done.

Three new sentences, each born as a key in both bundles. Four red-proofs seen failing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 20:46:50 +02:00
admin 811f75736e v0.261.0 — the controller no longer swaps itself out from under an app update (R-608, R-609)
gates / gates (push) Successful in 23s
The controller self-updates daily at 04:30 by default, and after any hub report
once a floor sits above the box. That swap restarts the controller container.
The window proposed for automatic app updates is 02:30-05:00. It contains 04:30.

R-608 — a two-way lock, wired in main.go (stacks never imports selfupdate):
- stacks.Manager.AnyUpdating() -> Updater.SetAppUpdatingCheck, consulted in the
  same three places as the existing backupRunning gate.
- Updater.IsUpdateRunning -> Manager.SetSelfUpdatingCheck; UpdatePreflight
  refuses `self_updating`.
- MEASURED: the gap was narrower than assumed. The update's `backing-up` phase
  already takes the backup single-flight, so that one phase was covered. The
  other six were not, and `starting`/`verifying` are where data may have moved.
- The lock must NOT latch: a held app does not block the controller's own
  updates, including the release that might fix the hold.

R-609 — the 409 carries `data.reason`, additively. transient (busy, updating,
deploying, migrating, self_updating) vs terminal (held, downgrade). Found while
writing the test: the router refuses a HELD app on its own line before the
preflight, so `held` would have been the one reason missing.

Five red-proofs, each seen to fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 14:23:59 +02:00
admin d0d431b42b Correct a stale count repeated in four places: 10 floating pins, not 23
gates / gates (push) Successful in 25s
Recounted at catalog 18a6d2d8: 66 unique pins — 48 full X.Y.Z, 6 two-part
lines, 4 major lines (10 float), 8 exact versions wearing a variant suffix.
The '23' carried since v0.233.0 matches no definition the catalog supports.
Definition written down beside the number so it can be rechecked.

Also: CONTEXT said the fleet floor was 0.257.0; the hub says 0.259.0.

Comment and doc only — no behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 13:02:48 +02:00
admin 8f8a64cad7 v0.260.0 — a box ahead of the catalog reads „Naprakész", and the pin never moves backwards (R-524)
gates / gates (push) Successful in 24s
MEASURED 2026-09-15 (BIGNIGHT Phase 6): privatebin updated 2.0.5 -> 2.0.6, catalog
reverted to 2.0.5, and the box read „Frissítés elérhető — ma" over an Update that
would have moved the pin BACKWARDS onto a possibly-migrated datadir.

- stacks.CatalogOrder: the comparison gains a fourth answer (Ahead) and moves out of
  web, so the badge and UpdatePreflight cannot drift apart.
- The badge: ahead reads „Naprakész"/"Up to date", tag-ok, with a title saying why.
- The refusal: UpdatePreflight returns `downgrade` (409), born as a bundle key; the
  API now renders update refusals through errText so it reaches English households.
- Ahead is narrow: every differing service must be orderable AND newer, else Behind.
- Ordering is util.Version.Compare behind a tag normaliser — no second comparator.
- Three red-proofs, each seen to fail.

R-589 was already fixed in v0.258.0; only its register row was stale.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 12:48:55 +02:00
admin 60f0a86bd4 v0.257.0: the app catalog can speak English (R-560, slice 5 Part A)
gates / gates (push) Successful in 25s
The READ PATH for a second language in `.felhom.yml`. An `i18n: {en: …}` sibling
block inside the same file; `Metadata.For(lang)` merges it FIELD BY FIELD over the
Hungarian, so a missing or blank English field shows the Hungarian one and a
half-translated app is a legal, shippable state.

`For("hu")` is the parsed struct with `I18n` cleared and nothing else — measured
against all 53 real catalog files, copied into `internal/stacks/testdata/catalog/`.
Lists replace whole; every other list is matched by its own key, never by position.
`For` never writes through the receiver: the metadata is the stack manager's, shared
by concurrent requests, and an in-place merge would leak one household's language
into another household's page.

Pages reach catalog copy only through `LocalizeStacks`/`LocalizeStackPtr`/`MetaFor`,
and `TestNoDirectMetaCopyReadOnPages` keeps a named, reasoned allow-list of every
direct `.Meta.<copy>` read in `internal/web` so the NEXT page to read one fails the
suite instead of quietly rendering Hungarian to an English household.

Eight red-proofs. Two of them convicted a hollow TEST rather than the code: a struct
copy shares its slices' backing arrays, so the obvious DeepEqual mutation check
passed a deliberately broken merge; and a one-entry fixture cannot tell key matching
from position matching. Both rewritten, both then seen to fail.

MinAgent: 0.131.0 (unchanged). Older controllers are unaffected — `LoadMetadata`
uses non-strict `yaml.Unmarshal`, so a pre-0.257.0 box drops the whole block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-20 14:31:52 +02:00
admin 7c05b59708 v0.253.0 — errors carry the key of the sentence they are (R-557 slice 2 release B)
gates / gates (push) Successful in 24s
179 Hungarian sentences were built deep inside a package with fmt.Errorf and printed by
whoever caught them: too late to translate where they are shown, too early where they are
made. Every one now carries its key across that gap. ZERO Hungarian error literals remain.

util.MsgError does three things at once, each earned:
  - Error() is the Hungarian, byte for byte, so every un-converted printer is unchanged;
  - errors.Is answers for the kind AND for a wrapped cause (KindErrorf dropped the cause);
  - an error ARGUMENT renders recursively, so "formázás sikertelen: %w" translates whole.
A foreign error — restic, docker, ssh, the stdlib — prints verbatim. It is not ours.

76 display sites go through errText, and TestNoErrErrorInPageOutput convicts any that do
not. memoryVerdict returns an error rather than a sentence, so the deploy's 409 and the
household's language come from one value; UpdateRefusal gained a Cause to carry it.

Plurals, one rule, stated once: a key with .one/.other takes its COUNT first. Not a
per-call-site flag — the producer somebody forgot would read "3 app is not running". The
guard caught a real key collision (alert.deadapp.one) the day the rule landed.

TWO DEFECTS FOUND IN MY OWN TOOLING, recorded rather than quietly fixed. The bulk converter
silently dropped multi-line concatenations, damaging 7 producers — and the parity gate could
not see it, because every surviving fragment WAS a real base literal while the CALL had lost
text; two behaviour tests caught it. And the counting script was case-sensitive, so it said
"0 left" while five remained.

MinAgent: 0.131.0 (unchanged). No hub release needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 11:44:30 +02:00
admin c00fed6db3 R-553: four decisions stop reading their own Hungarian words (sites 1-4)
Every Hungarian sentence is byte-identical; each decision now reads a signal set where the message is
made. util.KindErrorf builds the same bytes fmt.Errorf did while carrying a sentinel for errors.Is.

- Deploy status (api/router.go): deployStatusFor() by kind — stacks.ErrAlreadyDeployed (409),
  ErrRequiredField / ErrPathMissing / ErrNotEnoughMemory (400). The „kötelező" / „memória" /
  "does not exist" / "already deployed" text chain is gone.
- Off-site failure class (backup/offbox.go): ErrOffsiteQuota replaces the „tárhelykeretet" match. The
  restic/ssh signatures stay text matches on purpose — that output is not ours and is not translated.
- Alert placement (web/alerts.go): monitor.HealthReport carries WarningKinds parallel to Warnings;
  the "not on a separate drive" warning is inline by KIND. The hub report is untouched (builder.go
  copies Status/Issues/Warnings only) — pinned by a wire test.
- Stale off-site note (web/handlers.go): settings LastWarningKind + backup.OffboxWarnNoAppsSelected.
  The text test survives ONLY for kind == "" (a box whose last run predates 0.251.0) and is removed
  when R-570 closes; slice 2 must not translate that producer before then.

Tests (all red-proofed by restoring the pre-fix predicate — see the audit's redproofs.txt):
TestR553_Deploy_DecisionSurvivesWordingChange, TestR553_DeployHandlerUsesTheKind,
TestR553_DeployProducersCarryKindAndKeepTheirWords (through the real DeployStack),
TestR553_OffsiteQuota_{Decision,HeadLine}SurvivesWordingChange, TestR553_OffboxRunRecordsTheKind,
TestR553_StorageWarningsCarryKindsAndKeepTheirWords, TestR553_DiskWarningPlacementSurvivesWordingChange,
TestR553_HubReportWarningsAreUnchangedOnTheWire, TestR553_StaleNote*, TestR553_WarningKindIsPersistedAndCopied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 20:58:45 +02:00
admin 2f8ff2414c v0.244.0: the backup page stops promising what it does not hold (R-537/R-538/R-536)
gates / gates (push) Successful in 17s
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.

R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.

R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.

Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:55:55 +02:00
admin 843b319f35 v0.243.0: FileBrowser generated admin password (R-513); per-tier whole-guest backup truth (R-517); skip absent-storage tiers (R-518); OOM-killed worker visible (R-514)
gates / gates (push) Successful in 14s
MinAgent: 0.131.0

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:12:08 +02:00
admin d698ce343b controller v0.242.0: a removed app is listed with its kept backup; five small ones (R-487 R-491 R-490 R-489 R-476 R-456)
gates / gates (push) Successful in 14s
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
2026-09-13 22:50:18 +02:00
admin bdcbd50b42 controller v0.240.0: seven defects from the any-tier proof and the first nightly rotation
gates / gates (push) Successful in 13s
R-486 (P1): removing an app with its backups KEPT keeps its Tier-2 record,
so the second-drive restore is no longer refused over an intact mirror.
R-484: postgis/pgvector/timescaledb images are Postgres (logical dumps).
R-485: the backup card sizes the recovery unit and the mirror(s).
R-480: a held update's sentence leaves the card once the hold is lifted.
R-477: the update's off-site lookup is one snapshots call, no stats.
R-478: a copy older than this install's deploy does not count.
R-474: "delete backups" deletes the unit, the mirror(s) and the prefs.

Tests and red-proofs per row; evidence in felhom.eu
documentation/audits/v0240-2026-09-13/ and nightly-2026-09-13-adventurelog/.
2026-09-13 19:26:50 +02:00
admin b93c1543da controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.

Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 17:16:24 +02:00
admin 0d402f711d v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 11:41:31 +02:00
admin 42a73e667a v0.236.0: "delete my data too" deletes the data, or says that it could not (R-442)
gates / gates (push) Successful in 13s
Removal resolves the drive from the app's own app.yaml HDD_PATH (the 07 ~L437
rule), never the global cfg.Paths.HDDPath which no box sets. A data removal
that cannot be resolved, or whose drive is absent, is refused with a typed
RemoveRefusedError -> 409 + exact Hungarian sentence, before compose down, and
the app is kept. SSD app -> hdd_paths_removed: [] never null; missing folders
stated; backup-path refusals reach the response.

15 tests, two red-proofs run (pre-fix fallback -> C fails with err=nil and the
handler 200s; "no drive refuses" -> D fails).

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 08:52:25 +02:00
admin 2a56f557d0 v0.235.0: a delivered fix must also refresh the stored definition
gates / gates (push) Successful in 14s
Found by the LIVE validation on demo-hp, not by review. Scenario A passed - a
non-image catalog change reached the pinned app on the real 15-minute cycle - and
that is exactly what exposed the gap: the stored applied-compose.yml is written
when the PIN is written, so the fix landed in the live compose file and not in the
store. The first time the catalog then moved a version, the freeze would have
rendered the pre-fix definition and reverted every fix delivered since - silently
undoing the half of the operator's ruling that says fixes keep flowing.

The equal-images branch now refreshes the store as it delivers. The images cannot
move in that branch by construction, so no version moves and no intent is
rewritten. RenderPlan gains StackDir so the syncer can write it.

TestFixRefreshesTheStoredDefinition asserts both halves: the fix reaches the
store, and it survives the freeze that follows.
2026-09-06 10:04:21 +02:00
admin 8a0e0a59ad v0.235.0: freeze the version, keep the fixes flowing (operator ruling 2026-09-06)
gates / gates (push) Successful in 12s
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.

The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.

NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.

app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.

The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.

AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.

THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.

+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
2026-09-06 09:45:34 +02:00
admin 38d28b5b62 v0.234.0: seed installed_images at startup, so the label appears on an app nobody touched
gates / gates (push) Successful in 13s
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.

BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.

And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.

Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.

+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
2026-09-03 11:56:43 +02:00
admin 8025304acc v0.233.0: record what each compose service actually installed, and badge whether it is current
gates / gates (push) Successful in 12s
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no
auto-update, the three lifecycle buttons byte-identical.

Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each
entry carrying ref + repo digest + first-seen timestamp. Written by
Manager.recordInstalledImages after a successful compose up from StartStack,
RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never
from docker-compose.yml: the syncer overwrites a deployed app's compose on a
15-minute cycle and the two disagreed for 25 minutes in the spike's own
measurement. A failed write NEVER refuses the action - the deliberate opposite
of SetDesiredState, because this is an observation and that is an intent. Not
called from StartStackServices (the R-47 DB-only window). Its own docker seam
with a context and a 30s timeout, which neither existing exec helper has.

Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the
recorded ref per service against what the current template pins and returns a
*MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS.
NO RECORD RENDERS NOTHING: absent means unknown and never means current. No
version number reaches the customer and no registry is queried.

Known limitation, filed not hidden: 23 catalog pins float, so those apps can read
Naprakesz when the image behind the tag has moved.

+17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real
RestartStack plus an AST walk of the four call sites. Three companion red-proofs
run and reverted.
2026-09-02 20:18:01 +02:00
admin 5da11c4480 v0.222.0: ask whether a supervised member is DEAD before whether one is UNHEALTHY (R-384), and stop promising an undo copy nobody looked for (R-383)
gates / gates (push) Successful in 11s
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the
R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A
two-container app whose database exits goes unhealthy BECAUSE it cannot reach
that database - so the symptom the dead database causes was what suppressed the
alarm for it. unhealthy is not a down state, so classifyRunStates never marked
the app down and app_start_failed never fired.

Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the
F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich
failure, back through a different door.

Two things moved, and either alone leaves the defect standing: the supervised
test is hoisted above the unhealthy/starting/restarting returns, and "some
members are up" now counts ANY member not in the down bucket. The old guard was
running > 0, which made the R-51 block unreachable in exactly the case it was
written for.

IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy
container is running and folding it in reintroduces the flapping that exclusion
exists to stop. No new state was minted. Only the ORDER changed. The priority
comment was rewritten because it asserted an ordering the code no longer has.

Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they
asserted an unhealthy/starting/restarting member beat an exited peer on
unless-stopped, which pinned the defect as settled behaviour. They keep their
intent with the down member given a benign policy.

R-383. The double-failure message said the previous state's backup EXISTS,
built from the returned path without asking the filesystem - and a missing file
is one of the two ways that rollback fails. undoCopyPhrase now describes the
copy from disk: present, partial, missing (still naming where it should be), or
never written. Zero-length counts as missing.

Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two
halves of R-384 convict independently.
2026-08-23 07:25:00 +02:00
admin f94543ee5c v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.

PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.

PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.

PART 4 - measured before theorising, on the live off-site target:
  snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
  => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.

TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.

RED-PROOFS, mutation asserted applied then reverted to 0:
  A   three template guards dropped (count asserted 3) -> the blank form returned
  P4  inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential

DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.

NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
2026-08-21 21:29:01 +02:00
admin a96c3d9473 R-203: the export-mount resolver takes the namespace root too (its own commit)
gates / gates (push) Successful in 9s
ExportDataMounts lives in delete.go, which reads as a destructive path. IT IS NOT: its single
production caller is the .fab export adapter, and nothing deletes based on its result. The
delete path's own guard, ProtectedHDDPaths, is layout-agnostic by construction -- it protects
BOTH <hdd>/... and <hdd>/felhom-data/... -- so deletion was never affected by the
namespace-root defect. That scope note is now in the function's doc comment, because the file
placement will mislead the next reader exactly as it misled the spec for this change.

Separated into its own commit anyway, so a change to a function whose filename says "delete"
is reviewable on its own.

An empty nsRoot falls back to hddPath -- the pre-R-203 shape -- so any caller not yet updated
keeps working on enrolled drives.

Tests cover both drive kinds and assert the NEGATIVE: no emitted path lies outside the app's
own data roots. Red-proof: leaving the site bare fails the system-drive row, emitting
/mnt/sys_drive/userdata where the canonical root is /mnt/sys_drive/felhom-data/userdata.
2026-08-04 18:21:17 +02:00
admin 73efb091d9 R-203: the app and its backup look in the same directory — one resolver, every caller
gates / gates (push) Successful in 9s
appbackup's path helpers take a NAMESPACE ROOT. Five call sites passed a bare DRIVE path.
On an enrolled drive the two coincide, so nothing showed; on the system-data fallback they
differ by exactly the felhom-data segment, and the app then bound a directory the off-site
capture set never looked at -- while the run reported ok. Measured live on demo-hp: the app
wrote to /mnt/sys_drive/userdata/media/books, the capture set looked for
/mnt/sys_drive/felhom-data/userdata/media/books.

THE RULE NOW HAS ONE EXPRESSION. appbackup.NamespaceRootFor / IsEnrolledDrive encode the
drive-kind comparison; backup.Manager.namespaceRoot and stacks.Manager.inGuest delegate to
it. There were already TWO copies and they differed -- the backup package's compared without
filepath.Clean, the stacks package's with it, so a trailing slash from config would have
flipped the mode in one and not the other.

Sites routed through it:
  - stacks/deploy.go withPathVars -> ${USERDATA_PATH}   (the live defect)
  - appexport/fabplan.go + export.go                     (via a new provider method)
  - web/handlers.go FileBrowser mounts                   (latent: the system drive is
    deliberately never a registered StoragePath, so this is the identity today)

ComputeFabBuckets now receives the namespace root, which is what ComputeCaptureSet has always
received -- so the export's classified paths and the backup's capture set describe the same
directories by construction instead of by coincidence.

Tests are table-driven over BOTH drive kinds, because this survived by being invisible on the
kind that already worked. Red-proofs observed: restoring the bare-path call fails the
system-drive row with the two paths differing by /felhom-data; inverting the drive-kind
comparison fails every enrolled row.
2026-08-04 18:17:05 +02:00
admin 582135f861 v0.190.0 — the boot settle window, both gates on intent, and R-171
gates / gates (push) Successful in 8s
R-171 (a regression v0.189.0 introduced, CONFIRMED on hardware before any fix
was written). Replacing isBootOrphan's container-count term with recorded intent
made a drive-gate-stopped app read as a boot orphan: the gate stops apps with
`compose down` (zero containers) and never touches desired_state, because it is
not the customer. Observed on 9201 with the drive held unmounted — the sweep
found and started it, burned both attempts, and handed it to the dead-app alarm.
The write hazard did not materialise (the unbound mountpoint is host-root-owned
and the guest is unprivileged) but that protection is accidental and untested.
New consumer-side seam bootrecon.StartGate, fail-safe (cannot determine ⇒ do not
start), wired in main.go. The rule is not new: the API's startGatedByMissingDrive
already refuses this; the sweep bypassed it.

R-157 mechanism A. The sweep looked once at T+5s, deriving candidates from a
fleet docker was still restoring — three of six hard resets. Now a settle-then-
sweep window: sample every 5s, settled after 3 identical samples, sweep ONCE at
the end; ends on settled or a 50s budget, and the log says which. The budget is
50s because settle+budget+one retry must stay under the 90s dead-app grace — a
test rejected 60s at 95s. A window that overruns emits a LATE RECOVERY warn
rather than the grace being widened to hide it.

Widening the window made two more holders reachable, so the one gate covers all
three: an absent drive, a quiesce, and an in-flight app-data operation — reusing
quiesce.SuppressedStacks() and a new read-only AppStopGuard.HeldStacks().

R-170. shouldRecreateOnBoot now reads desired_state with the identical three-way
table; absent keeps the old hasContainers behaviour exactly. Its comment argued
for the container count and was rewritten. presentStable is untouched. The two
gates' agreement is pinned from both sides against one fixture table.

27/27 packages green; 6 red-proofs observed FAIL then restored.
2026-08-02 19:56:20 +02:00
admin dbcb306fcf v0.189.0 — desired state + the app-stop crash marker (R-166 / D-b)
gates / gates (push) Successful in 8s
The box stops inferring the customer's intent from a container count and reads
what they actually asked for.

Part 1 — desired state. AppConfig gains a tri-state `desired_state`
(""/running/stopped), written ONLY by the customer's own action: the API action
switch, DeployStack, UpdateOptionalConfig's redeploy branch, and the .fab
import. Intent is written BEFORE the act and a failed write REFUSES the act.
StartStack/StopStack are deliberately not writers — 14 callers, only 2 are the
customer. bootrecon.isBootOrphan now reads intent instead of len(Containers)>0,
which closes R-157 mechanism B (a power cut or interrupted deploy left an app
with zero containers, read as a deliberate stop, and stranded silently).

ABSENT MEANS UNKNOWN, NEVER "running": every pre-v0.189.0 app.yaml reads absent,
so the legacy fallback is byte-identical to the old rule. A running-only startup
backfill converges the unambiguous cases; `stopped` is never inferred.

Part 2 — backup.AppStopGuard, a persisted marker over every stop→work→start
window (volume dump, offbox reconstitute, .fab export). Its own file, never
quiesce's. Written before the stop, cleared only after a restart that succeeded,
kept when one fails. Recover() completes before the boot reconciler is launched
and returns its outcome, which main.go reports on the existing backup_failed
event once the notifier exists. A defer is not the mechanism — a SIGKILL runs
none (Campaign 8 fault 10).

Also: SaveAppConfig rebuilt AppConfig field-by-field (the R-100 shape) and would
have dropped desired_state on every save across nine call sites. Replaced with
copy-and-overlay. Measured: app.yaml does not round-trip unknown YAML keys.

No hub change, no agent coupling, no user-visible string. 27/27 packages green;
7 red-proofs observed FAIL then restored.
2026-08-02 18:40:17 +02:00