Commit Graph

195 Commits

Author SHA1 Message Date
admin ff1758a21c R-585: the last customer-facing producers follow the household's language
backup_integrity_ok / backup_integrity_failed now take facts and push bundle
keys (Hungarian bytes unchanged - go-parity, pinned verbatim by
TestR585_IntegrityHungarianIsUnchanged). The interrupted-operation alert
(backup_failed, customer-enabled by default) used to send the operator's
ENGLISH sentence to every household; NotifyInterruptedOperation composes it
per language, the English byte-identical to the operator's log line. The
now-callerless NotifyBackupFailed is removed. local_api_endpoint_drift is
operator-only (no customer toggle) and already English by design - not
changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 02:01:12 +02:00
admin ca89e70d53 R-585 (part): three producers follow the household's language
db_dump_failed, the off-box backup_failed and offbox_enlarge_blocked
(whose sentence IS the household's mail - no hub customerMessages entry)
now take facts and push a bundle key, so an English household reads
English. Hungarian bytes unchanged (go-parity). Added to
convertedProducers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 23:12:59 +02:00
admin 4347983a72 R-682: a Remove cut off by a controller restart is finished at boot
RemoveStack journals itself before compose down and clears on every
return; at start a found journal finishes the remove through the same
RemoveStack (once, before the boot reconciler), keeping drive data and
backups even if the household had asked to delete them (no unattended
deletion at boot; logged).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 23:12:59 +02:00
admin 6fac245d24 R-270: the local_api drift detector also names a TOKEN-only divergence
A rotation that reached bootstrap.json but not controller.yaml left the
agent channel at 401 across restarts while the endpoint-only detector
stayed silent. The tokens are now compared (constant time, boolean only);
the token case gets its own banner key and operator message. Still
detection only - nothing is reconciled (R-78).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 23:12:59 +02:00
admin 7b8216abac R-492: delete the always-empty cfg.Paths.HDDPath (field, env binding, every fallback)
Readers (report builder, health check, /api/system, web primaryHDDPath, metrics collector,
AutoDiscoverStoragePaths' fallback parameter) now use the storage registry only. An old
controller.yaml still carrying paths.hdd_path keeps loading (non-strict YAML), pinned by
TestR492_OldConfigWithHDDPathStillLoads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 22:24:22 +02:00
admin 800b32ceec R-521: an app a lost drive stopped no longer mails app_start_failed one by one
One unplugged drive sent storage_disconnected plus an app_start_failed per app on it (five operator
mails). The dead-app job now hands the notifier the apps of a disconnected drive (its recorded
StoppedStacks plus apps whose HDD_PATH is that drive) as not down; the storage event already names
them. The dashboard's dead list is unchanged. The hub-side cooldown (F6/F7) and a household mail
for a lost drive are not in this change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:55:37 +02:00
admin f885100d29 R-271: the channel recovery after a controller restart is reported
The agent_channel_unauthorized alert's remedy is a re-bootstrap (a restart), and an unseeded->up
first observation was silent, so following the instruction guaranteed no recovery event. A down
alert now leaves a marker in the data dir; the first UP after a restart sends the recovery and
clears it. A restart with no alert outstanding stays silent.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:55:37 +02:00
admin 1453cfc69b v0.297.0: burn-down round 2 — 24 small rows (R-591 R-568 R-567 R-363 R-547 R-10 R-552 R-251 R-104 R-619 R-362 R-675 R-256 R-257 R-240 R-365 R-425 R-565 R-564 R-603 R-454 R-208 R-457-swept) + the banner countdown and deepCopyStack twins; MinAgent 0.131.0
gates / gates (push) Failing after 50s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 20:12:56 +02:00
admin ff69074514 controller v0.296.0: a backup run cut off by a power cut or restart is said on the backup pages (R-519); the whole-system backup text states today's measurement (R-518)
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 12:14:54 +02:00
admin 635c33d381 v0.295.0: a box that was off at its backup time catches up once (R-871, decision 109); the missed-backup banner (decision 110); a late daily timer after a host suspend is skipped
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 09:27:16 +02:00
admin 7861bf9dde v0.294.0: off-site clean-up guard follows the policy's own constants (R-867); no image clean-up while compose pulls (R-863); stderr tail (R-864); move-aside destination logged (R-869)
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 07:13:02 +02:00
admin 2a7f6c6c42 v0.293.0: the controller heals itself after the guest's Docker socket is re-created (R-860)
gates / gates (push) Successful in 29s
internal/sockheal: 60 s of refusals (never a timeout, only after Docker answered once) → exit 75 so
Docker's restart policy brings the controller back on the current socket; every 5 min it restarts any
other socket user (traefik) holding an older inode. Measured on 9202: only a docker.socket restart
re-creates the file; dockerd crash / docker.service restart keep it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 19:48:29 +02:00
admin c1a73b24b3 v0.290.0: the clean-up guard skips same-day superseded young snapshots instead of refusing (R-824), refuses above the weekly cap; a due set-aside deletion is handed to the hub's 7-day wait (decision 74, R-823)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 07:27:38 +02:00
admin 55bb6c3d32 v0.289.0: off-site key cannot delete — append-only rclone transport, box sends only its public key (hub registrar), retention only inside a hub window behind the fake-snapshot guard (decisions 68-69, R-820, R-822)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-03 17:01:36 +02:00
admin 977665d8c0 The family gate (decisions 63/64, R-780): family members with their own logins, a permanent forwardAuth door per family app, anchored exceptions, min_controller
gates / gates (push) Successful in 27s
- internal/family: the family list (bcrypt, generated 4x4 passwords shown once) + 30-day sessions in family.json
  (0600, atomic); a reset (generation), a removal or a logout ends sessions at the next request.
- internal/stacks/family_gate.go: family_gate / family_gate_except / min_controller in .felhom.yml; the door is written
  BEFORE the first start (install and a removed app's restore), a life record in app.yaml, reconciled by the gate loop;
  priority below the install hold, setup gate and sign-up block; every exception anchored ^/prefix(/|$) (finding F1).
- internal/web/family_gate.go: forwardAuth /__felhom_gate/family (app cookie felhom_famgate, host-only, names a store
  session); /__family/start|login|logout on the dashboard host (session cookie felhom_family, Path=/__family);
  sign-in counted per visitor (clientIP) AND per name, short windows; the household's dashboard session vouches.
  RequireAuth never reads a family cookie. The "Család" card on the security page: add / new password / remove.
Red-proofs RP-F1..RP-F7 (felhom.eu audits/family-gate-2026-10-02/A/).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-02 07:42:47 +02:00
admin ac8ea72025 v0.285.0 — a box keeps two controller versions (decision 56, R-745); the update clean-up's nil-stack crash (R-751)
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-01 07:36:26 +02:00
admin 5a3437669f v0.284.0 — a box deletes old app images (decision 53, R-736); an after_install app is held until its known login is replaced (R-741)
gates / gates (push) Successful in 27s
Image retention: after a done/undone guarded Update and at remove, an app's images older than its running
and previous one are deleted — never an image any container, installed compose or installed/previous record
names (box-wide keep set read at delete time); exact id, never forced or pruned; paused while any update runs;
a one-time sweep of catalog app images at the first start. Install hold: an after_install app is installed
behind the setup gate's door and opens when after_install succeeds or the household says it changed the login.
Tests TestImageRetention_* and TestInstallHold_* with red-proofs; parity fixture for the held card.

MinAgent: 0.131.0 (unchanged).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 22:29:33 +02:00
admin 29d3c533aa v0.283.1: a Stop holds during the volume dump in production and at the crash recovery (R-721, found live on v0.283.0)
gates / gates (push) Successful in 28s
Red-proofs RP43, RP44. MinAgent 0.131.0 (unchanged).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:52:19 +02:00
admin 6be6c53e29 v0.283.0: apps go off-site by themselves (decision 50) with a size warning; a Stop holds during a backup (R-721); page slips (R-724/R-725)
gates / gates (push) Successful in 25s
Red-proofs RP31-RP38. MinAgent 0.131.0 (unchanged).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:31:55 +02:00
admin 7fa8768cfd v0.280.0: the setup gate (decision 46); R-710 'I changed it' + absent-record window; R-709 password fields off the page; password:N:special generator
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-29 08:51:46 +02:00
admin 0c702f834a v0.279.0: after_install (decision 45), known default logins on the page, Part D empty-backup alarm, night chain (R-705), R-706
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-28 18:38:30 +02:00
admin b6810f14ff v0.275.0: a backup's data and its version travel together (R-696, 07 §6.6, D4 option A); R-695, R-691, R-694
gates / gates (push) Successful in 23s
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-26 10:35:22 +02:00
admin 43e99d160c kept data: the choice at reinstall, the list, the read-only view, the load (09 decision 36); R-690 fixed
gates / gates (push) Successful in 26s
An install over an app's kept drive folder (appdata/<app> non-empty) asks the household:
"use my kept data" (a load from the newest copy of THIS drive's install, own unit or
second-drive mirror, then the template's after_load) or "start fresh" (the folder is
renamed into <drive>/kept/<app>/<date>/ with the removed app's unit; nothing deleted).
The install API answers 409 kept_data_choice until one is chosen; DeployStack refuses
too. New page Megorzott adatok / Kept data (/kept-data): Load / Look / Delete (typed
confirmation, the only deletion of kept data). FileBrowser gets a read-only source.
The drive-full warning names the kept folders. <drive>/kept is protected and outside
every backup leg.

R-690: the removed-app restore (R-487) never found a unit on a DATA drive — it asked
GetStackComposePath (true for every catalog app) and restored nextcloud with no env.
Now isStackDeployed; pinned with a production-shaped provider.

Red-proofs: audits/night-2026-09-26/E/redproofs/.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 13:32:30 +02:00
admin 2caae38a71 stacks: the box converts a PostgreSQL major as a guarded-update step (09 6.4 part 10, decisions 35/37/38)
gates / gates (push) Successful in 25s
A step whose ladder entry carries engine_conversion {service, engine, from, to}
converts the database: the old engine alone, the check (owners, roles,
extensions, per-table row counts), pg_dumpall validated by its completion line,
the volume emptied only after the undo copy's marker is validated again, the new
engine alone, the load with ON_ERROR_STOP, the check again + PG_VERSION. Any
failure goes to the existing undo; a restart during converting is undone.
A PostgreSQL major move without the mark is refused before anything moves.
The old datadir's copy is kept until a backup is proven after the conversion.
17 tests, 9 red-proofs (audits/night-2026-09-26/B/).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 12:57:03 +02:00
admin 44ae4dea70 controller v0.272.0: the backup page says when a whole-box backup does not fit (R-685); R-671, R-670, R-677
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 11:19:27 +02:00
admin 9cf13a3add controller v0.271.0: automatic app updates — the update leg after the off-site copy, the backup gate waits, the switch (09 6.4 part 7; R-680, R-678, R-643)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:07:24 +02:00
admin 504eae018b v0.270.0: no update for a current app (R-679); an interrupted install is reported (R-681); a restore brings back the pinned version's health check (R-669); R-674
gates / gates (push) Successful in 24s
Five red-proofs. MinAgent 0.131.0 unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:49:43 +02:00
admin 3c6b49b31c controller v0.269.0: whole restore from the second drive; crash loops stopped; exact image digests; steps judged by their own .felhom.yml (decisions 26-28, R-661 R-666 R-667 R-668 R-664 R-665 R-662, 09 6.4 part 6)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 12:18:39 +02:00
admin 206b0357d1 controller v0.268.0: the undo finds volumes by definition; a held app names only a whole copy; one press = one tested step (R-658, R-659, R-660, R-651; 09 §6.4 part 5)
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 08:15:27 +02:00
admin 80e6ad8c47 controller v0.267.0: tests off DooPlex's Docker, cut-off copies refused, two pages true
gates / gates (push) Successful in 26s
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 20:25:28 +02:00
admin 0054d4bd69 controller v0.265.0: R-634 cause fixed, held apps say so, OOM storm alarm, R-647 leftovers
gates / gates (push) Successful in 27s
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 17:16:57 +02:00
admin bc278944a3 controller v0.264.0: the household is told when an update is undone or held, in its language
gates / gates (push) Successful in 25s
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 13:51:21 +02:00
admin 8fc2b4a1a9 v0.263.0: a failed update puts the app back by itself (09 decision 15, R-637)
gates / gates (push) Successful in 26s
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.

- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
  (pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.

MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 11:12:49 +02:00
admin 811f75736e v0.261.0 — the controller no longer swaps itself out from under an app update (R-608, R-609)
gates / gates (push) Successful in 23s
The controller self-updates daily at 04:30 by default, and after any hub report
once a floor sits above the box. That swap restarts the controller container.
The window proposed for automatic app updates is 02:30-05:00. It contains 04:30.

R-608 — a two-way lock, wired in main.go (stacks never imports selfupdate):
- stacks.Manager.AnyUpdating() -> Updater.SetAppUpdatingCheck, consulted in the
  same three places as the existing backupRunning gate.
- Updater.IsUpdateRunning -> Manager.SetSelfUpdatingCheck; UpdatePreflight
  refuses `self_updating`.
- MEASURED: the gap was narrower than assumed. The update's `backing-up` phase
  already takes the backup single-flight, so that one phase was covered. The
  other six were not, and `starting`/`verifying` are where data may have moved.
- The lock must NOT latch: a held app does not block the controller's own
  updates, including the release that might fix the hold.

R-609 — the 409 carries `data.reason`, additively. transient (busy, updating,
deploying, migrating, self_updating) vs terminal (held, downgrade). Found while
writing the test: the router refuses a HELD app on its own line before the
preflight, so `held` would have been the one reason missing.

Five red-proofs, each seen to fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 14:23:59 +02:00
admin 7c4a33b6d1 v0.258.0: the last four Hungarian things an English household met
gates / gates (push) Successful in 24s
R-589 the update badge, R-590 the data-folder backup sentence, R-573 the two channel
banners, R-572 two dead helpers. All four were built in Go, which is why neither the
template parity fixtures nor TestI18nEnglishPages could see them; slice 5's LIVE proof
is what found them.

R-590 is the one that matters most: it is a promise about the customer's files. It
said, in Hungarian and under an already-English folder label, that a drop-zone is
temporary and unbacked. The test now asserts the CONSEQUENCE in both languages — and
that the two languages do not produce the same string, which would mean the English
fell back.

R-572 is not what its row said. The row claimed a template renders "vasárnap" on an
English page; measured, NO template and no Go file called pruneLabel or
nextPruneLabel. They were dead func-map entries returning Hungarian, so they are
deleted rather than translated — translating dead code would add machinery with no
reader and a test pinning a fiction. Deletion is fail-loud and that was proven: a
template naming the removed function panics loadTemplates at startup.

Five red-proofs. One of them says something about the gate rather than the code: the
Go-parity gate does not measure localeFuncs keys against the base capture, so the
citation to TestLocaleFuncsHungarianBundleMatchesFuncMap is what carries them — the
test was extended to make that citation true.

MinAgent: 0.131.0 (unchanged). Hungarian byte-identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-20 18:10:53 +02:00
admin bef39598d0 v0.256.0: the box sends its own sentence in the household's language (R-558 Part B)
gates / gates (push) Successful in 24s
MinAgent: 0.131.0 (unchanged). Needs hub v0.118.0+, which shipped first and
tolerates a box that sends none of this - every box in the fleet is that box
until this release reaches it.

The hub writes a household's e-mails in their language now, but about a third
of those mails carry a sentence the BOX composed, naming a drive, an app or a
number. The hub cannot translate one. So the box sends it twice.

- message_customer on POST /api/v1/event, omitempty. A HUNGARIAN household
  sends nothing extra at all, so its payload stays byte-for-byte what every box
  sends today and the hub's fallback path keeps being the one production
  exercises rather than a branch nobody takes.
- 19 producers render both sentences from ONE bundle key. `message` stays
  Hungarian always: it is what the operator is mailed and what the hub logs.
- customer.language bootstraps a new box - stored choice, then config, then
  Hungarian. The config value is NEVER written into settings.json: that would
  record a choice the household never made.

The Hungarian did not move, measured twice: the wire golden from the slice-2
base commit, and the Go parity gate over all 19 new keys.

Three guards had to learn the change and one caught me: the test seam now
carries the new field; the R-329 severity register reported two dynamic sites
as no longer existing the moment they moved off PushEvent (the walk now checks
36 severity literals, up from 20); and TestConfigLanguageIsWiredInMain reads
main.go, because cmd/ is gitignored and ripgrep does not.

A mistake, named: the first pass dropped displayName from three producers,
which would have mailed customers "Alkalmazás telepítve: %!s(MISSING)". Caught
reading the diff; now pinned by a test that refuses %!/MISSING/%s/%d in either
language.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 16:54:20 +02:00
admin 612c417024 v0.247.0: i18n spike — the dashboard can speak English, Hungarian byte-identical
gates / gates (push) Successful in 19s
Message bundles (internal/i18n) expanded into templates before parsing, one
template set per language. Launcher, /backups, /apps/<slug> and the layout
converted; household language setting, POST /settings/language, ?lang= override,
report field. Parity test against fixtures captured from unconverted templates;
copy gates read templates expanded; new i18n_missing_gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 14:50:57 +02:00
admin 0fe315b759 v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s
MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted.

R-550 (operator ruling: fix). A design reversed and recorded: the restore
op-status was in memory by choice. Now restore-status.json in DataDir, written
atomically at both ends of an op. At startup a record still marked running
becomes a failed, interrupted result kept per app until that app's next
restore, shown on /backups/restore and the off-site wizard, and raised once as
restore_interrupted. Cooldowns stay in memory.

R-546. The R-543 reminder bar consults the agent's own preflight ok (every
blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while
paused. /backup/escrow shows a waiting card that polls and reloads instead of
red crosses and English diagnostics. POST /api/escrow/start refuses 409 before
staging or starting - the direct path chaos night used. Unknown readiness keeps
the bar.

Red-proofs (each seen failing): restore record across restart; main() calls
both startup functions; startup helper with loading skipped; restore page card;
bar held back; waiting card; start refusal. go build/vet/test ./... green, 28
packages; controller_gates --fast all OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:45:29 +02:00
admin 2f8ff2414c v0.244.0: the backup page stops promising what it does not hold (R-537/R-538/R-536)
gates / gates (push) Successful in 17s
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.

R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.

R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.

Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:55:55 +02:00
admin 843b319f35 v0.243.0: FileBrowser generated admin password (R-513); per-tier whole-guest backup truth (R-517); skip absent-storage tiers (R-518); OOM-killed worker visible (R-514)
gates / gates (push) Successful in 14s
MinAgent: 0.131.0

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:12:08 +02:00
admin d698ce343b controller v0.242.0: a removed app is listed with its kept backup; five small ones (R-487 R-491 R-490 R-489 R-476 R-456)
gates / gates (push) Successful in 14s
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
2026-09-13 22:50:18 +02:00
admin 3e813307cc controller v0.241.0: a bind-data app leans on off-site before its own unit; the hold names what the copy holds (R-479)
gates / gates (push) Successful in 13s
Operator ruling 2026-09-13. An app with classified binds walks second
drive -> off-site -> own unit (its unit holds no files); volume apps keep
2 -> 1 -> 3. RestoreHold.CopyHolds records what the chosen copy holds and
the sentence ends with it; older holds keep their tier-only sentence.
Tests on both halves; red-proof: a layout-blind order fails the bind case.
2026-09-13 21:47:33 +02:00
admin b93c1543da controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.

Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 17:16:24 +02:00
admin cbcca03061 v0.238.1: the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (slice 4 follow-up)
gates / gates (push) Successful in 13s
Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not
yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition
(alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names
survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored
the broken definition over the copy the customer is told to restore from.

backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now
also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to
stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:25:09 +02:00
admin 0d402f711d v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 11:41:31 +02:00
admin 8a0e0a59ad v0.235.0: freeze the version, keep the fixes flowing (operator ruling 2026-09-06)
gates / gates (push) Successful in 12s
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.

The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.

NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.

app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.

The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.

AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.

THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.

+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
2026-09-06 09:45:34 +02:00
admin 38d28b5b62 v0.234.0: seed installed_images at startup, so the label appears on an app nobody touched
gates / gates (push) Successful in 13s
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.

BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.

And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.

Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.

+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
2026-09-03 11:56:43 +02:00
admin 303129e3af v0.231.0: the off-site proof gets a by-hand trigger, like its integrity sibling (R-87)
gates / gates (push) Successful in 12s
Without it the only way to see the job work is to wait for 05:30, which makes live
validation and any future diagnosis a next-day exercise. Same function as the scheduled
job - no second code path.

ONE deliberate difference from the integrity button: due-ness is NOT bypassed. There,
forcing means "check the store again", which is always answerable. Here due-ness IS the
target selection - an app is due when its newest snapshot has not been proved - so
ignoring it would mean inventing a second way to choose an app, exactly what having one
function prevents. When nothing is due the button says so, honestly.

Every other guard intact, including the single-writer flag: a hand-run during a backup
SKIPS exactly as the scheduled one would.

POST /api/debug/backup/offsite-proof, button beside "Restic integritas" on the debug page.
debug_route_gate pairs the two, so a button with no dispatch (R-400's shape) cannot ship.
2026-08-31 21:09:09 +02:00
admin e43b5ec07d v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
2026-08-31 20:55:34 +02:00
admin 3c49dc8ea4 v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").

R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.

Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
2026-08-31 10:24:29 +02:00