controller v0.269.0: whole restore from the second drive; crash loops stopped; exact image digests; steps judged by their own .felhom.yml (decisions 26-28, R-661 R-666 R-667 R-668 R-664 R-665 R-662, 09 6.4 part 6)
gates / gates (push) Successful in 27s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-24 12:18:39 +02:00
parent 7c3b3a9694
commit 3c6b49b31c
141 changed files with 3401 additions and 237 deletions
+30
View File
@@ -1,3 +1,33 @@
## v0.269.0 — the second drive brings a file app back whole; a crash loop is stopped; exact image fingerprints; each step judged by its own file (2026-09-24 night, `09` §3 decisions 26–28, R-661, R-666, R-667, R-668, R-664, R-665, R-662, §6.4 part 6)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.123.0 for `app_stopped_unhealthy` (older hubs answer 400 and the
event is lost; the stop itself does not depend on it). New strings: yes (hu + en).
- **Decision 26 (R-661):** `backup.RestoreTier2Whole` — the second drive's „Teljes visszaállítás" for a file app
brings back the drive files by four rules (never delete; never overwrite a newer live file; bring back every
missing one; an older differing live file is replaced and KEPT beside as `<name>.felhom-<ts>`), then the unit
(settings + database) from the mirror. Refuses before anything moves without a proven, openable mirror with
file legs, or below the 2 GB floor. `WholeOnTier` counts the second drive as whole for such an app. R-538's
guard stays for every other caller.
- **R-668:** the same-disk check fails CLOSED — a storage path it cannot read is the SAME disk (a removed folder
became a file app's "second drive" on its own disk, measured on 9202).
- **Decision 27 (R-666):** while a held app's page says support is informed, the remove dialog offers only
"keep my data" and the API refuses data deletion (409). The no-whole-copy sentence is informal now.
- **Decision 28 + R-667:** crash loop = ≥ 6 restarts in 10 min (Docker's back-off caps a steady loop at ~1/min —
gokapi measured 7 in 7 min), counted from `RestartCount`; OOM storm = ≥ 20 kernel kills in 30 min. The box stops
the app, records an `unhealthy_stop` hold (same store as every hold), shows the sentence with a Start button,
and sends `app_stopped_unhealthy`. Start lifts it (one more try); a second stop within 24 h says support is
informed. Never inside a deploy, an update, a restore, a backup's stop or a quiesce. Default-on event, seeded
add-only.
- **R-664 / R-665:** the update judges the new version by ITS OWN `.felhom.yml` — the ladder step's
`steps/<key>.felhom.yml` or the catalog's, journaled — never the stack dir's, which a restore rewrites. The
step's memory request and applied record come from the same file.
- **R-662:** the dead „database and settings only" second step is removed (decision 25 ruled that restore out).
- **`09` §6.4 part 6, box half:** the compose the app runs pins `name:tag@sha256:…` from the ladder entry that
tested those refs (update and sync); pins and records stay digest-free; the badge reads „Frissítés elérhető"
for a floating tag only when the catalog's tested digest differs AND was tested after the install.
- Red-proofs: fourteen, each seen failing (REPORT).
## v0.268.0 — the undo finds the storage after a restore; a held app's page tells the truth; one press = one tested step (2026-09-24, R-658, R-659, R-660, R-651; `09` §6.4 part 5)
**MinAgent: 0.131.0** (unchanged). Needs hub v0.122.0 for `app_hold_no_whole_copy` (older hubs answer 400 and