3d49df1e5a
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
3073 lines
264 KiB
Markdown
3073 lines
264 KiB
Markdown
# CONTEXT.md — Project Memory
|
||
|
||
> This file serves as persistent project memory across Claude Code sessions.
|
||
> It replaces the auto-generated "Memory" from the claude.ai Project.
|
||
> **Update this file at the end of each working session** with current state,
|
||
> recent decisions, and anything the next session needs to know.
|
||
>
|
||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||
|
||
Last updated: 2026-09-25 night (v0.271.0 — automatic app updates, `09` §6.4 part 7; R-680, R-678; floor 0.271.0)
|
||
|
||
> **2026-09-25 midday — v0.272.0, floor 0.272.0.** `web.noSpaceLine` (agent prefix `skipped: not enough space: `,
|
||
> key `backup.tier.no_space`); `backup.SetUndoCopyRemover` → `stacks.RemoveUndoCopies` from `clearUpdateHoldAfterRestore`
|
||
> (R-671); `stacks.LoadProbeMetadata` (R-670); `stacks.BehindSinceAge` in both badge producers (R-677). Rulings: Peti's
|
||
> box retired; `09` decision 34 — no resume of the update leg after a restart.
|
||
|
||
> **2026-09-25 night — v0.271.0, floor 0.271.0 (MinAgent 0.131.0).** `stacks/unattended.go`: `RunUpdateLeg` — one
|
||
> app at a time, ONE step per app per night (decision 33), only `proven` entries, never `needs_person`,
|
||
> `files_may_change` only with `backup.FreshWholeCopy` (the hold's truth table), never the step in `app.yaml`
|
||
> `failed_update_step` while the ladder print is unchanged (R-680), transient refusals retried once, no step at or
|
||
> after W+5h (`backupwindow.UpdateLegStopOffsetMin`). Chained by `chainUpdateLeg` inside the `offbox-backup` job
|
||
> (every path incl. panic). `quiesce.Options.UpdateLegFn` = decision 20 + 31 (the gate waits until W+5h, then only
|
||
> for a step in flight, cap W+5h30m). Self-update waits for the whole leg (decision 32). Switch `settings.json`
|
||
> `app_update.unattended`, absent = ON, card on /settings (`POST /settings/app-update`). Page line
|
||
> `last_auto_update` („Automatikus frissítés %s-kor — sikeres."), report `update_leg`. R-678: `ScanStacks` before
|
||
> `finishUpdate` on done/undone. `stacks.update_window` REMOVED. Decisions 31–33 taken unattended (operator may
|
||
> reverse). Open: R-686 (no resume after a restart), R-687 (live-proof gaps), R-685 controller half (the backup
|
||
> page line for a space skip). Record: `felhom.eu/documentation/audits/DRILL-night-2026-09-25.md`.
|
||
|
||
> **2026-09-24 evening — v0.270.0, floor 0.270.0.** `UpdatePreflight` refuses `already_current` (R-679).
|
||
> `stacks/install_interrupted.go`: `.felhom-install-pending` marker + `RecoverInterruptedInstalls` at start (after the
|
||
> deploy-done hook) → compose down, stale pins cleared, `app_deploy_failed`, `.felhom-install-interrupted` → the
|
||
> apps page sentence (R-681). The recovery unit captures `stacks.AppliedMetaFile`; a restore calls
|
||
> `RecordRestoredAppliedMeta` (R-669). Next: part 7 (the update leg, `09` §6.4.2), R-678, R-680, R-682.
|
||
|
||
|
||
> **2026-09-24 night — v0.269.0 + v0.269.1. Fleet floor 0.269.1 (MinAgent 0.131.0; needs hub v0.123.0).**
|
||
> `backup.RestoreTier2Whole` (`tier2_whole.go`, decision 26): `mergeRestoreFiles` by the four file rules
|
||
> (never delete; never an older file over a newer one; bring back missing; an older differing live file kept
|
||
> beside as `.felhom-<UTC>`), then `RestoreFromRecoveryUnitAtWith(AcceptMissingFiles)` from the MIRROR; refuses
|
||
> before anything moves (not a file app / no proven openable mirror / < 2 GB). `sameDevice` fails CLOSED (R-668).
|
||
> Decision 27: `RemoveKeepsDataOnly` → 409 `err.stacks.remove_keep_data_while_support`; `keep_data_only` on
|
||
> hdd-data. Decision 28/29: `stacks.ObserveUnhealthy` (RestartCount delta ≥ 6 in 10 min, OOM ≥ 20 in 30 min)
|
||
> → `stopUnhealthyApps` → `HoldUnhealthy` (`unhealthy_stop`, trip 2 within 24 h) → `app_stopped_unhealthy`;
|
||
> Start lifts it. R-664/R-665: `entry.NewMeta` = the step's own `.felhom.yml`. §6.4 part 6 box half:
|
||
> `stacks/digest.go` — `RenderWithLadderDigests` for update + fresh install; **`CarryDigests` for an installed
|
||
> app's sync (v0.269.1, decision 30)**. Open next: R-669 (applied-meta after hold + restore), R-679/R-680 +
|
||
> part 7 (the update leg, `09` §6.4.2), R-681/R-682 (interrupted install/remove). Record:
|
||
> `felhom.eu/documentation/audits/DRILL-night-2026-09-24.md`.
|
||
|
||
> **2026-09-24 — v0.268.0 (R-658, R-659, R-660, R-651; `09` §6.4 part 5). Fleet floor 0.268.0.** The undo
|
||
> selects volumes with `DeclaredVolumeNames` (the compose definition), never the project label; the restore
|
||
> labels volumes (project/volume/version, no config-hash). The hold names the newest WHOLE copy
|
||
> (`backup.WholeOnTier` = the restores' refusal predicate `DeclaredDriveFileLegs`; off-site only for file apps)
|
||
> or, with none, `hold.update.no_whole_copy` + no Mentések button + `app_hold_no_whole_copy` (needs hub
|
||
> v0.122.0). `UpdateHeldStacks` is the fourth app-down suppression. Remove deletes `applied-compose.yml` +
|
||
> `applied-meta/`. The ladder: `stacks/ladder.go` — `nextLadderStep` picks the NEWEST entry whose `from` is
|
||
> the pin; a non-head step pins `catalog-cache/templates/<app>/steps/<StepKey(to)>.yml`; missing/wrong →
|
||
> `update.error.pin_failed` before anything moves; no match → template, logged. `LadderStepsLeft` on the
|
||
> page. Open next to it: R-661 (Tier-2 no whole action), R-664 (a step has no .felhom.yml), R-665, R-666,
|
||
> R-667 (crash-loop clock resets on a flapping container).
|
||
|
||
> **2026-09-23 (night) — v0.267.0 (R-650, R-640, R-499, R-518).** Every docker exec goes through
|
||
> `internal/dockerexec`; under `go test` a real docker is refused (opt-in `FELHOM_TEST_REAL_DOCKER=1`; a
|
||
> stub under `os.TempDir()` passes). `api`/`stacks`/`web` tests run under `RunWithStub` (TestMain). Restore:
|
||
> `appbackup.CheckDumpComplete` — a dump without its engine's end marker is refused before the first
|
||
> mutation (unit + off-site) and again before any load. Tier-2 page: `systemBackupFact` → four sentences.
|
||
> Backup button: the measured ~8 min. R-626 not reproduced (closed by measurement).
|
||
|
||
> **2026-09-23 (late evening) — v0.265.0 (R-634, R-625, R-636, R-647).** A whole-box backup no longer
|
||
> stops/restarts a DEPLOYING app (`ListDeployedStacks` skips `Deploying`; the volume leg re-asks via the
|
||
> optional `deployingReporter`; `StopStack`/`StartStack` → `ErrStackDeploying`). Held apps: badge
|
||
> `badge.update.held`, no Update button. OOM: the kernel `oom_kill` counter (the `OOMKilled` flag is sticky);
|
||
> ≥20 kills / 30 min per container run → one `app_oom_storm`. Web `readerFuncs` render hold/error per reader
|
||
> in EVERY language. Needs hub v0.121.0. Floor 0.265.0. Open: R-649 (operator), R-650 (tests on real docker).
|
||
|
||
> **2026-09-23 (evening) — v0.264.0 (`09` §6.4 parts 2–3; R-606, R-620, R-646).** Two events,
|
||
> `app_update_undone` / `app_update_held`, one per app per outcome, via `Manager.SetUpdateEventSink`
|
||
> (main.go) → `NotifyAppUpdateUndone` / `NotifyAppUpdateHeld`; on by default and seeded once on existing
|
||
> boxes (`app_update_events_seeded`), with two toggles on the notifications page (the page pushes the
|
||
> list to the hub on save). Update sentences are stored as key + args (`UpdateErrorKey/Args`,
|
||
> `update.error.*`, `update.refusal.*`, `update.phase.*`, `hold.*`) and rendered per reader
|
||
> (`updatePhaseText` / `updateErrorText` / `holdText`, `RestoreHoldForLang`, `UpdateErrorIn`); stored
|
||
> Hungarian unchanged. Startup `BackfillAppliedMeta` (R-646). Disabled notifier WARNs once per type
|
||
> (R-620). Needs hub v0.120.0. Floor 0.264.0. Leftovers: R-647. Evidence:
|
||
> `felhom.eu/documentation/audits/undo-fleet-2026-09-23/`.
|
||
|
||
> **2026-09-23 — v0.263.0 → v0.263.2 (R-637, `09` §3 decisions 15 + 19).** A failed health check after
|
||
> an update is now UNDONE: phase `copying` (after the pull, the app stopped anyway) copies every NAMED
|
||
> volume `cp -a` into `<vol>.pre-update-<stamp>` (label `felhom.undo-copy-of`, finished-marker last);
|
||
> on failure `undoing` validates every copy first, refills the volumes, puts back definition + pin +
|
||
> the pinned version's `.felhom.yml` from the job's OWN copies, and checks the old version with ITS
|
||
> probe → `undone` (`app.yaml` `last_update_undone`, one page line) or a HOLD prefixed with the undo's
|
||
> failure and the data state. Bind folders are never touched. Journal recovery: `copying` → old version
|
||
> back; `undoing` → resumed. **Two live-only lessons, both now in code:** (1) the periodic probe with
|
||
> the CURRENT `.felhom.yml` flips the app `unhealthy`, so the undo's health wait must probe an
|
||
> `unhealthy` app with its own override (`waitUpdateHealthyMeta`, v0.263.1); (2) `.felhom.yml` flows
|
||
> into the stack dir on every catalog SYNC, so the pinned version's file must be recorded AT PIN TIME
|
||
> — `<stack>/applied-meta/.felhom.yml`, written at deploy, adoption and pin advance (v0.263.2). Apps
|
||
> pinned earlier have no record until their next pin (R-646). R-642: start/restart never say
|
||
> "completed". **The fleet floor is still 0.262.1** — raising it is the operator's decision. Evidence:
|
||
> `felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/`, `…/undo-live-2026-09-23/`.
|
||
|
||
> **2026-09-21 (evening) — v0.261.0 (R-608 + R-609).** The controller self-updates daily at **04:30** by default and after ANY hub report once a floor sits above the box; that swap restarts the controller, which is the supervisor of a running app update. The window `09` §3b Q1 proposes for automatic app updates is 02:30-05:00. **It contains 04:30.** **A two-way lock, wired in `main.go` — `stacks` never imports `selfupdate`:** `Manager.AnyUpdating()` -> `Updater.SetAppUpdatingCheck` (a sibling of the existing `SetBackupRunningCheck`, consulted in the SAME three places), and `Updater.IsUpdateRunning` -> `Manager.SetSelfUpdatingCheck` with `UpdatePreflight` refusing `self_updating`. **MEASURED: the gap was NARROWER than assumed** — the update's `backing-up` phase already took the backup single-flight, so only the other six phases were exposed. The live probe landed in `safety-dump`, i.e. in the real gap. **THE PROPERTY THAT MATTERS: the lock must NOT latch** — `Stack.Updating` clears on done, failed AND held, so a held app does not block the controller's own updates for ever. **R-609:** the 409 now carries `data.reason` — transient (`busy`,`updating`,`deploying`,`migrating`,`self_updating`) vs terminal (`held`,`downgrade`). **Found while writing the test: the router refuses a HELD app on its OWN line BEFORE `UpdatePreflight`**, so `held` — the reason an unattended caller needs most — would have been the one missing. Five red-proofs, each seen to fail. **Proven live on guest 9202 with a negative control:** with no app update the manual swap gives the AGENT refusal; with one in flight it gives OURS. The sentence changing IS the proof. Deployed to **9202 only**; the fleet floor is 0.260.0 and 0.261.0 is a separate operator ask. Measurements: `felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/`.
|
||
|
||
> **2026-09-21 — v0.260.0 (R-524, update arc).** A box that updated before the catalog was reverted under it read „Frissítés elérhető" over an Update that would have moved its pin BACKWARDS (measured BIGNIGHT Phase 6, privatebin 2.0.6 vs catalog 2.0.5). **The comparison gains a fourth verdict and MOVES OUT OF `web`:** `stacks.CatalogOrder` — Unknown/Current/Behind/**Ahead** — is read by BOTH the badge and `UpdatePreflight`, because a comparison implemented twice drifts. Ahead reads „Naprakész"/"Up to date" (`tag-ok`, same word and class as level — nothing for the household to do) with a title saying why; the Update is refused `downgrade` 409. **The API now renders update refusals through `errText`** — otherwise the new key would be a seam built and never wired. **Ahead is NARROW:** every differing service must be orderable AND newer, else Behind — this gate can BLOCK an update, so it errs towards letting one run. **THE TRAP, and the fixture caught it, not the design:** the first tag rule accepted only bare `X.Y.Z`, so every REAL catalog tag was unorderable and the refusal test failed with `got nil`. The rule now takes the version at the FRONT and requires the trailing suffix to be IDENTICAL on both sides — `31.0.14-apache → 31.0.15-apache` orders; `26.05.2-ls310 → -ls311`, `postgres:16-alpine`, `apache-2.57.0`, a date stamp and a digest pin do not. Ordering is `util.Version.Compare` and nothing else (one comparator, house rule). Three red-proofs, each seen to fail. **R-589 was NOT open** — it shipped in v0.258.0 and only its row was stale; a reviewer who reads ONE producer cannot see a SECOND that overrides it. **Floor NOT raised — the operator's step; and CONTEXT's own 0.257.0 was STALE — the hub says 0.259.0 (read live from `/configs`, never from a doc).** Deployed on demo-felhom 9201, demo-hp 9201 and demo-hp 9202 (scratch, upgraded from 0.245.0 for the live proof). Seven questions for Slices 6 and 7 are in `09` §3b; the state of the whole arc with drift numbers is `audits/UPDATE-ARC-STATE-2026-09-21.md`.
|
||
|
||
> **2026-09-21 — v0.259.0 (R-596 P1 + R-598, the drill's blockers).** The claim page's answers and the Backup page's protection warnings follow the reader's language. **Fourteen** live sites carrying **nine** messages, not the sixteen literals R-596 counted — and the sixteenth, `data["Title"]`, was **DEAD** (`claim.html` is standalone; `.Title` belongs to `layout.html`) and was DELETED rather than translated. `backup_handlers.go` had **nine** code literals, not twelve; three were inside comments. `degradedMessageFor` now returns a **KEY**, so the decision stays language-free in one place while the words are chosen by whoever knows the reader; `buildTierViews`/`backupTargetLabel`/`loadGuestBackup` take `lang` the `buildDataPathCards` way. **The anonymous cookie-less claim page's language chain** (`langFor` → `settings.GetLanguage` → `configLanguage` ← `cfg.Customer.Language`) **was an unpinned assumption and is now a test.** Six existing copy-contract tests were kept rather than weakened — each resolves its key through the real bundle, so they still convict on a reworded Hungarian sentence. **Two things the tests caught and review did not:** an apostrophe in an English value is escaped to `'` and NEVER matches on the page (the failure reads exactly like an unwired handler — R-603), and the i18n gate convicted two of my English sentences for saying "please". **Proven live on guest 9201 through the `felhom_lang` cookie** — and the lockout proved itself unasked: Hungarian attempts locked out the English request from the same source, so the counter is per SOURCE, not per language. **The instrument trap that nearly cost a second fix (R-602): the cookie works only on ANONYMOUS pages** — `langFor` step 2 means a signed-in request reads the household's setting and ignores the cookie, so `/backups?felhom_lang=en` returns HUNGARIAN and reads like an unfixed defect. Use `?lang=` behind auth. **Floor NOT raised — the operator's step**; it still stands at 0.257.0. Guest 9201 is on **felhom-pve**, and **demo-hp answers on no route** (R-601).
|
||
|
||
> **2026-09-20 — v0.258.0 (R-589/R-590/R-573/R-572, slice 6 Part 0).** The four leftovers slice 5's LIVE proof found, all Go-side: the update badge and the lifecycle badge through `localeFuncs` (reusing `compareInstalledToTemplate`/`EffectiveLifecycle`, so the DECISION cannot drift between languages — only the words do); the data-folder card's backup sentence through `s.msgLang`, with the test asserting the CONSEQUENCE in both languages AND that the two differ; the two channel banners keyed by the checker's own CLASSIFICATION with the composed sentence as a FAIL-OPEN fallback. **R-572 was not what its row said** — nothing called `pruneLabel`/`nextPruneLabel`, so the two dead Hungarian-producing func-map entries were DELETED; deletion is fail-loud and that was proven (a template naming a removed func panics `loadTemplates`). Five red-proofs; one of them convicted the GATE rather than the code — `localeFuncs` keys sit under `_preexisting` and are not measured against the base capture, so the citation to `TestLocaleFuncsHungarianBundleMatchesFuncMap` is what carries them, and that test was extended to make the citation true.
|
||
> **THE WALK (slice 6 Part B) then judged it**, on a box installed from scratch that day by an English speaker: `felhom.eu/documentation/audits/DRILL-first-hour-en-0258-2026-09-20.md`. Verdict **NOT YET ready for an English-speaking tester** — `internal/web/claim.go` has **16 raw Hungarian literals**, every one a message on the one screen between a household and their box (**R-596, P1**), and `backup_handlers.go` + `backup_target_offer.go` carry the Backup page's two protection warnings (**R-598**). Both are the *composed-sentence-into-page-data* shape — **three instances on file now** (R-573, R-596, R-598), worth naming as a class. Everything else held with ZERO Hungarian: the download page, the bilingual console, all three customer mails, the bind page, the dashboard's first language from `customer.language` alone, both app pages, the catalog, the language switch. **R-214 closed as a side effect.** **Floor NOT raised — the operator's step.**
|
||
|
||
> **2026-09-20 — v0.257.0 (R-560 slice 5 Part A).** The catalog's `.felhom.yml` may now carry an `i18n: {en: …}` sibling block, and the dashboard reads it. `Metadata.For(lang)` merges FIELD BY FIELD (absent or blank English shows Hungarian, so a half-translated app ships); `For("hu")` is the parsed struct with `I18n` cleared and nothing else, deep-compared over all **53 real catalog files** copied into `internal/stacks/testdata/catalog/`; the three prose lists replace WHOLE and every other list is matched by its own key (`env_var`, option `value`, `match_group`, `target`, `path`) — **never by position**; `For` never writes through the manager's shared metadata. Pages reach copy only via `LocalizeStacks`/`LocalizeStackPtr`/`MetaFor`, and `TestNoDirectMetaCopyReadOnPages` AST-parses `internal/web` against a named allow-list so the NEXT page to read catalog copy fails the suite instead of rendering Hungarian to an English household. **Two of the nine red-proofs convicted a hollow TEST, not the code:** a struct copy shares its slices' backing arrays, so the obvious `DeepEqual` mutation check passed a deliberately broken merge; and a ONE-entry fixture cannot tell key matching from position matching. Both rewritten. Proven live on demo-hp (English pages English, seven Hungarian pages byte-identical apart from the CSRF token) and on demo-felhom **0.255.0** (block synced, pages identical, no parse warning in a 93-line log window). Gaps: R-589 (update badge still Hungarian in English), R-590 (the data-folder promise sentence), R-591 (`Stack.Copy()` does not deep-copy the new `Meta.I18n`), R-592 (three defects inside the new catalog gate, each found by its own decoy). **FLEET FLOOR RAISED to 0.257.0** the same session (operator asked), `min_agent` 0.131.0 declared — above the vouched golden 0.246.0, so the declaration carries it (R-472). Hub: `managed floor SERVED for demo-felhom … from declared`; **demo-felhom went 0.255.0 → 0.257.0 by itself in under 12 s**, healthy, `settle-gate: GO — at/above floor 0.257.0`, **and then rendered the English tagline** — the floor delivered the FEATURE, not a version string. Peti's box and tester-1 are DOWN and take it unattended when they return. The catalog was fully translated the same day (50 more apps in three pushes, 1 031 of 1 032 strings): **the English Apps list shows ZERO Hungarian app descriptions across all 53**. Gap added: R-594 (the catalog gate can convict a retrieval promise but cannot REGISTER a true one).
|
||
|
||
> **2026-09-18 — v0.255.0 (fixes the v0.254.0 globe).** The sign-in-flow pages requested `style.css` with **no `?v=`**, so a browser holding a pre-0.254.0 copy served CSS with no `.lang-globe` rules and the globe rendered as a bare unstyled `<details>` — **invisible to every test, because they all read markup and the fault was in which CSS file the browser fetched**. It was FIVE shells, not three (both guest share pages too), and `.Version` was missing from three of their data maps — now set once in `executeTemplateLang`. The globe also moved INSIDE the card, centred under the footer, menu opening upward via the shared rule. 15 shell fixtures re-captured, **zero dashboard ones**. R-579.
|
||
|
||
> **2026-09-18 — FLEET FLOOR RAISED to 0.254.0** (operator asked), `min_agent` 0.131.0 declared — above the vouched golden 0.246.0, so the declaration carries it (R-472). Hub: `managed floor SERVED … from declared`; demo-felhom went 0.253.0 -> 0.254.0 **by itself in ~40 s**, healthy, four other containers up, own log `settle-gate: GO — at/above floor 0.254.0`, and its sign-in page shows **one globe, zero old text links**. **Peti Proxmox (DOWN 65 d, 0.115.0) and Tester 1 (DOWN 1 d, 0.245.0) did NOT get it** and take it unattended when they return — untested on this version. Evidence: `felhom.eu/documentation/audits/i18n-slice2-2026-09-18/floor-raise-0.254.0.md`.
|
||
|
||
> **2026-09-18 — v0.254.0 (R-557 slice 2 release C — SLICE 2 CLOSED).** Saved notes follow the box language at WRITE time (§16 option 1): ~70 producers, and `EndRestoreOp` no longer takes a Hungarian literal from anywhere. The language switch is a **globe** (`<details>`, no script) in the sidebar footer and on the sign-in/claim/recovery pages; a visitor's choice lives in a display-only `felhom_lang` cookie that `langFor` reads ONLY when there is no session, and a claim carries it into the household setting on success. `POST /lang` is CSRF-exempt for a narrow reason written at the exemption; `safeBackPath` refuses `//evil.example` too. **Two parity exceptions, measured with a real diff: exactly two change shapes across 106 fixtures, 5 byte-identical.** **A DEADLOCK was introduced and caught by the suite hanging** — a note rendered inside `UpdateOffboxStatus`'s callback takes the settings read lock while the write lock is held; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` guards it now. 629 literals left, 0 errors, 0 saved notes. New row R-577.
|
||
|
||
> **2026-09-18 — FLEET FLOOR RAISED to 0.253.0** (operator asked), `min_agent` 0.131.0 declared — the floor is above the vouched golden 0.246.0, so the declaration is what carries it (R-472). Hub: `managed floor SERVED … from declared`; demo-felhom went 0.250.0 -> 0.253.0 **by itself in ~20 s**, healthy, its other four containers up, and its own log reads `settle-gate: GO — at/above floor 0.253.0`. Both live boxes run agent 0.132.0 (above the requirement). **Peti Proxmox (DOWN 65 d, 0.115.0) and Tester 1 (DOWN 1 d, 0.245.0) did NOT get it** and will take it unattended when they return — untested on this version. Evidence: `felhom.eu/documentation/audits/i18n-slice2-2026-09-18/floor-raise-0.253.0.md`.
|
||
|
||
> **2026-09-18 — v0.253.0 (R-557 slice 2 release B).** All **179** Hungarian error messages carry a key (`util.MsgError`): `Error()` is still the Hungarian byte for byte, `errors.Is` answers for the kind AND the wrapped cause, and an inner error argument renders recursively. 76 display sites go through `errText`, pinned by `TestNoErrErrorInPageOutput`. `memoryVerdict` returns an error, `UpdateRefusal` gained a `Cause`. Plurals: a key with `.one`/`.other` takes its count FIRST (bundle rule, not a call-site flag) — the guard caught a real key collision. **Two tooling defects found and recorded:** the bulk converter dropped multi-line concatenations (7 producers; the parity gate could not see it, two behaviour tests could), and my counter was case-sensitive and under-reported. 705 literals left, 0 of them errors.
|
||
|
||
> **2026-09-18 — v0.252.0 (R-557 slice 2 release A).** The sentences the program BUILDS now follow the language: 226 Go literals converted (flash lines, page data, API answers, alert banners, 237 country names, the four app-named page titles — R-566 closed). Flash travels as a KEY in the redirect URL with `fa` parameters; an old link's prose is still shown verbatim. New gate `i18n_go_parity.py` refuses any key whose Hungarian is not byte-identical to the base commit (7 467 literals frozen; 3 decoys). Wire goldens freeze the report warnings and the event messages the hub MAILS — those stay Hungarian until slice 3 (R-558). Remaining: 894 literals, 176 of them `fmt.Errorf` (release B); persisted text (release C, §16 option 1). New rows R-572/R-573/R-574.
|
||
|
||
> **2026-09-17 night — v0.251.0 (R-553 + R-563 CLOSED).** Five decisions that read their own Hungarian words now read a kind: deploy status sentinels (`internal/stacks/deploy_errors.go`), `backup.ErrOffsiteQuota`, `monitor.WarningKinds`, `settings.LastWarningKind` (legacy text fallback until R-570), and `data-status` on the remote-backup status. `util.KindErrorf` keeps message bytes identical. Slice 2 (R-557) is unblocked except the one producer named in R-570.
|
||
|
||
> **2026-09-17 night — v0.250.0 (R-556 release C, slice 1 CLOSED).** All templates converted; the language switch is on every dashboard page (decision 6 superseded); login/claim/guest/catch-all through `executeTemplateLang`; Go-side messages and three app-name titles stay Hungarian (slice 2).
|
||
|
||
> **2026-09-17 evening — v0.248.0 + v0.249.0 (R-556 releases A, B).** Apps/settings and backup pages
|
||
> converted; `TestI18nParityCoversEveryMarker` (every marker rendered by a case); fixtures only written
|
||
> when missing (`-update-i18n-golden`, `-i18n-golden-only`); `executeTemplateLang` for pages outside the
|
||
> dashboard chrome (recovery now). Words a page COMPARES stay unconverted (R-563). Split Hungarian verbs
|
||
> escape the retrieval gate's Hungarian stems (R-564).
|
||
|
||
> **2026-09-17 — v0.247.0 (localisation starter; design `felhom.eu/documentation/architecture/10-localisation.md`).**
|
||
> Message bundles `internal/i18n/locales/{hu,en}.json`; templates carry `{{T "key"}}`, EXPANDED
|
||
> before parsing — one template set per language (`loadTemplates` / `parseTemplateSet`); `s.tmpl` is
|
||
> the Hungarian set. Converted: launcher, `/backups`, `/apps/<slug>`, `layout.html`. **Release gate for
|
||
> every future slice: `TestI18nParity`** — Hungarian render vs fixtures captured from UNCONVERTED
|
||
> templates; never regenerate fixtures to make a conversion pass. `settings.json` `language`,
|
||
> `POST /settings/language`, `?lang=` override; switch shown only on non-Hungarian pages (decided by CC,
|
||
> operator may reverse). Report carries `"language"`. Copy gates now read templates expanded
|
||
> (`scripts/i18n_bundle.py`) — they had gone blind/stale when copy moved. New `i18n_missing_gate.py`
|
||
> (ratchets: English gap 0, formal „ön" forms 6). Converting a page: `scripts/i18n_extract.py`, then
|
||
> REVIEW what it left (ASCII-only Hungarian slips past it). Plan rows R-553..R-562 (R-553 first).
|
||
|
||
> **2026-09-13 night — v0.242.0 (R-487 / R-491 / R-490 / R-489 / R-476 / R-456).** The local backup
|
||
> lists are keyed on the DRIVES, not on what is deployed (`ListRemovedAppUnits`, the R-237 rule one
|
||
> tier down): a removed app whose unit was kept is listed with its restore, the picker and
|
||
> `/api/backup/snapshots` answer for it, and the restore opens the unit where it sits
|
||
> (`primaryUnitDirFor`). A removal clears an update hold. `/api/system/info` reaches the API router.
|
||
> `volumes_removed` is a before/after difference — **but a volume recreated by a unit restore has no
|
||
> compose label and is missed (R-489 open, residual: list by `<project>_` prefix too)**. Tier-2 copy
|
||
> date = newest dump (`UnitDataDate`) unless the leg was preserved. **The scratch guest 9202 on demo-hp
|
||
> is where releases are validated now** (hand-set image; hub/tunnel off) — `felhom.eu/CONTEXT.md`.
|
||
|
||
> **2026-09-13 — v0.237.0 (slice 4: R-448, R-443, R-439).** Update is a guarded 202 job:
|
||
> refusals (hold, busy, memory, disk, no restorable Tier-2 copy) → backup-first if the proven copy is
|
||
> older than `update.backup_max_age` (24h) → safety dump → pin → pull (failure: pin BACK) → up →
|
||
> health (`.felhom.yml` check or 60 s settle, `update.health_timeout` 5m; failure: stop + HOLD,
|
||
> `RestoreHold.Reason=update_failed`, pin stays). `UpdateStack` deleted. **The copy is aged by the last
|
||
> successful Tier-2 copy, NOT the manifest `created_at`** — measured on demo-hp that `created_at` moves
|
||
> only on definition changes. Journal `update-journal.json`; `RecoverUpdates` before the boot sweep.
|
||
> Three unattended paths that ignored a hold now honour it (drive-return gate ×2, nightly volume dump);
|
||
> capture and Tier-2 skip held apps. **No automatic rollback** — measured per-app; route back = restore.
|
||
|
||
> **2026-09-13 — v0.236.0 (R-442).** Removal resolves the drive from the app's OWN `app.yaml`
|
||
> `HDD_PATH` (the `07` ~L437 rule), never the global `cfg.Paths.HDDPath` (set on no box); a data
|
||
> removal it cannot resolve is REFUSED (409, exact Hungarian sentence, typed `RemoveRefusedError`)
|
||
> before anything is touched and the app is kept. SSD app → `hdd_paths_removed: []`, never `null`.
|
||
> Backup-path refusals reach the response. Proven live on demo-hp — see `REPORT.md`.
|
||
|
||
|
||
> **2026-09-06 — v0.235.0. THE RULING, AND THE TRAP IT SET.**
|
||
>
|
||
> **1. OPERATOR RULING: freeze the version, keep the fixes flowing.** R-447 sat `BLOCKED` because
|
||
> R-438 established that `RestartStack`'s `up -d` was a CHOSEN behaviour with its reason in its own
|
||
> comment. Option 1 was taken: an app's version is frozen to what the customer has and only a
|
||
> deliberate Update moves it, while health-check fixes, memory limits and self-healing keep arriving
|
||
> on the 15-minute cycle. **Both halves of the old behaviour were examined; only the version change
|
||
> was unwanted.**
|
||
>
|
||
> **2. THE MECHANISM IS A RENDER, NOT A GATE — and that distinction is the whole design.** Nothing was
|
||
> added to any of the thirteen `compose up -d` call sites. Most of them are REPAIRS (boot reconciler,
|
||
> drive-return gate, app-stop guard), and a repair that refuses to repair leaves a customer's app
|
||
> down. They are made safe by removing the reason: the file they act on no longer changes version.
|
||
>
|
||
> **3. `pinned_images` IS INTENT; `installed_images` IS AN OBSERVATION. NEVER FEED ONE FROM THE
|
||
> OTHER.** Letting a reading become a deployment is the R-166 category error one field over. They will
|
||
> normally agree; when they disagree that is a signal.
|
||
>
|
||
> **4. THE TRAP THIS RELEASE SET FOR ITSELF, and it would have shipped silently.** `Stack.TemplateImages`
|
||
> is read from the app's LIVE compose file — which is now the RENDERED one. On a frozen app that file
|
||
> names the OLD version, so the update badge would have found installed == template and answered
|
||
> **„Naprakész" on exactly the apps that are behind**, with every test still green, because the new
|
||
> field has the same type and shape. The badge now reads `Stack.CatalogImages`, from the syncer's own
|
||
> clone. **A feature that silently inverts a previous feature is the failure mode to look for whenever
|
||
> a file changes meaning.**
|
||
|
||
> **2026-09-03 — v0.234.0. ONE GAP CLOSED, ONE TEST DEFECT OF MY OWN.**
|
||
>
|
||
> **1. A RECORD THAT ONLY THE BRING-UP PATHS WRITE NEVER REACHES A QUIET BOX.** v0.233.0 shipped with
|
||
> "the record appears after the next lifecycle action" written down as a known limitation. One day
|
||
> later the operator looked at demo-felhom and saw OpenGist — up 15 hours, running exactly the catalog
|
||
> pin, **no badge at all**. The limitation WAS the feature not working. `BackfillInstalledImages` now
|
||
> seeds the absences at startup by READING containers. **The general lesson: a feature that only fills
|
||
> itself in on an event nobody triggers is, on the quiet installations, not shipped.**
|
||
>
|
||
> **2. THE BACKFILL IS STRICTER THAN THE BRING-UP PATHS, AND THE ASYMMETRY IS THE DESIGN.** The badge
|
||
> reads a service-count mismatch as BEHIND. The bring-up recorder runs right after a SUCCESSFUL
|
||
> `up -d`, where a missing container is real news; a backfill meets a box in whatever state it is in,
|
||
> so a partial seed would render „Frissítés elérhető" over an app that is perfectly current. It
|
||
> therefore refuses to seed anything it cannot observe COMPLETELY. **Same data, two writers, two
|
||
> different admission rules — do not "make them consistent".**
|
||
>
|
||
> **3. A TEST THAT HARDCODES A DATE AND ASSERTS AN AGE IS GREEN ONLY ON THE DAY IT IS WRITTEN.**
|
||
> `TestGroupD_BadgeRendersOnBothSurfaces` pinned `catalog_since: "2026-07-18"` and the string
|
||
> "46 napja". The pure tests inject a clock; the RENDER path goes through the funcmap and reads
|
||
> `time.Now()`. It passed on 2026-09-02 and was red on 2026-09-03. Now derived from the same clock the
|
||
> code reads. **R-457** names six other test files that mix a literal date with `time.Now()` — as
|
||
> unchecked candidates, not accusations.
|
||
|
||
> **2026-09-02 — v0.233.0. TWO DECISIONS, AND ONE LIMITATION THAT IS NOT A DEFECT.**
|
||
>
|
||
> **1. `installed_images` is an OBSERVATION, so a failed write NEVER refuses the action — deliberately
|
||
> the opposite of `desired_state`.** `SetDesiredState` refuses the act when the record fails, because
|
||
> intent that could not be recorded recreates the exact ambiguity R-166 closed. `recordInstalledImages`
|
||
> does the reverse: refusing to start a customer's app because we could not write down which version
|
||
> it is trades a real outage for a bookkeeping gap. **The rule is "refuse on intent, log on
|
||
> observation", and the reason is in both code comments** so neither gets "made consistent" later.
|
||
>
|
||
> **2. THE RECORD READS THE CONTAINER, NEVER THE COMPOSE FILE — and the file is not a lesser source,
|
||
> it is a WRONG one.** The catalog syncer overwrites a deployed app's `docker-compose.yml` on a
|
||
> 15-minute cycle with no deployed check at all; the spike measured the file saying `v2.8.5` while the
|
||
> container ran `v2.8.6` for 25 minutes. Anything derived from that file answers "what will happen
|
||
> next time something runs `up -d`", which is a different question from "what is running".
|
||
>
|
||
> **3. THE LIMITATION, STATED: 23 catalog pins FLOAT, so „Naprakész" CAN BE FALSE.** The comparison is
|
||
> reference-to-reference and queries no registry (a box must not need eight upstream registries to
|
||
> render a page). For `postgres:16-alpine`, `mariadb:11.6` and 21 others the ref can be identical while
|
||
> the image behind it has moved — the spike caught `mariadb:11.4` and `mariadb:12.3` already moved.
|
||
> **This is a known gap with a register row, not an oversight.** Digest-level comparison needs a
|
||
> registry query and is deferred.
|
||
>
|
||
> **Nothing about updating changed.** No behaviour, no new endpoint, no auto-update, and the three
|
||
> lifecycle buttons are byte-identical (`TestScenarioE_TheUpdateButtonIsUntouched`). The behaviour work
|
||
> is slice 3 and needs an operator ruling. The whole arc's reasoning now has a home:
|
||
> `felhom.eu/documentation/architecture/09-update-architecture.md` (R-438 — its absence was a finding).
|
||
|
||
|
||
> **2026-09-01 — v0.232.0. THREE RULINGS.**
|
||
>
|
||
> **1. `restic stats` TAKES A REPOSITORY LOCK, and that is the fact the whole R-411 chain rested on.**
|
||
> Nobody had it. Clean-room measured on demo-hp 2026-08-31: nothing else running, four invocations,
|
||
> the sampler reads `locks=1`. A customer FULL restore shells `stats` in its size probe, so it holds a
|
||
> lock — and until v0.232.0 it held no single-writer flag, so the integrity check was not blocked, ran,
|
||
> met that lock, and `resticStep` removed it with `unlock --remove-all` while logging *"a stale
|
||
> exclusive lock left by a previous crash"*. There was no crash. Also measured, and recorded so the
|
||
> next reader does not re-derive it: `restic check` takes a lock; `restic snapshots` and `restic list`
|
||
> do **not**.
|
||
>
|
||
> **2. THE INVARIANT IS PINNED BY A WALK, NOT BY A COMMENT — and the walk is the deliverable, not the
|
||
> acquire.** `offbox_integrity.go:28` asserted *"Every off-site operation takes `acquireRunning`"* from
|
||
> v0.227.0 and it was false for months, which is the ninth instance of this project's most-repeated
|
||
> class. `TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered` is an AST pass over
|
||
> `internal/backup` — deliberately not `strings.Contains`, because a commented-out call still contains
|
||
> the string. **On its first run it found three entry points nobody had named**:
|
||
> `OffboxRestorePrepareFull` (the request the customer's UI reaches FIRST, and the one that shells
|
||
> `stats`), `RestoreSharesScratch` (R-411's exact shape on the shares tier, with a live caller) and
|
||
> `RestoreOffbox` (no caller today). All four now take the flag. `OffsiteInventoryList` is registered
|
||
> EXEMPT with its reason — `snapshots` only, measured not to lock, and flagging it would make a page
|
||
> refuse to load during a backup for no safety gain. **Adding a line to `offsiteExempt` is a deliberate
|
||
> act and belongs in the commit that adds it.**
|
||
>
|
||
> **3. R-414's BRANCH, AND THE EVIDENCE FOR IT.** The question was whether `offboxRestoreScratchDir`
|
||
> was MISSED by R-356's unification or EXCLUDED on purpose. **Neither label fits: it was consciously
|
||
> OUT OF SCOPE.** R-356's own commit (`08eb1a6`) says so in its test comments — *"the prepared scratch
|
||
> still resolves to the registered storage path … only the DESTINATION moves, which is precisely what
|
||
> this change is about"* — and every one of its fixtures assumed a registered storage path exists.
|
||
> `demo-felhom`, with `storage_paths: []`, is the case it never had. It was **never ruled out on
|
||
> state-only grounds**: the one comment about a `systemDataPath` fallback belonged to
|
||
> `PlaceOffsiteRestore`, concerned bulk **userdata**, and R-356 deleted it deliberately. This
|
||
> function's own documented exclusion is `cfg.Paths.DataDir` — the **rootfs** — a different filesystem.
|
||
>
|
||
> **So §6.3's `[DESIGN]` rule applies and now has a FOURTH consumer.** But the fallback is **SCOPED**,
|
||
> because the two callers ask different questions and one predicate answering both is the R-356 defect
|
||
> itself: **unit-only** may fall back (§7 records as `[FACT]` that a driveless app's unit already lives
|
||
> on `systemDataPath` indefinitely and that the same-device placement is *"intended, not a defect"*);
|
||
> **full** keeps the R-252 refusal, because it pulls bulk userdata onto a state-only tier (§2.2).
|
||
>
|
||
> **AND THE SILENCE ENDS EITHER WAY.** `ProofResultCannotRun` is recorded through
|
||
> `RecordProofVerdict`, so `last_proof_result` is never ABSENT — absent already means *"controller
|
||
> older than v0.231.0"*, and giving one field two meanings is the `StatsKnown` trap one level up. It
|
||
> does **not** advance per-snapshot due-ness: nothing was proved, and marking one proved would stop the
|
||
> app being retried once a drive is finally registered.
|
||
>
|
||
> **A MISTAKE OF MINE, RECORDED BECAUSE LIVE VALIDATION IS WHAT CAUGHT IT.** The fallback resolved a
|
||
> scratch that `removeProofScratch` then refused to delete — its accepted-roots list is built from
|
||
> REGISTERED drives, and a driveless box has none. Observed on `demo-felhom`: *"refusing to remove …
|
||
> it is not inside a proof root"*, with the copy still on disk. Every nightly proof would have left one
|
||
> behind, on exactly the boxes the fallback exists for. **The unit tests all registered a drive, so
|
||
> none of them could see it.** Fixed, and pinned by a pair — one that the copy IS removed on a
|
||
> driveless box, one that a path outside every proof root is still REFUSED, so the fix is not a
|
||
> widening into uselessness.
|
||
|
||
|
||
> **2026-08-31 — v0.231.0. FOUR RULINGS, recorded so none is re-litigated.**
|
||
>
|
||
> **1. THE ACCEPTANCE RULE IS NOT "EVERY DECLARED FILE IS PRESENT".** That is the spike's own one-line
|
||
> summary and taken literally it is worthless: a hollow unit declares nothing, so everything it
|
||
> declares is present, and the check passes on exactly the shape it exists to catch. The rule has TWO
|
||
> parts and needs both — (1) everything declared is present, AND (2) the manifest declares what the app
|
||
> is SUPPOSED to have. **Part 2 is the whole value; part 1 alone is the trap.**
|
||
> `TestR87_HollowUnitForAnAppWithADatabaseFAILS` is the fence and its red-proof models the naive rule.
|
||
>
|
||
> **2. THE EXPECTATION COMES FROM INSIDE THE UNIT, NEVER FROM THE LIVE BOX.** R-403's guard could tell
|
||
> hollow from legitimately-empty because it had TWO copies to compare; this has ONE. The snapshot may
|
||
> predate the app's current shape, and the point is to judge the snapshot on its own terms — so the
|
||
> source is the unit's own `compose/docker-compose.yml`, and `GetDockerVolumes` (live Docker,
|
||
> `backup.go`) is explicitly NOT it. Database half: `DBServiceNames`, the same discriminator
|
||
> `RestoreFromRecoveryUnit` already uses, so this cannot disagree with the restore path about what an
|
||
> app is. Volume half: `ParseComposeNamedVolumes`.
|
||
>
|
||
> **THE VOLUME HALF IS AN EXISTENCE CHECK AND NOT A NAME MATCH, and that half-rule is deliberate.**
|
||
> Volume tars are `<project>_<volume>.tar`; `ResolveDockerVolumeNames` derives the project from
|
||
> `filepath.Base(filepath.Dir(composePath))`, which inside a unit is the literal string `compose`, not
|
||
> the stack. Measured 2026-08-31 on all eight real units on demo-hp: the counts match exactly
|
||
> (bookstack 2/2, docmost 3/3, kimai 2/2, opengist 1/1, privatebin 1/1, calibre-web 1/1, paperless-ngx
|
||
> 3/3, romm 3/3) and `<stack>_<volume>.tar` held in every case. **"Held on eight" is not "derivable"**
|
||
> — R-355's standing rule is that a claim about the app must never be inferred from a counter. Half a
|
||
> rule that is true beats a whole rule that is invented. Name-level matching is available the moment
|
||
> the capture records the project, and is not worth inventing before then.
|
||
>
|
||
> **3. THE PROOF'S RESTORE IS A READ-ONLY VARIANT, NOT A CHANGE TO THE CUSTOMER'S PATH.**
|
||
> `RestoreOffboxScratch` is the customer's restore, it is proven, and a customer restore taking a lock
|
||
> is correct — so it was NOT changed. What was shared instead of forked: the scratch-dir resolver is
|
||
> now parameterised on its ROOT builder (`offboxScratchDirIn`), so the drive-preference rules, the
|
||
> network-storage refusal and the R-252 wording have exactly one implementation; and the unit-only
|
||
> headroom gate is extracted to `unitOnlyHeadroom` so both paths refuse at the same floor with the same
|
||
> Hungarian sentence. The proof's own three differences are the ones that must differ: `--no-lock`, no
|
||
> `unlockStale`, and `m.runner()` instead of `resticStep` so the `unlock --remove-all` escalation is
|
||
> unreachable rather than unlikely (REUSE.md's rule: replacing `resticStep` would hide the escalation
|
||
> from the assertion that must see it).
|
||
>
|
||
> **THE PROOF SCRATCH IS A SEPARATE ROOT (`backups/offsite-proof`) AND THAT IS A SAFETY DECISION, NOT
|
||
> TIDINESS.** The proof deletes its copy on every path including failure. Sharing
|
||
> `backups/offsite-restore/<app>` would mean a nightly background job deleting the verification copy a
|
||
> CUSTOMER made and is looking at — a poorer actor destroying a richer one, R-403's shape in different
|
||
> clothes. The separate root also keeps the proof copy invisible to `DeleteOffsiteRestoreCopy`, the
|
||
> copy listing and `OffboxFullScratchReady`, so it can never be offered for placement into a live app.
|
||
>
|
||
> **4. THE ALARM IS A NEW EVENT TYPE, AND REUSING `backup_integrity_failed` WOULD HAVE BEEN WRONG.**
|
||
> That type is the nearest existing one and it carries a hub-side Hungarian template saying the
|
||
> integrity check found an error — i.e. **the store is damaged**. Here the store is sound and the
|
||
> CONTENT is missing: a different fact, a different cause, a different customer action, and telling
|
||
> someone their backups are damaged when they are not is the more expensive mistake (the same asymmetry
|
||
> `looksLikeRepositoryDamage` is shaped around). So `offsite_proof_empty` was minted, severity `error`,
|
||
> with NO `customerMessages` entry so the controller's dynamic Hungarian survives, and **the hub half —
|
||
> `allowedEventTypes` plus `operatorOnlyEvents` — ships in the same commit**, because an unallowlisted
|
||
> type is answered 400 and vanishes, and a missing customerMessages entry is not a routing block.
|
||
> **This widened the task's stated scope to `felhom.eu/hub/`** and the reason is recorded here rather
|
||
> than left as an unexplained diff.
|
||
>
|
||
> **DELIBERATELY NOT DONE:** `07` §8 matrix row 4 was NOT moved. This proves the snapshot CONTAINS a
|
||
> recoverable unit; it does not prove a restore puts data back into a running app. R-408 (the missing
|
||
> `acquireRunning` on `RestoreOffboxScratch`) was NOT fixed — the job takes the flag itself and the row
|
||
> stays open.
|
||
|
||
|
||
> **2026-08-31 — v0.230.0. THREE RULINGS, recorded so none is re-litigated.**
|
||
>
|
||
> **1. HOLLOWNESS IS A MANIFEST QUESTION, NEVER A SIZE QUESTION.** `unitCarriesData` asks whether the
|
||
> unit's manifest lists any database dump or any volume tar, and nothing else. `dirSizeBytes` lives
|
||
> two files away and is the obvious wrong answer: a unit with a fat compose capture and no dumps is
|
||
> exactly the shape that deleted 120 MB on `demo-hp`, and a 360-byte unit belonging to a tiny app is
|
||
> perfectly healthy. Size answers *how big*; the question is *is there anything to recover*. Absent or
|
||
> unparseable manifest ⇒ hollow, fail closed: a unit whose contents cannot be vouched for must never
|
||
> authorise a delete of one whose contents can. `TestR403_SizeIsNeverConsulted` is the fence.
|
||
>
|
||
> **2. THE GUARD FENCES ONE SHAPE, NOT SHRINKING — because the derived-copy rebuild is a DESIGN
|
||
> DECISION.** `07-backup-architecture.md` §8 row 5 records that the secondary is a derived copy,
|
||
> rebuilt on the next run, and `tier2.go`'s own header records that a classified app's copy
|
||
> legitimately shrinks as `export` drops out of its class set. `rsync -a --delete` stays, the data legs
|
||
> are untouched, and complete→hollow and hollow→hollow both still mirror. The ONLY refusal is a source
|
||
> carrying no data over a destination that carries some. Widening this to "the secondary never shrinks"
|
||
> would be calling a decision a defect; `TestR403_DataLegShrinkIsUnaffected` is the guard on the guard.
|
||
>
|
||
> **3. THE REHYDRATE HAPPENS INSIDE THE RESTORE, BECAUSE A FOLLOW-UP JOB RACES THE CAPTURE.** The
|
||
> hollow manifest was written **two seconds** after a Tier-2 unit restore, by the 5-minute
|
||
> `backup-cache` job. A goroutine, a scheduled refresh or a "do it on the next run" would each lose
|
||
> that race some of the time, and the failure mode is silent. `RestoreTier2Unit` refills the primary
|
||
> before it returns, and the test asserts ORDERING rather than sleeping.
|
||
> **And the capture is deliberately NOT guarded:** a capture that describes an empty drive as empty is
|
||
> CORRECT. With the primary refilled there is no hollow state left to describe. Guarding the capture
|
||
> would have made the manifest lie, which is the opposite of every other fix this week.
|
||
>
|
||
> **A defect the LIVE run caught and the unit tests did not, worth remembering as a shape.** The first
|
||
> draft of `UnitRestoreDate` also flagged "the package is older than the run" by comparing their dates
|
||
> — and a unit is ALWAYS captured shortly before the run that mirrors it, so it was true for every
|
||
> healthy app on the box. Four apps would have been told their package was stale. **A warning that
|
||
> fires on everything costs the same as the comforting lie it replaces.** The flag is now
|
||
> `UnitLegPreserved` and nothing else.
|
||
>
|
||
> **The credential rider.** `felhom.eu/scripts/read_credential.py` is now the one place that value is
|
||
> read. Three occurrences (2026-07-20, and twice on 2026-08-31, the third of which rewrote a live box's
|
||
> password hash) happened while the project already had a memory file, a worked recipe and a session
|
||
> report describing the mistake. **A note is read by whoever thinks to look; a check runs whether or
|
||
> not anyone remembers.**
|
||
|
||
> **2026-08-31 — v0.229.0. THREE RULINGS, recorded so none is re-litigated.**
|
||
>
|
||
> **1. The source moves; the destination does not.** `RestoreFromRecoveryUnitAt(stack, unitDir)` takes
|
||
> the recovery-unit DIRECTORY, so the same restore reads a unit from the primary drive or from the
|
||
> Tier-2 mirror on the second drive. What it must NEVER take is a destination: data still lands in the
|
||
> live Docker volumes and the live database container, and the definition in the guest, resolved by
|
||
> `GetAppDrivePath` exactly as the capture is. A restore that also relocated an app's data would be a
|
||
> migration wearing a restore's label, and the customer pressed a button that said neither.
|
||
>
|
||
> **2. Two predicates, never one wider one — and it is the SECOND time this is written down.**
|
||
> `Tier2Coverage.CanRestore()` answers *"can the additive file restore run?"* and nothing else;
|
||
> `CanRestoreUnit()` answers *"can the unit restore open this copy?"*. `HasUnit` keeps its third,
|
||
> distinct meaning: *"is there captured data the file restore is not looking at?"* — true even for a
|
||
> half-copied mirror the unit restore refuses, because the disclosure is still owed. The temptation is
|
||
> always to widen the predicate already there. **R-356 is what that costs:** one predicate meaning both
|
||
> *"has this app a drive?"* and *"is this app installed?"* refused 40 running apps for months, while
|
||
> they were running, with a message telling their owners to reinstall them somewhere those apps never
|
||
> offer.
|
||
>
|
||
> **3. A destructive operation reached from a non-destructive surface must carry the difference in the
|
||
> CONFIRM, not in the label.** „Teljes visszaállítás a másolatból" sits beside „Fájlok
|
||
> visszaállítása" on the same row; one overwrites the app's database and internal volumes, the other
|
||
> only adds files that are missing and never overwrites anything. The confirm says exactly that, names
|
||
> the copy's date, and says so DIFFERENTLY when that date is only an attempt clock (R-101). It is built
|
||
> from named Go constants (`tier2UnitConfirmBase` / `…DateFmt` / `…DateUnprovenFmt` / `…Contrast`) and
|
||
> asserted verbatim, because a sentence assembled inside an HTML attribute cannot be pinned and R-364
|
||
> makes grepping accented Hungarian out of rendered markup unreliable on top of that.
|
||
>
|
||
> **The count is settled and must not be re-derived.** At catalogue `459766cb1639`, by the production
|
||
> rule: **A = 7 · B = 45 · C = 1**. The C9-F1 Phase-0 count (9/43/1) was wrong by two — **radarr and
|
||
> sonarr**, whose `${USERDATA_PATH}` binds are WRITABLE (so the `:ro` default rule Phase 0 applied does
|
||
> not catch them) and are excluded by an explicit `class: excluded` entry instead. C is **bentopdf**.
|
||
>
|
||
> **R-403, filed and NOT fixed.** Two seconds after a restore that ran with the primary unit absent, the
|
||
> 5-minute status refresh (`captureAllRecoveryUnits`) rewrote the primary unit from a drive with no
|
||
> dumps, yielding `"db_dumps": []` / `"volume_dumps": null`. Measured on demo-hp 2026-08-31. The
|
||
> dangerous half — that the next Tier-2 run would mirror that hollow unit over the good secondary copy,
|
||
> `rsyncMirror` carrying `--delete` — **was not tested and is recorded as unverified.**
|
||
|
||
> **2026-08-31 — v0.228.0. TWO RULINGS, recorded so neither is re-litigated.**
|
||
>
|
||
> **1. The off-site integrity check ships at FULL depth — `--read-data-subset=100%` — and `off` is the
|
||
> way back.** Viktor's ruling, 2026-08-31, taken on a measurement rather than a claim: on `demo-hp`,
|
||
> 2026-08-30, a pack damaged **without changing its size** made plain `restic check` report
|
||
> `no errors were found` and exit clean; every read-data form caught it. Cost on that store
|
||
> (140 829 678 B / 2 651 blobs / 67 snapshots): structure 35.0 s, 10% 35.9 s, 50% 37.3 s, 100% 39.2 s.
|
||
>
|
||
> Three consequences that are decided, not open:
|
||
> - **Empty means "not configured", therefore the default.** It does NOT mean off. `off` (any case) is
|
||
> the off token, and it exists because a setting with no off switch is not a setting.
|
||
> - **A malformed value falls back to the DEFAULT, never to structure.** Falling back to structure
|
||
> would silently remove the protection on a typo — R-357's shape, a guard that opens quietly.
|
||
> - **The default lives in `internal/backup/offbox_integrity.go`, NOT in `config.applyDefaults`.** Both
|
||
> integrity defaults are resolved in one accessor each, beside the argument that justifies them;
|
||
> symmetry with the other `Monitoring` defaults is worth less than that.
|
||
>
|
||
> **The thing a future session will get wrong: there is exactly ONE data point, on a 134 MB store.**
|
||
> The cost curves are governed by different quantities — structure tracks the index, read-data tracks
|
||
> the data — so nothing here extrapolates. That is why v0.228.0 ships a *notice* (a WARN over 5 minutes
|
||
> naming R-401) and NOT a rotation schedule, a size threshold or a bandwidth budget. Every one of those
|
||
> would be a number invented from one measurement, which is the shape of the four production designs
|
||
> this project has already specced against nothing. **R-401's trigger is that WARN firing on any box,
|
||
> not a calendar date.**
|
||
>
|
||
> **2. Implement or delete FIRST, register the gate SECOND.** R-400 found seven debug-page controls
|
||
> with no handler. The gate that makes that impossible (`controller/scripts/debug_route_gate.py`) was
|
||
> written and registered only after all seven were resolved — a registered-but-failing gate refuses
|
||
> every push, exactly as `instructions_gate` established. The gate is deliberately ten lines: two lists
|
||
> and a difference, in both directions, because a cleverer gate needs maintaining and an unmaintained
|
||
> gate is how the class hides in the first place. **Keep `handleDebugAPI`'s exact-match switch with its
|
||
> `NotFound` default** — a prefix match would have made the original defect invisible instead of merely
|
||
> silent.
|
||
>
|
||
> **The shape worth remembering is worse than "seven dead buttons":** three of the seven fetched on
|
||
> page LOAD, so those panels were permanently blank on the page an operator opens when something is
|
||
> already wrong.
|
||
|
||
|
||
> **2026-08-30 — v0.227.0/v0.227.1. THREE RULINGS, recorded so none is re-litigated.**
|
||
>
|
||
> **1. The integrity check TAKES the single-writer flag and SKIPS rather than waits.** `resticStep`
|
||
> self-heals a crash lock by running `unlock --remove-all` and retrying, and its own comment records
|
||
> why that is safe: every caller holds the in-process single-flight mutex, so any lock it meets is
|
||
> stale. A check that did not take the flag could meet a **live** `forget --prune`'s lock from this
|
||
> same box, remove it, and retry over the top of it. It skips rather than waits because waiting would
|
||
> pin the nightly backup behind a check, and a skip costs nothing — due-ness makes tomorrow try again.
|
||
>
|
||
> **2. Due-ness, not a weekday.** A daily job asking "is the last successful check older than 7 days?"
|
||
> catches up after downtime; a Sunday-gated job silently skips a week every time the box is off on a
|
||
> Sunday. **R-341 is that failure**, and no `Weekly` primitive was added to the scheduler.
|
||
>
|
||
> **3. The result is published on `OffboxReportStatus`, NOT on `report.BackupReport`'s `IntegrityOK` /
|
||
> `LastIntegrityCheck`.** R-331 retired those the day before, because the hub card rendering them read
|
||
> `Integrity Unknown` for every customer forever. Giving them a live value would resurrect a card that
|
||
> was deliberately removed and break `TestBackupReport_DeadFieldsStayZero`.
|
||
>
|
||
> **AND THE MEASUREMENT THAT MATTERS MOST, because it is the thing a future session will assume
|
||
> wrongly: the structure check that ships ON does NOT catch silent corruption.** Measured on demo-hp —
|
||
> a pack corrupted without changing its size returned `no errors were found`, exit 0. Only
|
||
> `--read-data*` caught it. The structure check does catch missing packs, broken indexes and unreadable
|
||
> snapshots, which are real; it does not re-hash pack contents. **R-399 is therefore not merely a
|
||
> bandwidth question.** The cost curve is measured and small at today's store size (100% costs +12%
|
||
> wall-clock over structure-only, 39.2 s vs 35.0 s on 134.3 MB) but does NOT extrapolate — the
|
||
> structure check's cost tracks the index, read-data's tracks the data.
|
||
|
||
|
||
> **2026-08-30 — v0.226.0. TWO RULINGS THIS SESSION MAKES, recorded so neither is re-litigated.**
|
||
>
|
||
> **1. A local restore's outcome is a claim about THE BACKUP, never about the app.** This is R-355
|
||
> extended from the off-site path to the Tier-1 path. On the off-site path `SafetyDump` is an honest
|
||
> discriminator for "does this app have a database" — a path is returned only when a live database was
|
||
> found AND dumped. **The local unit-restore path has no such discriminator at all**, so no claim about
|
||
> the app is available to it. „ennek az alkalmazásnak nincs adata" and every variant is forbidden in
|
||
> `unitRestoreOutcomeMsg`. It is not merely unproven but unprovable from a manifest:
|
||
> 07-backup-architecture §6.3 records that an absent dump has causes that say nothing about the app —
|
||
> R-361 destroyed apps' canonical `.sql` files for four months, and a restore in that window would have
|
||
> "proven" a database-bearing app had none.
|
||
>
|
||
> **2. The destructive reconstitute uses NO headroom margin**, matching `PlaceOffsiteRestore` and
|
||
> deliberately NOT `OffboxRestorePrepareFull`'s ×1.1. The ×1.1 exists because that gate is *predicting*
|
||
> the size of a download it has not made. The reconstitute copies a tree that already exists on disk,
|
||
> so its size is measured, not estimated, and a margin over a measured value is a refusal with no fault
|
||
> behind it. Stated in a comment at the gate as well, because the two neighbouring gates disagreeing
|
||
> looks like an oversight to a reader who does not know which is predicting.
|
||
>
|
||
> Also this session: three earlier-day releases stacked in front of this one — v0.224.0 (R-330, the
|
||
> nightly backup alarming about the apps it was holding down) and v0.225.0 (R-331, `stats_known` on the
|
||
> wire). **The task specifying v0.226.0's work was written against v0.223.0 and targeted v0.224.0**; the
|
||
> drift was re-confirmed against live Gitea before the first edit rather than assumed.
|
||
|
||
|
||
> **2026-08-23 — v0.223.0 (R-329 + R-386), and a defect that only became visible once another was fixed.**
|
||
>
|
||
> **[RULING] The severity a controller sends is the HUB's vocabulary: exactly
|
||
> `{info, warning, error, critical}`.** Anything else is **coerced to `info` at ingest, silently**, and
|
||
> `severityNotifies` drops `info` **before both** delivery legs. `app_start_failed` emitted `"warn"`.
|
||
> **Measured on the live hub DB: 91 such events stored all-time, ZERO notification rows ever.**
|
||
>
|
||
> **[FACT] This was the SECOND occurrence, and the first one's comment had recorded the lesson.**
|
||
> `DiskAlertKind.Severity` emitted `"warn"` until v0.215.0. **A comment is not a guard** — the guard is
|
||
> now an AST walk over the whole controller. **grep cannot do this job:** `"warn"` is a legitimate
|
||
> *healthcheck status* in `internal/monitor` and `internal/selftest`; the sweep hit nine such strings
|
||
> and exactly one defect. The walk cannot follow a variable, so the **six** dynamic call sites are
|
||
> registered by name with the values each can take — **an unlisted limit is not a limit, it is a hole**.
|
||
> The guard found two of those six that the hand sweep had missed.
|
||
>
|
||
> **[FACT] It hid because another defect hid it.** R-384's ordering bug meant `app_start_failed` could
|
||
> not fire at all, so a broken severity had nothing to break. **Fixing one defect made another
|
||
> reachable** — and the same shape appeared again downstream: the operator cooldown key carries no app
|
||
> identifier, so **only the first app-down per hour now e-mails the operator** (R-182's shape, newly
|
||
> load-bearing, filed not fixed).
|
||
>
|
||
> **[RULING] `app_start_failed`: operator always, customer OFF by default.** `processOperator` never
|
||
> consults customer preferences, so one word fixed the operator leg and left the customer leg where the
|
||
> ruling wanted it. **Deliberately NOT in `operatorOnlyEvents`** — that would make the new toggle
|
||
> visible, flickable and structurally incapable of delivering.
|
||
>
|
||
> **[RULING, R-386] "The customer stopped this" is a RECORD, never an inference.** `aggregateState`
|
||
> folds `StateExited` into the stopped counter, so an out-of-band stop and a customer's Stop are
|
||
> byte-identical on the Docker side — no state test can separate them. Ask `DesiredState`, which has
|
||
> exactly one writer. `Stopped` → no alarm; `Running` → **alarm**; **absent → UNKNOWN, keep the old
|
||
> behaviour AND announce it**, because reading absent as "nobody asked" would e-mail about every app
|
||
> anyone ever stopped, fleet-wide, on the first cycle after upgrade. **The backfill cannot help — it
|
||
> seeds `Running` only from an observed-UP reading.**
|
||
>
|
||
> **[MECHANISM] `IntentUnknown` + an INFO line naming the apps.** A rule without a mechanism is a wish.
|
||
> Measured on `demo-hp`: **0 of 8** deployed apps carry an absent intent.
|
||
>
|
||
> **[FENCE] Adding a `DesiredState` WRITER is the fenced act; reading is fine.** And `failedRestart`
|
||
> must still lift a `Stopped` intent, or F-CRIT-1 re-opens.
|
||
>
|
||
> **[TRAP, cost three attempts] An HTTP 200 can be a REFUSAL.** The settings save answers 200 while
|
||
> rendering the empty-email wipe-guard error. Scenario G's before/after hashes matched twice because
|
||
> **nothing was saved**, not because nothing changed. And the email `<input>` spans three lines, so a
|
||
> single-line grep reads it empty. **Assert the refusal banner is ABSENT before believing a save.**
|
||
>
|
||
> **[TRAP] A red-proof that passes may mean an INERT mutation.** `if next <= prev` → `if next < prev`
|
||
> in fillwatch changes nothing, because an earlier `if next == prev { continue }` already removed the
|
||
> equal case. Check the mutation applied before believing either verdict.
|
||
>
|
||
> **[RULING] The compound-toggle split's risk was the MIGRATION, not the split.** A save whose event
|
||
> SET is unchanged now stores the existing slice **verbatim**, so byte-identity is by construction —
|
||
> without that guard the defaults case reorders, and the red-proof caught it.
|
||
|
||
|
||
> **2026-08-23 — v0.222.0 (R-384 + R-383), and a bigger hole found by a measurement that was told not to fix it.**
|
||
>
|
||
> **[DECISION] A dead SUPERVISED member is asked about BEFORE a failing healthcheck, because they are
|
||
> different questions and the second was answering the first.** `aggregateState` returned
|
||
> `StateUnhealthy` the moment `unhealthy > 0`, and the R-51 mixed-case block that asks "is a supervised
|
||
> member dead?" sat below it. A two-container app whose database exits goes `unhealthy` seconds later
|
||
> *because it cannot reach that database* — so **the symptom the fault causes was what suppressed the
|
||
> alarm for it.** The supervised test is now hoisted above the unhealthy/starting/restarting returns.
|
||
>
|
||
> **[DECISION] "Some members are up" means ANY member not in the down bucket** — running, unhealthy,
|
||
> starting or restarting. The old guard was `running > 0` counting `StateRunning` alone, which made the
|
||
> R-51 block **unreachable in precisely the case it was written for**: an unhealthy survivor beside a
|
||
> dead database counted as nothing being up. Either half alone leaves the defect standing, and the two
|
||
> red-proofs convict independently.
|
||
>
|
||
> **[FENCE, unchanged] `IsDownState` is byte-identical and `unhealthy` stays excluded.** An unhealthy
|
||
> container is RUNNING; folding it in reintroduces the flapping that exclusion exists to stop. **No new
|
||
> state was minted** — `StateDegraded` already meant this and every consumer already handled it. The
|
||
> fix is an ORDER, not a widening.
|
||
>
|
||
> **[FACT] The register's own suggested remedy was wrong.** R-384's row proposed a sustained-`unhealthy`
|
||
> threshold on the `crashLoopAfter` model. The defect needed no threshold at all. **A register's
|
||
> "recommended fix" is a hypothesis written before the diagnosis, and must be re-derived from source.**
|
||
>
|
||
> **[TRAP] Three existing subtests pinned the DEFECT as settled behaviour.**
|
||
> `TestAggregateState_UnchangedBranches` asserted an unhealthy/starting/restarting member beat an
|
||
> `exited` peer that was on `unless-stopped`. They were amended (down member given a benign policy,
|
||
> which is the only case where that sentence was ever true) and the change is reported, not buried.
|
||
> **A green suite can be green about the wrong thing.**
|
||
>
|
||
> **[FINDING — R-386, OPEN, NOT FIXED] A single-container app stopped out of band raises NO alarm, and
|
||
> a comment states the opposite.** `aggregateState` folds `StateExited` into the `stopped` counter, so
|
||
> an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**;
|
||
> `classifyRunStates` then whitelists `stopped` as a deliberate user stop. The comment at
|
||
> `cmd/controller/main.go` claiming an out-of-band `docker compose stop` "still alerts" is **measured
|
||
> false** — `privatebin`, 9 scans, 0 events, 0 banner. **Case #10 of "a comment asserting an invariant
|
||
> the code does not provide".** The task asked for this as a MEASUREMENT and forbade a fix; it is filed.
|
||
>
|
||
> **[DECISION, R-383] A message may not assert a file exists without asking the disk.** The
|
||
> double-failure sentence named the undo copy as present, built from the returned path — and a missing
|
||
> file is one of the two ways that rollback fails. `undoCopyPhrase` now reads from disk; a zero-length
|
||
> dump counts as MISSING; and the absent case still names WHERE the file should have been, because
|
||
> R-351's lesson is that a refusal naming nothing forces someone to remember what the product knows.
|
||
>
|
||
> **[DECISION] The alarm ladder now has an owning document** —
|
||
> `felhom.eu/documentation/architecture/08-alarm-ladder.md`. Until 2026-08-23 no document owned it; the
|
||
> rules lived as comments in four packages, each locally correct, with the ordering between them legible
|
||
> only by reading one function top to bottom. **That absence is why R-384 survived review.**
|
||
>
|
||
> **[GOTCHA] `app_start_failed` still ships severity `warn` (R-329)**, which is not in the hub's
|
||
> vocabulary and coerces silently to `info`, e-mailing nobody, while the POST returns 200. R-384 moved
|
||
> this event from unreachable to load-bearing, so the severity bug now matters.
|
||
|
||
|
||
> **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.**
|
||
>
|
||
> **[DECISION] `db_dumps` lists the app's OWN dumps, not the `pre-restore-*` undo copies.** They are
|
||
> local material for a restore that went wrong, not part of the app's recovery set. **Every consumer
|
||
> of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`** (the
|
||
> declaration, the enumeration, the change-detection compare); none reads it for recovery, and no hub
|
||
> or agent consumer exists. Three copies per app were being pushed off-site permanently for no
|
||
> recovery value. **The files are neither deleted nor hidden** — their visibility is a separate
|
||
> recorded decision and it stands.
|
||
>
|
||
> **[TRAP, and it bit within minutes] A stable `db_dumps` lets `CaptureRecoveryUnit`'s already-current
|
||
> early return fire.** Anything that must happen on EVERY capture — bounding the undo copies — has to
|
||
> sit ABOVE that check. It did not, and the cap silently stopped applying: four copies against a cap
|
||
> of three, counted on the box. Fixed in v0.221.1. **One change made another unreachable, and only
|
||
> counting files on a real machine showed it.**
|
||
>
|
||
> **[FACT] The comment was the defect.** `writeSafetyDump` called `DumpOne` into the app's own unit
|
||
> and renamed afterwards; `DumpOne` writes the canonical `<stack>-<dbtype>.sql`, so every safety dump
|
||
> destroyed the app's real backup. The comment said the rename meant it "can never overwrite the app's
|
||
> real dump" — false as written, for four months. The fix is a DESTINATION (`DumpOneTo`), not a
|
||
> rename, and the `.tmp` derives from the final path so a nightly dump beside it cannot collide.
|
||
> **`DumpOne`'s signature did not move.**
|
||
>
|
||
> **[NEGATIVE — do not re-derive this] A HELD app does NOT raise the dead-app alarm.** It was read
|
||
> from source that it would, because it keeps its database container and so is not `StateStopped`.
|
||
> Measured on the shipped v0.220.2: it aggregates to `unhealthy`, `aggregateState` checks
|
||
> `unhealthy > 0` before the mixed-case degraded branch, and `IsDownState` excludes `unhealthy`.
|
||
> Heartbeat read `0 currently down` throughout. **No suppression was built.** The same measurement
|
||
> exposed **R-384**: an app whose database has died is `unhealthy` too, and is likewise silent.
|
||
>
|
||
> **Proven live on `demo-hp`:** the canonical dump's sha256 unchanged across a restore on both engines
|
||
> — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`. Evidence:
|
||
> `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
|
||
|
||
|
||
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
|
||
>
|
||
> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the
|
||
> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A
|
||
> running app on a half-written database lets the customer type into it, and that turns a recoverable
|
||
> state into a permanent one. The alternative — start it and mark it — was put to the operator and
|
||
> declined. If that judgement is ever revisited, this is the sentence to revisit.
|
||
>
|
||
> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored
|
||
> database; the only difference was whether it looked broken (Postgres emptied and crash-looping,
|
||
> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not
|
||
> transactional. Putting the customer's own copy back is what removes the half state, and it is the
|
||
> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps.
|
||
>
|
||
> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such
|
||
> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older
|
||
> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a
|
||
> two-database app would have restored one and left the other half-written.
|
||
>
|
||
> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found
|
||
> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as
|
||
> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec`
|
||
> against the dead id timed out — so the app was held for an infrastructure reason while its data was
|
||
> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and
|
||
> never look at container identity.**
|
||
>
|
||
> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure
|
||
> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either —
|
||
> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start
|
||
> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared
|
||
> `driveStartGate` **above** its driveless early return, because the apps this exists for have no
|
||
> drive.
|
||
>
|
||
> **The way out is `--clear-restore-hold <app>`, and it REQUIRES A CONTROLLER RESTART** — it runs as a
|
||
> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command
|
||
> say so. Clearing through the running controller is the right shape later; it needs an operator tier
|
||
> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own
|
||
> hold is what the hold prevents).
|
||
>
|
||
> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state
|
||
> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was
|
||
> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`.
|
||
|
||
|
||
> **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.**
|
||
>
|
||
> **[DESIGN] The restore destination is resolved by the SAME rule as the capture destination.** The
|
||
> drive if the app declares one, the system data path otherwise — `Manager.GetAppDrivePath`, one
|
||
> expression, now used by `CaptureRecoveryUnit`, `ReconstituteFromOffsite` and `PlaceOffsiteRestore`
|
||
> alike. Anything else and the restore aims somewhere the backup never came from, which surfaces as a
|
||
> placement-mismatch prompt on a box where nothing actually moved.
|
||
>
|
||
> **The refusal that protects a drive app from being restored onto the wrong disk (R-253, R-351)
|
||
> applies to apps that HAVE a drive to get wrong.** It used to be reached by `HDD_PATH == ""`, which
|
||
> also stood in for "is this app installed?". Measured in the catalogue at `459766cb1639`: **53
|
||
> templates, 13 `needs_hdd: true`, 40 `false`** — so for 40 apps that test was permanently true and the
|
||
> off-site restore refused them forever, while they were running, telling the customer to reinstall
|
||
> them "in the same place", which those apps never offer. **An app with no drive is not misconfigured**
|
||
> (`01-topology-and-trust.md` §8, `[DESIGN]`); it is the majority case, and between 19 and 22 August it
|
||
> was called a defect four times.
|
||
>
|
||
> **Now:** *installed?* is asked of `ListDeployedStacks()` via `Manager.isStackDeployed`, which **fails
|
||
> CLOSED on a nil provider** — "cannot tell" must not become "go ahead" when the next act is a write.
|
||
> *Where?* is asked of `GetAppDrivePath`. A third refusal, with its own sentence and its own route,
|
||
> covers installed-but-no-resolvable-data-root: widening `nincs telepítve` to cover that would send a
|
||
> customer to reinstall a running app and hide the real fault.
|
||
>
|
||
> **FENCED, and not changed:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`).
|
||
> Capture resolves an app's declared `userdata`/`import` file legs against that value; a system-data
|
||
> fallback there would write a snapshot claiming to hold the customer's files and not holding them. The
|
||
> fenced ACT is "introduce a fallback into capture-side path resolution" — reading the value elsewhere
|
||
> is fine.
|
||
>
|
||
> **Proven live on `demo-hp`, 2026-08-22.** `privatebin` (driveless): planted through the app's own
|
||
> HTTP API plus a direct file plant, off-sited, **deleted**, restored through
|
||
> `POST /backup/offbox/reconstitute` — **15/15 files back byte for byte**, two Hungarian accented names
|
||
> included, message „0 fájl és 1 adatkötet visszaállítva". The scratch was unit-only, exactly the shape
|
||
> nobody had ever driven to completion before, and every downstream leg held. `calibre-web` (drive
|
||
> app) walked the same way and did not move. Evidence:
|
||
> `felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`.
|
||
|
||
|
||
> **2026-08-22 — v0.218.0 (R-354/R-355). The database nobody backed up, and the restore that
|
||
> returned most apps nothing.**
|
||
>
|
||
>
|
||
> Both fixes came out of the 2026-08-21 backup-truth drill. **R-355 went first because it is the only
|
||
> place in the product where one customer action causes permanent total loss:** paperless-ngx's database
|
||
> was dumped into a directory for a stack that does not exist, so it never entered the recovery unit, the
|
||
> off-site copy or the restore — and the same misattribution meant a destructive restore of that app took
|
||
> NO undo copy, then told the customer the app has no database. Fixed by reading the compose project
|
||
> label, which is the stack name by construction. **One app of 53 affected**, established with a sweep
|
||
> proven able to convict by planting a second mismatch.
|
||
>
|
||
> **R-354:** the off-site restore had no named-volume leg at all. The archives live inside the unit, whose
|
||
> placement is correctly skipped, and the comment beside that skip said the dump is replayed from the
|
||
> scratch "so nothing is lost" — true of the database, false of the volumes. `restoreDockerVolumesFrom`
|
||
> now replays them from the scratch unit; `VolumesReplayed` reaches the message.
|
||
>
|
||
> **WHAT THIS DOES NOT FIX, and it is the blocking item for the apps that need it most:** the off-site
|
||
> restore still REFUSES outright for the 40 of 53 apps that declare no data drive (**R-356**), saying a
|
||
> running app „nincs telepítve". Those are exactly the apps whose entire dataset is a named volume, so
|
||
> R-354's fix cannot reach them until R-356 is closed. Proven again on hardware 2026-08-21. The live
|
||
> confirmation of R-354 was therefore done on `calibre-web` and `paperless-ngx`, which declare a drive
|
||
> and can reach the restore.
|
||
>
|
||
> Also still open from the drill: the empty-restore success message (R-353), where the 40 apps' data
|
||
> lives (R-352), and the remaining rows R-357..R-366.
|
||
>
|
||
> ## THE RESTORE'S OWN MEMORY (v0.217.0, 2026-08-21) — R-351 / R-352 / R-353
|
||
>
|
||
> > **Both sides of a written fact must be checked, not just the writing side.** Every recovery unit
|
||
> > has recorded `drive` and `namespace_root` since schema 1. **No non-test code in the repository ever
|
||
> > read either back.** A restore into a different destination than the backup recorded therefore
|
||
> > succeeded silently under a green message. This is the same shape as several defects closed this
|
||
> > month, and the cheap test for it is one grep: *who reads this field?*
|
||
> >
|
||
> > **`IsRunning()` is still the wrong flag, in one more place than we knew.** v0.154.0 fixed the wizard
|
||
> > and left a comment explaining why. The **seven handlers** were never moved over, so a second press
|
||
> > genuinely started a second run and reported „…elindult". A comment explaining a trap does not fix
|
||
> > the other call sites — grep for them.
|
||
> >
|
||
> > **A result nobody can see is the same defect as no result.** The banner gated its terminal state on
|
||
> > a page-local `sawRunning`. The 8.666 s OpenGist restore finished before any poll saw it, so no
|
||
> > screen said it had completed. Fixed with `RestoreOpStatus.LastRecent` — and the window now lives in
|
||
> > `internal/backup` as ONE expression that both surfaces read.
|
||
>
|
||
> **State, and what is next.**
|
||
>
|
||
> - **Shipped:** placement comparison + named mismatch + `ack_placement`; not-installed refusal names
|
||
> the recorded drive; deploy prefill from the app's own backup; `restoreOpBlocked()`; `LastRecent`;
|
||
> off-site listing bounded-concurrent (measured 16.1 s → two waves).
|
||
> - **R-352 partly closed.** 40 of 53 catalogue templates declare no data path, and
|
||
> `GetDefaultStoragePath()` is read by nothing that places data — its comment `// new apps use this by
|
||
> default` has never been true. **Only visibility shipped**; the deploy page now states where the data
|
||
> will live. **No placement changed, nothing migrated.** Specification:
|
||
> `felhom.eu/documentation/backlog/SPEC-app-data-placement-2026-08-21.md`.
|
||
> - **R-353 is the next session's first item.** A restore whose unit carries no `db_dumps` and no
|
||
> `volume_dumps` reports a bare completion. OpenGist's unit held configuration and nothing else, and
|
||
> the restore said only that it had finished. Fix the outcome first; *then* prove the off-site
|
||
> coverage of a named-volume app by running a dump cycle — do not close the first on the second.
|
||
> - **Open and unproven:** whether the 40-class reaches the off-site tier at all has **not been
|
||
> observed**. `runVolumeDumps` covers them on paper; every unit on the box read `volume_dumps: None`
|
||
> because no nightly run had happened yet.
|
||
>
|
||
> ---
|
||
>
|
||
> ## THE TWO RULES THE RECOVERY JOURNEY LEANS ON (v0.203.0, 2026-08-06)
|
||
>
|
||
> > **1. A credential the hub stages is collected by the box, not waited for.** The reconcile that
|
||
> > collects runs on a tick for exactly as long as the box's own declaration says it needs one — and
|
||
> > stops the instant a target exists. It is driven from `OffboxReportStatus().State`, the same statement
|
||
> > the hub acts on, so the two can never disagree about whether a retry is wanted.
|
||
> >
|
||
> > **2. A mount Felhom itself made is not "something else".** Enrolment mounts a drive twice — the
|
||
> > managed path and a raw `/mnt/<name>` on the host — and the host survives a guest rebuild while the
|
||
> > guest's registry does not. The claimed check forgives a non-managed mount **only when corroborated**
|
||
> > by the same device also being mounted under the managed path. **A genuinely foreign mount is still
|
||
> > refused, and that fence has its own test.**
|
||
>
|
||
> **Why both are stated here rather than left in the code:** each was a dead end that kept the unaided
|
||
> recovery journey failing, and each looked correct in isolation. R-218's declaration half shipped and
|
||
> worked while nothing consumed what it asked for; R-220's check was right about foreign disks and wrong
|
||
> about our own. **Neither is a bug in the thing it guards — both are about what runs, and when.**
|
||
>
|
||
> Two things that must not be "simplified" back:
|
||
> - **The settle gate stays.** The retry goes through `ReconcileWhenSettled`, so the day-0 floor race is
|
||
> unchanged. A retry that skipped it would trade one defect for another.
|
||
> - **The R-220 exemption is corroborated, never a prefix.** Widening it to any `/mnt/*` path offers a
|
||
> disk another system is using for formatting — the red-proof shows exactly that.
|
||
>
|
||
> ## THE UNLOCK PATH'S RULE (v0.202.0, 2026-08-06) — state it before changing anything there
|
||
>
|
||
> > **On the recovery unlock path the customer is blamed only after a real attempt REFUSED their code.
|
||
> > Every other outcome — including one that cannot be classified — says something else.**
|
||
>
|
||
> This is the rule, and it outlives the bug that produced it. It was learned twice, because fixing it
|
||
> once was not enough:
|
||
>
|
||
> - **v0.201.0** stopped an agent that is too OLD from being reported as a wrong code (R-216).
|
||
> - **v0.202.0** found the same defect through a different door: an agent that is **stopped**, and a hub
|
||
> that cannot be **reached**, still fell through to a message about the code. Measured with a
|
||
> **correct** code at 0.0299 s and 0.0556 s, against ~1.0 s for a real unseal — the machine accused the
|
||
> customer of something it had not tried (R-224).
|
||
> - And the inverse: the one message that says *"check your ten words"* was unreachable on any box that
|
||
> had re-escrowed, which is exactly the box a customer has just recovered (R-226).
|
||
>
|
||
> **How it is enforced.** `agentapi.ClassifyRecoveryFailure` maps the failure to one of five classes
|
||
> **from the value, never the text**; the typing message is reachable from **one** of them
|
||
> (`RecoveryAskedAndRefused`, i.e. HTTP 400, i.e. the bundle was fetched and `age` refused it); and the
|
||
> zero value is `RecoveryUnknown`, which renders **neutral**. **The safe default is the load-bearing
|
||
> part** — an unrecognised status must not fall into an accusation.
|
||
>
|
||
> **Two things that are deliberately NOT how it works, and must not be "fixed" into it:**
|
||
>
|
||
> 1. **Elapsed time is never a classifier.** It is what diagnosed this, it is logged for the operator,
|
||
> and that is all. A duration guard would be a second thing that can be wrong.
|
||
> 2. **The error's TEXT is never read.** A string match is a defect waiting for a rewording. When the
|
||
> distinction was not available as a value, the **agent was changed to provide one**
|
||
> (`escrow.ErrBundleFetch` → HTTP 502, agent v0.126.0, `MinAgent 0.126.0`) rather than parsed for.
|
||
>
|
||
> **The coupling degrades safely and silently:** an agent below 0.126.0 answers 400 for both causes, so
|
||
> `FeatureRecoveryFailureClass` withholds the refusal reading and the 400 becomes neutral. The gate
|
||
> blocks nothing; it only decides whether the customer may be told to check their typing.
|
||
>
|
||
> ## CAMPAIGN 11 — what changed in v0.201.0 (2026-08-05)
|
||
>
|
||
> **The off-site key recovery is a COUPLED feature and now declares it.** It needs agent **0.125.0**
|
||
> (`POST /escrow/recover-offsite-password`). `FeatureOffsiteKeyRecovery` has a `featureProbes` row, a
|
||
> `featureMinAgent` row and a `Supports` gate at the unlock entry point.
|
||
>
|
||
> ⚠ **That gate FAILS CLOSED — alone in that table.** The package default is fail-open, and that default
|
||
> is what produced R-216: an agent that could not answer 404'd, the unlock was attempted anyway, and the
|
||
> customer was told their correct recovery code was wrong. Anything but `SupportYes` now says *the
|
||
> machine* cannot ask yet, and no attempt is made. Do not "fix" it back to the package default.
|
||
>
|
||
> **The box declares `needs_credential` until the TIER WORKS, not until a key exists.** The old
|
||
> short-circuit on "a repository password is present" is deleted: installing one is the recovery
|
||
> screen's whole job, so it made succeeding at recovery switch off the mechanism that delivers the
|
||
> coordinates to use it. A disabled target still short-circuits at the first line (Scenario E).
|
||
>
|
||
> **The unlock finishes the job**: place the key → bring the tier up (`offsiteapply.Bridge.Reconcile`,
|
||
> wired via `SetRecoveryTierUp`) → list. Without the middle step the promised listing can never render on
|
||
> shape (a), because no key ⇒ no target ⇒ no inventory.
|
||
>
|
||
> **Four messages, not one.** Wrong code (the only one mentioning typing) · the machine cannot ask ·
|
||
> the store could not be read / the connection details have not arrived · the code belongs to a RETAINED
|
||
> earlier package. The last one is driven by the ACK's `superseded_present`/`superseded_at` (hub
|
||
> v0.97.0) and **promises nothing** — no read path for a superseded package exists.
|
||
>
|
||
> **Still open from the campaign:** R-214 (console pairing banner), R-220 (drives unenrollable after a
|
||
> rebuild), R-221 (a rebuilt box cannot run the escrow ceremony), R-223 (the Day-0 manifest still vouches
|
||
> agent 0.120.0 — operator decision).
|
||
>
|
||
> ## R-203 (v0.197.0) — the namespace-root contract, and what `ok` now means
|
||
>
|
||
> **The contract, in one line:** `appbackup`'s path helpers (`UserdataDir`, `PrimaryBackupPath`,
|
||
> `RecoveryUnitPath`, `AppDataDir`) take a **NAMESPACE ROOT**. Anything that came out of `HDD_PATH` or a
|
||
> `StoragePath` is a **DRIVE path** — put it through `appbackup.NamespaceRootFor(drive, systemDataPath)`
|
||
> first. `UserdataDir(bareDrivePath)` still compiles and is still wrong; five callers proved it.
|
||
>
|
||
> **The rule now has ONE expression.** `NamespaceRootFor` / `IsEnrolledDrive` in `appbackup`;
|
||
> `backup.Manager.namespaceRoot` and `stacks.Manager.inGuest` delegate. There were two copies before and
|
||
> **they differed** — one compared without `filepath.Clean`, the other with it.
|
||
>
|
||
> **Why it was invisible:** on an enrolled drive the namespace root IS the drive path. The two diverge
|
||
> only on the system-data fallback, which `paths.go:26` names as a supported arrangement.
|
||
>
|
||
> **`last_status` gains `incomplete`.** A run that could not capture a directory an app declares
|
||
> MANDATORY is not a successful run. **Not `error`** — the rest of the run worked, so `SnapshotCount`
|
||
> and `LastSuccess` still record what WAS captured. It reaches the operator via the existing
|
||
> `backup_run_failures` digest (a new event type is a two-repo change; the hub drops unlisted types).
|
||
> The Hungarian customer warning is unchanged; the page renders `! Hiányos`.
|
||
>
|
||
> **Still open, and NOT fixed here:** `resolveAbs` resolves `RootHDD` and `RootUserdata` against the same
|
||
> root. Both callers now pass the namespace root so the export and the backup agree with each other, but
|
||
> whether `${HDD_PATH}` should mean the namespace root on the system drive touches every deployed app's
|
||
> binds and needs a decision, not a patch.
|
||
>
|
||
> **Blast radius, measured before changing anything:** exactly one app in the fleet had
|
||
> `HDD_PATH == system_data_path` (`calibre-web` on demo-hp, the R-201 drill fixture). Its data was
|
||
> migrated and its sentinel re-verified byte-identical.
|
||
>
|
||
> ## R-203 (2026-08-04) — a MANDATORY userdata directory can be absent from the off-site snapshot while the run says `ok`
|
||
>
|
||
> Found live on demo-hp while staging the R-201 drill, and it **halted that drill**.
|
||
>
|
||
> `NamespaceRoot(drivePath, inGuestDrive)` (`appbackup/paths.go:28-33`) appends the `felhom-data` segment
|
||
> **when the drive IS the system data path** — `m.namespaceRoot` = `NamespaceRoot(drivePath, drivePath !=
|
||
> m.systemDataPath)` (`backup/backup.go:331`). The deploy-time bind does not: `${USERDATA_PATH}` is
|
||
> `<HDD_PATH>/userdata` (`stacks/classify_binds.go:14`).
|
||
>
|
||
> With `system_data_path: /mnt/sys_drive` and `calibre-web` deployed at `HDD_PATH=/mnt/sys_drive`:
|
||
>
|
||
> live bind (files land here): /mnt/sys_drive/userdata/media/books ← exists
|
||
> capture set looked for: /mnt/sys_drive/felhom-data/userdata/media/books ← does not
|
||
>
|
||
> **The same compose used BOTH roots** — `${IMPORT_PATH}` resolved *with* the segment,
|
||
> `${USERDATA_PATH}` *without*. The run logged one `[WARN] mandatory data path missing on disk, skipped
|
||
> from offsite`, then `0 mandatory path(s)` and **`backup OK: 3 app(s), 3 snapshot(s)`**, with
|
||
> `last_status: ok`. Nothing customer-visible or hub-visible said the directory was dropped.
|
||
>
|
||
> **Not established:** whether `HDD_PATH == system_data_path` is a supported deploy. It was accepted
|
||
> (HTTP 202) one call after the NAS path was correctly refused (R-108). **Either branch is a defect** —
|
||
> broken resolution, or a missing refusal.
|
||
>
|
||
> **Two things a fix must do:** make the two roots one function, and make a skipped **MANDATORY** path a
|
||
> customer/hub-visible signal rather than a container-log WARN. `opengist`/`privatebin` declare no
|
||
> mandatory userdata paths and are unaffected.
|
||
>
|
||
> ## R-200 (v0.195.0) — the offsite key recovery diagnostic
|
||
>
|
||
> `--recover-offsite-check` is a `docker exec` escape hatch (the `--print-reset-code` shape), NOT a page
|
||
> or an API a browser can reach. R comes from **STDIN** — never argv, never `ps`, never shell history,
|
||
> never a transcript. It asks the agent (>= v0.125.0) to fetch this host's sealed bundle and open it,
|
||
> then reports whether the recovered repository password matches the on-disk one **by sha256**.
|
||
>
|
||
> docker exec -i felhom-controller /usr/local/bin/felhom-controller --recover-offsite-check < /path/to/code
|
||
>
|
||
> **IT COMPARES AND NEVER INSTALLS.** `CheckOffsiteKeyRecoverable` must stay free of any write — if a
|
||
> future change makes it place the recovered password, it stops being a diagnostic and needs the drill's
|
||
> supervision (that is link 9, R-200's remaining half). Pinned by
|
||
> `TestCheckOffsiteKeyRecoverable_WritesNothing`, whose red-proof is adding the install call.
|
||
>
|
||
> Exit codes are load-bearing: **0** match, **2** a clean MISMATCH, **1** a step failed. A mismatch is a
|
||
> finding about the system; a failure is a finding about the run, and they must never share a status.
|
||
>
|
||
> **Proven live on demo-felhom 2026-08-04** — recovered sha256 == on-disk sha256 == the hub's stored
|
||
> hash. Nothing customer-facing ships with it: no card, no form, no preview.
|
||
>
|
||
> ## About Viktor (project owner)
|
||
>
|
||
> - Works at Deutsche Telekom (Budapest), building Felhom.eu as a side business
|
||
> - Felhom.eu: managed home-server service for Hungarian households
|
||
> - Technical but prefers pragmatic solutions over over-engineering
|
||
> - Runs all infrastructure on Gitea (gitea.dooplex.hu), k3s cluster for management
|
||
> - Customer deployments use Docker Compose (not Kubernetes) for simplicity
|
||
>
|
||
> ### felhom-controller (this repo)
|
||
> - **Version:** v0.16.1
|
||
> - **Phase 1:** ✅ COMPLETE — Stack Manager + Deploy Flow
|
||
> - **Phase 2:** ✅ COMPLETE — Monitoring & Health (scheduler, CPU/temp, healthchecks.io pings)
|
||
> - **Phase 3:** ✅ COMPLETE — Backups (DB dumps, restic integration, manual trigger, **dedicated backup page**)
|
||
> - **Phase 4:** ✅ COMPLETE — Monitoring Page with Metrics Store (SQLite, Chart.js, system + container metrics)
|
||
> - **Phase 5:** ✅ COMPLETE — Authentication, Persistence & Settings Page (settings.json, password change, session management)
|
||
> - **Phase 6:** ✅ COMPLETE — Monitoring Warnings, Dashboard Alerts & Notification System
|
||
> - **Phase 7:** ✅ COMPLETE — Storage Overview, Per-App Backup Toggles & Limited Restore
|
||
> - **Phase A:** ✅ COMPLETE — Storage Paths Foundation (registry, auto-discovery, per-app HDD_PATH, deploy dropdown, health monitoring)
|
||
> - **Phase B:** ✅ COMPLETE — Storage Management UI Polish & Health Severity Fix (flash messages, label editing, app details, FS info, deploy free space, backup context)
|
||
> - **Phase C:** ✅ COMPLETE — Storage Init Wizard, Data Migration & Startup Fix (disk scan/format/mount wizard, rsync-based migration, startup pings)
|
||
> - **v0.11.1 bugfix:** ✅ COMPLETE — Storage Scan: system disk detection via host fstab + blkid UUID resolution; FSType enrichment via `blkid -o export`
|
||
> - **v0.11.2 bugfix:** ✅ COMPLETE — /host-dev mount for block device access; `HostDevicePath()` helper; all format/scan/safety ops use /host-dev
|
||
> - **v0.11.3 bugfix:** ✅ COMPLETE — Added `fdisk` package to Dockerfile (provides `sfdisk`; not in `util-linux` on Debian bookworm)
|
||
> - **v0.11.4 bugfix:** ✅ COMPLETE — FormatAndMount: fixed sfdisk (wipefs+force+`,,`), mount (explicit device path), mount propagation (rshared), ASCII label, smart partition skip, findmnt verification
|
||
> - **v0.11.6:** ✅ COMPLETE — FileBrowser auto-mount sync (`syncFileBrowserMounts()`) + 3 UI fixes (badge color, progress bar, button text)
|
||
> - **v0.11.7:** ✅ COMPLETE — Stale data cleanup + FileBrowser sync after migration + deploy page title fix
|
||
> - **v0.11.8:** ✅ COMPLETE — Per-App Cross-Drive Backup (3-2-1 rule): rsync/restic to secondary drive, deploy page UI, backup page summary, scheduler jobs, API endpoints
|
||
> - **v0.11.9:** ✅ COMPLETE — UI Polish Fixes: spacing, tooltip on "Módszer", status dot instead of disabled checkbox, progressive disclosure, emoji cleanup
|
||
> - **First app deployed:** Paperless-ngx on demo-felhom.eu (2026-02-13)
|
||
> - **Running on:** demo-felhom (N100 mini PC) at 192.168.0.162:8080, felhotest (Proxmox VM) at router.abonet.hu:33022
|
||
> - **All Phase 1-5 features working:** deploy, start/stop/restart/update, logs, health-aware states, auth, monitoring, backups, backup detail page, system monitoring page, settings page
|
||
>
|
||
> ## Architecture decisions
|
||
>
|
||
> | Decision | Rationale |
|
||
> |----------|-----------|
|
||
> | Go stdlib for web (no Gin/Echo) | Minimal dependencies, single binary, easy to embed templates |
|
||
> | Templates as go:embed HTML/CSS files | Zero runtime file dependencies (compiled into binary), but each template is a separate editable file |
|
||
> | Docker Compose for customers (not k8s) | Simpler troubleshooting, customers don't need k8s knowledge |
|
||
> | k3s for management infra only | Viktor's own services (gitea, monitoring, website) run on k3s |
|
||
> | Cloudflare Tunnel for remote access | No port forwarding needed, works behind any NAT |
|
||
> | app.yaml per stack | Separates deploy config from compose files, survives git pulls |
|
||
> | Password fields require explicit input | Prevents accidental empty-password deployments |
|
||
> | Health-aware state from Docker Status field | Docker's State says "running" even for unhealthy containers |
|
||
> | Memory limits via deploy.resources.limits | Prevents runaway containers; ~50% headroom over expected usage |
|
||
> | System info from /proc/meminfo + statfs | No external dependencies, cheap to read on each page load |
|
||
> | mem_request vs mem_limit (K8s-inspired) | Requests = expected usage (hard block), limits = peak (overcommit OK) |
|
||
> | 384MB reserved for system | Prevents deploying apps that would starve the OS/controller |
|
||
> | Logo SVG embedded as Go constant | Same approach as CSS/HTML — zero external file deps |
|
||
> | Git sync via os/exec git CLI | No Go git library needed, git is in the container image |
|
||
> | SHA-256 for content comparison | Only copy changed files, avoid unnecessary disk writes |
|
||
> | 30s debounce on manual sync | Prevents spamming the git server |
|
||
> | Orphan = deployed but not in catalog | Safe lifecycle: remove from catalog → mark orphaned → user deletes via UI |
|
||
> | FileBrowser as infra (not catalog) | Needed even after apps deleted (user browses HDD data); deployed by setup script |
|
||
> | Protected HDD paths | Safety net: never delete top-level HDD dirs (media, storage, Dokumentumok, appdata) |
|
||
> | Central scheduler (not ad-hoc goroutines) | Single place to register/monitor all periodic tasks, graceful shutdown, skip-if-running |
|
||
> | CPU sampling via background goroutine | /proc/stat delta needs two readings — collector runs every 5s, GetInfo() reads cached value |
|
||
> | Temperature from /host/sys (Docker mount) | Container can't read host /sys directly — mount /sys:/host/sys:ro, try /host/sys first |
|
||
> | Restic password auto-generated | No manual setup needed — generated on first backup run, stored in named volume |
|
||
> | DB discovery via docker inspect | No config needed — discovers postgres/mariadb containers by image name + env vars |
|
||
> | Backup orchestrator with running flag | Prevents concurrent backups, supports both scheduled and manual trigger |
|
||
> | modernc.org/sqlite (pure Go) | No CGO/gcc needed in Docker build stage — keeps `CGO_ENABLED=0` static binary |
|
||
> | AlertManager state-based refresh | Alerts regenerated every 5min from health report — no persistent storage needed, always reflects current state |
|
||
> | Notification relay via hub | Controller → hub → Resend → email. Hub acts as central relay: knows customer email, handles Resend API. Controller only needs hub URL + API key |
|
||
> | In-memory notification cooldowns | Per-event-type cooldown map (default 6h). Lost on restart = acceptable (better to re-notify than miss). No persistence needed |
|
||
> | Health status change detection | Only notify on degradation (ok→warn, ok→fail, warn→fail). Avoids spam on flapping. First run records baseline, doesn't notify |
|
||
> | Resend HTTP API (no SMTP) | Direct POST to api.resend.com — same pattern as website contact-mailer. Simpler than SMTP setup, good deliverability |
|
||
> | Preferences sync on save + startup | Controller pushes prefs to hub (not pull). Startup sync handles hub DB rebuild. Local save always succeeds even if sync fails |
|
||
> | Chart.js embedded locally | Customer hardware may not have internet — CDN not reliable for offline environments |
|
||
> | StackDataProvider interface | backup package needs stack data but can't import stacks (circular). Interface in backup, thin adapter in main.go |
|
||
> | Password sync to hub via report | Restic password in Docker named volume on SSD. Hub sync provides redundancy for disaster recovery |
|
||
> | App backup via HDD mounts only | Docker volumes at /var/lib/docker/volumes/ not mounted in controller. HDD data is the important user data; DB in volumes covered by nightly dump |
|
||
> | Restore uses running mutex | Prevents concurrent backup+restore on same restic repo. Reuses existing `m.running` flag |
|
||
> | Storage paths registry in settings.json | Multi-storage support: each app's HDD_PATH from app.yaml is authoritative. Auto-discovery on startup avoids manual config. Registry enables UI management + health monitoring per path |
|
||
> | /mnt:/mnt:rw mount in controller | Replaces per-path HDD_PATH mount. Enables multi-storage + restore writes. All customer HDD mounts are under /mnt/ by convention |
|
||
> | Per-app HDD_PATH resolution (app.yaml > global) | App's own env HDD_PATH is Priority 1, registered storage paths as fallback. Eliminates dependency on global controller.yaml hdd_path |
|
||
> | Mount-point detection via syscall.Stat_t.Dev | Compares device ID of path vs parent dir — reliable check that path is on separate filesystem. Prevents data writes to SSD |
|
||
> | Health severity: mount-point = warning | Non-mount-point is informational, not a service failure. FAIL reserved for genuinely broken things. Avoids false alarms on demo/test environments |
|
||
> | FS info via findmnt + sysfs | `findmnt -n -o SOURCE,FSTYPE --target <path>` for filesystem type/device. `/sys/block/<dev>/device/model` for disk model. Best-effort, returns nil on failure |
|
||
> | Query param flash messages | Stateless, no session store needed. Consistent with backup page pattern. `?storage_msg=success&storage_detail=...` |
|
||
> | StorageLabels map on stacks page | Separate map passed to template (not modifying Stack struct). Built from deployed apps' HDD_PATH → registered path label lookup |
|
||
> | Metrics downsampling via SQL | Bucket-based AVG in GROUP BY keeps Chart.js responsive with up to 30 days of data |
|
||
> | 60s metrics collection interval | Good balance of resolution vs. storage — ~44K rows/month for system metrics |
|
||
> | /etc/os-release mounted read-only | Container can't read host OS info directly — mount to /host/etc/os-release:ro |
|
||
>
|
||
> ## Key file locations on demo-felhom
|
||
>
|
||
> ```
|
||
> /opt/docker/felhom-controller/ # Controller compose + config
|
||
> ├── controller.yaml # Customer config (domain, auth, paths)
|
||
> ├── docker-compose.yml # Controller's own compose
|
||
> └── data/ # Controller persistent data (named volume)
|
||
>
|
||
> /opt/docker/stacks/ # All app stacks
|
||
> ├── traefik/ # Reverse proxy (protected)
|
||
> ├── cloudflared/ # Tunnel (protected)
|
||
> ├── paperless-ngx/ # First deployed app ✅
|
||
> │ ├── docker-compose.yml
|
||
> │ ├── .felhom.yml # App metadata
|
||
> │ └── app.yaml # Deploy config (env vars, locked fields)
|
||
> └── whoami/ # Test stack (not deployed)
|
||
>
|
||
> /mnt/hdd_placeholder/storage/ # HDD storage for apps
|
||
> └── paperless/
|
||
> ├── consume/ # Drop files here for OCR
|
||
> ├── media/ # Processed documents
|
||
> └── export/ # Backup exports
|
||
> ```
|
||
>
|
||
> ## Related repositories and their state
|
||
>
|
||
> | Repository | Status | Notes |
|
||
> |------------|--------|-------|
|
||
> | felhom-controller | Active | This repo. Controller code + deploy scripts |
|
||
> | app-catalog-felhom.eu | Active | 10 app templates, all with .felhom.yml metadata + memory limits |
|
||
> | felhom.eu | Active | Website + hub/ subfolder (felhom-hub service) + k8s manifests |
|
||
> | homelab-manifests | Stable | k3s cluster running (dooplex.hu services) |
|
||
> | misc-scripts | Utility | collect-repo.sh, backup helpers |
|
||
>
|
||
> ## Gotchas & lessons learned
|
||
>
|
||
> - `docker compose restart` ≠ `docker compose up -d` — restart doesn't pick up new images
|
||
> - Go maps have random iteration order — always sort slices before displaying
|
||
> - Docker `.State`="running" doesn't mean healthy — check `.Status` for "(health: starting)" / "(unhealthy)"
|
||
> - Paperless-ngx needs `PAPERLESS_OCR_LANGUAGES` (plural) to install language packs, `PAPERLESS_OCR_LANGUAGE` (singular) to select
|
||
> - In-memory Deployed flag must be set BEFORE `docker compose up -d` (not after) — compose can take 30-60s for image pulls, during which the UI would show a stale "Telepítés" button
|
||
> - Cloudflare Tunnel handles *.demo-felhom.eu → Traefik handles Host()-based routing to containers
|
||
> - BIOS "AC Power Recovery" must be enabled on N100 for auto-restart after power outage
|
||
> - `docker compose up -d` returns exit 0 even when containers immediately crash-loop — need post-start status check to detect this
|
||
> - When logging env vars for debugging, only log keys (not values) to avoid leaking secrets in log files
|
||
> - Mealie image (`ghcr.io/mealie-recipes/mealie`) doesn't include wget/curl — use Python TCP socket check for healthcheck
|
||
> - Mealie DB migrations on first start take ~40s (alembic) — use `start_period: 60s` to avoid premature unhealthy status
|
||
> - Alpine-based images (filebrowser, vaultwarden) have wget via BusyBox — healthchecks with `wget --spider` work fine
|
||
> - Deploy `sed` command to update image version must target only the `image:` line — naive `sed 's|name:OLD|name:NEW|'` also matches the service name line (e.g., `felhom-controller:` → `felhom-controller:0.2.12`), breaking YAML. Use `sudo sed -i 's|image:.*felhom-controller:[^ ]*|image: ...felhom-controller:NEW|'` or similar scoped pattern
|
||
> - Hungarian quotation marks `„"` in YAML: `„` (U+201E) is safe inside YAML double-quoted strings, but the closing `"` must NOT be ASCII `"` (0x22) — it terminates the YAML string. Use `\"` escape or Unicode `"` (U+201D). This caused a silent parse failure for the entire `.felhom.yml` file
|
||
> - Never silently swallow parse errors — always log them. Silent failures make debugging impossible (took a dedicated debug session to find a simple quoting issue)
|
||
|
||
> **2026-08-14 — v0.215.0 (R-328..R-333). Disk health, phase 1: the alert that reached nobody.**
|
||
>
|
||
> ### A severity string is a WIRE CONTRACT with the hub, not a label we choose.
|
||
>
|
||
> The hub accepts exactly `{info, warning, error, critical}` and **silently coerces anything else to
|
||
> `info`**, which `severityNotifies` then drops. `disk_health_degraded` shipped `"warn"` — one letter
|
||
> short of the contract — so **every Figyelmeztetés-level disk alert this product ever produced was
|
||
> emailed to nobody, on both legs.** Proven live side by side on 2026-08-14: `"warning"` →
|
||
> `notification_log` status **`sent`**; `"warn"` → stored `info`, **no row at all**.
|
||
> `app_start_failed` (`notifier.go` ~L546) carries the identical defect and was deliberately NOT
|
||
> changed here — it needs its own decision on whether it should notify (**R-329**).
|
||
>
|
||
> ### A drive's own PASSED verdict cannot fail on bad sectors. Do not build on it.
|
||
>
|
||
> Attributes 187/197/198 all carry `thresh: 0`; a normalized SMART value floors at 1 and can never drop
|
||
> to or below the threshold. The real drive (ST3000VX010, S/N Z6A07P2G) read `PASSED` at **352** pending
|
||
> sectors and **1001** reported-uncorrectable reads. Felhom already read the raw counters, which is the
|
||
> only reason it would have noticed at all. Evidence + fixtures:
|
||
> `felhom.eu/documentation/audits/DIAG-smart-passed-trap-2026-08-14.md`.
|
||
>
|
||
> **DECISIONS MADE, so they are not re-litigated:**
|
||
>
|
||
> - **Predicted failure is labelled „Hiba" — there is no fourth verdict word.** A fourth Hungarian word
|
||
> sharing a root with „Figyelmeztetés" would make the MORE severe state read as the milder one. Four
|
||
> labels, final: Rendben / Figyelmeztetés / Hiba / Nincs adat.
|
||
> - **Sustain is the primary rule; the count is the backstop.** Truth-table row 6 (unreadable sectors
|
||
> present again at the next check) sits ABOVE row 8 (count >= 64) because on the real drive sustain
|
||
> fires 12 Aug and the count not until 13 Aug. Row 8 exists only for a box powered off across the
|
||
> sustain window.
|
||
> - **The numbers and where they come from.** 64: the benign excursion peaked at 16 and cleared inside
|
||
> an hour; the terminal run passed 64 at 13 Aug 11:28 and never returned. **It is a judgement from ONE
|
||
> drive** — a static backstop, expected to be replaced by growth-rate detection in Phase 3. 55/60 °C:
|
||
> the operator's existing Prometheus bands on DooPlex, adopted unchanged so the two systems cannot
|
||
> disagree about the same drive. **These bands are SPINNING-DISK bands and are questionable for NVMe**
|
||
> — demo-hp's healthy Toshiba NVMe idles at **53 °C**, 2 °C below Figyelmeztetés (**R-333**).
|
||
> - **Phase 2 owns the new SMART attributes (187 Reported_Uncorrect, 199, 188).** They are a declared
|
||
> wire change, so under the G-1 gate the hub must model them in the same session. Putting them here
|
||
> would have turned a one-word severity fix into a three-repo change (**R-330**). Everything v0.215.0
|
||
> needs was already on the wire.
|
||
> - **Phase 1 state is one small record per disk, NOT a sample series.** `metrics.MetricsStore` is the
|
||
> right home for Phase 2/3 history; using it now would have put a schema migration on the critical
|
||
> path of the severity fix.
|
||
>
|
||
> **The trap this change nearly shipped, caught by a test and not by review:** the card and the check
|
||
> share one verdict function so the chip and the email can never disagree — but the check CONSUMES the
|
||
> prior and then overwrites it, so a card rendering afterwards read its own check's write and showed one
|
||
> level MORE severe than the alert. Fixed by `diskRecord.PriorSawUncorrectable`, which replays the prior
|
||
> that produced the stored verdict. The guarantee was previously asserted in a comment only.
|
||
>
|
||
> **Cadence is hourly, and it was MEASURED:** demo-hp `/disks` costs median 0.821s (min 0.805 / max
|
||
> 0.841, 10 calls, 3 physical rows) — 6x under the 5s bar. Open question deliberately NOT acted on: the
|
||
> agent runs bare `smartctl -a -j` with **no `-n standby`**, so an hourly poll would wake a spun-down
|
||
> HDD. demo-hp is all-flash so the measurement could not show it (**R-333**).
|
||
>
|
||
> **NOT live-validated:** the Fail-from-counters path has never fired on real hardware — only against
|
||
> the fixture's values in unit tests (**R-332**).
|
||
|
||
> **2026-08-08 — v0.208.0 (R-254). THE RULE, stated so it outlives this session:**
|
||
>
|
||
> ### A secret is never in a page's response body. It is fetched by an explicit act, and the act is recorded.
|
||
>
|
||
> Three instances of one pattern shipped in two days, each found by hand: the retrieval passphrase
|
||
> (R-249), an app's real first-login password (R-254 site one), and an already-deployed app's generated
|
||
> secret field (R-254 site two). Every one was "hidden" with `display:none`, `hidden`, or
|
||
> `type="password"` — **instructions a browser honours when DRAWING and nothing else.** The plaintext
|
||
> was in the bytes; a `curl` returned it; caches, history, saved pages and screen-shares had it.
|
||
>
|
||
> **The shape of the fix, now used three times:** the page carries a BOOLEAN; the value comes from a
|
||
> **POST** (so CSRF covers it and it is not re-fetchable from history) with **`Cache-Control:
|
||
> no-store`**; the reveal is **LOGGED as an act** — reading a value off markup left no trace anywhere,
|
||
> which is why nobody can say whether any of these was ever read. **Per-secret endpoints, never one
|
||
> generic "reveal any named secret"** — that would turn three narrow exposures into one lever.
|
||
>
|
||
> **And the test must assert the RAW RESPONSE BODY.** Every test that asked what the customer *sees*
|
||
> passed while the bytes carried the secret. That is precisely how this survived three times.
|
||
>
|
||
> **What is NOT this defect:** a form must carry what it submits. The pre-deploy hidden input round-trips
|
||
> a generated secret deliberately (README §318) so the saved value is the one the customer wrote down.
|
||
> The defect there was the neighbouring READONLY input on an already-deployed app, where nothing is
|
||
> submitted at all.
|
||
>
|
||
> **The gate:** `scripts/secret_in_markup_gate.py`. Name-based, all 36 templates, **blind to a secret
|
||
> arriving under a neutral page-data key** — measured, not assumed. The complementary runtime
|
||
> body-assertion covers 4 of 27 page templates; the other 23 are **R-255**.
|
||
>
|
||
> **A correction to v0.207.0's report:** it said HTML comments ship in the response body. They do not
|
||
> here — `html/template` strips them (`text/template` does not). Measured.
|
||
|
||
> **2026-08-08 — v0.207.0 (R-249, R-252, R-253). Three things the fifth walk exposed BY PASSING.**
|
||
> The walk closed R-201 (both halves) on 2026-08-07; none of the below touches the recovery path it
|
||
> proved.
|
||
>
|
||
> **R-249 — a secret was living in the page source.** `settings_security.html` rendered the retrieval
|
||
> passphrase into a `display:none` span behind a „Megjelenít" button. That toggle stops a browser
|
||
> DRAWING it and nothing else: the plaintext was in the response body of every render. Found by doing
|
||
> exactly that — it landed in a session transcript while driving the documented rebuild path.
|
||
> **THE RULE, which the codebase already stated for R and this page did not follow:** a secret is
|
||
> revealed by an XHR, never templated server-side into HTML (`escrow_handlers.go`). The page now
|
||
> carries only `HasRetrievalPassword`; the value comes from `POST /settings/retrieval-password/reveal`
|
||
> — CSRF-covered, `no-store`, and **logged as an act**, which reading it off the markup never was.
|
||
> **The test asserts the RAW RESPONSE BODY** — every test that asked what the customer *sees* passed
|
||
> while the bytes carried the secret, and that is why it survived.
|
||
> **The census found two more instances** (`app_info.html`, a real per-install app password in a
|
||
> `hidden` span; `deploy.html`, a generated secret in a `value=`) — **filed as R-254, not fixed.**
|
||
>
|
||
> **R-252 / R-253 — the two obstacles, and the rule they share.** A rebuilt box keeps its drives but
|
||
> loses their REGISTRATION, so every restore refused with a sentence naming no next step; and the
|
||
> restore list promised „a visszaállítás előbb újratelepíti" three lines above a refusal that fired
|
||
> *because* the app was not installed. **The promise was the wrong half:** reconstitution writes to
|
||
> the app's own `GetStackHDDPath`, which exists only once the CUSTOMER has chosen a drive at deploy
|
||
> time — an automatic reinstall would mean the product making that choice for them, which is the one
|
||
> decision this recovery path exists to leave with them. Both now name a reason and route to the step
|
||
> that clears it, and both notices are conditional (a healthy box is byte-identical, pinned by a test
|
||
> that fails if either becomes unconditional).
|
||
>
|
||
> **The page and the resolver ask ONE question:** `HasRestoreDestination()` reads the same
|
||
> `GetSchedulableStoragePaths()` the scratch resolver reads. A second copy of that predicate is
|
||
> exactly how a page ends up promising what the handler refuses — which is R-253 itself.
|
||
|
||
> **2026-08-07 — v0.206.0 (R-241). THE RULING, and it reversed the fix: this was a MINTING defect,
|
||
> not a screen-predicate defect.** The recovery screen was telling the truth — there genuinely was
|
||
> nothing recoverable under the key the box held, because **the box minted that key itself over the
|
||
> top of a sealed package it already knew the hub was holding**. Fixing the predicate would have
|
||
> papered over a machine quietly making its own backups unopenable.
|
||
>
|
||
> **THE RULE: a box does not create a repository key while the hub holds a sealed package for it.**
|
||
> The guard is a conjunction (package held AND no key), so a first-time box is untouched, and the
|
||
> refusal is a HOLDING state rather than a failure — the transport is still configured so the
|
||
> recovery screen can bring the tier up the moment the key arrives.
|
||
>
|
||
> **THE SECOND RULE: the fact that answers a question must be kept where the question is asked.** The
|
||
> hub-vs-local key comparison had been computed on every ACK since SLICE 3 and persisted nowhere; on
|
||
> the venue it logged the right answer thirty-five minutes before the customer looked at a screen
|
||
> that could not see it. It is now persisted and drives shape (c) of the offer.
|
||
>
|
||
> **THE THIRD RULE (the operator's, and it generalises): fix the state, do not remember that it is
|
||
> wrong.** Abandoning the old history now starts a 14-day countdown that removes the set-aside store
|
||
> and its sealed package TOGETHER, after which the offer falls silent on its own because there is
|
||
> nothing left to compare — rather than a "they decided" flag suppressing a screen over a state that
|
||
> is still wrong. The recovery offer stays reachable for the whole grace; a grace in which recovery
|
||
> is impossible is decorative.
|
||
>
|
||
> **Surface:** the full page appears once per ENTRY into the offered state, not once ever — a box
|
||
> rebuilt months later is a new situation. Three dismissal levers with three scopes, and none of them
|
||
> removes the entry point on the backups page.
|
||
>
|
||
> **Needs hub v0.98.0** for the superseded-package purge. `felhom-agent` untouched.
|
||
>
|
||
> **Two real bugs were caught by tests rather than by review** — a missing `t.Enabled` (an existing
|
||
> test) and a missing falling-edge sync that reintroduced the very defect the epoch exists to fix.
|
||
>
|
||
> **NOT built, deliberately:** the automatic 30-day abandonment (R-245, with the operator's reasoning
|
||
> recorded), and R-242's release-to-golden gate.
|
||
|
||
> **2026-08-06 — v0.205.0 (R-234).** THE RULE: **a run that skipped an app the customer selected is
|
||
> not a successful run.** The R-203 verdict block already said *"a warning beside a success is read
|
||
> as a success"* and applied it to one of the two shapes it describes — a missing declared FOLDER
|
||
> made the run `incomplete`, an app skipped ENTIRELY did not. Now both do. A selected-but-UNDEPLOYED
|
||
> app is named with what to do but does NOT move the verdict, because a box left permanently amber by
|
||
> an app somebody removed is a status nobody reads.
|
||
>
|
||
> **§7.3, MEASURED rather than assumed — and the answer was "already done".** `CaptureRecoveryUnit`
|
||
> writes compose config + a manifest (a few KB), only ENUMERATES dumps rather than creating them, is
|
||
> idempotent, and does NOT stop the app; the off-site run already calls it for every deployed stack in
|
||
> its own pre-dump phase, through `admitApp`. So there is no wait to remove for a deployed app, and
|
||
> **nothing was built**. Proven on demo-hp: a unit moved aside was RECREATED by the run.
|
||
>
|
||
> **AND THE FILED MECHANISM WAS NOT THE MEASURED CAUSE.** R-234 was filed as "the first run after a
|
||
> toggle finds no bundle and skips the app". That cannot happen for a deployed app (above). What did
|
||
> happen on 2026-08-06: the manual run was dropped by the **single-flight** while an earlier run was
|
||
> still going; `runOffboxBackup` returned nil; the handler had already said „elindult”; and the card
|
||
> then showed the PREVIOUS run's „✓ Rendben”. Fixed by taking that decision synchronously in the
|
||
> handler. **The nightly path deliberately still returns nil** — nobody asked, and it retries.
|
||
|
||
> **2026-08-05 — v0.200.0 (R-193 CLOSED).** The customer-facing recovery screen. Until now a customer
|
||
> whose machine was rebuilt had everything needed to get their data back and no way to find out — the
|
||
> only route was a command line.
|
||
>
|
||
> **IT UNLOCKS AND ONLY UNLOCKS** (operator ruling). Explains, takes the recovery code, opens the
|
||
> repository, lists what is in it (apps, dates, sizes). **Restores nothing** — restore is per-app and
|
||
> lives in the backups area; the put-back is **R-213** and its stated requirement is a
|
||
> live-versus-backup comparison.
|
||
>
|
||
> **ONE CORE, TWO CALLERS.** `backup.RecoverInstallCore` is the only fetch→unseal→compare→install path.
|
||
> `RecoverAndInstall` is now a thin CLI wrapper — exit codes and printed lines byte-identical, every
|
||
> pre-existing CLI test passed unchanged — and the handler calls the same function. Asserted from
|
||
> source by AST on BOTH sides, plus a test that the routes and the landing-page interception exist.
|
||
>
|
||
> **THE TRIGGER HAS TWO SHAPES and the second is the one that matters.** `OffsiteRecoveryOffer` = the
|
||
> hub holds a package AND (no repository password OR the tier is orphaned). The literal "no repository
|
||
> password" alone is a window that CLOSES BY ITSELF — `WriteOffboxSecrets` auto-generates one on
|
||
> re-apply (R-193's own orphaning mechanism) and hub v0.96.0's self-heal re-applies within ~15–30 min.
|
||
> Shape (b) is also what the shipped move-aside requires, which is why the discard choice can reach it.
|
||
>
|
||
> **CLAIMED is part of the predicate** — a legacy-open box passes through `RequireAuth`, so without an
|
||
> explicit `authEnabled()` check the interception fired for an unauthenticated visitor. A test caught it.
|
||
>
|
||
> **„Most nem" suppresses the FULL PAGE ONLY.** The backups-area entry point is bound to
|
||
> `recoveryOffer`, never to the postpone flag.
|
||
>
|
||
> **The code:** POST body only, never logged/persisted/echoed, cleared on every path, `no-store`,
|
||
> `autocomplete=off`. **No lockout** — a ten-word phrase is not guessable and locking a customer out of
|
||
> their own data for a typo is worse; failures are logged locally without the code, and NO operator
|
||
> alert is raised (reasoning in REPORT.md §4).
|
||
>
|
||
> *Live:* demo-felhom is genuinely in shape (b), so validation needed no arrangement — `/launcher` →
|
||
> 302 `/recovery`, both mandatory sentences rendered, three wrong codes refused with the `offbox/`
|
||
> listing byte-identical and no lockout, and the code found in no file, log or ring **with a
|
||
> planted-copy positive control that first exposed a mis-aimed sweep**. **NOT proven live: a CORRECT
|
||
> code** — none was kept for demo-felhom's orphaned history and demo-hp's is operator-held.
|
||
|
||
|
||
> **2026-08-05 — v0.199.0 (R-204 item 4 / R-193).** The last of the four manual interventions the
|
||
> 2026-08-04 drill needed. **Operator ruling: automate it, and the trigger is a state the BOX
|
||
> DECLARES.** From the hub an absent off-site object has FOUR meanings — never configured,
|
||
> mid-restart, a transient config read failure, rebuilt-and-stranded — and the hub cannot tell them
|
||
> apart. The box can.
|
||
>
|
||
> **The declaration needs BOTH halves** (`backup.needsOffsiteCredential`): a fresh data area (no
|
||
> repository password) AND a hub-held recovery package (the ACK's `identity_blob_present`). Freshness
|
||
> alone is a box that never had off-site backups; dropping that condition makes the whole fleet ask
|
||
> for credentials, which is what `TestOffsiteDeclare_NeverHadOffsiteSaysNothing` catches. A merely
|
||
> DISABLED target is the customer's own choice and never declares.
|
||
>
|
||
> **The ACK field stopped being discarded.** `EscrowAutoConfirmer.Reconcile` returns early when the box
|
||
> is neither pending nor escrowed — exactly a rebuilt box — so the fact was thrown away every cycle. It
|
||
> is recorded FIRST, before every gate, via `RecordPresence`, wired in main.go and asserted by
|
||
> `TestMainWiresRecordPresence` (AST, comments dropped). Last-write-wins, not set-only, so a customer
|
||
> RESET turns the declaration back off; a nil ACK escrow records nothing.
|
||
>
|
||
> **Inert to every existing reader:** `enabled:false` + zero sizes, so the hub's `isStale` and
|
||
> `fillBand` both short-circuit; an unknown `state` string is ignored by encoding/json. **A configured
|
||
> box's report JSON is byte-identical to v0.198.0's.** The one reader that would have misread it is the
|
||
> hub's `reportHasOffsite`, tightened in hub v0.96.0 to require `enabled:true`.
|
||
>
|
||
> *Live:* both demo boxes now record `hub_escrow_identity_present=true` in settings.json (the recorder
|
||
> working on a HEALTHY box). demo-felhom 9201, arranged reversibly into the stranded shape, produced
|
||
> report id=16743 carrying `{enabled:false, state:needs_credential, quota_gb:0, repo_size_bytes:0}`;
|
||
> the single declaration was absorbed by the hub's debounce (no self-heal event) and the box was
|
||
> restored the same minute. **The hub half is felhom.eu v0.96.0.**
|
||
>
|
||
> *Rider:* `.githooks/pre-push` in all four repos now refuses a push from a clone outside
|
||
> `/mnt/5_hdd/felhom.eu`. Proven both ways against a scratch clone.
|
||
|
||
|
||
> **2026-08-05 — v0.198.0 (R-204 items 1 & 3).** The 2026-08-04 drill (R-201) passed only because a
|
||
> person was there; four manual interventions stood between a recovered key and a restored file. Two
|
||
> of the three defects are in this repo.
|
||
>
|
||
> **Item 1 — the reset code needed a restart.** `--print-reset-code` is a SEPARATE process; it
|
||
> persisted a new code while the running server kept the old one cached, so the code the customer was
|
||
> told to type was refused until the controller restarted, and nothing said so. `effectiveClaimCode`
|
||
> now calls `settings.ReloadClaimCode()` first. **The settings-vs-config precedence is unchanged** —
|
||
> the defect was freshness, not precedence. **Read-through, not a TTL, and that is the point:** a TTL
|
||
> makes the new code visible AND leaves a window in which the superseded one still works, which is
|
||
> worse than the bug. That is the mutation `TestClaimCode_SupersededByASecondMint_RefusedImmediately`
|
||
> exists to kill, and its red-proof produced exactly *"the SUPERSEDED code was accepted"*. The
|
||
> function now returns an error and **every caller fails closed**; an absent settings file is NOT an
|
||
> error. `ClaimConsumedGeneration` is deliberately NOT re-read — this process is its only writer and
|
||
> re-reading could move it BACKWARDS if a save had failed, resurrecting a consumed code.
|
||
>
|
||
> **Item 3 — the restore's default returned the wrong thing silently.** `mode=unit` restores the
|
||
> recovery unit (definition + config + DB dumps) and not the customer's files. `restoreScratchOutcomeMsg`
|
||
> now names what came back, what did not, and the next step; the wizard's intent card states its scope
|
||
> before the choice. **The size gate is untouched** and pinned unchanged by
|
||
> `TestOffboxRestore_FullPathUnchanged`. **The default stays `unit`** — all three wizard forms set
|
||
> `mode` explicitly, so a change would alter nothing visible while silently changing a mode-less POST.
|
||
>
|
||
> *Live-validated endpoint-level (no browser on DooPlex):* on demo-felhom 9201 with `restarts=0`
|
||
> across both mints, a superseded code returned „Hibás vagy lejárt kód" and the current one was
|
||
> accepted first time; on demo-hp 9201 a `privatebin` unit restore produced the scoped Hungarian
|
||
> outcome and `mode=full` without confirm revealed `full_size=6.8+KB` without restoring anything.
|
||
> demo-hp's drill scratch (`calibre-web`) was not touched.
|
||
>
|
||
> **Item 2 is the hub's** (felhom.eu v0.95.0, R-196). **Item 4 — a rebuilt box cannot obtain an
|
||
> off-site credential unaided — remains OPEN (R-193)** and was deliberately not begun.
|
||
|
||
|
||
> **2026-08-02 — v0.190.0 (R-157 mechanism A · R-170 · R-171).** Three items, one live validation
|
||
> cycle, because all three are boot behaviour and all three are proven by hard-resetting the box.
|
||
>
|
||
> **DIAGNOSE BEFORE THEORISING — and the first diagnosis was a FALSE NEGATIVE.** A hole was reasoned
|
||
> out of the v0.189.0 diff (a drive-gate-stopped app has zero containers and `desired_state: running`,
|
||
> so it now reads as a boot orphan) and confirmed on hardware BEFORE any fix was written. **Attempt 1
|
||
> produced `no boot-orphaned apps` and would have been reported as a disproof.** It was a race:
|
||
> unmounting only the parent bind is healed by the agent within ~60 s, so the drive gate's startup
|
||
> reconcile restarted the apps **one second before** the sweep looked. Holding the drive genuinely
|
||
> absent reproduced the defect immediately. **"It didn't happen this time" is not a mechanism.**
|
||
>
|
||
> **The confirmation moved the severity in BOTH directions.** The write hazard did not materialise —
|
||
> compose failed `mkdir …/userdata: permission denied` because the unbound mountpoint is
|
||
> host-root-owned and the guest is unprivileged. **That protection is ACCIDENTAL**: no code chose it,
|
||
> no test pinned it, and it is one `chown` or one privileged guest away from gone. But the harm that
|
||
> DID occur was not in the hypothesis and is real on every box: two wasted attempts and a **false
|
||
> dead-app alarm for an app the drive gate is deliberately holding**.
|
||
>
|
||
> **The fix already existed one path over.** `startGatedByMissingDrive` (the API) refuses a customer's
|
||
> start on an absent drive; the sweep bypassed it by calling `Manager.StartStack` directly.
|
||
> **`StartStack` HAS NO GATE OF ITS OWN** — carry this: every caller that is not the customer must
|
||
> decide for itself whether the app may run. New consumer-side `bootrecon.StartGate`, fail-safe
|
||
> (cannot determine ⇒ do not start).
|
||
>
|
||
> **Widening a window makes previously-unreachable overlaps reachable — a design input, not an
|
||
> afterthought.** The old T+5 s sweep never met a quiesce or an in-flight app-data operation; a 50 s
|
||
> window can. All three holders answer ONE seam because they differ only in the reason string.
|
||
>
|
||
> **A TEST REJECTED MY FIRST CONSTANT, and the comment says so.** `settle + budget + one retry` must
|
||
> fit inside `deadAppBootGrace`; 60 s gave 95 s against 90 s. The budget is 50 s **because a test said
|
||
> so** — recorded in the code rather than presented as taste. Widening the grace was rejected: it
|
||
> hides a late recovery instead of reporting one (`recordLateRecovery`).
|
||
>
|
||
> **THE FIX HAD ITS OWN DEFECT, FOUND LIVE AND NOT BY REVIEW.** The window sampled `GetStacks()` — the
|
||
> Manager's map, refreshed by the scheduler every **10 s** — every 5 s, so two identical samples could
|
||
> mean *the cache did not update*. Observed: a container removed ~5 s before the window closed was
|
||
> still in the sampled fleet and the sweep logged `no boot-orphaned apps` for an app that had none.
|
||
> `sampleBootFleet` now refreshes first. **Generalise: a settle detector is only as good as the
|
||
> freshness of what it samples — if the source is cached, refresh it, or you are watching the cache
|
||
> settle rather than the system.**
|
||
>
|
||
> **R-170:** `shouldRecreateOnBoot` reads intent with the identical three-way table; absent keeps the
|
||
> old `hasContainers` behaviour exactly; `presentStable` untouched and still load-bearing. Its comment
|
||
> argued at length FOR the count and was rewritten. Agreement pinned from BOTH sides against one
|
||
> fixture table (an import cycle prevents testing the two gates together).
|
||
>
|
||
> **Live: 6/6 hard resets** (every app back; the customer-stopped app down all six), settle times
|
||
> 10/40/10/10/15/15 s. Sharpest evidence: same app, same box — missed at 18:08:35, recovered at
|
||
> 18:18:50. R-170 proven in one reboot (calibre-web recreated, immich left stopped). 27/27 packages;
|
||
> 7 red-proofs. Detail: `REPORT.md`.
|
||
|
||
Last updated: 2026-08-02 (v0.189.0 — R-166 / D-b: the box stops guessing what the customer wanted)
|
||
|
||
> **2026-08-02 — v0.189.0 (R-166, operator decision D-b).** When an app was not running the box had
|
||
> to work out *why*, and it did so **by counting containers**: zero meant "the customer stopped it",
|
||
> some meant "something broke". A **power cut mid-compose** and an **interrupted deploy** also leave
|
||
> zero containers, so both were read as deliberate stops and stranded **silently** (R-157 mechanism
|
||
> B) — and a backup that stopped an app and died left it stopped with **nothing on disk** recording
|
||
> that it was owed a restart. The settling fact — what the customer asked for — **was written down
|
||
> nowhere**: `app.yaml` recorded *installed*, never *meant to be running*.
|
||
>
|
||
> **DECISION — one owner: the customer's action, and nothing else.** A census found **14 callers of
|
||
> `StartStack`/`StopStack`, of which exactly 2 are the customer**; the rest are quiesce, the volume
|
||
> dump, offbox reconstitution, app export/restore, the storage gate, migration and the boot
|
||
> reconciler. So the primitives are deliberately **not** writers — intent there would make a nightly
|
||
> backup indistinguishable from the customer pressing Stop. Writers: the API action switch,
|
||
> `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, the `.fab` import. Intent is written
|
||
> **BEFORE** the act and a failed write **REFUSES** the act.
|
||
>
|
||
> **DECISION — absent means UNKNOWN, never "running", and this is the whole safety property.** Every
|
||
> `app.yaml` on every box predates the field, so absent is what the fleet reads on upgrade; reading
|
||
> it as running would start every deliberately-stopped app on the first boot after the upgrade. The
|
||
> legacy branch of `isBootOrphan` keeps the old container-count rule **byte-for-byte**, and its test
|
||
> asserts BOTH legacy rows together because the safety property is the pair. Backfill is
|
||
> **running-only** — "zero containers ⇒ stopped" IS the defect, so an ambiguous app stays ambiguous.
|
||
>
|
||
> **Part 2 — `backup.AppStopGuard`**, a persisted marker over every stop→work→start window (volume
|
||
> dump, offbox reconstitute, `.fab` export), in its **own** file (one file, one writer). Written
|
||
> before the stop, cleared only after a restart that **succeeded**, kept when one fails. `Recover`
|
||
> **returns** its outcome instead of using a notifier seam, because it must complete before the boot
|
||
> reconciler (`main.go` ~236) while the notifier is not built until ~307 — a seam wired after the
|
||
> fact is a seam that never fires.
|
||
>
|
||
> **THE TEST LESSON, and it is the one worth carrying:** Scenario E's first version called
|
||
> `appStop.Begin` itself, and **survived the red-proof that deleted the production call**. It proved
|
||
> the marker type, not that `DumpAppVolumesSafe` uses it. Rewritten to drive the real function with a
|
||
> simulated hard abort (an unwind that skips the restart statement, since a `defer` is not
|
||
> crash-safety — Campaign 8 fault 10). **A test that constructs the thing it is meant to prove the
|
||
> caller constructs is hollow, and its red-proof will say so if you run it.**
|
||
>
|
||
> **FOUND EN ROUTE — `SaveAppConfig` rebuilt `AppConfig` field-by-field**, the R-100 shape (v0.181.0
|
||
> shipped two live instances). The literal named five fields, so `desired_state` would have been
|
||
> dropped on **every** save across nine call sites — a customer's Stop erased by the next unrelated
|
||
> `app.yaml` write. Copy-and-overlay (`saveCfg := *cfg`) is safe by construction. **Generalise it:
|
||
> treat any field-by-field struct rebuild in a save path as a defect on sight.** Measured, not
|
||
> assumed: `app.yaml` does NOT round-trip YAML keys the struct does not model (pinned by test).
|
||
>
|
||
> **R-157: mechanism B closed, mechanism A untouched** (the sweep observes ~5 s after start and never
|
||
> re-checks) — and B's fix makes A cost more, since the sweep now has more it could recover.
|
||
> **NEW R-170:** `shouldRecreateOnBoot` (`internal/web/intermediary.go:131`) still infers a Stop from
|
||
> `hasContainers` — the same defect one gate over, for drive-backed apps. Left deliberately.
|
||
>
|
||
> **Live on 9201, three flows** (stop survives a restart; a zero-container `running` app recovered by
|
||
> name; a legacy app.yaml skipped and never inferred stopped). The **interrupted-operation half is
|
||
> IMPLEMENTED, not PROVEN-LIVE** — nobody killed the controller mid-backup on metal. 27/27 packages;
|
||
> 7 red-proofs observed FAIL then restored. Detail: `REPORT.md`.
|
||
|
||
Last updated: 2026-07-28 (v0.182.0 — R-101 + F-DIAG: the restore dialog names the last SUCCESSFUL copy)
|
||
|
||
> **2026-07-28 — v0.182.0 (R-101 + F-DIAG).** `Tier2LastRun` is the ATTEMPT clock (written on failure)
|
||
> and was rendered as „Legutóbbi másolat" in the **restore confirm dialog** — misinformation at a
|
||
> decision point: the restore fills in MISSING files, so a customer with a failing Tier-2 restored and
|
||
> silently got OLDER files. New `CrossDriveBackup.LastSuccess` + **`SuccessTracked`**; the marker is
|
||
> load-bearing because **all 7 fleet rows were pre-anchor at deploy** — without it every customer sees
|
||
> „Még nincs sikeres másolat" at once. Legacy rows migrate on first touch (`ok` adopts its time,
|
||
> `error` seeds nothing). **PART 2 — the three `record*` helpers rebuilt the WHOLE struct with only 2
|
||
> fields carried over; the naive fix would have had `recordTier2Failure` CLEAR the anchor.** Replaced
|
||
> by `tier2Update` (copy-and-overlay = safe by construction). New `fmtTimeStr` → Budapest-local dates
|
||
> in the dialog instead of raw UTC RFC3339. **F-DIAG:** 6 classes incl. an honest `unknown`, and the
|
||
> notification no longer passes `err.Error()` through raw — **LESSON: my first sanitiser was regex-only
|
||
> and leaked a bare hostname; its own test caught it. Redact KNOWN values, don't guess at shapes.**
|
||
> Live on demo-hp: rendered dialog read in the failed, healthy AND legacy states. F-OPS documented at
|
||
> `felhom.eu/documentation/runbooks/RUNBOOK-manual-guest-restore.md`.
|
||
|
||
> **2026-07-28 — v0.181.0 (R-100).** `OffboxTarget.LastSuccess` + wire field `last_success`; the hub
|
||
> (v0.80.0) anchors offsite staleness on it. **`LastRun` is written unconditionally on every run
|
||
> INCLUDING failures** — it records an ATTEMPT — so the hub's "how long since LastRun" verdict read a
|
||
> nightly-failing tier as perfectly fresh forever. The rule is the pure `offboxAnchorAfterRun(prev, at,
|
||
> runErr)`: a failure neither ADVANCES nor CLEARS the anchor (both are distinct bugs; clearing it would
|
||
> make one bad night look like never-succeeded). `LastStatus == "error" ⇒ stale` was rejected — it pages
|
||
> on every blip, the F-A1 noise mode. **TWO SILENT-WIPE SITES CLOSED** (`offboxConfigHandler` and
|
||
> `ApplyOffsiteTarget` both rebuild the target and copy runtime status field-by-field — omitting
|
||
> LastSuccess would erase the anchor on any settings save or hub re-apply). **LESSON: my first test
|
||
> modelled the rule in a local closure and stayed GREEN when production was mutated — hollow; the
|
||
> extraction to a pure function is what made the red-proof bite.** Live on demo-hp: failing run advanced
|
||
> `last_run` to 11:25:48Z while `last_success` HELD at 11:24:20Z; demo-felhom healthy → advanced. The
|
||
> settings-save preservation was proven live too. Detail: `REPORT.md` + `felhom.eu/REPORT-r100.md`.
|
||
|
||
> **2026-07-28 — v0.180.0 (F-OBS).** Source: `audits/CAMPAIGN-8-backup-restore-2026-07-27.md`.
|
||
> On a default `logging.level: info` box there was **no positive observable that `deadapp-check` had
|
||
> run**: its per-cycle line goes through `Scheduler.dbg()`, gated on `level==debug`, so on a default
|
||
> box it was never *produced* and could not even reach the always-DEBUG ring. "No alarms" was
|
||
> therefore indistinguishable from "the detector never ran" — standing rule 3's exact fallacy, and it
|
||
> undermines F-CRIT-1's fix, which is a fix to **this same detector**.
|
||
> `noteDeadAppScan()` now emits an INFO line every **20th** scan (10 min at the 30 s cadence) carrying
|
||
> scans-since-boot / evaluated / currently-down. It reports **what it saw**, not that it ran, and it
|
||
> summarises rather than floods — one line per run is 2880/day, which is what made silence attractive
|
||
> in the first place. Both bounds are pinned by test in the direction that would break them.
|
||
> **The same shape then turned up in the agent's brand-new guest-power watchdog** (v0.107.0, shipped
|
||
> hours earlier): it logged only at startup and when it acted. Fixed in agent v0.109.0 with the same
|
||
> pattern. The anti-pattern reproduces itself — which is the argument for not having dropped this part.
|
||
> Live on demo-hp at INFO on a default-level box; deployed on both boxes. Detail: `REPORT.md`.
|
||
|
||
> **2026-07-26 — v0.173.0 (R-77).** Source: `audits/DIAG-agent-channel-2026-07-26.md`.
|
||
>
|
||
> **UNRESOLVED AND DELIBERATELY DEFERRED — which file is authoritative for `local_api`?** R-77 ships
|
||
> DETECTION ONLY. `controller.yaml` and `bootstrap.json` can disagree; the controller dials
|
||
> `controller.yaml`. The obvious "fix" — reconcile from `bootstrap.json` on every boot — has a failure
|
||
> mode **as severe as the bug it fixes**: on a guest whose `controller.yaml` is correct and whose
|
||
> `bootstrap.json` is stale (a re-provision that half-completed, a hand-repaired guest, a
|
||
> setup-wizard box), auto-reconcile would clobber a WORKING channel on the next restart — fleet-wide,
|
||
> silently, at the moment of a routine deploy. R-77's position is that **naming the drift is enough**:
|
||
> it would have converted the 17.5 h outage into a specific alert on the first health cycle. The
|
||
> authority ruling is **R-78** and needs its own spike — do not resolve it opportunistically.
|
||
>
|
||
> Corollary for anyone editing `bootstrap.MaybeIngest`/`ensureLocalAPI`: `ensureLocalAPI` is the ONLY
|
||
> writer, it fires only when the endpoint is EMPTY, and `DetectEndpointDrift` must stay write-free.
|
||
> Scenario A's test asserts `controller.yaml` is byte-identical after the check, and its red-proof
|
||
> covers the auto-correcting variant precisely because that is the tempting wrong turn.
|
||
>
|
||
> **Also settled here:** the samba protected-set must mirror EVERY early return in
|
||
> `reconcileSambaAt` (currently two: `!smb.Enabled`, `!smb.UserSet`). A third would need the same
|
||
> mirror, and the doc comment above `EffectiveProtected` must be updated with it.
|
||
|
||
|
||
> **2026-07-26 — v0.172.0 (R-75).** Spike `felhom.eu/documentation/audits/SPIKE-catalog-data-paths-2026-07-26.md`;
|
||
> feature doc `felhom.eu/documentation/controller/import-and-data-paths.md`.
|
||
>
|
||
> **RULING — the import root is CANONICAL on the system drive, overriding the spike's Fork-1
|
||
> recommendation of per-drive roots.** The spike weighed sidebar clutter and per-app link ambiguity and
|
||
> concluded per-drive; the operator overruled it on an argument the spike missed: each drop-zone app has
|
||
> exactly ONE ingest bind, so on a two-drive box every import folder except the app's own would look like
|
||
> a drop-zone and silently do nothing — and because `import/*` is `class: excluded`, files stranded there
|
||
> are never backed up either. A canonical root is the only shape with no dead drop-zone. Recorded as a
|
||
> deliberate deviation, not an oversight.
|
||
>
|
||
> **Phase-0 probe changed the shape of Part 6.** The system drive is NOT a registered `StoragePath` on
|
||
> either demo box (`/mnt/felhom-drives/hdd_1` on demo-felhom; `nvme-1tb` + `Felhom-Share` on demo-hp),
|
||
> so `sharingResolvePath` REFUSES `<sysroot>/userdata/import` — verified against the real guard with a
|
||
> passing control. Registering the drive was rejected (it would make the 50 GB volume holding the
|
||
> recovery units a customer-visible drive, deploy target and wipe candidate, and `SharingDeniedRoots`
|
||
> would then deny the namespace-consistent shape anyway). **Chosen: leave it unregistered and have the
|
||
> controller write the `beolvasas` share directly** — the picker guard validates CUSTOMER-supplied paths,
|
||
> a controller-generated constant is a different trust class. No guard was weakened.
|
||
>
|
||
> Also note: `withUserdataPath` computes `USERDATA_PATH` as `<hdd>/userdata`, NOT
|
||
> `NamespaceRoot(hdd)/userdata`. For an app on the system drive those disagree
|
||
> (`/mnt/sys_drive/userdata` vs the `felhom-data` namespace). Latent — no app with a userdata bind has
|
||
> ever been deployed there — but it is a real inconsistency, left untouched here.
|
||
>
|
||
> The other three forks followed the spike unchanged: all-apps skeleton / deployed-only in the UI;
|
||
> unknown role fails OPEN while a malformed path whole-block rejects; drop-zone copy driven by the
|
||
> derived backup class.
|
||
|
||
|
||
> **2026-07-24 — v0.169.0 (disk-health card + degradation alert).** Consumes the agent's new `smart`
|
||
> field (agent v0.94.0; MinAgent floor unchanged — feature-detect by presence). **Rulings:** (1) ONE
|
||
> pure verdict fn `agentapi.DiskVerdictFor` is the shared truth for the card chip AND the 6h check — they
|
||
> can never disagree. Thresholds: FAILING→Hiba; PASSED + any(reallocated>0/pending>0/offline_unc>0/
|
||
> critical_warning>0/media_errors>0/percentage_used **≥90**)→Figyelmeztetés; PASSED clean→Rendben;
|
||
> nil/UNKNOWN→Nincs adat (never alarms). (2) **No global alert banner** — the card + email carry disk
|
||
> health; banner fatigue is a real cost, so this is deliberately NOT wired into the dead-app/alert-banner
|
||
> machinery. (3) Degradation-only notification with an in-memory baseline: first run baselines silently,
|
||
> recovery never notifies, **UNKNOWN excluded both directions** (a transient blip neither fires nor erases
|
||
> history). (4) **Controller restart re-baselines silently** (in-memory baseline lost on restart) — an
|
||
> accepted trade consistent with the health-change pattern (a real post-restart degradation still fires on
|
||
> the following 6h check once a baseline exists). (5) A **60s TTL cache** wraps the card's /disks call so
|
||
> dashboard refresh-spam can't smartctl-storm the host; the 6h check fetches FRESH (cache-independent).
|
||
> Pairs with hub +1 (allowlist `disk_health_degraded`). No new smartctl load — serialization only.
|
||
|
||
Last updated: 2026-07-24 (v0.168.0 — customer-configurable backup window "Mentési időablak")
|
||
|
||
> **2026-07-24 — v0.168.0 (customer-configurable backup window).** ONE customer setting — the window
|
||
> start W ("Mentési időablak kezdete") — drives every nightly leg at FIXED, never-stored offsets so
|
||
> misordering is impossible: DB dump at W, tier-2 at W+60m, off-box at W+105m (wrap-safe). **Design
|
||
> rulings:** offsets are DERIVED and computed everywhere, never persisted and never exposed in the UI;
|
||
> precedence is settings > controller.yaml `db_dump_schedule` > "02:30" (mirrors PasswordHash); a change
|
||
> applies WITHOUT restart via the new scheduler seam `UpdateDaily` (per-daily-job buffered `resched`
|
||
> chan + a select case in `runDailyJob`). New pure package `internal/backupwindow` holds all the time
|
||
> math (ParseHHMM/FmtHHMM/LegTimes/GateWindow/EffectiveWindow). **Disk-tier (whole-guest PBS/vzdump)
|
||
> gate:** the quiesce loop's SCHEDULED cycles run only inside [W+2h, W+6h) (wall-clock Europe/Budapest),
|
||
> with a safety valve — last successful backup older than cadence+24h (or none) runs regardless, so a
|
||
> box only ever on outside its window never starves. **Manual "Mentés most"/TriggerNow is NEVER gated**
|
||
> (bypasses runOnce). The `quiesce.Backend.Due` seam now also returns the backup age (from the agent's
|
||
> own `/backup/due`); the agent, its cadence, and `/backup/due` are untouched. Window read fresh each
|
||
> poll (WindowStartFn) so runtime changes take effect. Cadence defaults to 24h controller-side (the
|
||
> response carries no cadence). Backup page gets a "Mentési időablak" card (time input + derived rows +
|
||
> the "kb. W+2h–W+6h között" rendszermentés line); POST /backups/window (RequireAuth+CsrfProtect).
|
||
|
||
|
||
> **2026-07-24 — v0.167.0 (outlined logo + favicon — Part 4 unblocked).** Viktor pushed the
|
||
> text-outlined `logo.svg` to felhom.eu `main` (`be9edb4`); the wordmark is now 17 real `<path>`
|
||
> glyphs. `FelhomLogoSVG` swapped to it; Inkscape's leftover **empty `<text/>` shells + font-* leftovers
|
||
> on the paths** were stripped via an lxml DOM pass (glyphs untouched — CC did NOT do text-to-path),
|
||
> editor `<sodipodi:namedview>` dropped. `FelhomFaviconSVG` vestigial `<text>` removed. Both constants:
|
||
> **0 `<text`, 0 `font-family`**; viewBoxes unchanged; palette + 14 gradients preserved. **Gotcha logged:
|
||
> Inkscape "Object→Path" leaves empty `<text/>` shells AND copies `style="…font-family:…"` onto the
|
||
> resulting `<path>`s — a search for `svg:text` misses them (elements are `<text>`, no prefix); grep
|
||
> `<text` and `font-family`.** Also: `serveLogoHandler`/`serveFaviconHandler` serve a **hub-synced file
|
||
> first** (`assetsSyncer.Resolve`) and fall back to the constant only if none is on disk — on 9201 the
|
||
> constant is what's live (verified). Still open (separate follow-up): website + hub serve their own
|
||
> non-outlined logo copies; login.html stylesheet link still unversioned.
|
||
|
||
|
||
> **2026-07-24 — v0.166.0 (mobile nav = off-canvas drawer; sidebar cleanup; ?v= on logo/favicon).**
|
||
> Mobile nav was broken: the ≤768px block predated the v0.146.0 accordion and flattened `.nav-links`
|
||
> into a horizontal `overflow-x` strip, clipping the accordion's nested sub-lists (they share the
|
||
> `.nav-links` class). **Decision: mobile nav = a sticky top bar + off-canvas left drawer that REUSES
|
||
> the vertical sidebar (Option A).** The accordion handler is untouched and works inside the drawer;
|
||
> a `no-js` html-class fallback renders the sidebar static inline so nothing dead-ends without JS.
|
||
> Options B (separate mobile menu) and C (exclude nested lists from the strip) were rejected. z-index
|
||
> ladder topbar 800 < backdrop 900 < drawer 950 < modal 1000; `100dvh`; reduced-motion disables the
|
||
> slide; focus-trap deliberately omitted (navigations reset state). **Sidebar customer-name removed**
|
||
> (logo only); `{{.CustomerName}}` stays in base data + login subtitle. **Logo policy decision: the
|
||
> wordmark must be OUTLINED paths, never live `<text>`** — under `<img>` secure static mode only
|
||
> locally-installed fonts resolve, so `font-family` in the SVG renders a fallback font everywhere.
|
||
> **Part 4 (swap `FelhomLogoSVG`/`FelhomFaviconSVG` to the outlined master) is GATED OUT** — §3a check
|
||
> against live felhom.eu `main` (`be9edb44`) found `website/assets/logo.svg` still has `<text>`/
|
||
> `font-family`; the outlined master is Viktor's manual Inkscape push, still pending. Only the `?v=`
|
||
> cache-bust (logo/favicon/login-logo, Cloudflare 4h edge-cache — the 0.126.1 failure mode) shipped
|
||
> from the logo work. Follow-up: when Viktor pushes the outlined asset, ship Part 4 (swap constants +
|
||
> clean the favicon's vestigial `<text>` nodes). Separately, the website + hub still serve their own
|
||
> non-outlined logo copies — propagation is a distinct follow-up.
|
||
|
||
|
||
> **2026-07-24 — v0.165.1 (native "Megosztás…" in the share modal, Web Share API).** The share modal
|
||
> gains a feature-detected `navigator.share` button (OS share sheet → Messenger/WhatsApp/email),
|
||
> sending **title + text + URL only**. Hidden unless supported; "Link másolása" stays the universal
|
||
> fallback (and catches the non-cancel rejection); `AbortError` (user cancel) is silent. **Ruling: the
|
||
> QR is NOT attached** (no Web Share Level-2 `files:`) — file-share support is narrow and several
|
||
> targets drop the URL when handed file+URL, leaving an unscannable QR picture in a chat; the QR's job
|
||
> (physical cross-device scanning) is already served by the modal image (mobile long-press). Template
|
||
> JS + tests only; the OS sheet interaction is an operator manual check (not endpoint-testable).
|
||
|
||
> **2026-07-24 — v0.165.0 (Indítópult megosztása — guest launcher via capability URL).** The admin
|
||
> launcher gets an "Indítópult megosztása" button that mints a **capability URL**
|
||
> (`https://<host>/s/<token>`, 160-bit `crypto/rand` token) serving a standalone, read-only guest
|
||
> launcher — same tiles, opens apps in new tabs — with **no account and no admin session**. **Security
|
||
> ruling: the link grants INFORMATION ONLY, ZERO CONTROL** — app names + public URLs; every privilege
|
||
> stays behind each app's own auth and the controller admin password. The token IS the secret (160-bit
|
||
> entropy is the whole defence for the GET — never rate-limited, never logged, `subtle.ConstantTimeCompare`
|
||
> only; an empty stored token = sharing OFF, matches nothing, so a wrong/disabled token is byte-identical
|
||
> to the mux default 404). Optional per-share password is a SEPARATE credential (own bcrypt hash, own
|
||
> attempt map — NEVER the admin ones); one pass mints a cookie = HMAC(`token|passwordHash`) keyed with
|
||
> the persisted `web.session_secret`, so rotate-token OR change-password invalidates all cookies for free.
|
||
> **Part-2 secret decision: REUSED `web.session_secret`** (persisted + box-scoped + stable — the SAME
|
||
> secret the claim pre-auth CSRF already trusts; not per-boot, not claim-generation-scoped → the reuse
|
||
> branch), so no `ShareCookieSecret` field was added. **Design rulings recorded:** member accounts are
|
||
> **superseded** by this capability-URL model; **per-member tile visibility is PARKED under the SSO arc.**
|
||
> Guest state labels ride the v0.164.0 invariants: `StateStopped` ⇒ "A tulajdonos leállította"; any
|
||
> other non-clickable state ⇒ "Átmenetileg nem elérhető" (guests never see stopped/exited/degraded/
|
||
> unhealthy). Accepted residuals (documented, no code action): link-preview crawlers fetch once and see
|
||
> app names (noindex prevents indexing); reverse-proxy/CF access logs may hold the path (ops-tier); the
|
||
> modal link carries the request Host, so a LAN-IP admin session yields a LAN-IP link. New dep:
|
||
> `github.com/skip2/go-qrcode`. Tests: Groups A–G (14 tests) + 3 red-proofs verified red.
|
||
|
||
> **2026-07-24 — v0.164.0 (stopped ≠ fault).** Operator finding on 9201: a UI stop (Leállítás) raised
|
||
> the global "Telepített alkalmazás nem fut: … (stopped)" banner on every page AND fired the
|
||
> `app_start_failed` email. RULING: **a deliberate user action must not alarm anywhere.** One-line
|
||
> filter at the single fix-3 derivation point — `scanDeployedAppRunStates`'s pure core extracted to
|
||
> `classifyRunStates([]stacks.Stack)`, down predicate now
|
||
> `stacks.IsDownState(st.State) && st.State != stacks.StateStopped`. `StateStopped` is dropped from
|
||
> BOTH the banner dead-list and the notifier Down-set (⇒ no banner, no event, clean tracker). Rests on
|
||
> **two invariants that MUST both hold for this suppression to be correct:** **I1** — the UI stop path
|
||
> `Manager.StopStack` runs `docker compose down` → containers removed → a deployed stack with zero
|
||
> containers aggregates to `StateStopped` (refreshStatusLocked). **I2** — the P2 restart-policy census
|
||
> (2026-07-21, 53 templates / 78 services) found every catalog service on `unless-stopped`, so a crash
|
||
> never rests at `stopped` — faults surface as `exited`/`degraded`/`restarting`/`unhealthy`. **If
|
||
> either invariant changes, revisit this suppression.** `IsDownState` UNCHANGED (other callers rely on
|
||
> stopped=down). Out-of-band `docker compose stop` (containers remain → `StateExited`) still alerts —
|
||
> correct, tampering is reportable. The `stopped_by_user` intent flag was considered and PARKED (only
|
||
> adds value against out-of-band stops, which should keep alerting). Tests +4 (notify 3→4, main 4→7),
|
||
> both red-proofs verified. No template/funcmap/notifier/counter/copy change.
|
||
|
||
> **2026-07-24 — v0.163.1 (launcher polish).** Two v0.163.0 live findings fixed. RULE recorded:
|
||
> **every app-logo surface ends in a visible placeholder** (`SVG → PNG → /static/app-placeholder.svg`,
|
||
> infra rows → `infra-logo.svg`) — the four sibling `onerror` chains (`backups_apps`, `stacks`,
|
||
> `app_info` hero, `deploy`) now match `app_row.html`; `app_info` screenshots deliberately still
|
||
> vanish on error. And the **launcher monogram is launcher-only AND failure-only**: hidden by default,
|
||
> revealed when the tile's img chain fails (`onerror` adds `.launch-tile--noimg`) — it was bleeding
|
||
> through every transparent white glyph. Template/CSS only; no handler/funcmap change. 5 tests + 2
|
||
> red-proofs. [[launcher-v0163-2026-07-24]]
|
||
|
||
> **2026-07-24 — v0.163.0 (Indítópult app launcher + universal placeholder icon).** New
|
||
> customer-facing `/launcher` page: the FIRST sidebar item (above Vezérlőpult), a grid of large
|
||
> tappable tiles for openable deployed apps. `/` stays the Vezérlőpult — the launcher is ADDITIVE.
|
||
> Design rulings recorded here:
|
||
> - **(a) The felhom brand mark is NEVER an app placeholder** — brand = platform identity only. The
|
||
> logo-less fallback everywhere is the new generic `AppPlaceholderSVG` (a 2×2 app-grid glyph,
|
||
> `/static/app-placeholder.svg`), now the DEFAULT `FallbackIcon` on `app_list_row` (was
|
||
> `visibility:hidden`). On the launcher tile the fallback is the **monogram**, not the placeholder.
|
||
> - **(b) A launcher tile exists ⟺ a „Megnyitás" button would** — subdomain presence (env `SUBDOMAIN`
|
||
> > `.felhom.yml` subdomain > `protectedStackSubdomains`) is the single openability criterion. The
|
||
> controller stack is excluded by name. The subdomain assembly was extracted to
|
||
> `Server.subdomainMap` (3 callers: dashboard, Alkalmazások, launcher; priority byte-unchanged).
|
||
> - **(c) Colored-tile + mono-glyph design.** `tileColor` = validated `.felhom.yml` `brand_color`
|
||
> (`#rgb`/`#rrggbb`, new `Metadata.BrandColor`, omitempty) OR a deterministic FNV-1a-of-slug HSL
|
||
> (fixed S/L, hue per app). Invalid `brand_color` silently falls back to the hash color (the one
|
||
> §8 exception to no-silent-failure — cosmetic). `tileColor` returns `template.CSS` (we
|
||
> validate/compute in Go; html/template's CSS filter mangles a legit `hsl()` from a func pipeline).
|
||
> - **(d) `/` remains the Vezérlőpult.** No role/auth gating — member-role gating is a future arc
|
||
> (ROADMAP: member role → launcher becomes the member landing page). No catalog app sets
|
||
> `brand_color` yet (curation parked).
|
||
> No agent coupling; MinAgent unchanged. 10 new test functions + 4 red-proofs (all observed FAIL then
|
||
> restored). Gates green (app_row_dedup / template_id / emoji).
|
||
|
||
|
||
> **2026-07-24 — v0.162.0 (R-71a), SHIPPED + deployed BOTH boxes (demo-felhom 9201 + demo-hp 9201
|
||
> via G1 break-glass), clean+healthy, settle-gate GO line captured on both.** B′ live note: both
|
||
> above-floor boxes GOed correctly but NOT literally first-poll — the floor is in-memory (not
|
||
> persisted), unknown at t=0, so the gate logged `awaiting floor knowledge` then GOed ~10 s later the
|
||
> instant the report ACK landed (report-ACK latency = exactly what the 90 s sub-bound is sized to;
|
||
> zero-wait-when-floor-known is unit-proven, test E). The gate correctly did NOT burn the one-time
|
||
> password before the update picture was clear.
|
||
> The structural fix for the F10 day-0 race (DIAG-f10): the apply-bridge no longer consumes the
|
||
> single-use offsite password while a managed floor-update is in flight or imminent (below floor).
|
||
> New seam `offsiteapply.SettleProvider.SettleState()` + `SettleFunc` adapter over the updater's own
|
||
> `GetFloor()`/`IsUpdateRunning()` (no second floor path); `Bridge.AwaitSettle` polls 10 s BEFORE the
|
||
> 3-min Reconcile ctx (deferral never eats the reconcile budget), bounds 90 s floor sub-bound / 5 min
|
||
> overall (both GO+WARN — the "hub that can't serve a floor can't serve a consume → no burn" argument,
|
||
> R-71c is the belt). At/above floor → GO first poll, zero wait (B′). Bridge goroutine MOVED after the
|
||
> updater in main.go; wired only when an updater exists. **Ordering-only** — consume/persist/404
|
||
> contract untouched; R-71(b) rejected-by-design. **FINDING:** the floor is in-memory
|
||
> (report-ACK-derived ~5–10 s), NOT persisted → unknown on any restart until the first ACK (sized the
|
||
> 90 s sub-bound to that). 5 test scenarios (A–E) + nil-provider + cancelled-gate; **4 red-proofs all
|
||
> observed FAIL then restored** (gate/updateRunning/sub-bound/overall-bound). Deferral paths NOT
|
||
> live-fired (precondition now structurally prevented by the v1.25.0 build gate). **Layering: gate
|
||
> prevents, (a) defers, (c) heals.** ROADMAP R-71 → SHIPPED (a)+(c). Live leg = the B′ first-poll GO
|
||
> line on both above-floor boxes.
|
||
|
||
> **2026-07-23 — v0.161.0 (R-70 controller leg), SHIPPED + deployed BOTH boxes.** When
|
||
> `offsite.enabled` is in controller.yaml but no `offbox` target exists (pre-apply window / burned
|
||
> credential — the F10 shape), Távoli mentés now shows „Felhom offsite tárhely kiépítve — a
|
||
> beállítás automatikus, folyamatban…" on BOTH empty surfaces (status card + target line) instead
|
||
> of „igényelhető" / „Még nincs beállítva". Data key `OffsiteHubEnabled` (from `Server.cfg`, no new
|
||
> wiring); render tests per gate branch; banner leg is unit-proven/live-pending (no healthy box
|
||
> occupies the window; next fresh onboarding is the natural live leg). Hub sibling v0.72.0 carries
|
||
> the detector + `offsite_delivery_stuck` + the R-71c self-heal. Origin + rulings:
|
||
> `felhom.eu/documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`.
|
||
|
||
> **2026-07-22 — v0.160.0 (R-67), SHIPPED + deployed BOTH boxes, full live leg on demo-hp.**
|
||
> Network shares now bind their share ROOT into FileBrowser (`…/<name>:/srv/<name>:rslave`) — no
|
||
> skeleton/userdata toward the NAS, ever. Pure assembly = `buildFileBrowserPaths` + `fbPathDeps`
|
||
> (handlers.go), returning mounts AND config sources together so they can't disagree.
|
||
>
|
||
> **DECISION — two classes, two gates:** drives keep the drive-absent gate (byte-identical,
|
||
> tested + observed live: demo-felhom logged a no-op sync); network shares use the STUB classifier
|
||
> gate instead (stub ⇒ excluded from both lists + WARN — an exposed stub swallows uploads the real
|
||
> mount later shadows; idle autofs is HEALTHY and included; unknown fails open). Never force-wake
|
||
> in the sync (doctrine).
|
||
>
|
||
> **Phase-0 probe = GO:** in-container access through an rslave bind WAKES an idle autofs trigger
|
||
> (proved on demo-hp against the real Felhom-Share). Live leg: upload from demo-hp's filebrowser
|
||
> container (uid 1000) landed on demo-felhom's share dir and deleted clean; dead-NAS gave
|
||
> `Host is down` in seconds (no hang) and recovered unaided after samba restart. RESIDUAL for the
|
||
> operator: the FileBrowser HTTP click-through — its admin credential is customer-held (CC got 401
|
||
> on admin/admin and the demo password; by design). ROADMAP R-67 SHIPPED (coupled to R-64).
|
||
|
||
> **2026-07-22 — v0.159.0 (R-66), SHIPPED + deployed to BOTH boxes.** Three legs: „Hálózat" card on
|
||
> Beállítások → Rendszer (Helyi cím / Hálózati név only-while-Megosztás / Átjáró; „—" fallback),
|
||
> `network` section in the Debug dump (best-effort per item), and the NetBIOS trap named on the NAS
|
||
> add form (Szerver helper text + a purely lexical hint on `unreachable` for single-label non-IP
|
||
> names).
|
||
>
|
||
> **DECISION (the load-bearing one): all guest-net reads go through the samba netns door.** The
|
||
> controller is bridge-netns'd, so `/proc/net/route`/resolv.conf/net.Interfaces in-process answer
|
||
> for the CONTAINER (172.x / 127.0.0.11) — the S-2 trap. `internal/stacks/guestnet.go` docker-execs
|
||
> into host-networked felhom-samba (one `guestNetExecFn` seam); Megosztás off ⇒ door closed ⇒ „—" /
|
||
> in-place error strings, never a plausible-wrong substitute (S-5). Nothing stored anywhere.
|
||
>
|
||
> Deploy: 0.159.0 on demo-felhom 9201 (open-door path live: .104/.1/\\FELHOM) AND demo-hp 9201 via
|
||
> G1 break-glass (closed-door path live: dashes, no name row, in-place dump errors; secret shredded).
|
||
> demo-hp gotcha worth keeping: the controller 404s on direct container-IP probes without the
|
||
> customer-domain Host header (`felhom.enkisfelhom.hu` there). Red-proofs A2 + C2 run and recorded.
|
||
> ROADMAP: R-66 SHIPPED; R-64 (pairing blessed, drill = evidence leg) + R-65 (buddy-box replication,
|
||
> post-alpha spike-first) minted. NAS doc gained the naming-caveat paragraph.
|
||
|
||
> **2026-07-21 — v0.155.0.** v0.154.0's wizard sourced "is an op running" from `Manager.IsRunning()`
|
||
> — the CONCURRENCY single-flight, acquired inside the goroutine, and **`RestoreOffboxScratch` never
|
||
> acquires it**. So the execution step was unreachable for „Ellenőrzés" and the full-restore
|
||
> preparation: live buttons while a restore downloaded, with the progress banner contradicting the
|
||
> phase strip on the same screen. Found by the operator on the first live click-through.
|
||
>
|
||
> **DECISION: display reads `RestoreStatus()` (the `opRunning` flag), never `IsRunning()`**, through
|
||
> the named `restoreOpInFlight` seam, and the handler reads the status ONCE per render so the strip,
|
||
> the suppression and the running-op name cannot diverge. The lesson generalises: `opstatus.go` is the
|
||
> DISPLAY surface and says so in its own header — the concurrency flag is not a substitute.
|
||
>
|
||
> **The test lesson:** a table test over a pure function proves the function, not the caller. Scenario
|
||
> E passed throughout because it injected `OpRunning=true` directly. The new test drives a real
|
||
> `Manager` through `BeginRestoreOp` and asserts the render.
|
||
>
|
||
> **DECISION: „Eredmény" earns its place.** The strip's highlight is now `Phase`, derived separately
|
||
> from `Step`: a finished restore is back on the intent step while the strip reads „Eredmény" and an
|
||
> outcome card shows the result — window-bounded (10 min) and app-bound.
|
||
|
||
|
||
> **2026-07-21 — v0.154.0 (R-48).** Collapses the offsite restore controls to a single
|
||
> „Visszaállítás…" entry per app row plus a per-app wizard at `GET /backups/restore/app?name=<app>`.
|
||
> The defect it closes is the CAUSE of the round-2 incident: the list rendered up to five inline
|
||
> forms per row, two of which — the missing-only merge and the true reconstitution — were sibling
|
||
> buttons whose difference is whether the data comes back. The rule it establishes: *two adjacent
|
||
> controls whose difference is "your data comes back" vs "your data cannot come back" must never be
|
||
> distinguishable only by layout.*
|
||
>
|
||
> **DECISION: the wizard is server-rendered on the EXISTING endpoints.** No new mutation endpoint,
|
||
> no JSON state API, no client router. Every card is a real form POST to
|
||
> `/backup/offbox/{restore,place,reconstitute}` with the same field names and gates, and the server
|
||
> renders the next step — so it works with JavaScript disabled. `TestRestoreWizard_NoNewMutationEndpoints`
|
||
> makes that structural: adding a form that posts somewhere new fails the suite by design.
|
||
>
|
||
> **DECISION: R-45 stays its own item.** The wizard polls the two existing status surfaces as-is; the
|
||
> generalized job registry (and with it a real per-phase progress feed) is not built here.
|
||
>
|
||
> **DECISION: the step is derived, never requested.** `deriveWizardStep` is pure over (op running,
|
||
> size-gate flash, scratch ready). Precedence is load-bearing — a running op outranks a stale
|
||
> `?full_prep=` in the URL, or a commit button reappears mid-restore. While ANY op runs every
|
||
> mutation form is suppressed server-side rather than offered and then refused with a 409.
|
||
>
|
||
> Latent bug found and fixed on the way: `offboxRedirectTo` hardcoded `"?"` when appending its flash,
|
||
> which would have buried the flash inside `?name=<app>`. **No agent coupling — MinAgent stays
|
||
> 0.90.0.** 9 new tests + the Group-B red-proof; full suite green.
|
||
>
|
||
> **NOT live-validated at commit time by design:** v0.154.0 is published but deliberately NOT
|
||
> hand-deployed — the operator's hub floor save (0.153.0 → 0.154.0) pulls it via the self-update
|
||
> path, and that swap IS the R-23(a) single-fire validation (STOP-1).
|
||
|
||
|
||
> **2026-07-20 — v0.153.0 (R-47).** Closes the H4 race on **BOTH** restore paths. The replay needs a
|
||
> running DB container, so both paths started the WHOLE stack first — giving the application a window
|
||
> to rebuild the schema objects the dump was about to create. Measured at 8 s on 2026-07-19
|
||
> (`DIAG-immich-restore-round2-2026-07-19`): immich-server rebuilt `clip_index` two seconds before
|
||
> the dump's `CREATE INDEX`, the replay aborted `already exists` under `ON_ERROR_STOP=1`, and immich
|
||
> then reported schema drift. The photos came back **by accident** — `pg_dump` emits COPY before
|
||
> CREATE INDEX, so the abort landed after the rows; a collision earlier in the script would have left
|
||
> a genuinely half-restored database, reported identically.
|
||
>
|
||
> **DECISION: the DB-only bring-up is done by compose SERVICE scoping**, not by container tricks —
|
||
> `StartStackServices(name, []string{svc})` → `compose up -d <svc>`. Every catalog template's
|
||
> dependency direction is app→db, so naming the DB starts the DB and nothing else. `docker start
|
||
> <ctr>` was never an option: `StopStack` is `compose down`, so the containers no longer exist.
|
||
> `RestartStack`/`RedeployFromEnv` are traps here — both end in a full `up -d`.
|
||
>
|
||
> **DECISION: fail-closed.** A `.sql` dump with no identifiable DB service refuses BEFORE the first
|
||
> mutation, on both paths (one Hungarian string, shared). The alternative would be to start everything
|
||
> and replay into the race. It should be structurally unreachable — `dbTypeForImage` is now shared by
|
||
> `DiscoverDatabases` and `DBServiceNames`, and a dump can only exist because discovery matched the
|
||
> container's image, which IS the compose `image:` value — so this is the belt for template drift.
|
||
>
|
||
> Enablers: `RedeployFromEnv` split into `PersistUnitRedeployConfig` (persist, starts nothing) + the
|
||
> unchanged tail; `StackDataProvider.RecreateStackFromUnit` renamed to
|
||
> `RecreateStackDefinitionFromUnit` because the old name promised less than the method did — the
|
||
> hidden `up -d` inside it is what carried the defect on the local path. `StartStackServices` REFUSES
|
||
> an empty list (argument-less `up -d` is a full start). **No agent coupling — MinAgent stays 0.90.0.**
|
||
> 19 new tests, 3 red-proofs, 23/23 green. **NOT live-validated yet:** STOP-1 supervised reconstitute,
|
||
> golden 0.153.0 bake (P3 registry-reachability probe from the vacation site is load-bearing), Viktor's
|
||
> two hub saves, and his C6 customer-restore UI run.
|
||
|
||
|
||
> **2026-07-20 — v0.152.0 + felhom-samba 1.1.0 (Megosztás on a Mac).** Closes **S-3**. **A capture
|
||
> on the box overturned the earlier guess:** macOS DOES send a correct NBNS query for `<NÉV><20>` and
|
||
> nmbd DOES answer it correctly in 140 µs (flags `0x8580`, RCODE=0, right address) — macOS simply
|
||
> never acts on it. NetBIOS there feeds legacy browsing, not `smb://` URL resolution, so **the bare
|
||
> `smb://<NÉV>` can never work from a Mac** and nmbd was never the broken part (it is what serves
|
||
> Windows). felhom-samba 1.1.0 adds **avahi + dbus**, templating `avahi-daemon.conf` and the
|
||
> `_smb._tcp` service file from `FELHOM_SERVER_NAME` so a rename re-advertises; both daemons are
|
||
> non-fatal on failure. v0.151.0's card had offered `smb://<NÉV>` for Mac — the one dead form — now
|
||
> `smb://<NÉV>.local`; Windows keeps flat `\\<NÉV>`. Spiked live by hand and confirmed from the
|
||
> operator's Mac BEFORE publishing the image (the operator's call, and it chose the design too).
|
||
> **STILL OPEN: Finder-sidebar discovery is NOT shipped** — the record is published and answers
|
||
> browse queries, but was never observed working; likely a Finder Settings → Sidebar toggle, but
|
||
> unverified. **Windows was not retested.** Two test bugs fixed en route, neither a production
|
||
> defect: `TestRenderSambaCompose` pinned a literal image tag, and `TestFabUpload_GCAndIdleTimeout`
|
||
> asserted an async unlink synchronously (it passed alone, failed in the full package once the new
|
||
> render tests made `web` heavier). 23/23 green twice; 2 red-proofs.
|
||
|
||
> **2026-07-20 — v0.151.0 (Megosztás).** Closes **S-1/S-2/S-4-core/S-5** of
|
||
> `felhom.eu/documentation/audits/DIAG-sharing-2026-07-20.md`; **S-3 (no mDNS/Bonjour) stays OPEN**,
|
||
> awaiting Viktor's `smbutil lookup FELHOM` + `dns-sd -B _smb._tcp` from the Mac. **The `/sharing`
|
||
> page had been reload-looping at ~1.2 s for every customer with sharing enabled since v0.147.0** —
|
||
> `/sharing/status` coerced `idle`→`running` on the JOB phase channel, and the client answers a
|
||
> terminal `running` with a one-shot `location.reload()`, so the first poll of every steady-state
|
||
> page load re-armed it. The rule this leaves behind, now recorded against R-45 too: **a phase a
|
||
> client answers with a one-shot action is an EDGE — never synthesise it from a level, and serve it
|
||
> exactly once.** Both halves are server-side; `sharing.html`'s `<script>` is byte-identical to
|
||
> v0.150.0. `consumeIfRunning` is the serve-once half (`failed`/`needs_password`/in-flight are never
|
||
> consumed). Also new: the „Csatlakozás a megosztáshoz" card — Windows + Mac forms plus the direct
|
||
> `smb://<IP>`, read from the SAMBA container's netns because the controller is on a docker bridge
|
||
> and would answer `172.x`; **derived per render, cached nowhere** (the guest holds it by DHCP).
|
||
> Live-verified endpoint-level on 9201: `phase:"idle"` on 3/3 polls, and the card renders the real
|
||
> `192.168.0.104`. 23/23 green twice; 3 red-proofs. **The human "does the page sit still" check is
|
||
> still Viktor's** — no browser here.
|
||
|
||
> **2026-07-20 — v0.150.0.** **`go test ./...` on DooPlex is fully green again (23/23 packages, run
|
||
> twice) — no surviving `t.Skip`s, no weakened assertions, no deleted tests.** The 7 red
|
||
> `TestTier2V2_*` / `TestSharesTier2*` were all one environmental class: Tier-2's off-drive guard
|
||
> uses `system.SamePhysicalDevice` (st_dev equality), and every `t.TempDir()` here shares one
|
||
> filesystem, so the fixture's "two drives" looked identical and the guard correctly refused the
|
||
> target (`nincs másik fizikai meghajtó`). Fixed with a nil-defaulted `Manager.samePhysicalDevice`
|
||
> seam + `sameDevice` wrapper — **production behaviour is byte-for-byte unchanged** (nil → the real
|
||
> check); only the two fixtures inject a fake. All 7 mutation-proved. **F7/R-53 shipped:**
|
||
> `app_export.html` built the app URL from `{{$.CSRFToken}}`; now `{{$.Domain}}`, with
|
||
> `exportPageHandler` supplying the key (it bypasses `baseData`). Host: the orphaned `dhclient` on
|
||
> the non-existent `eth0` was killed and did not respawn. Still open from the arc: R-50 (durable F1,
|
||
> spike-first — its ROADMAP note about needing a new cert SAN was **corrected**: the pin is a raw
|
||
> leaf-DER SHA-256 with `InsecureSkipVerify`, so SAN never enters it), R-51, R-52, R-39(b)/F6.
|
||
|
||
> **2026-07-20 — remote-site remediation + v0.149.0.** **F1 is MITIGATED FOR THE WINDOW, not durably
|
||
> fixed:** `vmbr0` on the demo host is now **static `192.168.0.162/24`** (was DHCP; the remote router
|
||
> had handed it `.147`, and the agent binds that literal), applied with `ifreload -a`; the agent came
|
||
> up clean and the red „a tárolókezelő ügynök nem elérhető" banner is gone. The control plane is
|
||
> still pinned to a LAN literal — the durable fix (host-internal island bridge) is a separate
|
||
> spike-first arc, **R-50**. Restoring the agent immediately let the quiesce loop run the overdue
|
||
> whole-guest backup by itself (**F2 closed**), a manual app-data run followed (2 DBs, 3 volume dumps,
|
||
> 43 s), and **Immich is back** (`photos.demo-felhom.eu` → 200; it had been left `Exited` by the
|
||
> pre-transport shutdown, not by the offsite-restore test). **v0.149.0 fixes F3** — the dashboard card
|
||
> said „Utolsó mentés: Még nem futott" on every box because `dashboardHandler` never passed the
|
||
> `BackupStatus` key the template branches on. F4/F5/F6/F7 are roadmap-only (**R-51/R-52/R-53/R-54**).
|
||
> Evidence: `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md` §Remediation.
|
||
|
||
> **2026-07-20 — demo box moved to a remote site until ~2026-08-02; `ssh felhom-pve` = tailnet
|
||
> 100.70.170.35 (direct, ~37 ms). THE HOST AGENT IS DOWN THERE:** its `localapi` binds the literal
|
||
> `192.168.0.162`, the host now DHCPs `192.168.0.147` → `bind: cannot assign requested address`, so
|
||
> the service has never started at the remote site and the controller's `agentapi` dials of the same
|
||
> literal get `no route to host` — that is the whole "A tárolókezelő ügynök nem elérhető" banner, and
|
||
> it kills storage/PBS-backup/quiesce/restore-test/DR until fixed (Viktor GO: config **and** guest
|
||
> bootstrap state). Calibre-Web and `immich-server` were left `Exited` by the pre-transport shutdown
|
||
> and never came back despite `unless-stopped`; Calibre-Web was restarted via the real UI endpoint,
|
||
> Immich deliberately left down. Two code defects found and NOT fixed: **the dashboard's "Utolsó
|
||
> mentés: Még nem futott" is a display bug** (`dashboardHandler` never sets `BackupStatus`, so
|
||
> `dashboard.html:116`'s `{{if}}` branch is unreachable — it renders on every box regardless of
|
||
> history), and **multi-container apps under-alert** (`IsDownState` excludes `unhealthy`, so Immich's
|
||
> dead primary container produced no banner and no `app_start_failed` event for 18 h).
|
||
> Full evidence + ranked findings: `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md`
|
||
|
||
> **2026-07-19 — v0.148.0: R-43 + R-44.** Viktor deleted 11 immich photos to test offsite restore;
|
||
> both runs flashed success and the photos stayed gone (`DIAG-immich-restore-2026-07-19`). Two
|
||
> defects. **R-43:** no offsite path could restore a DATABASE — all three buttons were file-only, so
|
||
> a DB-indexed app got its bytes back and still could not see them. **R-44:** a manual push shipped
|
||
> whatever dump the 02:30 local run left; that day's predated the customer's account by four hours
|
||
> and held zero users and zero assets inside 52MB of shipped geodata.
|
||
>
|
||
> Now: every run (manual AND nightly) dumps FIRST, then captures, so each snapshot is a coherent
|
||
> `{DB@T, files@T}` pair stamped with `offsite_run_id`; and „Teljes visszaállítás (fájlok +
|
||
> adatbázis)" does safety-dump → stop → overwrite files → start → replay the snapshot's dump.
|
||
> **Two invariants: nothing is ever deleted, and the undo is verified on disk before the act.**
|
||
> Warn-level honesty surfaces for stale/empty-looking pairs — never gates.
|
||
>
|
||
> **The live acceptance has NOT run yet** (§9: upload → push → empty trash for real → one button →
|
||
> photos back). Until it does, the capability-map offsite row stays PARTIAL/scope-contested, the
|
||
> customer-restore row stays MISSING, and R-3 stays DRAFT. Floor raise = Viktor's click, BEFORE the
|
||
> acceptance run. Still open: the `00-capability-map.md:61` ruling — did CAMPAIGN-6D's "immich
|
||
> end-to-end from offsite alone" exercise the DB half, or only the file half?
|
||
|
||
> **2026-07-19 — v0.147.0 → v0.147.3: feedback slice 1.** The systemic complaint, twice in one
|
||
> evening: you press a button and nothing happens. Three worst offenders fixed on the two patterns
|
||
> that already existed (deploy 3-step panel; storage-init status poll). **Deliberately NOT a
|
||
> framework** — that is ROADMAP **R-45**, and the two lessons it must encode are already written
|
||
> down there: a terminal state must be **probed, not inferred** (`compose up -d` exits 0 on a
|
||
> crash-loop), and a progress source reporting nothing is **normal, not broken** (restic reports 0
|
||
> bytes for a whole incremental run).
|
||
>
|
||
> **4a** — the offsite verification restore names its **full path** in the flash, and
|
||
> `/backups/restore` lists existing verification copies (app · size · date · path) with a
|
||
> double-confirmed per-copy delete. That delete takes a **stack name, never a path**; red-proofed
|
||
> (neutralise the name guard and `stack:""` resolves to the offsite-restore ROOT and takes every copy
|
||
> with it). `backups/offsite-restore` now has ONE home, `offsiteRestoreRootFor`.
|
||
> **4b** — Megosztás enable/password no longer reconcile inside the POST; detached job + poll, with
|
||
> „képfájl letöltése" vs „indítás" decided BEFORE the work starts (afterwards the image is always
|
||
> present and the distinction is unrecoverable).
|
||
> **4c** — „Távoli mentés most" streams restic `--json`. **Manual only**; the nightly stays silent,
|
||
> pinned by a test.
|
||
>
|
||
> **Three of the four versions exist because the cards were watched against real runs on the demo
|
||
> box** — none of these would have surfaced from unit tests: (.1) an incremental run reports 0 bytes
|
||
> for its whole duration, so a byte-only bar looks hung in the COMMON case; (.2) restic 0.14 counts a
|
||
> file only when it completes, so one big archive freezes the file counters too — fall back to
|
||
> current file + elapsed; (.3) the run does not end with the last app — the shares leg and
|
||
> `forget --prune` took 40 of a 57-second run, and the card froze on the last app until phases were
|
||
> added.
|
||
>
|
||
> Also: `infra.Images()` + `--print-infra-images` close the golden/controller infra-image drift at
|
||
> the source. The golden's own copy had already drifted (felhom-samba missing → 3 of 4 baked), which
|
||
> is **why** enabling Megosztás pulled at runtime in the first place. Effective at the next bake
|
||
> (`felhom-agent` build-golden v2.1.0); no golden rebuilt. **Floor NOT raised — Viktor decides.**
|
||
>
|
||
> Earlier: 2026-07-18 (v0.145.0 — R-7b: share data enters the live backup runs, Model B′)
|
||
|
||
> **2026-07-18 — v0.145.0: R-7b — the „Felhőmentés" toggle is now TRUE (Model B′), + samba liveness.**
|
||
> Until v0.144.0 a share could be marked „Felhőmentés: bekapcsolva" while its files were in NO backup:
|
||
> both engines are recovery-unit shaped (`RunTier2` short-circuits on `os.Stat(unitDir)`; the offsite
|
||
> runner enumerates `GetOffboxApps()`) and a share-only infra stack has neither. **Viktor's ruling was
|
||
> Model B′: a SIBLING shares source** — additive job/leg code reusing the proven primitives (tier-2
|
||
> mirror seam, restic wrappers, quota gate, status recorders) while every per-app engine path stays
|
||
> **byte-identical**. That invariant is enforced by test in both tiers, red-proofed.
|
||
>
|
||
> Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares/<driveKey>/…`,
|
||
> payload at `_payload/`, marker LAST). Tier 3 → `runOffboxSharesLeg`, ONE extra restic call tagged
|
||
> `_shares`, placed after the app loop and BEFORE retention so `--group-by host,tags` covers it for
|
||
> free. Restore → „Megosztások" on /backups/restore: scratch, then a missing-only merge whose every
|
||
> destination is PREFIX-ASSERTED against live storage roots; definitions merge existing-wins; then
|
||
> `ReconcileSamba`; then the credential, best-effort.
|
||
>
|
||
> The **payload** is the point: a byte-deterministic `_shares-manifest.json` + a best-effort
|
||
> secret-bearing `passdb.tar`, so DR returns files + configuration + password, not loose bytes. A
|
||
> quota-blocked offsite push degrades to the MANIFEST ONLY — never to nothing.
|
||
>
|
||
> **Findings:** (a) the reserved-name assumption was FALSE — `nbNameRe` accepted „_shares" as a share
|
||
> name; `ValidateSMBShareName` now refuses a leading underscore and both run loops skip a `_shares`
|
||
> stack loudly. (b) The samba-liveness fold-in needed NO alert/e-mail pipeline change and adds no new
|
||
> event type (so the `allowedEventTypes` gotcha does not apply) — `EffectiveProtected` just gains a
|
||
> settings-backed dynamic extra watching the CONTAINER name (`felhom-samba` ≠ stack name `samba`).
|
||
> (c) A real bug surfaced in review-by-test: `shareSourceDrive` returned a slash-normalised path,
|
||
> making the target selector's equality check miss so a share group could target its own source drive.
|
||
>
|
||
> **OPEN:** Viktor raises the managed-update floor to **v0.145.0** (supersedes the 0.144 note) so the
|
||
> N100 rehearsal's day-0 box converges onto the honest version. Pre-existing, untouched:
|
||
> `docker_run_volume_path_gate` fails on `internal/appexport/estimate.go` (predates this work).
|
||
|
||
> **2026-07-18 — v0.144.0: „Megosztás" LAN SMB sharing (R-7 slice 1), LIVE on demo.** SMB ships as an
|
||
> **embedded controller feature** — the FOURTH protected infra stack (traefik/cloudflared/filebrowser/
|
||
> **samba**), NOT a catalog app (it needs `network_mode: host` per the R-6 spike, its config is a
|
||
> generated share list, and its roots ride the backup classification). New own image
|
||
> **`felhom-samba:1.0.0`** (pinned alpine + smbd + **nmbd** + wsdd + tini; smb.conf bind-mounted
|
||
> read-only, nothing templated inside, passdb on a volume). nmbd is REQUIRED alongside wsdd — the R-6
|
||
> spike proved wsdd-only leaves the box visible but the Explorer double-click fails `0x80070035`.
|
||
> New top-nav category „Megosztás" → „Hálózati megosztás": enable + one household SMB password
|
||
> (STDIN→smbpasswd, NEVER persisted — only `user_set`), shares table, and a create flow (new folder
|
||
> under `<storage>/shares/` or an existing folder via a guarded browse modal). Every customer path goes
|
||
> through `sharingResolvePath` (absolute → EvalSymlinks → containment in a registered LIVE storage root
|
||
> → deny-listed system subtree → is-a-dir) with **uniform** refusals so the picker is never a
|
||
> filesystem oracle; the deny-list is DERIVED from `ProtectedHDDPaths` (provably a subset).
|
||
> `ClassifiedBinds("samba")` resolves from the shares registry: Felhőmentés ON → mandatory
|
||
> (offsite+tier-2), OFF → optional (tier-2 only), with ZERO backup-engine edits.
|
||
> **OPEN / needs a Viktor ruling (suggested R-7b):** the classification seam is correct but share data
|
||
> is **not in any live backup run** — `RunTier2` short-circuits on the missing recovery unit before it
|
||
> reaches the seam, and the offsite runner enumerates `GetOffboxApps()`. Both engines are recovery-unit
|
||
> shaped; teaching them about a share-only stack is a structural change, so it was reported as a design
|
||
> fork rather than improvised (task STOP clause). Live-validated end-to-end through the real endpoints
|
||
> + a Windows 11 workstation (445, NetBIOS `FELHOM` resolves, write/read byte-compare PASS, write to a
|
||
> read-only share REFUSED with no effect, SMB writes land as uid 1000). **Explorer leg PASSED (Viktor): both
|
||
> shares open, an interactive Explorer save landed as uid 1000, a write into the read-only share was
|
||
> refused. Slice 1 fully PROVEN-LIVE.** Also: samba is protected in CODE (`config.alwaysProtectedStacks`) because
|
||
> controller.yaml is golden-generated — a side effect is that `monitor.EffectiveProtected` does NOT
|
||
> monitor samba liveness (deliberate: no false alarms while off; see REPORT §9).
|
||
|
||
|
||
|
||
> **2026-07-17 — v0.143.0: guest RAM resize UI (R-24), LIVE on demo. MinAgent: 0.90.0.** The customer
|
||
> sees the guest's current memory + allowed range on the **Rendszer** settings page and resizes it
|
||
> ("Szerver memória (RAM)" card). The controller proxies + maps the agent's machine `code` to Hungarian;
|
||
> the AGENT (felhom-agent v0.90.0) enforces every bound and applies live (no reboot — R-24's old
|
||
> hub-desired-state framing is SUPERSEDED by Viktor's controller-direct ruling). agentapi
|
||
> `GuestMemory`/`ResizeMemory` (`*MemoryRefusedError` carries the code); capability
|
||
> `FeatureGuestMemoryResize` (featureMinAgent 0.90.0, probe type-asserts GuestMemory so the shared
|
||
> SupportProber/netAgent stay untouched); `POST /api/system/memory/resize`; JS confirm only on a shrink;
|
||
> agent-outdated hides the control, agent-unreachable falls back to the guest's `/proc/meminfo`. Deployed
|
||
> to BOTH demo guests (felhom-pve 9201 + nested demo-vm-felhom-4846bc 9201). **LIVE-validated** through the
|
||
> real endpoint on the nested demo (above_max + below_min refusals render the Hungarian message; the agent
|
||
> English never leaks; the capability gate resolves SupportYes via the 0.90.0 version header). A successful
|
||
> grow couldn't be shown on the tiny 4 GB nested host (max<current, correctly refused); the apply is
|
||
> Phase-0-proven at the agent layer. See REPORT.md.
|
||
|
||
> **2026-07-17 — v0.142.0: offsite repo continuity (Parts A + C), LIVE on demo.** Closes the
|
||
> reinstall-orphaned-repo incident (a recreated data volume mints a new repo passphrase → the offsite
|
||
> repo, keyed under the old one, errors nightly with `wrong password or no key found`). Part A:
|
||
> `ensureOffboxRepo` classifies the `cat config` failure → ORPHANED state + calm Hungarian card
|
||
> (exception color) + `offbox_repo_orphaned` event (once, not nightly); reset = move-aside (never
|
||
> delete, `mv <repo> <repo>.orphaned-<date>`) + init — UNCLAIMED auto, CLAIMED reveal-then-confirm.
|
||
> Part C: `GET /backup/offbox/status` + poll on backups_remote flips Fut→Rendben/Hiba without a manual
|
||
> reload. Pairs with hub v0.60.0 (superseded-escrow retention). Live leg staged for the rehearsal
|
||
> (scratch-target swap would disturb the live escrow state; live repo untouchable) — mechanism covered
|
||
> by 3 fake-based scenarios + 2 red-proofs. Details: REPORT.md.
|
||
|
||
> **2026-07-17 — v0.141.0: N100 polish (F6 + F7), LIVE on demo.** F6 (MEDIUM): drive "initialize"
|
||
> now ends in a mounted+registered drive even on a client disconnect. `POST /api/storage/init` runs
|
||
> the format→mount→register chain as a DETACHED single-flight job (`web/storage_init_job.go`,
|
||
> `context.Background()`, netAddState shape) the wizard polls via `GET /api/storage/init/status`
|
||
> (3-step Hungarian progress); register is the last step (marker-last crash-safety). Fork verdict:
|
||
> controller-side, NO agent change (the chain must reach FileBrowser sync = controller-only).
|
||
> **Deeper half found on the live leg:** a slow mkfs (64 GB USB, ~27 s) outruns the agentapi client's
|
||
> 15 s timeout → the controller now polls the agent's `GET /disks/format/status`
|
||
> (`agentapi.FormatStatus` → `awaitAgentFormat`) then continues. Live-validated on `/dev/sdd` →
|
||
> `/mnt/felhom-drives/scratch1` (mounted+registered). F7 (LOW): storage init/attach Vissza → `/storage`.
|
||
> Red-proofs for both F6 halves + F7. Security review of the commit flagged the pre-existing
|
||
> format→resolve→assign device-node TOCTOU (agent-guarded destructive step, benign fs-UUID mount) —
|
||
> acknowledged as an Observation, not expanded. Fork/landmarks/live evidence: REPORT.md.
|
||
|
||
> **2026-07-16 — v0.140.0: Direction-2 immediate-sync (hub→box) SHIPPED.** The reverse of v0.139.0:
|
||
> an OPERATOR action on the hub now reaches the box in seconds. `report.Waiter`
|
||
> (`internal/report/waiter.go`) holds a hanging `GET {hub}/api/v1/wait?gen=N` (same hub URL+key as
|
||
> the pusher — no new config keys) against hub ≥ v0.58.0's in-memory operator-intent generation
|
||
> counter; on a generation CHANGE it fires the v0.139.0 `report.Trigger` and nothing else, so the
|
||
> report ACK delivers everything through the UNCHANGED machinery (the box pulls even the wake-up).
|
||
> No overall client timeout (held GET); first-observation records-not-fires (no restart echo);
|
||
> same-gen timeout fires nothing; errors incl. a 404 from a pre-v0.58.0 hub back off 5s→5min while
|
||
> the 15-min cycle reconciles. Gated on the SAME `hubPusher!=nil && Hub.Enabled` as the trigger.
|
||
> Copy soften: backups_remote/escrow "néhány **másodperc**, legfeljebb 15 perc" (15-min bound stays
|
||
> as the honest worst case). Red-proof: baseline-branch-off → first-obs fires (reverted). Agent-plane
|
||
> poke is PARKED in the OOB arc (spike P4). Deployed to 9201; live-validated (hold + immediacy).
|
||
> Detail: `CHANGELOG.md` v0.140.0, `controller/README.md` §9, `REPORT.md`.
|
||
|
||
Last updated (prior): 2026-07-16 (v0.139.0 — immediate out-of-cycle hub report, Direction 1)
|
||
|
||
> **2026-07-16 — v0.139.0: immediate out-of-cycle hub report on user actions (Direction 1
|
||
> SHIPPED; Direction 2 pending SPIKE-immediate-sync-transport).** Viktor's ruling: user actions
|
||
> with hub-side effects round-trip in seconds. New `report.Trigger` (`internal/report/trigger.go`):
|
||
> buffered-1 chan + worker, quiet 2 s → drain → min-interval 15 s → ONE full BuildReport+Claimed+
|
||
> Push; trailing-edge coalescing (burst ≤ 1+ceil(burst/15 s) pushes, last state always lands),
|
||
> no own retries, failures degrade to the UNTOUCHED 15-min cycle. Generalizes the v0.70.0 geo
|
||
> `reportPushNow` seam (raw goroutine in main.go replaced by the debounced trigger). Wired: geo
|
||
> save/sync + app deploy/remove/delete (api), and via `web.SetReportTrigger`/`reportTriggerNow`
|
||
> (nil-safe, AFTER successful local commit only): escrow recovery-code claim (headline — the
|
||
> v0.138.0 "megerősítésre vár" card now collapses in seconds via the unchanged EscrowAutoConfirmer
|
||
> ACK hash-match), notification-prefs save, app-email toggle, offsite config + per-app toggle,
|
||
> customer claim. `hub.enabled:false` → seams nil → strict no-op. NO hub change, NO UI copy change
|
||
> ("legfeljebb 15 perc" stays the honest worst case; post-live-proof a soften to "általában néhány
|
||
> másodperc" is a later one-liner). Tests: trigger_test.go (2 red-proofs recorded),
|
||
> report_trigger_seam_test.go, report_trigger_nilsafe_test.go. Known pre-existing Windows-only
|
||
> test failures (appexport df=0, stacks paperless, web fab pipelines) verified failing on base
|
||
> 8f3564c too.
|
||
|
||
> **2026-07-16 — v0.138.0: escrow "awaiting hub confirmation" waiting state.** Fixes the customer-zero
|
||
> (N100) UX gap: after a completed escrow ceremony the Távoli mentés page kept showing the yellow
|
||
> "Helyreállítási kód szükséges" card for ~15 min until the next hub-report ACK flipped
|
||
> pending→escrowed. **Phase-0 diagnosis (read-only) = verdict A (report-cycle lag), already resolved on
|
||
> the box:** demo logs show the ceremony claimed `16:13:39`, the next `hub-report` ACK at `16:27:58`
|
||
> auto-confirmed via hash-match (`d517ce7f…`); `settings.json` = `escrow_state:"escrowed"`. **Hub
|
||
> Hypothesis B verified FALSE → no hub change:** `SaveHostEscrow`'s `ON CONFLICT … stale_at = NULL`
|
||
> already clears stale on upload (store.go:2084); withhold only fires while `stale_at != ""`
|
||
> (store.go:2150). **Part 1 (code):** new persisted `OffboxTarget.CeremonyCompletedAt` (stamped on the
|
||
> recovery-code claim, zeroed on the flip + manual confirm); `offboxCeremonyWaitState` +
|
||
> `escrowCeremonyGraceWindow`=35m; `backups_remote.html` gains an info "megerősítésre vár, legfeljebb 15
|
||
> perc" card → warn "a megerősítés nem érkezett meg" past the window; `backups_escrow.html` final step
|
||
> gains a "Mi történik ezután?" note. Test `web/escrow_wait_state_test.go` + red-proof. No
|
||
> scheduler/agent/hub/endpoint changes. Deploy 0.136.0→0.138.0 to 9201. First "Távoli mentés most" =
|
||
> Viktor's click (NOT done). Note: demo host key changed (box reprovisioned for N100) → known_hosts
|
||
> refreshed.
|
||
|
||
|
||
> **2026-07-15 — v0.137.0: cleanup bundle (email-wipe guard + carried hygiene).** Closes the arc's
|
||
> carried micro-queue. **Part 1 (code):** `settingsNotificationsHandler` now REFUSES a save with a
|
||
> blank email box while events are enabled (it would push empty to the hub → wipe the customer's
|
||
> provisioning-seeded alert address — the 2026-07-15 demo incident). Returns before
|
||
> SetNotificationPrefs + sync, Hungarian error, repaints submitted checkboxes; empty+zero-events
|
||
> clear-all still allowed. No HTML `required` (it would block the legit clear-all). Tests + red-proof
|
||
> in `web/notifications_guard_test.go`. Deployed 0.137.0 to 9201 (healthy). **Part 2 (hygiene):**
|
||
> removed the confirmed-older `felhom-flash/backups/primary/immich` recovery unit (CreatedAt 06-23 <
|
||
> live usb 07-15, 44M); STOPPED audiobookshelf/komga/romm on flash (CreatedAt TIED with usb →
|
||
> tie-break is drive-order-dependent, not confirmable — manual disposition pending). **Part 3:**
|
||
> campaign6 is a bare empty leftover dir (not a live mount); safe `rmdir` refused (Permission
|
||
> denied — autofs-ghost/immutable); no mount disturbed → left for Viktor's reboot window. **Part 4:**
|
||
> tagged the campaign6 6D-audit finding track-only (felhom.eu `dee72cd`).
|
||
|
||
> **2026-07-15 — v0.136.0: `.fab` exclusion scoping (Task 4).** Architecture §2 `.fab` row + SQ5
|
||
> verdict + R1-C. SQ6 over-capture FIXED for classified apps: the userdata root tar is exclude-scoped
|
||
> (keeps only ancestors/descendants of a SELECTED bind relpath — the tier2Reconcile keep-rule); no
|
||
> selected userdata bind ⇒ no root tar (radarr state-only). New `appbackup.ComputeFabBuckets` (shared
|
||
> `resolveGuardCollapse` pipeline; class buckets; guards over ALL classes; no cross-bucket containment).
|
||
> New `appexport/fabplan.go`: `computeFabPlan` (SkipMounts/SkipUserdataTar/UserdataExcludeRels) +
|
||
> `tarDirectoryExcluding` + `fabEstimateSplit`. `ExportRequest` += DeselectOptional/OptInExcluded (both
|
||
> start handlers — two-call-site); mandatory is a SERVER-SIDE floor. Manifest v1 + import UNTOUCHED;
|
||
> legacy apps byte-identical to v0.130.0. Estimate gains an additive class split; export UI shows
|
||
> locked-mandatory / optional-checkboxes / excluded-opt-in + the two-number warning. All 6 §10
|
||
> red-proofs verified. CAMPAIGN-6D Accept #1 (≥1 GiB .fab full circle) now runs against this shape.
|
||
|
||
> **2026-07-15 — v0.135.0: tier-2 engine rework (Task 3b).** Architecture §2/§8. New
|
||
> `tier2_capture.go`: classified apps get the `TierSecondary` per-bind legs (paperless copy shrinks —
|
||
> export drops); legacy apps keep the byte-identical resolver set. v2 relpath-mirroring layout
|
||
> (`backups/secondary/<stack>/{.felhom-tier2-layout marker LAST, recovery-unit/, hdd/<rel>/,
|
||
> userdata/<rel>/}`) — N>1 native (flat-appdata refusal + `errTier2MultiDir`/`tier2AppDataName`
|
||
> deleted). Migration=delete-and-rebuild + reconcile (prunes dest dirs a bind no longer covers); all
|
||
> `os.RemoveAll` via `tier2SafeRemove` (refuses outside backups/secondary/). SSD=state-only tier
|
||
> (unit+mandatory). `selectTier2Target` never picks NETWORK storage (pinned+auto, F-6C-1). Restore
|
||
> reads v2 behind a marker gate (pre-v2 refused); two-subtree missing-only merge. **Part 0:**
|
||
> offbox_enlarge_blocked is now a persisted one-time Load seed (`OffboxEnlargeNoticeSeeded`), NOT a
|
||
> getter append — opt-out STICKS (fixes the 3a-fix un-disableable checkbox). **Part 0.5:** offsite
|
||
> restore scratch prefers a local (non-network) path. Full v2 test suite + all 10 §10 red-proofs
|
||
> verified. Every destructive write bounded to backups/secondary/.
|
||
|
||
> **2026-07-15 — v0.134.1 (+ hub v0.55.0): Task 3a-fix.** Placement hardening in
|
||
> `offbox_restore.go`: F-3a-1a live target uses raw `GetStackHDDPath` (not `AppNamespaceRoot` — its
|
||
> systemDataPath fallback would merge userdata onto the SSD; empty ⇒ undeployed ⇒ refuse); F-3a-1b
|
||
> placement headroom gate; F-3a-4 stat PRE-PASS over all placements before any copy (no partial
|
||
> writes); F-3a-3 `mapOffsiteRestorePaths` refuses the namespace root itself; F-3a-2 scratch removed
|
||
> on success (place button then gone), kept on failure. Enlarge-blocked notification delivery chain:
|
||
> `DefaultEnabledEvents` + `GetNotificationPrefs` append-if-absent migration + settings checkbox +
|
||
> handler slice (controller), and hub v0.55.0 allowlists `offbox_enlarge_blocked` — NO
|
||
> customerMessages entry (raw dynamic message must survive). +8 controller tests, +2 hub; all §10
|
||
> red-proofs verified. Hub LIVE (ArgoCD synced, :0.55.0). Migration trade-off noted (getter re-enables
|
||
> on opt-out — future persisted marker). 3a deferred list shrinks after §13 live legs.
|
||
|
||
> **2026-07-14 — v0.134.0: offsite tier policy engine (Task 3a — FIRST behavior change).**
|
||
> Implements architecture §2/§6/§7/§9. Each toggled app's offsite push = ONE multi-path restic
|
||
> snapshot (recovery unit + TierOffsite mandatory userdata via `ComputeCaptureSet`); legacy/undeployed
|
||
> stay unit-only. New `offbox_capture.go` (`offboxCaptureSet` + loud gaps: restic 0.14.0 silently
|
||
> skips missing paths, SP-3.4) + `offbox_restore.go` (ID-first snapshot introspection, scratch off the
|
||
> rootfs + headroom gate F-A1, `RestoreOffboxScratch(full)` unit-only default via `--include`,
|
||
> `PlaceOffsiteRestore` missing-only merge, pure `mapOffsiteRestorePaths`). Quota → `stats --mode
|
||
> raw-data` (SP-1; **displayed size drops once after deploy**). Pre-push enlargement gate blocks the
|
||
> userdata enlargement over-quota (unit-only push continues; `OffboxTarget.EnlargedBlocked`;
|
||
> edge-triggered notify). `forget --group-by host,tags` on both sites (SP-2). UI: /backups/restore
|
||
> three actions (unit / full two-step / place-to-live); /backups/remote per-app blocked note. New
|
||
> route `POST /backup/offbox/place`. **HUB FLAG:** `offbox_enlarge_blocked` event needs hub
|
||
> allowlist+customerMessages for push delivery (in-dashboard LastWarning works now). +13 tests, all 10
|
||
> §10 red-proofs verified. NOT-live-yet (6D): PlaceOffsiteRestore, large full restore, live
|
||
> enlarge-block, notification delivery, SQ3 immich full-circle. Tier-2 (3b) + .fab (Task 4) untouched.
|
||
|
||
> **2026-07-14 — v0.133.0: capture-set computation (Task 3-core, INERT).** Task 3-core of the
|
||
> backup-classification-redesign arc (architecture `felhom.eu/documentation/architecture/07-backup-architecture.md`
|
||
> §3; spike verdicts `SPIKE-restic-snapshot-shape-2026-07-14.md`). New
|
||
> `appbackup/captureset.go`: pure `ComputeCaptureSet(binds, hasClassification, tier, hddPath)` →
|
||
> `CaptureSet{HasClassification, Paths []CapturePath, Skipped []SkippedPath}`. Pipeline: legacy
|
||
> short-circuit → tier filter (`TierOffsite`=mandatory only, `TierSecondary`=mandatory+optional,
|
||
> excluded dropped) → structural guards (traversal / bare HDD drive-root / reserved `backups/` →
|
||
> `Skipped` with English reasons; bare userdata allowed) → equal-Abs collapse (mandatory>optional) →
|
||
> containment dedup (keep ancestor) → sort by Abs. Slash algebra only (no `filepath`). Pure
|
||
> `CrossAppOverlaps` advisory (WARN wiring deferred to 3a/3b). **Deliberately INERT — no engine
|
||
> consumes it yet; 3a (offsite policy) and 3b (tier-2 rework) are the consumers.** ARCHITECTURE
|
||
> IMPACT from the spike (SP-3.4, already in §2.5): restic 0.14.0 does NOT error on a missing source
|
||
> path (exit 0, silent partial snapshot) → the stat-filter in 3a/3b is load-bearing. Wiring test
|
||
> through a real Manager (F-S3 no-seam); all 6 §10 red-proofs verified. felhom.eu §3 docs aligned
|
||
> (`8d85da7`).
|
||
|
||
> **2026-07-14 — v0.132.0: backup classification (Task 2, INERT).** Task 2 of the
|
||
> backup-classification-redesign arc (spike: `felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md`).
|
||
> Ships the referential-coupling classification as DATA + PARSER + PURE CLASSIFIER, **deliberately
|
||
> inert** — no backup tier changes behavior. New `appbackup/classify.go`: `BackupSpec`/`BindSpec`
|
||
> (the `.felhom.yml` `backup:` block), `ComposeBind` (`${VAR}`-relative + `:ro`), `ClassifyBinds`
|
||
> (SQ5 two-level default: explicit beats `:ro`; unlisted writable→mandatory, unlisted `:ro`→excluded;
|
||
> **no block → legacy/false**), `ValidateBackupSpec` (whole-block-reject on any defect). New
|
||
> `stacks/classify_binds.go` `ParseComposeClassifiableBinds` (relative-space, keeps `:ro` — NOT
|
||
> `ParseComposeHDDMounts`/`ExportDataMounts`, the classifier-input traps). `LoadMetadata` is the
|
||
> SINGLE validation choke point (bad catalog block → nil + one `[ERROR]` within one sync cycle →
|
||
> legacy). Wired seam `Manager.ClassifiedBinds` + `StackDataProvider.GetStackClassifiedBinds`
|
||
> (nil-stubbed in every fake) so **Task 3 (tier policy engine)** consumes a tested seam, not a fresh
|
||
> one. Inertness proven: full pre-existing suite green with ZERO test-logic edits. The 13 catalog
|
||
> `backup:` blocks ship in the same `app-catalog-felhom.eu` change (controller deployed FIRST so the
|
||
> parser validates on first sync). audiobookshelf PENDING-VETO: media/audiobooks ruled **optional**
|
||
> (consistency with komga/romm) pending a Viktor veto to excluded. +14 tests, RP-1..RP-4 confirmed.
|
||
> **Next: Task 3** consumes `ClassifiedBinds` to scope offsite/tier-2/`.fab` capture by class.
|
||
|
||
> **2026-07-14 — v0.131.0: F-S2 + F-S3 (compose-derived appdata dir resolution).** Task 1 of the
|
||
> backup-classification-redesign arc (spike: `felhom.eu/documentation/audits/SPIKE-backup-classification-2026-07-14.md`).
|
||
> The controller assumed `appdata/<stackName>`; paperless-ngx writes `appdata/paperless` (stack
|
||
> `paperless-ngx`). ONE canonical resolver `appbackup.AppDataDirNames(hddPath, stackName, mounts)`
|
||
> derives the real dir name(s) from compose `${HDD_PATH}` binds (deduped/sorted; fallback `[stackName]`);
|
||
> all consumers use it. **F-S2** (spike-proven): `RunTier2`/`Tier2Info`/`RestoreTier2Files` now hit the
|
||
> resolved dir (paperless documents got NO tier-2 copy before — the appdata leg stat-skipped a dir that
|
||
> never existed). **F-S3 (NEW, found this session):** `migrate.go` keyed all six per-app appdata legs by
|
||
> stack name; **scope="app"** has no merge walk, so migrating paperless-ngx copied nothing, verified
|
||
> vacuously, flipped HDD_PATH → **empty media dir** (scope="all" was saved by the merge walk — data safe,
|
||
> accounting off). All six legs now loop resolved names. **Multi-dir (N>1) refusal** is defensive (no
|
||
> catalog app hits it today: immich/nextcloud/romm match, paperless mismatches, each app = exactly ONE
|
||
> dir): tier-2 backup/info/restore refuse loudly (Hungarian); **migrate supports N naturally**. This
|
||
> limitation is **deferred to Task 3 (tier-policy engine)**, which owns the destination layout. Storage
|
||
> page sums resolved dirs. Truth repairs: the v0.130.0 CHANGELOG/CONTEXT "tier-2 copies the namespace
|
||
> wholesale" claim is FALSE — corrected in the v0.131.0 CHANGELOG entry + `main.go` export-adapter
|
||
> comment; tier-2 copies the recovery unit + resolved `appdata/<name>` ONLY (NOT userdata — F-S1,
|
||
> unaddressed here). New seam `tier2Mirror`; `migSeams.resolveNames`. +9 tests, RP-1..RP-5 all
|
||
> confirmed. Controller-only, no agent/hub coupling. **NOT live-validated here:** scope="app" migration
|
||
> of a real app between drives (F-S3 live proof — supervised leg, Viktor's session).
|
||
|
||
> **2026-07-14 — v0.130.0: CRITICAL C6B-F1 (hollow .fab export) + C6B-F2 (share-removal guard).**
|
||
> CAMPAIGN-6B proved `.fab` export shipped **config-only, data-free bundles** for 12/13 `needs_hdd`
|
||
> catalog apps (sonarr 4.17 GB → 2308 B, success, past the v0.125.0 guard). Three compounding fixes
|
||
> (all red-proofed run→fail→revert): (1) `stacks.ExportDataMounts` — export mount discovery unions
|
||
> `${HDD_PATH}` binds + the `${USERDATA_PATH}` **ROOT** (single `userdata` entry; root-not-per-bind
|
||
> is LOAD-BEARING: the manifest keys tars by basename and the untouched import maps basename →
|
||
> `<HDD_PATH>/<subdir>` — per-bind subpaths would restore to wrong places; this deviates from the
|
||
> task's literal per-bind+namespaced-names instruction, which could not round-trip without import
|
||
> changes the task forbade); (2) export + estimate are ADDITIVE for `needs_hdd` apps (HDD data AND
|
||
> named volumes — sonarr_config was silently dropped); (3) anti-hollow guard: `needs_hdd` manifest
|
||
> with zero data fails loudly. Plus §8: basename collision between mounts = loud Hungarian failure
|
||
> (was silent overwrite). **C6B-F2:** `netstorage/remove` refuses 409 while a DEPLOYED app's
|
||
> HDD_PATH is on the share (the orphaned-autofs trigger); resolves via the netAgent seam.
|
||
> **Residual flagged for a felhom-agent task:** RemoveNetworkMount's tolerate-and-continue stop
|
||
> (felhom-agent netmount.go:434-443) still deletes unit files under a busy mount if some non-product
|
||
> path calls it. Scheduled/tier-2 backup path was NOT affected and is untouched (`stackAdapter`
|
||
> deliberately unchanged). CAMPAIGN-6C's first acceptance test = the full-circle byte-compare this
|
||
> unblocks.
|
||
|
||
> **2026-07-13 night — v0.128.1 + demo storage hygiene (ruling F5).** `classTag` suppresses the
|
||
> rotational class hint for `type==='usb'` (card already carries the USB tag; hub `ClassHint`
|
||
> UNCHANGED; pinned by `TestStorageTemplate_USBClassBadgeSuppressed` + red-proof). Host op on
|
||
> demo-felhom: the two pre-intermediary legacy `dir:` storages (`felhom-usb`, `felhom-flash`,
|
||
> content=Backup, is_mountpoint) RETIRED via `pvesm remove` after G1/G2/G3 gates all PASSED
|
||
> (agent-owned UUID .mount units; both `enrolled` in drive-intents.json; zero /etc/pve refs, empty
|
||
> dump/, no customer app on either drive). Post-removal: mounts+binds intact (marker round-trip
|
||
> through the guest), `GET /disks` shows both registry-sourced (role+durable-id intact, class
|
||
> absent), `pvesm status` clean. **The demo node now matches the fresh-install storage shape** —
|
||
> drives are registry+units-sourced only, no legacy PVE dir: storages. v0.128.1 LIVE on demo 9201
|
||
> (drill guest skipped — optional, no behavioral dependency; it runs 0.128.0).
|
||
|
||
> **2026-07-13 night — v0.128.0: CHUNKED BROWSER .FAB UPLOAD on /import (ruling F3: chunked).**
|
||
> Step-0 probe on the real tunnel PROVED the Cloudflare edge cap (120 MiB POST → edge 413 with
|
||
> `Server: cloudflare` on Content-Length alone; 80 MiB → origin 302 /login; local DNS overrides
|
||
> the hostname to the LAN guest, probe needed `--resolve` onto CF's public IP). Design: JS
|
||
> `File.slice` 64 MiB strictly-sequential chunks → `POST /api/export/upload/{init,chunk,finalize,
|
||
> abort}` inside `ServeExportAPI` (inherits RequireAuth+CsrfProtect; single-flight; offset must
|
||
> equal received else 409+echo; 96 MiB request cap; free-space gate size+1 GiB; finalize =
|
||
> exact-size + fsync + atomic rename, collision → lowest-free `"name (N).fab"`). Lands in the
|
||
> DEFAULT drive's exports dir — scan/validate/import pipeline untouched. No client hash
|
||
> (deliberate: .fab self-validates). In-memory state: startup GC of `*.part-*`, 15-min idle
|
||
> abort. §7 A–F tested + 3 red-proofs. `appexport.DiskFree` exported (REUSE.md row).
|
||
> **NOT live-validated: the end-to-end multi-GB browser upload through the real tunnel needs a
|
||
> dashboard login → Viktor's 5-minute leg (export an app to .fab, download, re-upload, import —
|
||
> full circle).** Possible follow-up if it itches: per-drive target picker (v1 = default drive only).
|
||
|
||
> **2026-07-13 eve — v0.127.0: CUSTOMER-FACING ESCROW CEREMONY WIZARD (/backup/escrow) +
|
||
> Scenario-F stale-blob re-check. MinAgent 0.88.0 (wizard only). LIVE on demo 9201 + drill guest
|
||
> (both healthy).** The friend-alpha missing piece: preflight → warnings → password re-auth
|
||
> (login rate limiter) → **re-stage-first** (abort on failure — the UI can never mint a hash-less
|
||
> blob) → agent job (poll 2 s) → ONE-SHOT R reveal (no-store; R only in the claim XHR + page JS;
|
||
> 10-min TTL → void) → typed-back (two random words) → finish. Ruling F1: R over the CF tunnel
|
||
> once = accepted (threat model in felhom.eu RUNBOOK-escrow-ceremony.md). Scenario F: an ESCROWED
|
||
> box re-checks the ACK hash — mismatch/hash-less ⇒ stale flag (card warning + CTA) + one WARN
|
||
> per hash; never flips, never blocks; **fired LIVE on both boxes' hash-less blobs at first ACK**
|
||
> (drill = the spike's superseded blob, since REPAIRED via a real ceremony —
|
||
> `restic_pw_sha256` now covers; demo = its legacy blob, warning stays until a wizard run).
|
||
> Manual-confirm BUTTON removed (endpoint stays, deprecated). **OPEN: one supervised full-browser
|
||
> wizard pass with Viktor's login (re-auth needs the customer-owned password — CC validated
|
||
> everything beneath it endpoint-exact); demo wizard run to clear its stale warning.**
|
||
|
||
> **2026-07-13 — v0.126.0: UI UNIFORMITY BUNDLE (shared app-list rows + infra identity +
|
||
> restore-form polish + mojibake gate + honest stale line). Presentation-layer only — NO
|
||
> backup/toggle/engine behavior change. MinAgent 0.81 + floor unchanged.**
|
||
> (A) `templates/app_row.html` `app_list_row`/`app_list_row_end` is THE canonical list row
|
||
> (icon+name left, caller action right, compact 44px) — dashboard installed-apps, Távoli mentés
|
||
> toggles, Visszaállítás restore-to-verify + .fab lists render through it; the backups-apps
|
||
> expander header is ALIGNED (own markup, allowlisted); gate `scripts/app_row_dedup_gate.py`
|
||
> (red-proven). funcmap: `dict`/`appHref`/`infraMeta`. (B) `inframeta.go`: cloudflared →
|
||
> „Cloudflare Tunnel", traefik → „Traefik", filebrowser → „FileBrowser" + Hungarian descriptions
|
||
> + generic `/static/infra-logo.svg` fallback; filebrowser = the ONLY Linked infra
|
||
> (files.<domain>); render test counts exactly one customer link (red-proven). (C) .fab password
|
||
> field standard („Opcionális jelszó" + helper; import-page input got `.form-input`).
|
||
> (D) `scripts/mojibake_gate.py` — templates+Go strict UTF-8, zero Ã/Â/Ă-signature chars,
|
||
> allowlist ZERO (red-proven); source had NO mojibake — the live „Tárhely" is the felhom-usb
|
||
> drive-label DATA, repaired via the label-edit UI (live step). (E) `offboxWarningDisplay`
|
||
> display pick — stale „nincs mentésre jelölt alkalmazás" run-warning → „A kijelölés módosult…"
|
||
> note once ≥1 app toggled (neutral color); 0 toggled unchanged (red-proven).
|
||
> Housekeeping: one-shot `backups_split_move_check.py` RETIRED (served its purpose).
|
||
> **Live QA fixes:** v0.126.1 — `.form-input`/`.form-row` had NO CSS rule at all (root cause of
|
||
> the unstyled .fab password field; styled as the `.form-control` twin). v0.126.2 — CF edge
|
||
> caches /static/style.css 4h → stylesheet link now `?v={{.Version}}` (auto-bust per release).
|
||
> **0.126.2 LIVE drill+demo (both healthy).** §13 visual QA ran on the DRILL box via a
|
||
> reversible SSH gate-lift (hash restored byte-identical, gate verified back ON) — the demo is
|
||
> customer-claimed and CC does not enter credentials. OPEN human step: demo login → felhom-usb
|
||
> label repair via the label-edit UI (data = 'Tárhely (felhom-usb)' in storage_paths, documented
|
||
> read-only; the Part D gate closed the code side).
|
||
|
||
> **2026-07-13 — v0.125.0: .FAB VOLUME PATH-STRAND DATA LOSS FIXED (IA finding 1, HIGH).
|
||
> MinAgent 0.81 unchanged; floor may advance to 0.125.0 next train (must NOT halt above 0.124.0
|
||
> without this).** Both volume legs stream via docker cp (helper container + `dockerExec` seam —
|
||
> zero shared paths, correct bare-metal AND containerized; §3 live probe first). Export FAILS
|
||
> LOUD on any missing/empty claimed tar (`assertBundleDataComplete`); import VALIDATES BEFORE it
|
||
> destroys (`validateBundleData` in step 0 — hollow bundle → refusal, app untouched). Class
|
||
> extinguished by `scripts/docker_run_volume_path_gate.py` (every `"-v"` allowlisted with WHY;
|
||
> Tier-1/2 mounts documented host-visible). Live: the exact failed ActualBudget leg round-trips
|
||
> byte-identically (`ec8ea6cb…` before==after); engine-invalid volume → loud export failure.
|
||
> **ASYMMETRY (needs a customer-docs line):** .fab bundles exported by containerized ≤0.124.0
|
||
> controllers are hollow — re-export; the import guard refuses them loudly.
|
||
|
||
> **2026-07-13 — v0.124.0: BACKUPS IA RESTRUCTURE. MinAgent 0.81.0 + floor unchanged.
|
||
> Operator decisions (2026-07-13, treat as settled):** (1) single active offsite destination per
|
||
> box STANDS — the dual-destination `managed_by` model is the separate queued Task B;
|
||
> (2) the Felhom-offsite status card NEVER changes anything — display + opt-in pointers only;
|
||
> (3) .fab export/download is PORTABILITY, not a backup tier — no scheduling, no status surface,
|
||
> point-in-time framing. Mechanics: four sub-pages (`/backups{,/remote,/apps,/restore}`, sections
|
||
> moved VERBATIM — `scripts/backups_split_move_check.py` gates vs df7ad37), status card (3 states,
|
||
> display-only), .fab download exit (existing exporter + staging dir + guarded stream + 1h TTL;
|
||
> traversal guard red-proven). Live-validated on drill+demo incl. a supervised import round-trip.
|
||
> **NEW FINDINGS:** **(HIGH)** containerized .fab export strands Docker-VOLUME tars on the guest
|
||
> host (`docker run -v <container-tmp>` → host path) — bundle ships empty volumes, import brings
|
||
> the app up EMPTY; fix = host-visible staging + fail-loud post-export assertion. **(MEDIUM,
|
||
> agent)** legacy-boot PVE (LVM root, no ESP mount) → SystemDisks empty → sysKnown=false → drive
|
||
> wizard offers ZERO candidates ever. Follow-ups: .fab browser-upload; mega-zip parked.
|
||
|
||
> **2026-07-13 — v0.123.0: POLISH BATCH (take-two F-15/F-11 + rename + zero-toggle). MinAgent
|
||
> 0.81.0 unchanged; floor 0.122 unchanged. Requires hub v0.52.0 for F-15 (old hub = clean no-op).**
|
||
> (1) F-15: the reset-request RESPONSE carries the rotated code hash, applied via the ACK's
|
||
> generation-guarded ClaimSync — emailed codes work immediately (live: 1 s, first-try accept).
|
||
> (2) F-11: zero native `confirm()` — `felhomConfirm`/`data-confirm` inline Igen/Mégse (layout.html);
|
||
> gate `scripts/native_confirm_gate.py`. (3) Tier-3 customer branding is **"Távoli mentés"**
|
||
> (NAS-mentés gone; manual form generalized to any SFTP target; gate
|
||
> `scripts/offbox_rename_gate.py`; "Hálózati tárhely" feature untouched). (4) Zero-toggle honesty:
|
||
> hint + "Sikeres — nincs mentésre jelölt alkalmazás" run copy. Also: `atomicPromoteTar` O_RDWR
|
||
> fsync (Windows dev-box green gate was permanently red). Deployed drill qm300 + demo 9201.
|
||
> NOT started (separate queued task, operator fork pending): offbox `managed_by` coexistence.
|
||
|
||
> **2026-07-12 — v0.122.0: CUSTOMER-CLAIM PASSWORD GATE (closes DRILL-day0-vm F-4/F-5). MinAgent
|
||
> 0.81.0 unchanged. Requires hub v0.50.0.** The dashboard password is CUSTOMER-OWNED via a one-time
|
||
> claim code the hub emails to the registered address — the "no password → open dashboard" default
|
||
> is GONE. Unclaimed box (code hash delivered, no password) → serves ONLY `/claim`; every other route
|
||
> → claim page (302) or 401 (API). A set password disables the gate (auth wins). Reset rides the same
|
||
> code engine (login "Elfelejtett jelszó"). Legacy-open (no password + no hash) → red transition
|
||
> banner until the hub delivers a hash. `internal/web/claim.go` (gate + pages + HMAC pre-auth CSRF +
|
||
> 5-try/15-min lockout → `claim_lockout` event), `report/claim_sync.go` (ACK cache, idempotent by
|
||
> generation), settings `Claimed`/`ClaimCode*`/`ClaimConsumedGeneration`, `config.web.claim_code_*`
|
||
> (hub-baked), `--print-reset-code` root hatch. Gate-coverage signature test + 4 red-proofs.
|
||
> **LIVE-PROVEN on drill guest 9201 (0.122.0): gate ON via the real edge (/ → 302 claim page, /api →
|
||
> 401), code emailed to demo-vm-felhom's registered address.** Floor raise 0.120→0.122 = operator's
|
||
> LAST step (supervised). Details: felhom.eu/documentation/audits/DRILL-day0-vm-2026-07-12.md (F-4/F-5
|
||
> RESOLVED).
|
||
|
||
> **2026-07-12 — v0.121.0: BACKUPS PAGE TRUTH PASS. MinAgent 0.81.0 unchanged. Controller-only, no
|
||
> agent-API change.** Pure UI/data-plumbing on `/backups`; no backup-engine behavior change. Fixes the
|
||
> self-contradicting live page: (1) **removed the dead "Részletek" card** (operator decision — redundant;
|
||
> per-app rows + Adatbázisok section already carry the truth) — kills the last uses of the never-set
|
||
> template fields `Tier2DriveGroups`/`ResticPassword`, the `restic-pw` element, and the `toggleTier`/
|
||
> `toggleResticPw`/`copyResticPw` JS. (2) **per-app "3. mentés" row now shows real off-box state** via a
|
||
> new pure `tier3State` (configured→toggle→escrow precedence): unconfigured/off/escrow_pending/active —
|
||
> the "Hamarosan — B2/S3/SFTP" placeholder is gone. (3) **SQLite-honest DB messaging** via pure
|
||
> `dbSectionState(discovered,dumps)` → dumps/pending/embedded (embedded-only box shows "–" + "beágyazott
|
||
> DB-k a kötetmentésben", not a bare "0"). (4) **dead/raw fields fixed** — `Tier1LastRun`/`Tier1LastStatus`
|
||
> now populated from `ListRestorePoints`; Tier-1/Tier-2 labels via `timeAgoStr` (relative), confirm()
|
||
> dialog keeps raw. (5) **terminology split** — off-box section = "Távoli mentés (3. mentés)" (+
|
||
> `#offbox-section` anchor); whole-guest PBS card = "Távoli rendszermentés" (was both "Távoli mentés").
|
||
> (6) deploy page gains a "Mentési beállítások →" link. DECISIONS: Részletek removed as redundant
|
||
> (operator-approved); "Távoli mentés (3. mentés)" (app off-box) vs "Távoli rendszermentés" (PBS whole-CT)
|
||
> are two distinct customer-facing names. Pure helpers in `internal/web/backup_page_state.go`. +9 web
|
||
> tests, 4 red-proofs. Observations: orphaned style.css classes from the Részletek removal left in place
|
||
> (details-tier*, repo-encryption*, restic-pw-field, drive-detail-*, tier-empty-state, repo-info-row*,
|
||
> repo-tier-title) — noted, not cleaned. Backlog: felhom.eu backup-architecture.md offbox refresh (separate task).
|
||
|
||
> **2026-07-12 — v0.120.0: fix-3 + fix-6 → CAMPAIGN-3 CLOSED (LIVE on 9201 + hub 0.48.0).
|
||
> MinAgent 0.81.0 unchanged.** **fix-3:** a `deadapp-check` job (30s, 90s boot grace) flags a DEPLOYED
|
||
> app in stopped/exited state (`stacks.IsDownState`) → self-clearing WARN dashboard banner + one
|
||
> `app_start_failed` hub event per running→down transition (`Notifier.NotifyAppStartFailures`, in-memory
|
||
> tracker, hub owns cooldown). **fix-6:** ring cap 1000→5000 (display cap raised too); periodic
|
||
> scheduler/refresh success lines → `[TRACE]` (ring drops at write-time, failures never TRACE); atomic
|
||
> JSON-lines spill to `<DataDir>/debug-ring.log` (SSD, survives recreate) every 30s + shutdown, loaded
|
||
> on boot. **hub v0.48.0** accepts `app_start_failed` (allowlist + customerMessages). LIVE: docker stop
|
||
> seerr → banner + ONE hub event across 3 cycles (anti-spam) → docker start → banner self-cleared; ring
|
||
> 0 spam lines + restart PRESERVED the pre-restart window (oldest unchanged, 63KB spill on SSD volume).
|
||
> **CAMPAIGN-3 CLOSED** (F12/F11/F10/F9/F2/F1→agent 0.85; F7/F6/F5→0.118; F8/F4→0.119; fix-3/6→0.120).
|
||
> Follow-ups: agent-ring persistence; F13 (active-nfs-mp8 rc255); publish train (agent 0.85 + ctrl
|
||
> 0.118/0.119/0.120 + hub 0.48) to Peti. Seams: deadapp scanDeployedAppRunStates, notify.pushFn.
|
||
|
||
> **2026-07-12 — v0.119.0: STORAGE-HEALTH COHERENCE (LIVE on 9201). MinAgent 0.81.0 unchanged.** Fixes
|
||
> CAMPAIGN-3 F8+F4. **F8 (MED):** the share row's health came only from the agent's server-level TCP
|
||
> dial (blind to a single unexported share) → it showed benign "Készenlét" while the stacks cards
|
||
> showed the stub — a contradictory UI. `networkStorageItems`→`fuseNetHealth` now reuses the SAME
|
||
> `system.ClassifyPathFS` the stacks stub badge reads (§3 fork = option B, controller-only): a new
|
||
> `stub` health (badge "Hibás — az alkalmazások nem a NAS-t látják") overrides idle/ok when the
|
||
> namespace sees local disk; `unreachable` still wins over stub; autofs/network/unknown leave agent
|
||
> health intact (never force-mount). Row + stacks badge now share ONE classifier → can't contradict.
|
||
> **F4 (LOW):** `handleNetStorageAdd` range-checks container uid/gid 1..65533 (`validMappedID`) →
|
||
> friendly 400, nothing installed (was raw agent_error on 101000). LIVE: F8 row=stub matching stacks
|
||
> badge through an exportfs cut, cleared to ok on re-export; F4 uid 101000→400, uid 1000 passes.
|
||
> Seam: `s.classifyFSPath`. Task D (fix-3 alerting + ring revision) still queued.
|
||
|
||
> **2026-07-12 — v0.118.0: BACKUP INTEGRITY (LIVE on 9201). MinAgent 0.81.0 unchanged.** Fixes
|
||
> CAMPAIGN-3 backup findings (`documentation/audits/CAMPAIGN-3-2026-07-11.md`). **F7 (HIGH) atomic
|
||
> volume dumps:** `DumpAppVolumes` writes `<vol>.tar.tmp` → fsync → `os.Rename` over the `.tar` only on
|
||
> success (`atomicPromoteTar`), mirroring dbdump.go DumpOne; a mid-write NFS cut can no longer
|
||
> truncate the last good tar to 0 bytes. **F6 (LOW) no single-copy:** `RunAllTier2` no longer skips
|
||
> volume-only apps (they now get a cross-drive tier-2 copy); sys_drive restore-point label is clear
|
||
> ("Belső SSD (rendszer)"); single-drive box shows an honest `SingleCopyWarning` banner. **F5 (LOW)
|
||
> stale-primary sweep:** `pruneStalePrimaryDirs` removes an orphaned `backups/primary/<app>` dir on an
|
||
> OLD drive after an HDD_PATH move (guarded: deployed + different-current-drive only, never a restore
|
||
> point). **Part 4 locality fork → operator chose (A) keep locality, doc-only** (NAS tier-1 stays on
|
||
> the NAS; tier-2 is the off-NAS leg). LIVE: F7 money-shot (all NAS tars byte-identical through a
|
||
> mid-write cut, no 0-byte, success:false); F6 (actualbudget/seerr on felhom-usb/secondary); F5
|
||
> (seeded stale dir swept, current kept); restore round-trip byte-identical. Seams: `tarVolume`,
|
||
> `perAppTier2`. Task C (F8/F4) + Task D (ring/alerting) still queued; Peti reaches 0.118 + agent 0.85
|
||
> at his next train (agentless-on-proxmox2 gap noted).
|
||
|
||
> **2026-07-11 — v0.116.0/0.116.1: OBSERVABILITY PASS (LIVE on 9201; agent v0.83.0 + hub v0.46.0).
|
||
> MinAgent: 0.81.0 unchanged.** The debug ring (`LogBuffer`) now ALWAYS exists — logger =
|
||
> `MultiWriter(LevelFilterWriter(stdout, logging.level), ring)`, so DEBUG detail is remotely
|
||
> readable on an `info` box while docker logs keep the configured level. New `internal/logx`
|
||
> leveled helpers = the sweep standard (netstorage_job phases/verdicts/durations, netprobe,
|
||
> validation refusals, orphan WARN, `SupportsWithSource` gate line, agentapi per-call DEBUG,
|
||
> migrate phases, tier2/offbox unswallowed persists). Report ACK gains `controller_log_requested` →
|
||
> next report ships `controller_log_tail` (selftail.go, consume-once, 128 KB; the customer-visible
|
||
> `operator log pull served` INFO rides in the tail; app-tail wire byte-compatible). Debug page:
|
||
> `Vezérlő | Ügynök` tabs — the agent tab proxies agent `GET /debug/logs` (`Client.DebugLogs`;
|
||
> typed-404 → the "after the agent's next update" notice). **v0.116.1 (found by live validation):
|
||
> `/debug` + `/api/debug/*` + the nav item were STILL gated on logging.level=debug — ungated (auth
|
||
> unchanged), the incident's actual blind spot.** Live-proven at info: a real refused NAS add is
|
||
> fully reconstructable on both tabs (capability gate w/ source=version, phase lines, 502+duration,
|
||
> category). Conventions: `felhom.eu/documentation/runbooks/logging-conventions.md`. OPEN: operator
|
||
> clicks the hub's two "Request logs" buttons (hub UI password-gated) to close the live bundle
|
||
> round-trip; legacy `isDebug()` emission sites left for incremental migration. NOT published —
|
||
> Peti stays 0.113/0.81.
|
||
|
||
> **2026-07-11 — v0.115.0: version-aware Supports + DSM-validated NAS guidance (LIVE on 9201, pairs
|
||
> with agent v0.82.0 + hub v0.45.0). MinAgent: 0.81.0.** Capability detection now compares the agent
|
||
> version from agent v0.82.0's `X-Felhom-Agent-Version` header (`Client.noteAgentVersion` captures it
|
||
> on every response, strict semver; `features.go featureMinAgent` table + version-first `Supports`)
|
||
> instead of route-probing — the probe stays as the fallback for header-less (≤0.81) agents, so
|
||
> nothing changed for Peti's box. THE one comparator moved to `internal/util/version.go` (selfupdate
|
||
> aliases it). Part A DSM spike (real DSM 7.2 via virtual-dsm) validated the consumer recipes E2E; the
|
||
> NAS-page NFS guidance gained the verified Synology steps (File Services → NFS → **NFSv4.1**; "Map
|
||
> all users to admin"; `/volume1/<share>`); caveat narrowed to QNAP-only. Live-checked on demo: a real
|
||
> add shows `capability gate: netstorage_verify=yes` via the version compare, zero probes. **Q1c
|
||
> (Part E, supervised) FAILED**: a NAS automount trigger does NOT survive a guest reboot (guest sees
|
||
> an empty dir; agent has no network-mount reassert) — fix is felhom-agent's, spec'd at
|
||
> felhom.eu/documentation/backlog/FOLLOWUP-nas-automount-guest-reboot-reassert.md (controller
|
||
> health-cross-check follow-on noted there). NOT published (agent 0.82 demo-only; Peti 0.81).
|
||
|
||
> **2026-07-11 — v0.114.0: agent-capability gate (option-1) + publish-train rules (option-2).**
|
||
> Answer to the 0.81/0.113 train's 9-minute controller-before-agent skew (Peti's box): the box now
|
||
> protects itself. `internal/agentapi/features.go` — `Supports(Feature)` route-probes the agent
|
||
> (`GET /netstorage/verify-status` = the v0.81.0 coupling signal; typed `StatusError` 404 ⇒ No, 2xx
|
||
> ⇒ Yes, transport/5xx ⇒ Unknown NEVER refused; `SupportCache` TTL 5m both polarities, Unknown
|
||
> uncached). `handleNetStorageAdd` refuses on No BEFORE the single-flight claim (412 +
|
||
> `agent_outdated` + honest Hungarian message); settings page swaps the add form for a banner
|
||
> (list/remove untouched in every state). remove/list NOT gated. NO agent/hub changes; NOT
|
||
> published, floor untouched, Peti stays 0.113.0 — the gate is inert protection until the next
|
||
> train. Rules codified: `felhom.eu/documentation/runbooks/publish-train-rules.md` (manifest before
|
||
> floor; floor field LAST — hub_settings DB row overrides env + acts immediately; MinAgent fleet
|
||
> gate — CHANGELOG header convention starts with this release; the gate as box-level backstop).
|
||
> Tests T1–T6 + wire-level 404-typing; red-proofs RP1–RP5 in REPORT.md. Roadmap: agent
|
||
> version-in-envelope upgrade of `Supports`; hub floor-UI separation = its own task. The
|
||
> `agent_outdated` branch is test-proven only (demo agent is current — downgrade not justified).
|
||
|
||
> **2026-07-11 — v0.113.0: NAS verify-before-commit + page redesign (LIVE on 9201, pairs with agent
|
||
> v0.81.0 + host-install v1.13.0).** `POST /api/storage/netstorage/add` no longer registers blind
|
||
> (the bogus-share-at-Készenlét bug is dead): detached single-flight orchestration
|
||
> (`internal/web/netstorage_job.go`, migrate shape; poll `GET .../add/status`) = agent add (units +
|
||
> agent-side detached verify with journal classification + auto-rollback) → controller **uid-1000
|
||
> re-exec write probe** (`--netprobe`, SysProcAttr.Credential — catches the squash trap) → register
|
||
> LAST. Any failure = full rollback; verify-lost after agent restart ⇒ controller rollback; unregistered
|
||
> agent shares surface as remove-only "Árva megosztás" rows. §3.2 Hungarian error map server-side
|
||
> (`netAddMessage`; `nfs_export` MERGES not-found/not-permitted — NFSv4 identical strings).
|
||
> storage_network.html rebuilt on the storage_attach pattern (form-row/form-input killed), SMB listed
|
||
> first, NFS two-recipe guidance with live computed uid+100000. Agent v0.81.0: NFS `retry=0`
|
||
> (dead-NAS access 91 s→3.8 s), `ClassifyNetVerifyFailure` (Q4-verbatim), unprivileged journal read
|
||
> (systemd-journal group — host-install v1.13.0 adds it; NO new sudoers). Live-validated A–E on 9201
|
||
> vs an isolated sim NAS (all transcripts + red-proofs in REPORT.md); Route A proven in production
|
||
> (alien-uid squash → server-side 1060:1060). Authoritative doc:
|
||
> felhom.eu/documentation/controller/network-storage-nas.md. NOT published (0.81.0 not in Gitea /
|
||
> Day-0 manifest; Peti untouched — his rollout incl. the usermod one-liner comes with the floor bump).
|
||
> Gotcha for future sessions: the controller container is bridge-only — in-guest API tests need the
|
||
> CONTAINER IP + `Host: felhom.demo-felhom.eu` (127.0.0.1:8080 is stale advice).
|
||
|
||
> **2026-07-10 — v0.112.0: self-update without credentials (LIVE on 9201, pairs with hub v0.43.1).**
|
||
> Root cause on Peti's box: the updater refused without Git Sync creds, but the public package is
|
||
> anonymously pullable. `queryRegistry` now does the Docker v2 anonymous token dance when both creds are
|
||
> empty (realm/service parsed FROM the WWW-Authenticate header — never hardcoded); `pullImage` skips
|
||
> `docker login` credential-less; creds path byte-unchanged (private catalogs); half-configured pair =
|
||
> loud misconfig; denial = "registry denied anonymous access — a private registry requires Git Sync
|
||
> credentials". Settings panel gains the mode line "Registry: nyilvános (hitelesítés nélkül) /
|
||
> hitelesített" — credential-less is a supported mode, not an error. Red-proof green (old guard restored
|
||
> → anonymous tests fail with the old message). LIVE-PROVEN on the credential-less demo (git creds are
|
||
> quoted-empty): /api/selfupdate/check → ok, latest=0.112.0, no error; settings shows "nyilvános".
|
||
> PENDING OPERATOR: floor-bump Peti to 0.112.0, then delete his temp Git Sync creds → clean "nyilvános"
|
||
> check. runCommand/runCommandStdin are now package VARS (test seam).
|
||
|
||
> **2026-07-10 — v0.111.0: remote app-log diagnostics (LIVE on 9201, pairs with hub v0.43.0).** The
|
||
> telemetry scraper now attaches `LogIssue.Context` (±5 raw lines around the FIRST occurrence of each
|
||
> error-severity issue; ≤11 lines, ≤400 chars/line, 16KB/report budget dropping lowest-count first; warns
|
||
> carry none) and `metrics.RedactLine` sanitizes EVERY off-box context/tail line (password/token/api-key/
|
||
> authorization/bearer → `[REDACTED]`, 64-hex → `[REDACTED-HEX64]`). On-demand log tails ride the ACK pull
|
||
> pattern: hub ACK `log_tail_requests` → next report `log_tails` (200 lines via stacks.GetLogs /
|
||
> FetchContainerLogTail, ordered, ≤64KB/app newest-kept, redacted, consume-once drain). Hub v0.43.0 stores
|
||
> context (first-capture-wins + `context_customer` provenance), renders click-to-expand copyable issues,
|
||
> fixes the period filter on Known Issues, replaces issue deletion with DISMISSAL (`dismissed_at`,
|
||
> resurface only on `last_seen > dismissed_at`), adds `?customer=` filtered drill-down, and keeps the last
|
||
> 2 tails per app with an ordered viewer + .log download. All red-proofs green (capture, redaction,
|
||
> consume-once ×2, dismissal guard, range filter, context clobber). Live-proven on demo: synthetic error →
|
||
> hub row with ordered 11-line context and `password=[REDACTED]`. OPERATOR: one click ("Request log tail"
|
||
> on demo felhom-controller) completes the live tail round-trip — hub UI is password-gated, CC cannot.
|
||
|
||
> **2026-07-09 — v0.106.0: offsite provisioning SLICE 2 — the apply-bridge (pairs with hub v0.38.0).** On
|
||
> startup the controller reconciles the hub-served `offsite:` descriptor into a key-only offbox target:
|
||
> `internal/offsiteapply.Bridge.Reconcile` — verify-pin the box host key against `host_fingerprint` (NO blind
|
||
> TOFU) → generate keypair → consume the one-time password (`POST /api/v1/offsite/consume-password/{id}`,
|
||
> single-use, never logged) → `sshpass ssh-copy-id -s -f` install + verify → `Manager.ApplyOffsiteTarget`
|
||
> (fork-4 enable → `EscrowState="pending"`) → persist a descriptor-hash marker LAST. **Idempotent** (no
|
||
> re-consume of a spent password) + **fail-safe** (any step fails → nothing persisted, retry next restart;
|
||
> consumed-but-failed install = loud "reset on the hub"). Seams faked in tests; both red-proofs (no-TOFU,
|
||
> marker-after-success) green. `Dockerfile` + `sshpass`. **NOT yet live-applied** — supervised end-to-end
|
||
> (hub provision → controller apply) is the next runbook, gated on the hub's new scoped `HETZNER_TOKEN`.
|
||
> NEXT slices: SLICE 3 (escrow auto-confirm), SLICE 4 (soft-quota).
|
||
|
||
> **2026-07-09 — v0.105.0: fork-4 offsite password custody (pairs with agent v0.77.0).** The restic-offsite
|
||
> repo password now rides the **customer-R escrow** (age-under-R in the agent `IdentityBundle`; custody spike
|
||
> `febdc56`). Enable → controller pushes the password (`StageEscrowSecret` → agent `POST /escrow/stage-secret`)
|
||
> → `EscrowState="pending"`. **Atomicity gate:** no offsite RUN until `EscrowState="escrowed"` (operator
|
||
> `POST /backup/offbox/confirm-escrow` after the escrow ceremony) — so no un-recoverable offsite ciphertext
|
||
> exists. **DR:** `POST /backup/offbox/inject-password` pre-places the recovered password (honored by
|
||
> `WriteOffboxSecrets`). DR recipe gains non-secret `offsite_restic` coords (`DRResticCoord`); the SFTP key is
|
||
> regenerated at DR (not escrowed). Atomicity + inject companion red-proofs green. **NOT yet live-validated**
|
||
> — the supervised escrow ceremony (enable→stage→escrow-create→confirm→gated run) is operator-run; NEXT =
|
||
> hub-verified auto-confirm + customer-self-serve enable (provisioning task). Deployed to 9201; see REPORT.
|
||
|
||
> **2026-07-09 — v0.104.0: off-box discovery over inference + no-silent-success.** The Storage-Box spike
|
||
> found offbox reporting `ok`/0 snapshots while backing up nothing; DIAG pinned it: offbox resolved each
|
||
> toggled app's recovery unit via `AppNamespaceRoot`→`GetAppDrivePath`, which reads the app's *live*
|
||
> `app.yaml` `HDD_PATH` and **silently falls back to `systemDataPath`** when the app isn't deployed → it
|
||
> looked on the wrong drive. **Decision: DISCOVER, don't infer** — scan the durable storage registry
|
||
> (schedulable, non-decommissioned paths ∪ systemDataPath) for `backups/primary/<app>`, deployment-state
|
||
> independent; newest-by-manifest wins on drive churn. **Silent-success closed:** 0-of-N toggled → hard
|
||
> error + operator alert; partial → `ok` + customer `LastWarning`. Write paths + `AppNamespaceRoot`
|
||
> untouched. Unit suite + both companion red-proofs green. **NEXT:** supervised box re-provision + a real
|
||
> offbox→Storage-Box endpoint round-trip (this task did NOT re-point at the live box — spike creds were
|
||
> torn down). Deployed to 9201; see REPORT.md.
|
||
|
||
> **2026-07-07 — v0.103.0: F-C2-1 (LIVE on 9201).** The config loader ran `os.ExpandEnv` over the
|
||
> whole YAML before parse, silently corrupting a bcrypt `web.password_hash` (`$2a$10$…` → `"a0"`) — a
|
||
> silent auth-integrity bug. Removed both `ExpandEnv` calls (parse raw bytes); typed
|
||
> `FELHOM_WEB_PASSWORD_HASH` override unchanged. Live-proven: a bcrypt hash in controller.yaml now
|
||
> loads intact and login succeeds (pre-fix it corrupted → login fail). Behavior change: literal
|
||
> `${VAR}` in a value is now preserved verbatim (no repo config depends on the old expansion).
|
||
|
||
Last updated: 2026-07-06 (v0.102.0 — async restore family; F4 re-adjudicated + fixed)
|
||
|
||
> **2026-07-06 — v0.102.0: async restore family (F4 UX fix, LIVE on 9201).** All three restore surfaces
|
||
> (`/backup/restore`, `/backup/tier2/restore`, `/backup/offbox/restore`) blocked the HTTP request until
|
||
> completion → through cloudflared's 100s cap a customer got an error page while the restore succeeded
|
||
> (offbox worse: bounded on `r.Context()`, canceling the SFTP restore mid-flight). Now async (offboxRun
|
||
> shape): fast-path IsRunning refuse → background goroutine (offbox ctx off r.Context()→Background+30m) →
|
||
> instant redirect. New `GET /api/backup/restore-status` + mutex op-status (`opstatus.go`) + 3s-polling
|
||
> `backups.html` banner. Live-proven: restore POST 0.018s internal / **0.235s external (F4 tunnel)**, canary
|
||
> bit-identical, status transitions. Restore single-flight unchanged. OPEN: op-status is in-memory (no
|
||
> persistence, by design).
|
||
|
||
> **2026-07-06 — v0.101.0: no-mercy campaign findings.** F3: git subprocess deadline in
|
||
> `internal/sync/sync.go` (`gitCmdTimeout=120s`, `exec.CommandContext`) — a hung remote no longer
|
||
> wedges `syncing=true` until restart. F2 evidence gap: `agentapi.EjectDisk`/`Decommission` now use
|
||
> `postWithStatus` + `refusalError` so the agent's `"…refused (role: X)"` reaches the operator
|
||
> instead of a bare `HTTP 403`. Companion: catalog `d86e256` (F1 vaultwarden `_ENABLE_SMTP` boot-gate
|
||
> — fresh email-off deploys crash-looped; live-validated Scenarios A/B on 9201). F2 diagnosed to a
|
||
> verdict (REAL finding — `roleForMountPath` over-refuses an enrolled user-data drive that isn't a
|
||
> PVE storage; fail-safe direction; agent fix DEFERRED). Full triage:
|
||
> `felhom.eu/documentation/audits/CAMPAIGN-nomercy-2026-07-06.md` addendum. OPEN follow-ups: agent
|
||
> `roleForMountPath` fallback; the targeted P1–P3 campaign re-run for clean backup/restore coverage.
|
||
|
||
> **2026-07-05 — v0.100.0 (TASK C2): drill finding F2 CLOSED — one-click class-C file restore.**
|
||
> `POST /backup/tier2/restore` + "Fájlok visszaállítása" on the Tier-2 row: in-place, ADDITIVE-ONLY
|
||
> (`rsync -a --ignore-existing` from the recorded Tier-2 copy — never overwrites, never deletes).
|
||
> Serves "I deleted my files"; corruption/point-in-time stays offbox/operator. **The C-series
|
||
> (drill findings F1/F2/F3/O4) is now fully closed.** Reindex caveat (e.g. Nextcloud occ files:scan)
|
||
> documented in backup-architecture.md.
|
||
|
||
> **2026-07-05 — v0.99.0 restore-path fixes (TASK C1): drill findings F1/F3/O4 RESOLVED.**
|
||
> F1: `GET /api/backup/snapshots` implemented (`backup.ListRestorePoints`) — the restore panel
|
||
> populates and the restore button enables. F3: `runVolumeDumps` wired into the nightly/manual
|
||
> backup run (volume gate BEFORE DumpAppVolumesSafe; before unit capture). O4: unrecoverable
|
||
> resettable secrets get a generated replacement (`GenerateSecretForField` + `SetSecretGenerator`
|
||
> seam); data-key gate untouched. **F2 (one-click in-place class-C restore) remains OPEN → TASK C2.**
|
||
> O4 residual: restored volume tar with an OLD credential hash may still need a manual in-DB reset.
|
||
|
||
> **2026-07-04 — app-data restore drill** → see `felhom.eu/documentation/audits/DRILL-appdata-restore-2026-07-04.md`.
|
||
> Keep-side restore proven live on 9201 (class-A DB replay + fail-closed data-key gate + non-destruction). **F1 (HIGH): UI restore is dead — `/api/backup/snapshots` has no handler, so the restore button never enables.** F2: no one-click in-place class-C (HDD bind-mount) restore. F3: named-volume data never backed up (`DumpAppVolumes*` has no caller). *(F1/F3/O4 resolved in v0.99.0, see above.)*
|
||
|
||
> **2026-07-03 — CLAUDE.md slimmed to stable orientation** (full package map, verified 9201 deploy
|
||
> summary, no version-pinned state). Deep runbooks/design/testing doctrine now in the personal
|
||
> skills `felhom-build-deploy` / `felhom-ui-design` / `felhom-testing` (source: `felhom.eu/skills/`).
|
||
|
||
> **2026-07-03 — `REUSE.md` exists at the repo root** (canonical helpers / patterns / traps / seams, code-verified). Check it before writing new code; update it in the same commit that adds/changes a shared helper (maintenance rule now in CLAUDE.md).
|
||
|
||
> **2026-07-02 — v0.98.0 (deployed on 9201): Tárhely IA follow-up.** User feedback on D1: NAS-add and
|
||
> local-drive enrollment buttons sat side by side — confusing. `/storage` split into two subpages
|
||
> under the Tárhely nav item (nested sub-links, `.nav-links-nested`): **/storage = Meghajtók**
|
||
> (drives + agent view + wizards + manual add) and **/storage/network = Hálózati tárhely (NAS)**
|
||
> (NAS-megosztások list + add form; own openDialog copy). `networkStoragePageData` (page key
|
||
> `storage-network`) split out of `storagePageData`. No API/storage-semantics change.
|
||
|
||
> **2026-07-02 — v0.97.0 (deployed on 9201): TASK-D1 — settings split + Tárhely page (IA only).**
|
||
> The 1451-line `settings.html` monolith is split into four pages: `/storage` (new main-nav
|
||
> **Tárhely**: Adattárolók + NAS + unified agent drive view + wizard entry), `/settings` (Rendszer:
|
||
> konfig, verzió/frissítés, vezérlő+kiszolgáló újraindítás), `/settings/notifications` (Értesítések
|
||
> + Alkalmazás-email), `/settings/security` (Jelszó, Földrajzi korlátozás, Vészhelyzeti információk).
|
||
> Sidebar gains the Tárhely item + a "Beállítások" group with three sub-links. **No storage/agent/API
|
||
> behavior change** — pages moved, logic didn't; every `/api/storage/*` endpoint + POST action
|
||
> unchanged. **Decisions:** (a) storage flashes + wizards moved to `/storage`; old
|
||
> `/settings/storage/{init,attach}` **301** to `/storage/{init,attach}`. (b) The two duplicated drive
|
||
> views MERGED: registry cards render server-side, then JS enriches each connected user-data card in
|
||
> place from the agent `/api/disks` (role tag, durable-id, agent actions) joined on mount path, plus
|
||
> two groups — **Rendszermeghajtók** (system/backup, read-only, lock tag, NO actions) and **Nem
|
||
> regisztrált** (register only); agent-down → one warn note, registry cards still render. (c) Every
|
||
> native `confirm()`/`prompt()` on the four pages now routes through a light `.confirm-overlay`
|
||
> (`openDialog`; texts verbatim; type-to-confirm kept only where it existed) — the D0 tab-freeze is
|
||
> gone. **Contract for future UI:** new templates must pass `scripts/template_id_gate.py` (every JS
|
||
> element-ID resolves in its own template) and `scripts/emoji_gate.py` (0 emoji — the D0 grep gate
|
||
> false-negatived multibyte emoji on Windows; 8 survivors removed). `.badge-lock`/`.lock-ico` CSS
|
||
> deleted; other `.badge*` still used on secondary pages (D2 to finish). **Next: D2** (badge→tag on
|
||
> backups/secondary pages; storage.html is still ~700 lines — NAS-add partial candidate). NOT
|
||
> live-validated: agent-down degradation (don't stop the live agent — static/unit only); destructive
|
||
> storage ops via the moved overlay paths (endpoints unchanged; supervised session).
|
||
|
||
> **2026-07-02 — v0.96.0 (deployed on 9201): TASK-D0 — design system v2 re-skin (appearance only).**
|
||
> The whole customer UI (dashboard AND setup wizard) now renders in the approved v2 language: navy
|
||
> token palette (`--bg-0/1/2, --line, --text-1/2/3, --blue, --warn, --crit`), single 2px radius, no
|
||
> shadows; **exception-based status color** — nominal is blue/neutral, amber/red only on deviation;
|
||
> zone bars → the hairline `.meter` (neutral 70/85 ticks, warn/crit flag „Fogyóban a hely" /
|
||
> „Kritikusan kevés hely"); pills → `.tag` / `.metarow`. **Decisions:** (a) `stateColor` semantics
|
||
> changed — **stopped/exited = neutral, NOT red** (operator-approved; failures surface via unhealthy
|
||
> + alerts); restarting = warn; outputs are now `run/progress/warn/neutral/off` and
|
||
> `nominal/warn/crit` — any new template MUST use these suffixes (truth tables guarded by unit tests,
|
||
> `stateLabel` copy frozen byte-identical). (b) Fonts (PJS + JBM variable woff2, latin+latin-ext) and
|
||
> a 30-icon Lucide sprite are **vendored in the binary** (`/static/fonts/`, `templates/icons.html`)
|
||
> — no CDN; `fonts.googleapis.com`, `box-shadow`, 999px and the emoji set are banned strings (grep
|
||
> gate in REPORT §5, 143→0). (c) Setup wizard `handleCSS` bug fixed: it read a dataDir-derived path
|
||
> that never exists in the container → always served minimalCSS in production; now serves embedded
|
||
> `web.StyleCSS()`. (d) Found+fixed pre-existing v0.93.0 bug: `/backups` 500'd once an off-box backup
|
||
> had run (`OffboxTarget.LastRun` is an RFC3339 string fed to `timeAgo`; new `timeAgoStr`).
|
||
> Canonical reference: `felhom.eu/documentation/design/design-system.md`. **Next: D1** (settings IA
|
||
> split + Tárhely main-nav promotion; also migrate the remaining native `confirm()` dialogs to the
|
||
> type-to-confirm overlay and finish badge→tag markup on secondary pages). NOT live-validated: setup
|
||
> wizard visuals (unit-only), warn/crit meters on real hardware (demo healthy).
|
||
|
||
> **2026-07-01 — v0.95.0 (deployed on 9201): Impl-2b — raw-drive enrollment wizards.** Both enrollment
|
||
> wizards (`storage_init.html` / `storage_attach.html`) now fetch candidates from the agent's raw-device
|
||
> scan `GET /api/disks/candidates` (proxy `agentDiskCandidatesHandler` → agent Impl-2a) instead of the
|
||
> `Observe()`-based `/api/disks`, so a brand-new (non-PVE-storage) drive is discoverable + enrollable.
|
||
> `agentapi.ListCandidates` + `resolveEnrollUUID` (a raw candidate isn't in `/disks` → resolve its
|
||
> fs-UUID from the scan). Live end-to-end on felhom-pve: raw `/dev/sdd` (SD card) → wizard → format (Impl-1
|
||
> guard) / attach → mount → bind → intent → guest-bind → **"Aktív"** in the UI (3rd drive, no false
|
||
> detach). Needed a chain of AGENT fixes (v0.56–0.58: durableIDForMount / ReassertGuestBinds / HostReader
|
||
> wiring / /disks GuestPath+BoundUnderParent for registry rows) — the whole enroll+tracking machinery had
|
||
> assumed a PVE-storage drive. **Follow-ups:** `runStorageInit` should poll the agent's detached-format
|
||
> status (`/disks/format/status`) so a SLOW-device format completes in one flow (pre-existing); Impl-3
|
||
> shared-box operator format gate.
|
||
|
||
> **2026-06-30 — v0.94.0 (deployed on 9201): pull-based config-refresh.** The hub report ACK now also
|
||
> carries a per-customer **`config_version`** (hub v0.26.0; a stored counter bumped on every config save).
|
||
> `OnPushResponse` → `ConfigRefresher.Reconcile` (`internal/report/config_refresh.go`): on a change vs.
|
||
> `settings.applied_config_version`, `bootstrap.RefreshConfig` re-pulls `controller.yaml` (re-merging
|
||
> `local_api` from bootstrap.json; overwrites controller.yaml, never settings.json) → record → graceful
|
||
> self-restart (`api.GracefulSelfRestart`). First-run records baseline (no restart); unchanged = no-op
|
||
> (no storm); failed pull keeps config + retries. This is the box-pulls-config replacement for the hub's
|
||
> retired inbound "Push Config" — the hub never connects into the box. **Companion hub v0.26.0** also
|
||
> retired Trigger Update / Pull Config / Show Diff and dropped geo-disable's inbound notify (kept the
|
||
> hub→Cloudflare WAF removal), and fixed the stale `docker-setup.sh` setup command → host-install.
|
||
> Live-validated on 9201: first-run baseline (applied=1, no restart) → DB bump 1→2 → re-pull+self-restart
|
||
> (session_secret rotated, local_api preserved, RestartCount 0→1) → 4 cycles no-loop → apps stayed up.
|
||
|
||
> **2026-06-29 — Self-health arc complete (agent v0.48.0 + hub v0.22.1).** The hub now proactively
|
||
> watches every agent's served local-API **leaf fingerprint** fleet-wide: the agent reports
|
||
> `leaf_fingerprint` on its host report (v0.48.0), and the hub `HostLeafChecker` (`monitor/host_leaf.go`)
|
||
> raises `host_leaf_changed` when it changes (trust-on-first-report baseline; empty fp = unknown, never
|
||
> alerts; reads from report_json, no migration; hub-generated so no allowlist change). Independent of the
|
||
> controller channel-check. Live: a leaf regen raised `host_leaf_changed` (`82078fab…→60b5974d…`) AND the
|
||
> controller's `agent_channel_pin_mismatch` — complementary. NOTE: v0.22.0 shipped the checker but the
|
||
> main.go wiring never applied (concurrent file-touch); the LIVE TEST caught it (no event on regen) →
|
||
> fixed in v0.22.1. The original silent-multi-day-outage incident class is now caught from THREE angles
|
||
> (agent capabilities v0.44.0 / controller channel-check v0.90.0+F2 / hub leaf-fp v0.22.x) AND prevented
|
||
> (agent v0.46.0 loud-regenerate + host-install --preserve-state-from). Backlog: the served-fp-vs-pinned-fp
|
||
> authoritative cross-check; the test-run coverage gaps (capability→hub live, host-reboot doubling).
|
||
|
||
> **2026-06-29 — v0.91.0: F2 (channel-health born-down alerting).** The v0.90.0 channel-health check
|
||
> alerted only on a live up→down transition; a channel broken at **startup/reseed** (e.g. the
|
||
> controller boots right after a leaf regen → first observation is `pin_mismatch`) was dashboard-only,
|
||
> no operator email — forever. Fixed with an `alerted` flag in `channelhealth/checker.go`: a confirmed
|
||
> down coming from up/unseeded OR a reason-change re-arms then alerts once; steady-down no dup; recovery
|
||
> re-arms; debounce intact (transient born-down still N≥2; healthy first-obs silent). Live-validated:
|
||
> controller restarted into a regenerated-leaf channel → `[channel] DOWN (unseeded->down:pin_mismatch)`
|
||
> + `Event pushed: agent_channel_pin_mismatch` (~60s) — the capstone's missing half. Companions: hub
|
||
> v0.21.0 (same fix for HostCapabilityChecker/HostStalenessChecker — seed only healthy, born-degraded/
|
||
> stale emits on first Check, cooldown dedups), agent v0.46.0 (EnsureLeaf loud-WARNs a REGENERATED leaf
|
||
> + host-install `--preserve-state-from`/populated-host guard = the prevention for the original incident
|
||
> class). Self-health story now complete: agent watches its capabilities, controller watches its link,
|
||
> detection works at startup, and a reinstall can preserve the leaf. Backlog: F1 swap-rollback
|
||
> RestartCount hardening; full live reinstall (supervised); hub-side leaf-fp fleet comparison.
|
||
|
||
> **2026-06-29 — v0.90.0: controller→agent channel health-check (self-health slice).** A ~60s
|
||
> scheduler job (`internal/channelhealth`) probes the local-API channel via the PRODUCTION memoized
|
||
> client (`Server.ProbeAgentChannel` → `s.agentClient()` + GET /storage — NOT a fresh client, per the
|
||
> spike: self-heals, mirrors the disk UI, no transport leak). Classifies failures (spike Q1 map:
|
||
> pin_mismatch/unauthorized/unreachable/timeout/misconfigured/construction_error/unknown), **debounces**
|
||
> transient reasons (refused/timeout need N≥2 consecutive so the ~1s agent-restart blip doesn't page;
|
||
> pin/401/DNS/construction alert first-obs), seeds the first observation silently, and on a transition
|
||
> emits an **English operator-only** event (`Notifier.NotifyAgentChannelDown/Recovered`) + a Hungarian
|
||
> dashboard banner (`AlertManager.SetAgentChannelAlert`). Closes the R1-incident gap (channel was only
|
||
> checked once at startup). **Required a hub change** (felhom-hub v0.20.0): the `agent_channel_*` event
|
||
> types were rejected by the hub's `allowedEventTypes` allowlist (HTTP 400) — controller-pushed events
|
||
> are gated, unlike the hub-generated host_* ones. Live-validated: transient restart → no alert;
|
||
> sustained stop → debounce(1/2) then DOWN transition + dashboard banner; start → recovered. No agent
|
||
> change. Self-health story now: agent watches its own capabilities (v0.44.0) + controller watches its
|
||
> link to the agent (this). Backlog: hub-side leaf-fp comparison.
|
||
|
||
> **2026-06-29 — OPS (no code change): controller↔agent TLS leaf-pin mismatch RESOLVED via R1.**
|
||
> The 2026-06-28 root→non-root agent migration regenerated the agent's local-API leaf
|
||
> (`60b5974d…`→`ced34036…`) and truncated its token store, so the in-guest controller's pinned
|
||
> `agentapi` client failed every call (`TLS pin mismatch`) — disk UI dead AND the quiesce
|
||
> app-consistent-backup loop dead. Fix: restored the pre-migration `local-api.{crt,key}` **+**
|
||
> `local-tokens.log` from `/root/agent-backup-bundle-test/aside-var-lib/` on felhom-pve (R1 RUNBOOK,
|
||
> supervised). Agent now serves `60b5974d…` again (== controller pin + bootstrap.json seed) and the
|
||
> restored token store's 9201 entry matches the controller's current `local_api.token` — **zero
|
||
> in-guest change.** Gates passed: leaf loaded (not regenerated), wire fp `60b5974d…`, authenticated
|
||
> `GET /disks` 200/5-disks (bogus token 401), and the next quiesce cycle ran a real backup job.
|
||
> Rollback point kept at `/root/agent-rollback-20260629-150137/`. **OPEN (carried forward):** the
|
||
> migration/Day-0 path must carry these three files into the new `felhom-agent`-owned data dir (or
|
||
> treat their loss as a re-bootstrap trigger), else any re-migration reintroduces this. Full record:
|
||
> `felhom.eu/documentation/audits/SPIKE-agent-pin-mismatch-2026-06-29.md` (§9 Outcome).
|
||
|
||
> **2026-06-27 — v0.86.0 (deployed on 9201): Phase 2 managed updates — controller-version FLOOR.**
|
||
> On top of the Phase 1 opt-in "update to latest" button, the controller now honors an operator-enforced
|
||
> **minimum version (FLOOR)** delivered on the hub **report ACK** (`min_controller_version` +
|
||
> `latest_version`; `internal/report/pusher.go` `PushResponse`). `OnPushResponse` →
|
||
> `updater.SetFloor()` + `updater.MaybeAutoUpdate()` (rides the report cycle — no new timer). Below the
|
||
> floor → **auto-update to the floor** (reuses Phase 1 `performUpdate`: in-guest pull → agent swap →
|
||
> rollback; `initiatedBy="auto-floor"`). At/above floor → nothing (does NOT chase latest — that's the
|
||
> button). Guards: dev/no-agent/backup → skip; floor must be pullable (floor ≤ latest; floor>latest →
|
||
> warn+noop); one attempt per below-floor condition (in-mem + persisted state) → no flapping. Floor source
|
||
> + operator UI are hub-side (felhom-hub v0.15.0: per-customer override + global default + report ACK).
|
||
> **No agent change** (reuses Phase 1 `POST /controller/swap`). Day-0 now ships current (golden rebuilt at
|
||
> 0.85.1) AND stays current (floor) — the fleet-currency story is closed. Live: dogfood 0.85.1→0.86.0 via
|
||
> the button, then floor 0.87.0 → auto 0.86.0→0.87.0 (no click).
|
||
|
||
|
||
> **2026-06-26 — v0.84.0 (deployed on 9201): show an app's auto-generated first-login on its page.**
|
||
> General, catalog-driven mechanism: `.felhom.yml initial_credentials: {file, format json|regex|plain,
|
||
> username_key/password_key | username_pattern/password_pattern, note}`. The controller reads the file
|
||
> **live** from the container (`internal/stacks/initialcreds.go` `ReadInitialCredentials` + pure
|
||
> `parseInitialCreds`, unit-tested), never persists it, and renders a "Kezdeti belépési adatok" card on
|
||
> `/apps/{slug}` (masked password + reveal/copy). First consumer: crafty-controller (Crafty writes a
|
||
> random admin password to `/crafty/app/config/default-creds.txt`). Live-verified on 9201: card shows
|
||
> username `admin` + the real extracted password. Reuse for any future self-seeding app. Security: same
|
||
> exposure class as the existing post-deploy reveal / default_creds card — relies on the prod dashboard
|
||
> being auth-gated (demo public-unauth is the separate tracked issue). Backlog idea unchanged: a
|
||
> `backend_scheme` hint for TLS backends (v0.83.0 line).
|
||
|
||
> **2026-06-26 — v0.83.0 (deployed on 9201): scoped Traefik backend transport for self-signed HTTPS apps.**
|
||
> Crafty is the **first/only catalog app with an HTTPS backend** (self-signed TLS on `:8443`, no plain-HTTP
|
||
> port). The crafty healthcheck fix un-withheld its route, exposing a pre-existing 502 (Traefik proxied
|
||
> HTTP to the HTTPS backend). Since traefik v3 forbids `insecureSkipVerify` via Docker labels, the
|
||
> controller now renders a file-provider dynamic file `dynamic/serverstransports.yml` defining a **named**
|
||
> `insecure-skip-verify` transport (`infra.RenderServersTransports` + `stacks.ensureServersTransports`,
|
||
> written from `EnsureBaseStack` outside the traefik-running guard, write-if-changed, hot-loaded). Apps opt
|
||
> in per-service via catalog labels `scheme=https` + `serverstransport=insecure-skip-verify@file`.
|
||
> **No global `insecureSkipVerify`** — verification stays ON for every other backend (scoped Option B).
|
||
> Pattern to reuse for any future self-signed HTTPS backend. Live-verified: `minecraft.demo-felhom.eu`
|
||
> 502→302; filebrowser (HTTP) unaffected. Backlog: a `.felhom.yml backend_scheme` hint so the catalog
|
||
> convention sets these labels instead of hand-adding.
|
||
|
||
> **2026-06-22 — v0.76.0 (deployed on 9201): three campaign-#3 hardening fixes.**
|
||
> - **S1**: corrupt `settings.json` no longer crash-loops — `save()` writes a last-known-good `.bak`
|
||
> (after the primary rename); `Load()` recovers from `.bak`, else preserves the corrupt file as
|
||
> `*.corrupt-<ts>` + safe defaults. New `Settings.LoadWarning` → dashboard banner. (`main.go` Fatalf
|
||
> now only on the IO-unreadable path.)
|
||
> - **F2**: `web.validStackName` gates `backupRestoreHandler` + `apiExportStart` (reject `/ \ .. NUL`)
|
||
> before any restore/export — closes the traversal defense-in-depth gap (no escape had occurred, but
|
||
> it relied on downstream map-lookups).
|
||
> - **S3**: `quiesce.readMarker` logs `[WARN]` + quarantines a corrupt marker to `*.corrupt-<ts>`
|
||
> instead of silently dropping it.
|
||
> All three live-validated on 9201 (truncate settings → recover from .bak no crash-loop; traversal
|
||
> restore → rejected, /etc intact; corrupt quiesce marker → quarantined). Tests + red-proofs.
|
||
> Still-open campaign-#3 deferrals (NOT done): time-chaos on a dedicated VM, host reboot (supervised),
|
||
> `.fab` import compose-fuzzing. See `felhom.eu/documentation/tests/test-campaign-3-2026-06-22-findings.md`.
|
||
|
||
Last updated: 2026-06-22 (v0.75.0 — gate userdata MkdirAll on a live mountpoint)
|
||
|
||
> **2026-06-22 — v0.75.0 (deployed on 9201): userdata MkdirAll gated on a live mountpoint.**
|
||
> Both `MkdirAll`-into-`<drive>/userdata` sites — the deploy belt (`stacks.ensureUserdataMounts`) and
|
||
> the FileBrowser sync (`web.syncFileBrowserMounts`) — now skip when an **external** drive root (under
|
||
> `StableParentDir=/mnt/felhom-drives`, `!= sysDataPath`) is **not a live mountpoint** (`system.IsMountPoint`).
|
||
> Closes the campaign-#2 hazard where a drive-absent window produced `mkdir …/userdata: permission denied`
|
||
> + transient `Created` flapping AND could write app data onto the guest **rootfs** (shadowed when the
|
||
> drive returns). The app is held by `planDriveGates` instead. System/local path is never gated. New
|
||
> `Manager.isMountPoint` seam + pure `web.skipFileBrowserPath` helper for tests. Live-proven via a C6
|
||
> re-run: 0 permission-denied, 0 shadow dirs on the rootfs, apps recover on reconnect.
|
||
>
|
||
> **Boot-ordering design note (NOT implemented — decide separately):** the *boot-time* `mkdir …
|
||
> permission denied` is a different cause — **docker's** boot-restore auto-starts drive-backed
|
||
> containers (`restart: unless-stopped`) before the agent mounts the drives, so docker (not the belt)
|
||
> tries to create the bind source; `planDriveGates` recovers them after mount convergence (why a reboot
|
||
> ends healthy). Options: (A) accept + suppress (apps self-heal; lowest risk; maybe downgrade the
|
||
> boot-window log level) — recommended now; (B) drive-backed app containers don't docker-auto-start at
|
||
> boot (restart policy `no`/`on-failure`) so the controller's gate is the sole starter after mounts
|
||
> converge — cleaner but bigger (must cover crash-restart too). See
|
||
> `felhom.eu/documentation/tests/test-campaign-2-finding1-recovery-diagnosis.md` (boot-ordering observation).
|
||
|
||
Last updated: 2026-06-22 (v0.74.0 — controller→agent connection-leak fix)
|
||
|
||
> **2026-06-22 — v0.74.0 (deployed on demo guest 9201): fixed the controller→agent connection leak.**
|
||
> `Server.agentClient()` built a new `agentapi.Client` (new bare `http.Transport`, `IdleConnTimeout:0`)
|
||
> per call → one leaked idle ESTABLISHED socket per agent call to `192.168.0.162:8443`, exhausting the
|
||
> ephemeral source-port range after ~5 days of controller uptime → EADDRNOTAVAIL, which had taken down
|
||
> storage UI + host-metrics + whole-guest backup (found in the 2026-06-22 unattended campaign). Fix:
|
||
> memoize ONE shared client via `sync.Once` + harden the Transport (`MaxIdleConnsPerHost:2`,
|
||
> `IdleConnTimeout:90s`). Live-proven: idle sockets to `:8443` stay flat at 2 across a 120-call burst
|
||
> (pre-fix grew ~1/call → 132). Agent/firewall untouched. Diagnosis + campaign findings live in
|
||
> `felhom.eu/documentation/tests/unattended-test-campaign-2026-06-22-*.md`. **Separate open item:** the
|
||
> defense-in-depth host firewall rule scoping `:8443` to the guest bridge subnet is still absent
|
||
> (pve-firewall disabled) — to be closed independently.
|
||
|
||
Last updated: 2026-06-13 (v0.60.0 backlog-Medium cleanup)
|
||
|
||
> **Live version: controller v0.60.0** (deployed on demo guest 9201), agent **v0.30.0** (AGENT-001 deployed), hub **v0.11.0**.
|
||
> **NOTE:** the long-form sections far below this banner are stale (last full pass ~v0.16.1). Current
|
||
> state is the dated entries immediately below + `CHANGELOG.md` + the now-authoritative **central docs at
|
||
> `felhom.eu/documentation/controller/`** (code-verified) + auto-memory `MEMORY.md`.
|
||
>
|
||
> **2026-06-13 — v0.60.0 backlog-Medium cleanup (BUGHUNT reconcile):**
|
||
> - **M25** (Server.integrationMgr data race — constructor goroutine vs post-construction `SetIntegrationManager`):
|
||
> **FIXED** via `atomic.Pointer`; `-race`-verified. **M4/M5/M6** verified already FIXED (no action).
|
||
> - **M18** (dump re-validation every 5 min — perf) and **M19** (`deriveStackName` misattribution — low
|
||
> incidence) verified LIVE but cross-package-entangled → branches `fix/m18-dump-validation-cache` /
|
||
> `fix/m19-stackname-crossref` (notes + fix plan, PENDING REVIEW, not deployed).
|
||
>
|
||
> **2026-06-13 — v0.59.0 audit fixes + documentation centralization:**
|
||
> - Fixed the validated 2026-06-13 audit findings (records: `felhom.eu/documentation/audits/`):
|
||
> **CTRL-001** (`.fab` import path traversal — manifest segment validator, fail the parse);
|
||
> **CTRL-T2-1** (ghost-deployed on crash — `app.yaml` now persists `deployed:false` until `compose up -d`
|
||
> succeeds, flipped true only on success; in-memory flag still true during pull for UX);
|
||
> **H10** (plaintext secret on encrypt failure — `SaveAppConfig` now fail-closed, returns an error);
|
||
> **M2** (misleading lock on init-only `SetStackProvider` removed). All with regression tests.
|
||
> - **AGENT-001** (wrong-disk wipe TOCTOU) fixed on agent branch `fix/agent-001-wipe-durable-reresolve`
|
||
> — PENDING REVIEW, NOT deployed (supervised merge+golden-rebake reserved; see `AGENT-001-FIX-NOTES.md`).
|
||
> - Built + deployed v0.59.0 to demo guest 9201 (bootstrap mechanism); verified healthy, dashboard 200.
|
||
> - Documentation centralized under `felhom.eu/documentation/` (controller subtree + audits + top index),
|
||
> code-verified; `controller/README.md` banner points there; this CONTEXT banner refreshed.
|
||
>
|
||
> **2026-06-13 — v0.58.0 OS/Docker-data split prevention layer (Phase 2; Phase 1 = agent v0.29.0):**
|
||
> - OS rootfs + Docker data split onto separate local-lvm volumes (golden bakes 32G rootfs + 256G
|
||
> /var/lib/docker, backup=1, **overlay2** so images live on the data vol). Infra protected by
|
||
> PREVENTION not placement: `system.GetDockerVolumeHeadroom()` (statfs "/" = data vol) reserves
|
||
> max(5GB,10%); `deployStack` refuses HTTP 507 below buffer; deploy.html banner+disable; monitor
|
||
> (warn80/crit90) watches the same vol; log rotation baked into the golden daemon.json.
|
||
> - Live-validated by DESTROYING + RE-PROVISIONING 9201 from the split golden: split layout, overlay2,
|
||
> images on data vol, lean rootfs, OS isolation (data vol 100% → rootfs healthy), deploy gate 507,
|
||
> ActualBudget volume on data vol, hub config pull, external access via CF. RomM/USB re-enroll = the
|
||
> documented final restore step (RomM data safe on host USB). See memory [[os-data-split]].
|
||
>
|
||
> **2026-06-13 — v0.57.0 UI fixes (Part A of the UI-fixes/storage-spike spec):**
|
||
> - A1: fixed the RIGHT storage list — `#host-storage-bars` (the JS-filled, agent-PVE-storage list:
|
||
> `local`/`local-lvm`/`felhom-pbs`/`felhom-usb`), which reordered on every poll. Now
|
||
> `enrichHostStorageTargets` sorts `/api/host-metrics` server-side + adds friendly Hungarian
|
||
> labels/purpose. Display-only — PVE storage ids never renamed. (v0.56.0's 4C had sorted the OTHER,
|
||
> server-rendered user-data list.)
|
||
> - A2: per-app Tier-2 config panel at `GET/POST /stacks/{name}/backup`; the dead-end "Beállítás" button
|
||
> (was → deploy page) is repointed there. Pin a target drive / toggle Tier 2 off; prefs
|
||
> (`UserDisabled`/`PreferredTarget`) persist on `CrossDriveBackup` and survive the runner's status
|
||
> writes (`withTier2Prefs`). Always visible incl. single-SSD + non-HDD (PBS-context) apps.
|
||
> - Part B (storage OS/data split spike) = build-nothing; findings → `felhom-agent/REPORT-storage-split-spike.md`.
|
||
> - Live-validated on guest 9201; build/deploy = golden bootstrap (`/etc/felhom-controller-image` + restart `felhom-controller-bootstrap.service`).
|
||
>
|
||
> **2026-06-13 — v0.56.0 Phase 4: FileBrowser scoping + UI polish (SLICE COMPLETE):**
|
||
> - 4A: FileBrowser bind scoped to `<drive>/appdata` (recovery units + Tier 2 copies under `backups/`
|
||
> NOT mounted → customer can't browse/delete the restore source). 4B: deploy storage step states
|
||
> files-on-drive / DB-on-fast-SSD. 4C: `buildStorageBars` stable sort + purpose description on the
|
||
> monitoring list (user-data drives only; agent local/local-lvm/pbs live on the storage page, not here).
|
||
> - Live-validated (9201): FileBrowser mount `/mnt/felhom-usb/appdata -> /srv/felhom-usb` (backups hidden);
|
||
> deploy + monitoring text rendered. **All 5 phases (1, 2, 2b, 3, 4) shipped + live-validated, v0.52→v0.56.**
|
||
>
|
||
> **2026-06-13 — v0.55.0 Phase 3: auto off-drive Tier 2 (rootfs-headroom guard):**
|
||
> - `internal/backup/tier2.go`: rsync `-a --delete` of each HDD app's recovery unit + appdata → a
|
||
> DIFFERENT physical disk (`<target>/backups/secondary/<app>/`). Auto target: prefer another registered
|
||
> drive (off-disk via `system.SamePhysicalDevice`), else internal SSD for SMALL units only.
|
||
> - **Rootfs-headroom guard** (`tier2FitsHeadroom`, unit-tested): SSD = ~8G guest rootfs, so REFUSE
|
||
> unless the unit fits leaving reserve = max(2G, 20%) free; honest "needs 2nd HDD" status when nothing
|
||
> fits — never fills the rootfs. Status via surviving `settings.CrossDriveBackup`; "2. mentés" UI card
|
||
> now populated (`buildAppBackupRows`). Daily `tier2-backup` 03:30 + `POST /api/backup/tier2`.
|
||
> - **Live-validated (9201):** happy path (RomM → SSD, off felhom-usb, 77KB, "[SSD: DB/config only]");
|
||
> refuse path (1G userdata dummy → REFUSED with honest msg, rootfs not filled); UI card shows
|
||
> "Sikeres → belső SSD (csak DB/konfiguráció)". Demo cleaned.
|
||
> - Next: Phase 4 (FileBrowser scoping + deploy-UI DB-on-SSD note + monitoring sort).
|
||
>
|
||
> **2026-06-13 — v0.53.0/v0.53.1 Phase 2: per-app recovery unit (capture side, SECRET-FREE):**
|
||
> - Each app's `backups/primary/<app>/` becomes a self-contained recovery unit: `compose/`
|
||
> (docker-compose.yml + .felhom.yml + **secret-stripped** app.yaml) + db-dumps/ + volume-dumps/ +
|
||
> `manifest.json` (image pins, secret env-var NAMES, data_key names, checksums, secret_source note).
|
||
> - **Secret-free by design.** Decided after reading the ACTUAL hub code: hub is zero-knowledge (no app
|
||
> secrets); app.yaml + key live on the guest rootfs → in the PBS whole-guest snapshot. So the unit
|
||
> stores no secret/data-key/image; restore recovers secrets from the guest's app.yaml (live/PBS),
|
||
> regenerates nothing. `data_key` (DeployField.DataKey; AdventureLog SECRET_KEY marked) = fail-closed
|
||
> restore annotation only.
|
||
> - Capture needs no decryption (non-secret env is plaintext; excludes secret-named + encrypted keys).
|
||
> Wired into RunDBDumps AND the periodic RefreshCache (idempotent checksum-skip → no USB thrash).
|
||
> - **Deploy mechanism resolved:** controller in guest 9201 is golden/bootstrap-managed —
|
||
> `felhom-controller-bootstrap.service` docker-runs the tag from `/etc/felhom-controller-image`
|
||
> (gitea anon-pull). Deploy = build+push → anon-pull → update tag file → restart the service.
|
||
> - **Live-validated (9201):** RomM unit captured (images=3, secrets=3, data_keys=0), secret-leak grep
|
||
> = NO_LEAK.
|
||
> - **v0.54.0 Phase 2b (restore-from-unit + fail-closed gate):** `RestoreFromRecoveryUnit` recreates an
|
||
> app from its unit + secrets recovered from the GUEST's live app.yaml (`RecoverStackSecrets`,
|
||
> `stacks.RedeployFromEnv`), regenerating nothing. `reconcileRestoreSecrets` (pure, unit-tested) is the
|
||
> fail-closed gate: missing/empty data-key → REFUSE (needs PBS whole-guest restore); missing resettable
|
||
> secret → warn+proceed. Wired into `/backup/restore`. Gate + orchestration + data_key parsing
|
||
> unit/integration-tested; deployed v0.54.0 healthy.
|
||
> - **LIVE-validated (9201, AdventureLog):** unit manifest `data_key_env_vars:[SECRET_KEY]`
|
||
> (catalog→manifest live); with SECRET_KEY made unrecoverable, `POST /backup/restore` REFUSED with the
|
||
> exact fail-closed message BEFORE any compose-up. Demo has NO dashboard password → API open (auth+CSRF
|
||
> skipped), driven via public URL. NOTE: full deploy-with-data→restore e2e blocked because AdventureLog
|
||
> images don't fit the 8G guest rootfs ("no space left") — that's the Phase 3 rootfs-headroom concern
|
||
> seen live. Demo left clean (AdventureLog reverted to not-deployed).
|
||
> - Next: Phase 3 (Tier 2 auto off-drive, rootfs-headroom guard), Phase 4 (FileBrowser + UI).
|
||
>
|
||
> **2026-06-13 — v0.52.0 Phase 1 GATE: deploy-side double-nest fix (catalog) + path-agreement test:**
|
||
> - The `felhom-data` double-nest lived in the **app-catalog compose templates**
|
||
> (`${HDD_PATH}/felhom-data/appdata/<app>`), not in `deploy.go`. On a Model-A in-guest drive the mount
|
||
> already IS the `felhom-data` namespace, so it double-nested on disk while the v0.51.0 backup helpers
|
||
> resolved single-nested → divergence. Fixed all four HDD templates (romm, nextcloud, immich,
|
||
> paperless-ngx) → `${HDD_PATH}/appdata/<app>`.
|
||
> - New `internal/stacks/hddpath_agreement_test.go` locks deploy-resolver (`ParseComposeHDDMounts`) ==
|
||
> backup helper (`AppDataDir(NamespaceRoot(.,true))`). No controller runtime change → no image rebuild
|
||
> (deployed stays 0.51.0, functionally current; golden not rebaked for a no-op).
|
||
> - **Live (guest 9201):** git-sync auto-delivered the fix to all four stack files; RomM migrated
|
||
> (stop→move→verify→redeploy) from `/mnt/felhom-usb/felhom-data/appdata/romm` →
|
||
> `/mnt/felhom-usb/appdata/romm`, healthy + HTTP 200, no data loss, old namespace empty. **GATE PASSED.**
|
||
> - Next: Phase 2 (per-app recovery unit), Phase 3 (auto-enabled off-drive Tier 2 w/ rootfs-headroom
|
||
> guard), Phase 4 (FileBrowser scoping + deploy-UI DB-on-SSD note + monitoring sort).
|
||
>
|
||
> **2026-06-12 — storage UX polish (v0.45.0), pairs with felhom-agent v0.24.0:**
|
||
> - **Agent eject role-gate (Part A, felhom-agent v0.24.0):** `POST /disks/eject` now refuses to
|
||
> unmount system/backup storage *at the agent* (fail-safe to protected on ambiguity) — the UI hiding
|
||
> the button was never the control. Validated live on guest 9201 (eject `/var/lib/vz` → 403, no unmount).
|
||
> - **Controller (Part B, v0.45.0):** B1 deterministic `/api/disks` order (user-data→system→backup,
|
||
> alpha within); B2 init wizard excludes mounted drives; B3 **Regisztrálás** primary action for a
|
||
> mounted-but-unregistered user-data drive (`POST /api/storage/register`); B4 per-card purpose
|
||
> descriptions + app-backing tags + tiering note (`local` & `local-lvm` both kept); B5 eject already
|
||
> names affected apps. All validated live on guest 9201.
|
||
> - **Golden REBAKED** to controller 0.45.0 (`/root/build-golden.sh`; gitea allows anon pull, no creds);
|
||
> archive `local:backup/vzdump-lxc-9100-2026_06_12-09_53_03.tar.zst`, build guest 9100 purged.
|
||
|
||
---
|