Commit Graph

593 Commits

Author SHA1 Message Date
admin e81ad613ae v0.256.1: the push log says whether a household sentence was attached (R-558)
gates / gates (push) Successful in 25s
Found while trying to PROVE v0.256.0 on a live box: there was nothing to look
at. The box logs the Hungarian message only, the hub logs the Hungarian message
only, and the payload is inside TLS - so "did the box attach a second
sentence?" was unanswerable from either end. That is the fork in the diagnosis
when an English household reports a Hungarian mail.

  [INFO] Event pushed: app_deployed (info) [+household(en)] — Alkalmazás ...
  [INFO] Event pushed: app_deployed (info) [hu-only] — Alkalmazás ...

The sentence itself is not logged: it is the same sentence twice and one copy
is already on the line. Pinned in both languages.

A feature whose only failure mode is "the wrong language arrived" needs an
observable that says which branch was taken. Without one, every diagnosis is a
guess.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 17:06:20 +02:00
admin bef39598d0 v0.256.0: the box sends its own sentence in the household's language (R-558 Part B)
gates / gates (push) Successful in 24s
MinAgent: 0.131.0 (unchanged). Needs hub v0.118.0+, which shipped first and
tolerates a box that sends none of this - every box in the fleet is that box
until this release reaches it.

The hub writes a household's e-mails in their language now, but about a third
of those mails carry a sentence the BOX composed, naming a drive, an app or a
number. The hub cannot translate one. So the box sends it twice.

- message_customer on POST /api/v1/event, omitempty. A HUNGARIAN household
  sends nothing extra at all, so its payload stays byte-for-byte what every box
  sends today and the hub's fallback path keeps being the one production
  exercises rather than a branch nobody takes.
- 19 producers render both sentences from ONE bundle key. `message` stays
  Hungarian always: it is what the operator is mailed and what the hub logs.
- customer.language bootstraps a new box - stored choice, then config, then
  Hungarian. The config value is NEVER written into settings.json: that would
  record a choice the household never made.

The Hungarian did not move, measured twice: the wire golden from the slice-2
base commit, and the Go parity gate over all 19 new keys.

Three guards had to learn the change and one caught me: the test seam now
carries the new field; the R-329 severity register reported two dynamic sites
as no longer existing the moment they moved off PushEvent (the walk now checks
36 severity literals, up from 20); and TestConfigLanguageIsWiredInMain reads
main.go, because cmd/ is gitignored and ripgrep does not.

A mistake, named: the first pass dropped displayName from three producers,
which would have mailed customers "Alkalmazás telepítve: %!s(MISSING)". Caught
reading the diff; now pinned by a test that refuses %!/MISSING/%s/%d in either
language.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 16:54:20 +02:00
admin 8fb2f9ef9d v0.255.0 — the globe on the sign-in-flow pages: styled, and inside the card
gates / gates (push) Successful in 23s
Two defects in v0.254.0's globe, both plain on a browser and neither catchable by anything
that existed — every test read the MARKUP, and the fault was in which CSS file the browser
fetched.

The shells requested /static/style.css with NO ?v=, while layout.html has carried one since
v0.166.0. A browser holding a copy from before v0.254.0 kept serving CSS with no .lang-globe
rules, so the globe came out as a bare unstyled <details> — a stray triangle and two plain
words at the edge of the window. It was FIVE shells, not the three named: both guest share
pages have the same fault for any CSS change, and their visitor is the likeliest of all to be
holding an old copy. And .Version was missing from three of those five data maps, which is
exactly how the next one would be forgotten — it is now filled at the one choke point every
shell renders through.

The globe also floated outside the card, pinned to the corner of the VIEWPORT, reading as part
of the browser rather than the page. It now sits inside the card, centred under the footer, with
the menu opening upward via the shared rule — so the dashboard and the shells cannot drift.

AND A THIRD, caught by a test that already existed: putting the version on the guest share pages
would have printed the controller build onto a page a stranger with a capability URL can open.
TestShareGuest_HeadersTilesNoAdminChrome refused it. Those two now take an opaque per-build tag
— same cache-busting, no disclosure. The fill is ONE function shared with the parity harness,
because a fixture rendered through a different data path is a picture of a page nobody serves,
which the previous release got wrong twice.

15 shell fixtures re-captured; 91 identical, every dashboard page among them.

MinAgent: 0.131.0 (unchanged). No hub release needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 15:07:49 +02:00
admin 0b1486df8f fix: the recovery page's globe must write the HOUSEHOLD's setting, not a cookie nothing reads
gates / gates (push) Successful in 24s
/recovery is in the AUTHENTICATED route table — its reader is the household, not a visitor. The
first draft of v0.254.0 gave it the anonymous form, which sets the felhom_lang cookie that
langFor deliberately ignores once there is a session: the button would have appeared to work
and done nothing. Found by the live probe on demo-hp reporting no globe on /recovery (it 302s
to /login without a session) and then reading the route table.

executeTemplateLang now branches on hasSession: household form with its session CSRF, or the
visitor form without. The parity harness and TestI18nDirectRenderPagesFollowLanguage carry the
same branch, so the fixture is the form the real page serves — the trap this release already
walked into once with the shells.

A per-session CSRF token cannot be a fixture value, so it is blanked on both sides of every
parity comparison, exactly as relative ages already were. What stays pinned is that the field is
THERE and WHICH form it sits in — the half that says whether the globe writes the household's
setting or the visitor's cookie.

Evidence regenerated: 3 change shapes across 106 fixtures, 5 byte-identical (both guest share
pages and the catch-all — the three that must not change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 14:28:24 +02:00
admin 48f3336956 v0.254.0 — the saved notes follow the language, and the switch becomes a globe (R-557 slice 2 release C; SLICE 2 CLOSED)
gates / gates (push) Successful in 23s
The notes a background run SAVES — last night's backup line, the last error, the proof
result, the restore outcome — are written in the BOX's language at the moment they are
written. A household that switches sees the previous run's note in the old language until
the next run rewrites it: the operator's §16 option 1, stated rather than hidden.
EndRestoreOp no longer receives a Hungarian literal from anywhere.

The language switch is a globe. Two text links wrapped in the sidebar footer and asked the
reader to recognise "Magyar"/"English" as links; a globe is the one symbol every web user
already reads as "language", so nobody has to read Hungarian to escape Hungarian. It is
<details>/<summary> — a menu with no script, drawn inline because the icon sprite lives
only in layout.html and the visitor pages have their own shell.

Those visitor pages get the same globe, and a visitor's choice stays theirs: a display-only
felhom_lang cookie that langFor reads ONLY when there is no session. A signed-in household
can never inherit a language a previous visitor picked in the same browser. POST /lang is
CSRF-exempt for a narrow reason written at the exemption — its only achievable effect is the
language of the page the victim's own browser shows them — and safeBackPath refuses
//evil.example as well as https://, because "starts with /" alone is not the test. §16 taken:
a successful claim carries the cookie into the household's setting.

TWO PARITY EXCEPTIONS, MEASURED: 106 fixtures compared with a real diff — exactly two change
shapes (the dashboard footer, the globe in the shells) and 5 byte-identical, which are the
three pages that must not change.

I INTRODUCED A DEADLOCK AND THE SUITE CAUGHT IT BY HANGING. UpdateOffboxStatus holds the
settings write lock while running its callback; boxLang() wants the read lock; sync.RWMutex
is not reentrant. On a real box an off-site run would have hung forever HOLDING the settings
lock. Fixed by resolving the language before the callback, and guarded by a test that names
the file and line in a second instead of hanging for 25 minutes.

MinAgent: 0.131.0 (unchanged). No hub release needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 14:19:31 +02:00
admin 7c05b59708 v0.253.0 — errors carry the key of the sentence they are (R-557 slice 2 release B)
gates / gates (push) Successful in 24s
179 Hungarian sentences were built deep inside a package with fmt.Errorf and printed by
whoever caught them: too late to translate where they are shown, too early where they are
made. Every one now carries its key across that gap. ZERO Hungarian error literals remain.

util.MsgError does three things at once, each earned:
  - Error() is the Hungarian, byte for byte, so every un-converted printer is unchanged;
  - errors.Is answers for the kind AND for a wrapped cause (KindErrorf dropped the cause);
  - an error ARGUMENT renders recursively, so "formázás sikertelen: %w" translates whole.
A foreign error — restic, docker, ssh, the stdlib — prints verbatim. It is not ours.

76 display sites go through errText, and TestNoErrErrorInPageOutput convicts any that do
not. memoryVerdict returns an error rather than a sentence, so the deploy's 409 and the
household's language come from one value; UpdateRefusal gained a Cause to carry it.

Plurals, one rule, stated once: a key with .one/.other takes its COUNT first. Not a
per-call-site flag — the producer somebody forgot would read "3 app is not running". The
guard caught a real key collision (alert.deadapp.one) the day the rule landed.

TWO DEFECTS FOUND IN MY OWN TOOLING, recorded rather than quietly fixed. The bulk converter
silently dropped multi-line concatenations, damaging 7 producers — and the parity gate could
not see it, because every surviving fragment WAS a real base literal while the CALL had lost
text; two behaviour tests caught it. And the counting script was case-sensitive, so it said
"0 left" while five remained.

MinAgent: 0.131.0 (unchanged). No hub release needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 11:44:30 +02:00
admin 5270bad76e v0.252.0 — the sentences the program builds follow the language (R-557 slice 2 release A)
gates / gates (push) Successful in 23s
Slice 1 translated the dashboard's markup. The sentences the program BUILDS were still
Hungarian literals in Go, so an English household clicked an English button and was
answered in Hungarian. 226 of them move into the bundle here.

A flash was the hard part: it travels inside the redirect URL and is rendered by a
DIFFERENT request, so it now carries a bundle key plus its parameters. A link minted by
an older controller carries prose and is shown verbatim — never a raw key, never dropped.

Also converted: page data and view-model text, the internal/api JSON answers, the alert
banners (Alert.MessageKey, rendered on the way out of GetAlerts), 237 country names at
display, and the four page titles built around an app name (R-566 closed).

Hungarian is byte-identical, and that is measured rather than read:
scripts/i18n_go_parity.py freezes every Go literal at the base commit (7 467) and refuses
a key whose Hungarian is not that text, byte for byte. Three decoys, each seen to convict.
Its own first version filtered the capture through an ASCII-Hungarian word list and missed
seven real literals — the R-565 class. The filter is gone.

Nothing on the wire moved, and wire goldens now hold it there: the report's health
warnings and every notify event message stay Hungarian, because the hub MAILS the
controller's sentence when it has no entry of its own. Slice 3 (R-558) owns those.

MinAgent: 0.131.0 (unchanged). No hub release needed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 10:15:56 +02:00
admin f806baf4c7 R-563: the remote-backup page polls from a status attribute, and „Fut…" is finally translatable
- backups_remote.html: the status stat carries data-status="{{.Offbox.LastStatus}}"; the poll reads
  v.dataset.status === 'running' instead of v.textContent.indexOf('Fut'). The English page could never
  see a running backup before this (slice 1 shipped the English page with that word left Hungarian).
- The running label becomes {{T "backups_remote.fut"}} — hu „Fut…" unchanged, en „Running…".
- 12 backups_remote parity fixtures re-captured. Proven (audit switch/fixture diff): each differs from
  its predecessor ONLY by the data-status attribute and that one poll line; the other 94 are
  byte-identical. The displayed Hungarian is unchanged.
- Tests: TestR563_PollStartsFromAttribute (hu and en), TestR563_AttributeFollowsTheStatus.
  Red-proofed twice: poll back on textContent → both languages fail; word back inline → the English
  page shows Hungarian.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 21:02:47 +02:00
admin c00fed6db3 R-553: four decisions stop reading their own Hungarian words (sites 1-4)
Every Hungarian sentence is byte-identical; each decision now reads a signal set where the message is
made. util.KindErrorf builds the same bytes fmt.Errorf did while carrying a sentinel for errors.Is.

- Deploy status (api/router.go): deployStatusFor() by kind — stacks.ErrAlreadyDeployed (409),
  ErrRequiredField / ErrPathMissing / ErrNotEnoughMemory (400). The „kötelező" / „memória" /
  "does not exist" / "already deployed" text chain is gone.
- Off-site failure class (backup/offbox.go): ErrOffsiteQuota replaces the „tárhelykeretet" match. The
  restic/ssh signatures stay text matches on purpose — that output is not ours and is not translated.
- Alert placement (web/alerts.go): monitor.HealthReport carries WarningKinds parallel to Warnings;
  the "not on a separate drive" warning is inline by KIND. The hub report is untouched (builder.go
  copies Status/Issues/Warnings only) — pinned by a wire test.
- Stale off-site note (web/handlers.go): settings LastWarningKind + backup.OffboxWarnNoAppsSelected.
  The text test survives ONLY for kind == "" (a box whose last run predates 0.251.0) and is removed
  when R-570 closes; slice 2 must not translate that producer before then.

Tests (all red-proofed by restoring the pre-fix predicate — see the audit's redproofs.txt):
TestR553_Deploy_DecisionSurvivesWordingChange, TestR553_DeployHandlerUsesTheKind,
TestR553_DeployProducersCarryKindAndKeepTheirWords (through the real DeployStack),
TestR553_OffsiteQuota_{Decision,HeadLine}SurvivesWordingChange, TestR553_OffboxRunRecordsTheKind,
TestR553_StorageWarningsCarryKindsAndKeepTheirWords, TestR553_DiskWarningPlacementSurvivesWordingChange,
TestR553_HubReportWarningsAreUnchangedOnTheWire, TestR553_StaleNote*, TestR553_WarningKindIsPersistedAndCopied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 20:58:45 +02:00
admin be2efe5624 v0.250.0: i18n slice 1 release C — the language switch on every dashboard page (R-556)
gates / gates (push) Successful in 20s
- addLanguageData: the v0.247.0 hide condition (switch only on non-Hungarian pages or with ?lang=) is
  deleted; every template is converted, so the offer no longer leads to a half-English page.
- 89 layout parity fixtures re-captured; each equals its predecessor plus exactly one switch form after
  the version span (checked byte-for-byte, 89 of 89). The 17 standalone fixtures are unchanged.
- TestLanguageSwitch_EndToEnd: a Hungarian household with no ?lang= sees the form (red-proofed by
  restoring the condition: this test and 89 parity cases fail).
- CHANGELOG v0.250.0, README §18, CONTEXT.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 18:43:01 +02:00
admin ac149c2aa4 i18n slice 1 release C: storage, sharing, sign-in, guest, catch-all and debug pages in English
- 11 templates converted (storage, storage_network, storage_init, storage_attach, sharing, login,
  claim, launcher_shared, launcher_share_password, catchall, debug); every key translated.
  Six ASCII-only Hungarian JS fragments the extractor missed were found by eye and converted by hand
  (", majd a(z)", "FIGYELEM:", "jelenlegi:", "mp", "p", " db").
- login, claim, both guest share pages and the catch-all render through executeTemplateLang (the
  household language; no session CSRF, no escrow reminder). renderLogin now takes the request.
- Page titles: TitleKey for storage, network storage, the two drive wizards and sharing.
  TestHandlerTitleKeysMatchHungarianTitle pins handler literal == hu.json value for every TitleKey.
- Tests: TestDirectRenderHandlersFollowLanguage (real routes, en + hu),
  TestI18nDirectRenderPagesHaveNoAdminChrome (escrow reminder due; none on the guest page).
- Fixture correction: the release C cases had invented page titles; the cases now carry the handlers'
  real titles (and the wizard pages their real page name), and those 12 fixtures were RE-CAPTURED from
  the unconverted templates at 0555091. The diff to the old fixtures is the <title> line (and the nav
  highlight on the two wizard pages) only.
- i18n gate: HU_FORMAL_CEILING 12 -> 16 — four more „ön" forms in the converted copy, unchanged by rule (R-516).

Hungarian: byte-identical (TestI18nParity against fixtures from unconverted templates).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 18:29:53 +02:00
admin 0555091f83 i18n slice 1 release C: parity cases + fixtures captured from UNCONVERTED templates
storage, storage_network, storage_init, storage_attach, sharing, login, claim,
launcher_shared, launcher_share_password, catchall, debug. Case completeness checked with the
coverage probe on a trial conversion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 18:07:37 +02:00
admin 11b790e314 v0.249.0: i18n slice 1 release B — backup pages in English, Hungarian byte-identical (R-556)
gates / gates (push) Successful in 20s
Seven backup pages + the restore-progress JS converted against fixtures captured unconverted
(f8ebc47); recovery renders through executeTemplateLang; English retrieval claims registered;
the Fut comparison left unconverted (R-563).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 18:01:09 +02:00
admin f8ebc47728 i18n slice 1 release B: parity cases + fixtures captured from UNCONVERTED backup templates
backups_apps, backups_remote, backups_restore, backups_restore_wizard, backups_escrow,
tier2_config, recovery. Case completeness checked with the coverage probe on a trial conversion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 17:43:33 +02:00
admin ff68b0b433 v0.248.0: i18n slice 1 release A — apps and settings in English, Hungarian byte-identical (R-556)
gates / gates (push) Successful in 20s
Ten pages converted against fixtures captured from unconverted templates (6be55a0); marker
coverage test; English page mask narrowed; English retrieval-promise stems; common keys;
title keys and infraMeta from the bundle; old wording tests read the Hungarian expansion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 17:36:08 +02:00
admin 6be55a0eb3 i18n slice 1 release A: parity cases + fixtures captured from UNCONVERTED templates (25 states, 10 pages)
dashboard, stacks, logs, monitoring, app_import, app_export, deploy, settings_system,
settings_security, settings_notifications. Captured before any of these templates carries a
marker; case completeness was checked with the coverage probe on a trial conversion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 17:19:51 +02:00
admin 1af89a54af i18n slice 1 (A.2-A.4, A.6): narrowed English mask, marker coverage test, English retrieval stems, parse timing
Five v0.247.0 fixtures re-captured from UNCONVERTED v0.246.0 (89dd3e94) because their
one-word Hungarian data values had to become multi-word; two new cases (backups_tier_due,
app_info_installable) captured the same way — the coverage test found both gaps.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 17:15:23 +02:00
admin 612c417024 v0.247.0: i18n spike — the dashboard can speak English, Hungarian byte-identical
gates / gates (push) Successful in 19s
Message bundles (internal/i18n) expanded into templates before parsing, one
template set per language. Launcher, /backups, /apps/<slug> and the layout
converted; household language setting, POST /settings/language, ?lang= override,
report field. Parity test against fixtures captured from unconverted templates;
copy gates read templates expanded; new i18n_missing_gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 14:50:57 +02:00
admin 0fe315b759 v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s
MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted.

R-550 (operator ruling: fix). A design reversed and recorded: the restore
op-status was in memory by choice. Now restore-status.json in DataDir, written
atomically at both ends of an op. At startup a record still marked running
becomes a failed, interrupted result kept per app until that app's next
restore, shown on /backups/restore and the off-site wizard, and raised once as
restore_interrupted. Cooldowns stay in memory.

R-546. The R-543 reminder bar consults the agent's own preflight ok (every
blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while
paused. /backup/escrow shows a waiting card that polls and reloads instead of
red crosses and English diagnostics. POST /api/escrow/start refuses 409 before
staging or starting - the direct path chaos night used. Unknown readiness keeps
the bar.

Red-proofs (each seen failing): restore record across restart; main() calls
both startup functions; startup helper with loading skipped; restore page card;
bar held back; waiting card; start refusal. go build/vet/test ./... green, 28
packages; controller_gates --fast all OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:45:29 +02:00
admin ad398b60d9 v0.245.0 — R-543: the household is asked for the recovery code, the page says "szunetel" until then
gates / gates (push) Successful in 14s
Off-site backup is ON by default and does not RUN until the household creates its
recovery code. The pause is the zero-knowledge escrow design and is untouched here;
what was missing is that nothing ASKED, while the app-backup page promised the very
copy that had never run.

- a reminder bar on every authenticated page while the off-site tier is configured
  and its escrow is not complete, linking /backup/escrow. It is the R-241 bar, second
  instance: same session-cookie dismissal, back next visit, gone for good when
  escrowed. No second banner system. It hangs off executeTemplate, the single render
  choke point, so it cannot reach only the pages someone remembered.
- the tier-1 file sentence renders by tier3State's own vocabulary instead of the
  app's shape: active -> "vedi", escrow_pending -> "vedene ... szunetel" + the route,
  no copy at all -> says so and names both ways out.
- both fixes red-proofed: the bar test fails on BOTH pages with the hook removed; the
  sentence test quotes the exact v0.244.0 promise when the state is ignored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 21:02:51 +02:00
admin 2f8ff2414c v0.244.0: the backup page stops promising what it does not hold (R-537/R-538/R-536)
gates / gates (push) Successful in 17s
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.

R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.

R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.

Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:55:55 +02:00
admin d3eacbb7cc backup tile: unknown size shows a dash, not 0 B (R-517 follow-up, measured on 9201)
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:53:20 +02:00
admin 843b319f35 v0.243.0: FileBrowser generated admin password (R-513); per-tier whole-guest backup truth (R-517); skip absent-storage tiers (R-518); OOM-killed worker visible (R-514)
gates / gates (push) Successful in 14s
MinAgent: 0.131.0

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:12:08 +02:00
admin d698ce343b controller v0.242.0: a removed app is listed with its kept backup; five small ones (R-487 R-491 R-490 R-489 R-476 R-456)
gates / gates (push) Successful in 14s
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
2026-09-13 22:50:18 +02:00
admin 3e813307cc controller v0.241.0: a bind-data app leans on off-site before its own unit; the hold names what the copy holds (R-479)
gates / gates (push) Successful in 13s
Operator ruling 2026-09-13. An app with classified binds walks second
drive -> off-site -> own unit (its unit holds no files); volume apps keep
2 -> 1 -> 3. RestoreHold.CopyHolds records what the chosen copy holds and
the sentence ends with it; older holds keep their tier-only sentence.
Tests on both halves; red-proof: a layout-blind order fails the bind case.
2026-09-13 21:47:33 +02:00
admin bdcbd50b42 controller v0.240.0: seven defects from the any-tier proof and the first nightly rotation
gates / gates (push) Successful in 13s
R-486 (P1): removing an app with its backups KEPT keeps its Tier-2 record,
so the second-drive restore is no longer refused over an intact mirror.
R-484: postgis/pgvector/timescaledb images are Postgres (logical dumps).
R-485: the backup card sizes the recovery unit and the mirror(s).
R-480: a held update's sentence leaves the card once the hold is lifted.
R-477: the update's off-site lookup is one snapshots call, no stats.
R-478: a copy older than this install's deploy does not count.
R-474: "delete backups" deletes the unit, the mirror(s) and the prefs.

Tests and red-proofs per row; evidence in felhom.eu
documentation/audits/v0240-2026-09-13/ and nightly-2026-09-13-adventurelog/.
2026-09-13 19:26:50 +02:00
admin b93c1543da controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.

Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 17:16:24 +02:00
admin cbcca03061 v0.238.1: the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (slice 4 follow-up)
gates / gates (push) Successful in 13s
Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not
yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition
(alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names
survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored
the broken definition over the copy the customer is told to restore from.

backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now
also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to
stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:25:09 +02:00
admin 129201abab v0.238.0: the page follows the update, and a held app offers no way to start it (update arc slice 4 Part 4)
gates / gates (push) Successful in 13s
No behaviour change on the box — the surface only.

- Frissítés follows the job: the button shows the phase label (polling GET /api/stacks/{name}
  every 3 s) and the page reloads when updating goes false.
- An updating card offers no lifecycle button; a held card (failed update OR failed restore) shows
  the hold sentence with a Mentések link and nothing that would start it; a failed update that held
  nothing shows its sentence above the buttons. app_info shows the same three notices.
- The updating/held checks run BEFORE isOperational, which counts `restarting` as operational — how
  the 2026-09-01 spike saw a green Frissítés beside a crash loop. Pinned with StateRestarting
  fixtures; red-proofed by moving the checks after it (both tests fail).
- No new CSS, no version number.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:07:35 +02:00
admin 0d402f711d v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 11:41:31 +02:00
admin 42a73e667a v0.236.0: "delete my data too" deletes the data, or says that it could not (R-442)
gates / gates (push) Successful in 13s
Removal resolves the drive from the app's own app.yaml HDD_PATH (the 07 ~L437
rule), never the global cfg.Paths.HDDPath which no box sets. A data removal
that cannot be resolved, or whose drive is absent, is refused with a typed
RemoveRefusedError -> 409 + exact Hungarian sentence, before compose down, and
the app is kept. SSD app -> hdd_paths_removed: [] never null; missing folders
stated; backup-path refusals reach the response.

15 tests, two red-proofs run (pre-fix fallback -> C fails with err=nil and the
handler 200s; "no drive refuses" -> D fails).

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 08:52:25 +02:00
admin 2a56f557d0 v0.235.0: a delivered fix must also refresh the stored definition
gates / gates (push) Successful in 14s
Found by the LIVE validation on demo-hp, not by review. Scenario A passed - a
non-image catalog change reached the pinned app on the real 15-minute cycle - and
that is exactly what exposed the gap: the stored applied-compose.yml is written
when the PIN is written, so the fix landed in the live compose file and not in the
store. The first time the catalog then moved a version, the freeze would have
rendered the pre-fix definition and reverted every fix delivered since - silently
undoing the half of the operator's ruling that says fixes keep flowing.

The equal-images branch now refreshes the store as it delivers. The images cannot
move in that branch by construction, so no version moves and no intent is
rewritten. RenderPlan gains StackDir so the syncer can write it.

TestFixRefreshesTheStoredDefinition asserts both halves: the fix reaches the
store, and it survives the freeze that follows.
2026-09-06 10:04:21 +02:00
admin 8a0e0a59ad v0.235.0: freeze the version, keep the fixes flowing (operator ruling 2026-09-06)
gates / gates (push) Successful in 12s
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.

The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.

NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.

app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.

The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.

AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.

THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.

+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
2026-09-06 09:45:34 +02:00
admin 38d28b5b62 v0.234.0: seed installed_images at startup, so the label appears on an app nobody touched
gates / gates (push) Successful in 13s
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.

BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.

And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.

Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.

+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
2026-09-03 11:56:43 +02:00
admin 8025304acc v0.233.0: record what each compose service actually installed, and badge whether it is current
gates / gates (push) Successful in 12s
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no
auto-update, the three lifecycle buttons byte-identical.

Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each
entry carrying ref + repo digest + first-seen timestamp. Written by
Manager.recordInstalledImages after a successful compose up from StartStack,
RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never
from docker-compose.yml: the syncer overwrites a deployed app's compose on a
15-minute cycle and the two disagreed for 25 minutes in the spike's own
measurement. A failed write NEVER refuses the action - the deliberate opposite
of SetDesiredState, because this is an observation and that is an intent. Not
called from StartStackServices (the R-47 DB-only window). Its own docker seam
with a context and a 30s timeout, which neither existing exec helper has.

Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the
recorded ref per service against what the current template pins and returns a
*MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS.
NO RECORD RENDERS NOTHING: absent means unknown and never means current. No
version number reaches the customer and no registry is queried.

Known limitation, filed not hidden: 23 catalog pins float, so those apps can read
Naprakesz when the image behind the tag has moved.

+17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real
RestartStack plus an AST walk of the four call sites. Three companion red-proofs
run and reverted.
2026-09-02 20:18:01 +02:00
admin 8b55de734c R-414: the fallback scratch must also be DELETABLE - caught by live validation
gates / gates (push) Successful in 12s
The system-data fallback resolved a scratch fine and removeProofScratch then refused to
delete it: its accepted-roots list is built from REGISTERED drives, and a driveless box has
none. Observed on demo-felhom: 'refusing to remove ... it is not inside a proof root', with
the copy still on disk. Every nightly proof would have left one behind, growing forever, on
exactly the boxes the fallback exists for.

My defect, introduced with the fallback in the same session. The unit tests missed it
because every one of them registers a drive; the new pair deliberately does not, and the
second asserts the guard still REFUSES a path outside every proof root, so the fix is not a
widening into uselessness.
2026-09-01 10:29:08 +02:00
admin fcef8e069c one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:

  RestoreOffboxScratch      - the reported one
  OffboxRestorePrepareFull  - the SECOND request in the customer's own two-step full-restore
                              flow, and the one that actually shells `restic stats`. The UI
                              reaches it FIRST, so flagging only the restore would have left
                              the collision reachable by the ordinary path.
  RestoreSharesScratch      - R-411's exact shape on the shares tier: unlockStale + resticStep,
                              a live web caller, and its sibling PlaceSharesRestore has always
                              taken the flag.
  RestoreOffbox             - no production caller today, but the same dangerous pattern.
                              Flagged rather than left for a future caller to inherit.

OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.

THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.

R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.

R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.

The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.

And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.

R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.

16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
2026-09-01 10:19:36 +02:00
admin 303129e3af v0.231.0: the off-site proof gets a by-hand trigger, like its integrity sibling (R-87)
gates / gates (push) Successful in 12s
Without it the only way to see the job work is to wait for 05:30, which makes live
validation and any future diagnosis a next-day exercise. Same function as the scheduled
job - no second code path.

ONE deliberate difference from the integrity button: due-ness is NOT bypassed. There,
forcing means "check the store again", which is always answerable. Here due-ness IS the
target selection - an app is due when its newest snapshot has not been proved - so
ignoring it would mean inventing a second way to choose an app, exactly what having one
function prevents. When nothing is due the button says so, honestly.

Every other guard intact, including the single-writer flag: a hand-run during a backup
SKIPS exactly as the scheduled one would.

POST /api/debug/backup/offsite-proof, button beside "Restic integritas" on the debug page.
debug_route_gate pairs the two, so a button with no dispatch (R-400's shape) cannot ship.
2026-08-31 21:09:09 +02:00
admin e43b5ec07d v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
2026-08-31 20:55:34 +02:00
admin b48a7fa326 R-403: the restore OUTCOME names the package's date, not the run's
gates / gates (push) Successful in 11s
Live on demo-hp the confirm said 11:43 (the preserved package) and the outcome said 14:23 (the
copy's newest run) for the same restore. A customer reading both cannot tell which one they had, and
one of the two is the flattering sentence. Part 2.3's rule is 'not a plain green success ANYWHERE',
and the outcome is an anywhere.

tier2UnitSourceMsg now asks UnitRestoreDate, the same resolver the confirm uses, so the two cannot
disagree. TestR403_OutcomeNamesThePackageDateNotTheRunDate pins it, with a negative control for the
ordinary case.
2026-08-31 14:32:29 +02:00
admin 5429d651ee R-403 fix, caught by the live run: the stale flag fired for every healthy app
gates / gates (push) Successful in 11s
UnitRestoreDate also compared the package's date against the run's and flagged 'older'. A recovery
unit is ALWAYS captured shortly before the run that mirrors it, so that comparison is true for every
healthy app. Measured on demo-hp: bookstack, kimai, opengist and privatebin all had src and dest
manifests at 12:03:49Z against a run at 12:14:24Z - perfectly healthy, and all four would have been
told their package was stale.

A warning that fires on everything is a warning nobody reads, which costs the same as the comforting
lie it was meant to replace. The second return is now UnitLegPreserved and nothing else.

TestR403_AHealthyAppIsNeverCalledStale pins it; red-proof: reinstate the comparison -> it fails.
2026-08-31 14:22:32 +02:00
admin 2358e561b7 R-403: a poorer copy must never delete a richer one
gates / gates (push) Successful in 11s
MEASURED FIRST, then fixed. On the shipped v0.229.0, on demo-hp, an app's Tier-2 copy went from
120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either) in ONE nightly
run, and the run recorded itself a success: 'Tier 2 copied docmost -> ... (14.9 KB, 0 leg(s), 0s)'.
Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/.

The mechanism was three individually-correct lines: RunTier2 guards the unit leg with os.Stat only
(does the folder exist), rsyncMirror is rsync -a --delete, and nothing between them compared source
to destination. An EMPTY unit is a folder that exists.

THE GUARD. One predicate, unitCarriesData/unitIsHollow (r403_hollow.go), asking the MANIFEST and
never the byte size - a big compose tree with no dumps is dangerous, a tiny unit for a tiny app is
fine. Fail closed on an absent or unparseable manifest. RunTier2 skips the unit leg when the source
is hollow AND the destination is not; the other legs still run, the run is not failed, and the skip
is recorded for the SURFACE (CrossDriveBackup.UnitLegSkipped + UnitPackageDate) as well as logged.

--delete STAYS and shrinking stays legal. 07 section 8 row 5's derived-copy rule is unchanged; the
fence is exactly one shape. TestR403_DataLegShrinkIsUnaffected is the guard on the guard.

THE HONESTY. A preserved package is older than the run that preserved it, so the card carries a
notice and the unit-restore confirm names the PACKAGE's date - read from the mirrored manifest's own
created_at, not from the status record - plus a clause saying why it is older.

THE CAUSE. RestoreTier2Unit now refills a hollow or absent primary unit from the mirror it just
restored from, INSIDE the call before returning. The hollow manifest was written two seconds after
a restore by the 5-minute capture job; any follow-up job races it. The capture itself is NOT guarded:
a capture describing an empty drive as empty is correct, and with the primary refilled there is no
hollow state left to describe. Never over a complete primary, never after a failed restore.

recordTier2Success and tier2UnitConfirmMsg keep their old signatures as thin callers, so no existing
test needed editing. New seam unitRehydrate, separate from tier2Mirror on purpose.

22 new Go tests. Red-proofs run and reverted: A6 (predicate -> size threshold), B1 (guard removed ->
the copy's 3 files are DELETED and the seam is called), B6 (a general never-shrink rule -> the shrink
case fails), C2 (only-when-hollow dropped -> the complete primary is overwritten).
2026-08-31 14:02:13 +02:00
admin 4c8f0d2919 R-103: the Tier-2 refusal becomes an action
gates / gates (push) Successful in 12s
An app whose Tier-2 copy holds no file legs but a full recovery-unit mirror - 45 of the 53 catalog
templates - was told to press a button on a DIFFERENT page. Since R-102 the data it is asking for is
restorable from the copy it is looking at.

New POST /backup/tier2/unit-restore and backupTier2UnitRestoreHandler: same guards, same
restoreOpBlocked() refusal (R-351b), same async shape as the file restore beside it, plus a
fail-closed pre-flight so the app is never stopped for a mirror that could not be opened. The
outcome reuses unitRestoreOutcomeMsg and adds which copy overwrote the live data.

The row offers the action where the refusal was, in a danger style, as a SEPARATE button. The two
are not merged: one adds what is missing, the other overwrites. The confirm carries that difference
in words and names the copy's date - and says so differently when that date is only an ATTEMPT
(R-101). It is built from named Go constants rather than assembled inside an HTML attribute, so a
test can assert it verbatim; fmtTimeStr now delegates to a package-level fmtRFC3339Local so the
confirm and the outcome cannot render the same date two ways.

tier2NoCoverageMsg is NARROWED to the case that remains - no legs and no openable unit - and still
names the route that works. tier2UnitNotCoveredMsg is NOT deleted: it is appended where the FILE
restore ran and is still exactly true of it.

Tests C1-C2 and D1-D6 plus four more. Red-proofs: C1 (widen CanRestore to include HasUnit -> the
unit-only cases fail), D6 (drop EndRestoreOp from the handler goroutine -> 'the restore never
published a result').
2026-08-31 11:42:03 +02:00
admin 0f9b796615 R-102: the recovery unit on the second drive becomes a way back
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on
every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the
primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any
action in the product (07-backup-architecture 6.3, 7.2).

Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is
the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the
DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all
unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the
fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched.
reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now
named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second
one invented, which is what lets the acceptance test assert the volume leg's source directory.

Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a
parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken
inside RestoreFromRecoveryUnitAt, not beside it.

Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it
still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which
refused 40 running apps for months.

Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with
the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'.
Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the
primary -> fails with the mirror never reaching the redeploy, and with permission denied once the
primary tree is unreadable).
2026-08-31 11:30:34 +02:00
admin c732006d26 R-102 phase 1.1: split the recovery-unit path helpers, zero behaviour change
Every unit path helper took (nsRoot, stackName) and joined backups/primary/<stack>/... . That
hard-coded 'primary' is the mechanism of R-102: Tier-2 mirrors the whole unit directory to
<dest>/backups/secondary/<stack>/recovery-unit/ every night, and because no reader could NAME a
unit outside backups/primary/, that mirror has been captured for months and read by nothing.

Adds four unit-directory-relative primitives - UnitComposeDir, UnitManifestFile, UnitDBDumpDir,
UnitVolumeDumpDir - each taking the recovery-unit DIRECTORY itself. The four existing
(nsRoot, stackName) helpers become thin wrappers over them and keep their exact signatures and
their exact return values; every current caller compiles untouched.

ONE implementation, two callers - the rule restoreDockerVolumesFrom already follows in this repo.

TestR102_PathWrappersAreByteIdenticalToToday pins the wrappers against hand-written literals (not
re-derived from the helpers under test). Red-proof: UnitComposeDir join changed to 'compose2' ->
the test fails on all three fixtures.
2026-08-31 11:16:02 +02:00
admin 3c49dc8ea4 v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").

R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.

Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
2026-08-31 10:24:29 +02:00
admin 45770f2282 v0.227.1: the damage classifier matched restic's ordinary progress output
gates / gates (push) Successful in 11s
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and
re-pushing changed bytes under a live tag is the :latest hazard with extra steps.

looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A
HEALTHY restic check prints "check all packs" and "check snapshots, trees and
blobs" -- so any check that failed for a NON-damage reason, a connection dropped
mid-run for instance, would have been classified as a corrupted repository and
told the customer their backups may be damaged. That is the false alarm that
teaches an operator to ignore the true one.

Caught by the NEGATIVE control, built from the real bytes of a real passing
check on demo-hp. The spec made the negative control mandatory and this is what
it was for: a control that has only ever seen the failing case proves nothing.

Signatures are now phrases from restic's own error wording.

Also in this commit: CONTEXT.md records the three rulings (take the flag and
skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead
fields) plus the measurement a future session would otherwise assume wrongly --
THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job,
the route and the config, and corrects a line that listed four debug backup
routes when only two exist. REUSE gains three rows, including one that records
R-398 was my own mistake so nobody re-files it.
2026-08-30 21:22:23 +02:00
admin 0d52a42c17 R-359 + R-397: the off-site store gets checked, and the advertised check becomes real
gates / gates (push) Successful in 12s
Nothing ever verified that the off-site copies are still readable. The
whole-guest tier has verify jobs; the tier holding the customer's documents and
photos had none -- the complete set of restic verbs this controller used
contained no `check`. We would have found out at restore time, with a customer
waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain
`restic check`; we had never run it.

R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub
allowlists both event types and carries the Hungarian text for both, the
settings checkbox exists, and the debug button posts to /api/debug/backup/
integrity. Everything was built except the part that runs. SIXTH instance of
that shape in this project.

THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by
running `unlock --remove-all` and retrying, and its own comment records why that
is safe: every caller holds the in-process single-flight mutex, so any lock it
meets is stale. A check that did not take that flag could meet a LIVE prune's
lock from this same box, remove it, and retry over the top of it. So the check
TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly
backup behind it, and a skip costs nothing because due-ness makes tomorrow try
again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never
invoked, `unlock` never in any argv. Its red-proof prints the real thing --
restic running `check` while the flag was held.

DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful
check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape,
a dated check quietly missed and never caught up. No Weekly primitive added.

THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts.
"I could not look" is not "I looked and it is broken" -- R-339 already owns
reachability, and a second alarm for the same fact trains the operator to
discount the one alarm that means the backups are damaged. A timeout is
unreachable, never damage. A failure advances due-ness (a broken store must not
be re-checked nightly); a skip and an unreachable store do not.

Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by
design. A weekly success e-mail is how people stop reading their alerts.

The customer gets a SENTENCE; restic's words go to the log, truncated (R-379:
615 bytes of raw database text reached a customer once). read-data-subset ships
OFF and a malformed value is refused at read time rather than handed to restic,
where one typo would fail the whole check.

Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK --
those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero
still passes unmodified.

Also: the monitoring page stopped promising a Sunday job that never existed, and
the debug button got its dispatch case.

PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for
already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable
since the off-site tier shipped, and other tests drive restic-backed paths
through it. A resticStepFn seam would have been WORSE here -- it would replace
the `unlock --remove-all` escalation and hide it from the assertions that must
see it. R-358's AST ordering test is converted to a real execution test instead,
which immediately surfaced something the AST walk could not: unlockStale
legitimately runs before the restore.

Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages,
rc 0. All 12 controller gates OK.
2026-08-30 21:03:29 +02:00
admin c0c8fe67bf An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero
result, which is honest" exposed that the sentence was FALSE.

A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's
fallback to RestoreApp -- which returns only an error, and whose signature is
deliberately out of scope -- would have printed "ez a mentes csak a
beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the
app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure
direction this whole change exists to remove, re-introduced by the change.

UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a
fourth sentence claiming only what is known -- the restore ran, the app is back,
and we cannot say what came back. RestoreApp's signature is untouched.

Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam
test was corrected too: its fixture has no recovery unit, so it exercises exactly
this path and had been asserting the wrong sentence -- it now asserts the
unknown, which is what pins the fallback to it.

IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate
written to stop findings dying in an overwritten REPORT.md caught a live defect
instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the
monitoring page advertises a weekly integrity check that does not exist) and
R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST
test) rather than leaving them in a file that is overwritten every session.

REPORT.md is the full run record: baselines re-confirmed, per-test results, the
five red-proofs with their observed output, the live validation with verbatim
Hungarian messages, what was NOT validated and why, teardown across three
layers, and the register 165 -> 167 -> 161.

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
2026-08-30 19:51:16 +02:00
admin b8af72764d R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21
backup-truth drill, all still in shipped code. They share one acceptance idea: a
restore surface must state what it actually did, and must refuse what it cannot
do.

VERSION NOTE. The task specifying this targeted v0.224.0 against baseline
f8c9390. Both were consumed earlier the same day by R-330 (0.224.0) and R-331
(0.225.0). Drift re-confirmed against live Gitea before the first edit, operator
authorised proceeding, every symbol the spec named re-verified present at the
real baseline e5eee50.

R-353 -- a restore that gave back nothing still said it worked.
RestoreFromRecoveryUnit returned only error, so the surface printed
"<app> visszaallitva (<snapshot>)." -- equally true of a run that returned an
entire dataset and one that returned nothing. The count already existed and was
discarded one line deep: restoreDockerVolumesFrom always returned it, the
wrapper threw it away. Now (UnitRestoreResult, error), carrying replayed counts
AND what the manifest LISTED, because zero-replayed has two causes that are
opposite news. Three cases, three sentences, and EVERY one is a claim about the
BACKUP, never about the app -- this path has no SafetyDump discriminator, and
07-backup-architecture 6.3 records that an absent dump says nothing about the
app (R-361 destroyed canonical .sql files for four months).

R-357 -- the destructive restore had no free-space gate. offbox_reconstitute.go
contained ZERO references to offboxFree; all three existing gates guard
non-destructive paths. The gate now sits before mapOffsiteRestorePaths,
writeSafetyDump and StopStack, so a refusal costs nothing. Position IS the fix,
which is why the test asserts StopStack was never called. No headroom multiplier
(matches PlaceOffsiteRestore; the x1.1 elsewhere predicts a download). Fail
closed on either probe <= 0 -- otherwise `free < need` with need==0 is FALSE and
an unmeasurable scratch sails through: a gate present and inert.

R-358 -- a failed download was offered as a good one. The gate answered "the
directory exists and is non-empty", which is exactly what a part-way restic run
leaves. Now a completion marker written 0600 atomically AFTER restic returns
nil, with any stale one cleared BEFORE it starts; both orders pinned by an AST
test because resticStep is not a seam. Both handlers refuse server-side: the
wizard flags control a button, and a hidden button is not a guard.

SCENARIO F ANSWERED, and worse than the question assumed: a unit-only scratch IS
reachable through the real flow, by the most ordinary route. "Ellenorzo
visszaallitas" (mode=unit, advertised non-destructive) writes the SAME directory
-- offboxRestoreScratchDir ignores `full` and --include limits what restic
extracts, never where -- so a customer who ran the SAFE restore was then offered
the destructive one over a unit-only copy. Filed R-396; the marker closes it.

R-360 -- the delete refused only while a BACKUP ran. IsRunning() is FALSE for the
whole of a verification restore; the five sibling handlers all use
restoreOpBlocked(). Its doc comment claimed it already did this, which is why
nobody looked -- corrected in place.

Red-proofs, each printing the pre-fix behaviour, in CHANGELOG and REPORT. The
first R-357 red-proof exposed a hollow test OF MY OWN and it is recorded rather
than quietly fixed: the fixture refused earlier at the placement stat pre-pass,
so `stops == 0` passed against the pre-fix code. Fixture corrected, assertions
reordered so a removed gate reports the outage rather than "no error returned".

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
2026-08-30 19:31:31 +02:00