Use my kept data / Load consider the off-site snapshot when it is newer than every local copy or
the only one; the unit is downloaded alone, judged (drive, data, recorded data version) and only
then restored. The page names the copy and its date.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Recounted at catalog 18a6d2d8: 66 unique pins — 48 full X.Y.Z, 6 two-part
lines, 4 major lines (10 float), 8 exact versions wearing a variant suffix.
The '23' carried since v0.233.0 matches no definition the catalog supports.
Definition written down beside the number so it can be rechecked.
Also: CONTEXT said the fleet floor was 0.257.0; the hub says 0.259.0.
Comment and doc only — no behaviour change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The composed-sentence-into-page-data class is now five instances and has its
own warning block in the README's i18n section, with the two things that make
a live probe lie: the felhom_lang cookie is ignored behind auth (so an English
probe of /backups returns Hungarian and reads as unfixed), and an apostrophe in
an English value is escaped so no assertion on the rendered page ever matches.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The release fixed four Go-side leftovers. The walk that judged it on a box installed
from scratch that day found two more of the SAME SHAPE — a composed sentence handed
to a renderer as page data — on the claim page (R-596, P1) and the Backup page
(R-598). That shape now has three instances on file with R-573, which is worth naming
as a class rather than fixing one at a time.
Verdict recorded: not yet ready for an English-speaking tester.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The operator asked for the floor at the end of the translation work. Raised the same
session with the declared MinAgent 0.131.0 — above the vouched golden, so the
declaration carries it. demo-felhom went 0.255.0 -> 0.257.0 by itself in under 12
seconds and THEN rendered the English tagline, which is the observable worth having:
the floor delivered the feature, not a version string.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
README gains "Catalog copy in a second language" under §18: the format, the
field-by-field fallback, whole-list replacement vs key matching, the no-write-through
rule, the four templates that carry copy, and why an older controller is unaffected.
REPORT records the nine red-proofs — including the two that convicted a hollow TEST
rather than the code — and the live proof on both demo boxes, naming the method
(endpoint level, no browser on DooPlex) rather than implying a click-through.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Two defects in v0.254.0's globe, both plain on a browser and neither catchable by anything
that existed — every test read the MARKUP, and the fault was in which CSS file the browser
fetched.
The shells requested /static/style.css with NO ?v=, while layout.html has carried one since
v0.166.0. A browser holding a copy from before v0.254.0 kept serving CSS with no .lang-globe
rules, so the globe came out as a bare unstyled <details> — a stray triangle and two plain
words at the edge of the window. It was FIVE shells, not the three named: both guest share
pages have the same fault for any CSS change, and their visitor is the likeliest of all to be
holding an old copy. And .Version was missing from three of those five data maps, which is
exactly how the next one would be forgotten — it is now filled at the one choke point every
shell renders through.
The globe also floated outside the card, pinned to the corner of the VIEWPORT, reading as part
of the browser rather than the page. It now sits inside the card, centred under the footer, with
the menu opening upward via the shared rule — so the dashboard and the shells cannot drift.
AND A THIRD, caught by a test that already existed: putting the version on the guest share pages
would have printed the controller build onto a page a stranger with a capability URL can open.
TestShareGuest_HeadersTilesNoAdminChrome refused it. Those two now take an opaque per-build tag
— same cache-busting, no disclosure. The fill is ONE function shared with the parity harness,
because a fixture rendered through a different data path is a picture of a page nobody serves,
which the previous release got wrong twice.
15 shell fixtures re-captured; 91 identical, every dashboard page among them.
MinAgent: 0.131.0 (unchanged). No hub release needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
0.253.0 -> 0.254.0 with min_agent 0.131.0 declared — the floor sits above the vouched golden,
so the declaration is what carries it (R-472), and the agent versions were checked first
because a floor is HELD for a box whose agent is below the requirement.
demo-felhom took it by itself in ~40 s. The box's own log is the observable that counts
("settle-gate: GO — at/above floor 0.254.0"), not the hub's claim, and the release's visible
change is there too: one globe on the sign-in page, zero of the old text links.
The two DOWN customers did not get it and will take it unattended when they return, from
0.115.0 and 0.245.0 — said plainly rather than left to be inferred from a green dashboard.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The notes a background run SAVES — last night's backup line, the last error, the proof
result, the restore outcome — are written in the BOX's language at the moment they are
written. A household that switches sees the previous run's note in the old language until
the next run rewrites it: the operator's §16 option 1, stated rather than hidden.
EndRestoreOp no longer receives a Hungarian literal from anywhere.
The language switch is a globe. Two text links wrapped in the sidebar footer and asked the
reader to recognise "Magyar"/"English" as links; a globe is the one symbol every web user
already reads as "language", so nobody has to read Hungarian to escape Hungarian. It is
<details>/<summary> — a menu with no script, drawn inline because the icon sprite lives
only in layout.html and the visitor pages have their own shell.
Those visitor pages get the same globe, and a visitor's choice stays theirs: a display-only
felhom_lang cookie that langFor reads ONLY when there is no session. A signed-in household
can never inherit a language a previous visitor picked in the same browser. POST /lang is
CSRF-exempt for a narrow reason written at the exemption — its only achievable effect is the
language of the page the victim's own browser shows them — and safeBackPath refuses
//evil.example as well as https://, because "starts with /" alone is not the test. §16 taken:
a successful claim carries the cookie into the household's setting.
TWO PARITY EXCEPTIONS, MEASURED: 106 fixtures compared with a real diff — exactly two change
shapes (the dashboard footer, the globe in the shells) and 5 byte-identical, which are the
three pages that must not change.
I INTRODUCED A DEADLOCK AND THE SUITE CAUGHT IT BY HANGING. UpdateOffboxStatus holds the
settings write lock while running its callback; boxLang() wants the read lock; sync.RWMutex
is not reentrant. On a real box an off-site run would have hung forever HOLDING the settings
lock. Fixed by resolving the language before the callback, and guarded by a test that names
the file and line in a second instead of hanging for 25 minutes.
MinAgent: 0.131.0 (unchanged). No hub release needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
min_controller_version 0.250.0 -> 0.253.0 with min_agent 0.131.0 declared — the floor sits
above the vouched golden 0.246.0, so the declaration is what carries it (R-472).
demo-felhom took it by itself in ~20 s and is healthy; its own log is the observable that
matters ("settle-gate: GO — at/above floor 0.253.0"), not the hub's claim alone. The two DOWN
customers did not get it and will take it unattended when they return, on versions nobody has
carried forward from — said plainly rather than left to be inferred from a green dashboard.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
179 Hungarian sentences were built deep inside a package with fmt.Errorf and printed by
whoever caught them: too late to translate where they are shown, too early where they are
made. Every one now carries its key across that gap. ZERO Hungarian error literals remain.
util.MsgError does three things at once, each earned:
- Error() is the Hungarian, byte for byte, so every un-converted printer is unchanged;
- errors.Is answers for the kind AND for a wrapped cause (KindErrorf dropped the cause);
- an error ARGUMENT renders recursively, so "formázás sikertelen: %w" translates whole.
A foreign error — restic, docker, ssh, the stdlib — prints verbatim. It is not ours.
76 display sites go through errText, and TestNoErrErrorInPageOutput convicts any that do
not. memoryVerdict returns an error rather than a sentence, so the deploy's 409 and the
household's language come from one value; UpdateRefusal gained a Cause to carry it.
Plurals, one rule, stated once: a key with .one/.other takes its COUNT first. Not a
per-call-site flag — the producer somebody forgot would read "3 app is not running". The
guard caught a real key collision (alert.deadapp.one) the day the rule landed.
TWO DEFECTS FOUND IN MY OWN TOOLING, recorded rather than quietly fixed. The bulk converter
silently dropped multi-line concatenations, damaging 7 producers — and the parity gate could
not see it, because every surviving fragment WAS a real base literal while the CALL had lost
text; two behaviour tests caught it. And the counting script was case-sensitive, so it said
"0 left" while five remained.
MinAgent: 0.131.0 (unchanged). No hub release needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Slice 1 translated the dashboard's markup. The sentences the program BUILDS were still
Hungarian literals in Go, so an English household clicked an English button and was
answered in Hungarian. 226 of them move into the bundle here.
A flash was the hard part: it travels inside the redirect URL and is rendered by a
DIFFERENT request, so it now carries a bundle key plus its parameters. A link minted by
an older controller carries prose and is shown verbatim — never a raw key, never dropped.
Also converted: page data and view-model text, the internal/api JSON answers, the alert
banners (Alert.MessageKey, rendered on the way out of GetAlerts), 237 country names at
display, and the four page titles built around an app name (R-566 closed).
Hungarian is byte-identical, and that is measured rather than read:
scripts/i18n_go_parity.py freezes every Go literal at the base commit (7 467) and refuses
a key whose Hungarian is not that text, byte for byte. Three decoys, each seen to convict.
Its own first version filtered the capture through an ASCII-Hungarian word list and missed
seven real literals — the R-565 class. The filter is gone.
Nothing on the wire moved, and wire goldens now hold it there: the report's health
warnings and every notify event message stay Hungarian, because the hub MAILS the
controller's sentence when it has no entry of its own. Slice 3 (R-558) owns those.
MinAgent: 0.131.0 (unchanged). No hub release needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- addLanguageData: the v0.247.0 hide condition (switch only on non-Hungarian pages or with ?lang=) is
deleted; every template is converted, so the offer no longer leads to a half-English page.
- 89 layout parity fixtures re-captured; each equals its predecessor plus exactly one switch form after
the version span (checked byte-for-byte, 89 of 89). The 17 standalone fixtures are unchanged.
- TestLanguageSwitch_EndToEnd: a Hungarian household with no ?lang= sees the form (red-proofed by
restoring the condition: this test and 89 parity cases fail).
- CHANGELOG v0.250.0, README §18, CONTEXT.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Seven backup pages + the restore-progress JS converted against fixtures captured unconverted
(f8ebc47); recovery renders through executeTemplateLang; English retrieval claims registered;
the Fut comparison left unconverted (R-563).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Message bundles (internal/i18n) expanded into templates before parsing, one
template set per language. Launcher, /backups, /apps/<slug> and the layout
converted; household language setting, POST /settings/language, ?lang= override,
report field. Parity test against fixtures captured from unconverted templates;
copy gates read templates expanded; new i18n_missing_gate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Removal resolves the drive from the app's own app.yaml HDD_PATH (the 07 ~L437
rule), never the global cfg.Paths.HDDPath which no box sets. A data removal
that cannot be resolved, or whose drive is absent, is refused with a typed
RemoveRefusedError -> 409 + exact Hungarian sentence, before compose down, and
the app is kept. SSD app -> hdd_paths_removed: [] never null; missing folders
stated; backup-path refusals reach the response.
15 tests, two red-proofs run (pre-fix fallback -> C fails with err=nil and the
handler 200s; "no drive refuses" -> D fails).
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.
The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.
NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.
app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.
The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.
AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.
THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.
+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.
BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.
And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.
Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.
+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no
auto-update, the three lifecycle buttons byte-identical.
Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each
entry carrying ref + repo digest + first-seen timestamp. Written by
Manager.recordInstalledImages after a successful compose up from StartStack,
RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never
from docker-compose.yml: the syncer overwrites a deployed app's compose on a
15-minute cycle and the two disagreed for 25 minutes in the spike's own
measurement. A failed write NEVER refuses the action - the deliberate opposite
of SetDesiredState, because this is an observation and that is an intent. Not
called from StartStackServices (the R-47 DB-only window). Its own docker seam
with a context and a 30s timeout, which neither existing exec helper has.
Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the
recorded ref per service against what the current template pins and returns a
*MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS.
NO RECORD RENDERS NOTHING: absent means unknown and never means current. No
version number reaches the customer and no registry is queried.
Known limitation, filed not hidden: 23 catalog pins float, so those apps can read
Naprakesz when the image behind the tag has moved.
+17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real
RestartStack plus an AST walk of the four call sites. Three companion red-proofs
run and reverted.
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.
THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.
IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.
THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".
THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.
THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".
IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.
IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.
DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.
ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.
SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.
NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.
33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.
A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
CHANGELOG v0.230.0, leading with the measurement rather than the fix: 120 082 104 B -> 7 036 B on
the shipped v0.229.0, reproduced before anything was built.
CONTEXT records three rulings: hollowness is a MANIFEST question and never a size question; the
guard fences one shape and NOT shrinking, because the derived-copy rebuild is a design decision; and
the rehydrate happens inside the restore because a follow-up job races the 5-minute capture. Plus
the shape the live run taught: a warning that fires on everything costs the same as the comforting
lie it replaces.
README documents the refusal, what each surface says, and why the capture job is deliberately not
guarded. REPORT leads with Part 1's result, carries the six red-proofs, the per-row Scenario D table
with its seven-app control, and eight observations including R-404 filed-not-acted-on and three
mistakes of mine recorded rather than tidied away.
CHANGELOG v0.229.0. CONTEXT records three rulings: the source moves and the destination does not;
two predicates and not one wider one (R-356's cost restated); and a destructive operation reached
from a non-destructive surface carries the difference in the CONFIRM, not the label. README documents
the new action and route and corrects the coverage note to the measured count. REUSE maps the four
unit-directory-relative primitives and the new manager methods, with the traps.
REPORT covers the live drill on demo-hp (docmost, class B, primary unit moved aside — 3 volumes of 3
and 1 database of 1 in 28.65 s, accented filename byte-identical verified as hex, the app reading its
own row over TCP; Scenario D with the guest app.yaml also aside, secrets recovered=2/2), the settled
count (A=7 B=45 C=1, and why the earlier 9/43/1 was wrong), the five named red-proofs, and seven
observations including R-403 and a process error of mine that changed the box and is now in memory.
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").
R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.
Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and
re-pushing changed bytes under a live tag is the :latest hazard with extra steps.
looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A
HEALTHY restic check prints "check all packs" and "check snapshots, trees and
blobs" -- so any check that failed for a NON-damage reason, a connection dropped
mid-run for instance, would have been classified as a corrupted repository and
told the customer their backups may be damaged. That is the false alarm that
teaches an operator to ignore the true one.
Caught by the NEGATIVE control, built from the real bytes of a real passing
check on demo-hp. The spec made the negative control mandatory and this is what
it was for: a control that has only ever seen the failing case proves nothing.
Signatures are now phrases from restic's own error wording.
Also in this commit: CONTEXT.md records the three rulings (take the flag and
skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead
fields) plus the measurement a future session would otherwise assume wrongly --
THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job,
the route and the config, and corrects a line that listed four debug backup
routes when only two exist. REUSE gains three rows, including one that records
R-398 was my own mistake so nobody re-files it.
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21
backup-truth drill, all still in shipped code. They share one acceptance idea: a
restore surface must state what it actually did, and must refuse what it cannot
do.
VERSION NOTE. The task specifying this targeted v0.224.0 against baseline
f8c9390. Both were consumed earlier the same day by R-330 (0.224.0) and R-331
(0.225.0). Drift re-confirmed against live Gitea before the first edit, operator
authorised proceeding, every symbol the spec named re-verified present at the
real baseline e5eee50.
R-353 -- a restore that gave back nothing still said it worked.
RestoreFromRecoveryUnit returned only error, so the surface printed
"<app> visszaallitva (<snapshot>)." -- equally true of a run that returned an
entire dataset and one that returned nothing. The count already existed and was
discarded one line deep: restoreDockerVolumesFrom always returned it, the
wrapper threw it away. Now (UnitRestoreResult, error), carrying replayed counts
AND what the manifest LISTED, because zero-replayed has two causes that are
opposite news. Three cases, three sentences, and EVERY one is a claim about the
BACKUP, never about the app -- this path has no SafetyDump discriminator, and
07-backup-architecture 6.3 records that an absent dump says nothing about the
app (R-361 destroyed canonical .sql files for four months).
R-357 -- the destructive restore had no free-space gate. offbox_reconstitute.go
contained ZERO references to offboxFree; all three existing gates guard
non-destructive paths. The gate now sits before mapOffsiteRestorePaths,
writeSafetyDump and StopStack, so a refusal costs nothing. Position IS the fix,
which is why the test asserts StopStack was never called. No headroom multiplier
(matches PlaceOffsiteRestore; the x1.1 elsewhere predicts a download). Fail
closed on either probe <= 0 -- otherwise `free < need` with need==0 is FALSE and
an unmeasurable scratch sails through: a gate present and inert.
R-358 -- a failed download was offered as a good one. The gate answered "the
directory exists and is non-empty", which is exactly what a part-way restic run
leaves. Now a completion marker written 0600 atomically AFTER restic returns
nil, with any stale one cleared BEFORE it starts; both orders pinned by an AST
test because resticStep is not a seam. Both handlers refuse server-side: the
wizard flags control a button, and a hidden button is not a guard.
SCENARIO F ANSWERED, and worse than the question assumed: a unit-only scratch IS
reachable through the real flow, by the most ordinary route. "Ellenorzo
visszaallitas" (mode=unit, advertised non-destructive) writes the SAME directory
-- offboxRestoreScratchDir ignores `full` and --include limits what restic
extracts, never where -- so a customer who ran the SAFE restore was then offered
the destructive one over a unit-only copy. Filed R-396; the marker closes it.
R-360 -- the delete refused only while a BACKUP ran. IsRunning() is FALSE for the
whole of a verification restore; the five sibling handlers all use
restoreOpBlocked(). Its doc comment claimed it already did this, which is why
nobody looked -- corrected in place.
Red-proofs, each printing the pre-fix behaviour, in CHANGELOG and REPORT. The
first R-357 red-proof exposed a hollow test OF MY OWN and it is recorded rather
than quietly fixed: the fixture refused earlier at the placement stat pre-pass,
so `stops == 0` passed against the pre-fix code. Fixture corrected, assertions
reordered so a removed gate reports the outage rather than "no error returned".
Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate
"warn" strings that are healthcheck statuses), the hub manifest's real location
since the task's premise was wrong, all five red-proofs with the layer each
guard sits at, the live walk in six steps with the hub's own records quoted, and
the absent-intent count (0 of 8).
Three things are reported that a tidier account would omit: red-proof 5 passed
first time because the mutation was INERT; Scenario G was silently refused twice
behind an HTTP 200; and the live Scenario A does NOT prove the customer gate,
because demo-hp has no prefs row at all.
CONTEXT records the severity vocabulary as a ruling with its mechanism, the
intent ruling with its three-way handling of unknown, both fences, and two traps
worth more than the fixes: a 200 can be a refusal, and a passing red-proof can
mean an inert mutation.
README: the event table said `app_start_failed | warn` - the defect, written
down as if correct. Now `warning`, with the vocabulary contract and who receives
what. `disk_critical` also corrected from `error` to `critical`, which is what
fillwatch has always sent.
REPORT.md overwritten with the full run: baselines and the hub's four numbers,
the four red-proofs with the mutation and observed text for each, the five
IsDownState consumers walked and named, the live walk in full with the old and
new heartbeat lines quoted side by side, and the halt.
CONTEXT records the decision - a dead supervised member is asked about before a
failing healthcheck, because they are different questions and the second was
answering the first - plus the fence that IsDownState did not move, the trap
that three existing subtests pinned the defect, and R-386.
README gains the `degraded` row, which the state table never had, and a note
that the ORDER is load-bearing. Points at the new alarm-ladder architecture doc.
Records the db_dumps decision with every consumer named, the trap that a stable
db_dumps lets CaptureRecoveryUnit's already-current early return fire (so
per-capture housekeeping must sit above it), and the NEGATIVE that a held app
does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives
it.
The README's reconstitution sequence gains the volume replay it never had, and the AppBackup row
states that DiscoverDatabases now prefers the compose project label. CONTEXT records the thing an
operator most needs next: R-354's fix cannot reach the 40 apps that need it most until R-356 is
closed, because the off-site restore still refuses outright for every app that declares no data
drive. The live confirmation was therefore done on calibre-web and paperless-ngx.
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.
PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.
PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.
PART 4 - measured before theorising, on the live off-site target:
snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
=> 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.
TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.
RED-PROOFS, mutation asserted applied then reverted to 0:
A three template guards dropped (count asserted 3) -> the blank form returned
P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential
DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.
NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
- CHANGELOG leads with the severity fix and the live warning-vs-warn proof.
- CONTEXT records the settled decisions so they are not re-litigated: Hiba is
the label for predicted failure (no fourth word); sustain before count and
why; the provenance of 64/55/60; phase 2 owns the new SMART attributes
because they are a wire change under G-1; phase 1 state is one record per
disk, not a series.
- README documents the 14-row ladder, the persisted state, the hourly cadence
and the five message shapes.
- REUSE pins the severity wire contract on PushEvent — the defect's real home,
so the next typo'd severity is caught at the table rather than in production
— and records priorFor vs cardPriorFor, which differ by one observation and
make the chip disagree with the email if mixed up.