internal/sockheal: 60 s of refusals (never a timeout, only after Docker answered once) → exit 75 so
Docker's restart policy brings the controller back on the current socket; every 5 min it restarts any
other socket user (traefik) holding an older inode. Measured on 9202: only a docker.socket restart
re-creates the file; dockerd crash / docker.service restart keep it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- internal/family: the family list (bcrypt, generated 4x4 passwords shown once) + 30-day sessions in family.json
(0600, atomic); a reset (generation), a removal or a logout ends sessions at the next request.
- internal/stacks/family_gate.go: family_gate / family_gate_except / min_controller in .felhom.yml; the door is written
BEFORE the first start (install and a removed app's restore), a life record in app.yaml, reconciled by the gate loop;
priority below the install hold, setup gate and sign-up block; every exception anchored ^/prefix(/|$) (finding F1).
- internal/web/family_gate.go: forwardAuth /__felhom_gate/family (app cookie felhom_famgate, host-only, names a store
session); /__family/start|login|logout on the dashboard host (session cookie felhom_family, Path=/__family);
sign-in counted per visitor (clientIP) AND per name, short windows; the household's dashboard session vouches.
RequireAuth never reads a family cookie. The "Család" card on the security page: add / new password / remove.
Red-proofs RP-F1..RP-F7 (felhom.eu audits/family-gate-2026-10-02/A/).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Image retention: after a done/undone guarded Update and at remove, an app's images older than its running
and previous one are deleted — never an image any container, installed compose or installed/previous record
names (box-wide keep set read at delete time); exact id, never forced or pruned; paused while any update runs;
a one-time sweep of catalog app images at the first start. Install hold: an after_install app is installed
behind the setup gate's door and opens when after_install succeeds or the household says it changed the login.
Tests TestImageRetention_* and TestInstallHold_* with red-proofs; parity fixture for the held card.
MinAgent: 0.131.0 (unchanged).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The unit's data files are stamped with the versions that wrote them; the capture keeps the
definition the data belongs to; a restore never starts data under another version's
definition (unit restores refuse a mismatch; the off-site restore writes the snapshot's
definition); every tier's time is its data's; the conversion-copy release needs a dump on
the new engine. File-browser sync single-flight + no empty kept folder (R-695); the kept
view joins the folder's owning group, language switch resyncs (R-691); a restore-generated
login is not shown as the password (R-694). Red-proofs in
felhom.eu/documentation/audits/version-travel-2026-09-26/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
An install over an app's kept drive folder (appdata/<app> non-empty) asks the household:
"use my kept data" (a load from the newest copy of THIS drive's install, own unit or
second-drive mirror, then the template's after_load) or "start fresh" (the folder is
renamed into <drive>/kept/<app>/<date>/ with the removed app's unit; nothing deleted).
The install API answers 409 kept_data_choice until one is chosen; DeployStack refuses
too. New page Megorzott adatok / Kept data (/kept-data): Load / Look / Delete (typed
confirmation, the only deletion of kept data). FileBrowser gets a read-only source.
The drive-full warning names the kept folders. <drive>/kept is protected and outside
every backup leg.
R-690: the removed-app restore (R-487) never found a unit on a DATA drive — it asked
GetStackComposePath (true for every catalog app) and restored nextcloud with no env.
Now isStackDeployed; pinned with a production-shaped provider.
Red-proofs: audits/night-2026-09-26/E/redproofs/.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A step whose ladder entry carries engine_conversion {service, engine, from, to}
converts the database: the old engine alone, the check (owners, roles,
extensions, per-table row counts), pg_dumpall validated by its completion line,
the volume emptied only after the undo copy's marker is validated again, the new
engine alone, the load with ON_ERROR_STOP, the check again + PG_VERSION. Any
failure goes to the existing undo; a restart during converting is undone.
A PostgreSQL major move without the mark is refused before anything moves.
The old datadir's copy is kept until a backup is proven after the conversion.
17 tests, 9 red-proofs (audits/night-2026-09-26/B/).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-650: internal/dockerexec — every docker exec routed through it; under
go test a real docker is refused (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub
under the temp dir is allowed). api/stacks/web tests run under a silent
stub (TestMain). TestR650_NoBareDockerExec pins it repo-wide.
R-640: a dump without its engine's completion marker is refused before
the first mutation (unit + off-site restore) and again before any load.
R-499: the Tier-2 page's system-disk sentence has four true branches.
R-518: the backup button states the measured ~8 min stop.
R-626: measured on 9202, not reproduced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-634: a whole-box backup no longer stops/restarts a DEPLOYING app (the
measured cause of containers running under 'not deployed'); StopStack
and StartStack refuse a deploying stack for every caller.
R-625: held badge 'Stopped - restore needed', no Update button.
R-636: kernel oom_kill counter; 20+ in 30 min -> one app_oom_storm.
R-647: held error per reader, copy_holds key, two log wordings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
app_update_undone / app_update_held events (09 decision 15), on by
default and seeded once on existing boxes; R-606 update sentences as
key+args rendered per reader; R-646 startup applied-meta backfill for
apps current with the catalog; R-620 a disabled notifier WARNs once per
event type. Needs hub v0.120.0.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The guarded update gains a folder copy of the app's named volumes, taken
after the pull where the app stops anyway (decision 19, chosen by the
2026-09-23 bake-off). On a failed health check the box undoes: every copy
validated by its finished-marker first, volumes refilled, definition and pin
from the job's own pre-update copies, the old version checked with the OLD
.felhom.yml probe. It holds only if the undo fails, and the hold sentence
says so and what state the data is in. Bind-mounted folders are never
touched.
- R-637 built; R-638/R-640/R-641 do not arise with a folder copy; R-639
(pre-update copies incl. .felhom.yml kept until the undo is over).
- journal phases copying/undoing with power-cut recovery.
- app.yaml last_update_undone + one line on the app page (hu/en).
- R-642: start/restart never answer "completed".
- Removal deletes kept undo copies.
MinAgent unchanged (0.131.0). Nine red-proofs in REPORT.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The controller self-updates daily at 04:30 by default, and after any hub report
once a floor sits above the box. That swap restarts the controller container.
The window proposed for automatic app updates is 02:30-05:00. It contains 04:30.
R-608 — a two-way lock, wired in main.go (stacks never imports selfupdate):
- stacks.Manager.AnyUpdating() -> Updater.SetAppUpdatingCheck, consulted in the
same three places as the existing backupRunning gate.
- Updater.IsUpdateRunning -> Manager.SetSelfUpdatingCheck; UpdatePreflight
refuses `self_updating`.
- MEASURED: the gap was narrower than assumed. The update's `backing-up` phase
already takes the backup single-flight, so that one phase was covered. The
other six were not, and `starting`/`verifying` are where data may have moved.
- The lock must NOT latch: a held app does not block the controller's own
updates, including the release that might fix the hold.
R-609 — the 409 carries `data.reason`, additively. transient (busy, updating,
deploying, migrating, self_updating) vs terminal (held, downgrade). Found while
writing the test: the router refuses a HELD app on its own line before the
preflight, so `held` would have been the one reason missing.
Five red-proofs, each seen to fail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-589 the update badge, R-590 the data-folder backup sentence, R-573 the two channel
banners, R-572 two dead helpers. All four were built in Go, which is why neither the
template parity fixtures nor TestI18nEnglishPages could see them; slice 5's LIVE proof
is what found them.
R-590 is the one that matters most: it is a promise about the customer's files. It
said, in Hungarian and under an already-English folder label, that a drop-zone is
temporary and unbacked. The test now asserts the CONSEQUENCE in both languages — and
that the two languages do not produce the same string, which would mean the English
fell back.
R-572 is not what its row said. The row claimed a template renders "vasárnap" on an
English page; measured, NO template and no Go file called pruneLabel or
nextPruneLabel. They were dead func-map entries returning Hungarian, so they are
deleted rather than translated — translating dead code would add machinery with no
reader and a test pinning a fiction. Deletion is fail-loud and that was proven: a
template naming the removed function panics loadTemplates at startup.
Five red-proofs. One of them says something about the gate rather than the code: the
Go-parity gate does not measure localeFuncs keys against the base capture, so the
citation to TestLocaleFuncsHungarianBundleMatchesFuncMap is what carries them — the
test was extended to make that citation true.
MinAgent: 0.131.0 (unchanged). Hungarian byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
MinAgent: 0.131.0 (unchanged). Needs hub v0.118.0+, which shipped first and
tolerates a box that sends none of this - every box in the fleet is that box
until this release reaches it.
The hub writes a household's e-mails in their language now, but about a third
of those mails carry a sentence the BOX composed, naming a drive, an app or a
number. The hub cannot translate one. So the box sends it twice.
- message_customer on POST /api/v1/event, omitempty. A HUNGARIAN household
sends nothing extra at all, so its payload stays byte-for-byte what every box
sends today and the hub's fallback path keeps being the one production
exercises rather than a branch nobody takes.
- 19 producers render both sentences from ONE bundle key. `message` stays
Hungarian always: it is what the operator is mailed and what the hub logs.
- customer.language bootstraps a new box - stored choice, then config, then
Hungarian. The config value is NEVER written into settings.json: that would
record a choice the household never made.
The Hungarian did not move, measured twice: the wire golden from the slice-2
base commit, and the Go parity gate over all 19 new keys.
Three guards had to learn the change and one caught me: the test seam now
carries the new field; the R-329 severity register reported two dynamic sites
as no longer existing the moment they moved off PushEvent (the walk now checks
36 severity literals, up from 20); and TestConfigLanguageIsWiredInMain reads
main.go, because cmd/ is gitignored and ripgrep does not.
A mistake, named: the first pass dropped displayName from three producers,
which would have mailed customers "Alkalmazás telepítve: %!s(MISSING)". Caught
reading the diff; now pinned by a test that refuses %!/MISSING/%s/%d in either
language.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Message bundles (internal/i18n) expanded into templates before parsing, one
template set per language. Launcher, /backups, /apps/<slug> and the layout
converted; household language setting, POST /settings/language, ?lang= override,
report field. Parity test against fixtures captured from unconverted templates;
copy gates read templates expanded; new i18n_missing_gate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted.
R-550 (operator ruling: fix). A design reversed and recorded: the restore
op-status was in memory by choice. Now restore-status.json in DataDir, written
atomically at both ends of an op. At startup a record still marked running
becomes a failed, interrupted result kept per app until that app's next
restore, shown on /backups/restore and the off-site wizard, and raised once as
restore_interrupted. Cooldowns stay in memory.
R-546. The R-543 reminder bar consults the agent's own preflight ok (every
blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while
paused. /backup/escrow shows a waiting card that polls and reloads instead of
red crosses and English diagnostics. POST /api/escrow/start refuses 409 before
staging or starting - the direct path chaos night used. Unknown readiness keeps
the bar.
Red-proofs (each seen failing): restore record across restart; main() calls
both startup functions; startup helper with loading skipped; restore page card;
bar held back; waiting card; start refusal. go build/vet/test ./... green, 28
packages; controller_gates --fast all OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.
R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.
R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.
Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
Operator ruling 2026-09-13. An app with classified binds walks second
drive -> off-site -> own unit (its unit holds no files); volume apps keep
2 -> 1 -> 3. RestoreHold.CopyHolds records what the chosen copy holds and
the sentence ends with it; older holds keep their tier-only sentence.
Tests on both halves; red-proof: a layout-blind order fails the bind case.
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.
Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not
yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition
(alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names
survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored
the broken definition over the copy the customer is told to restore from.
backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now
also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to
stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.
The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.
NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.
app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.
The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.
AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.
THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.
+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.
BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.
And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.
Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.
+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
Without it the only way to see the job work is to wait for 05:30, which makes live
validation and any future diagnosis a next-day exercise. Same function as the scheduled
job - no second code path.
ONE deliberate difference from the integrity button: due-ness is NOT bypassed. There,
forcing means "check the store again", which is always answerable. Here due-ness IS the
target selection - an app is due when its newest snapshot has not been proved - so
ignoring it would mean inventing a second way to choose an app, exactly what having one
function prevents. When nothing is due the button says so, honestly.
Every other guard intact, including the single-writer flag: a hand-run during a backup
SKIPS exactly as the scheduled one would.
POST /api/debug/backup/offsite-proof, button beside "Restic integritas" on the debug page.
debug_route_gate pairs the two, so a button with no dispatch (R-400's shape) cannot ship.
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.
THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.
IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.
THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".
THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.
THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".
IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.
IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.
DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.
ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.
SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.
NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.
33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.
A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged
without changing its size made plain `restic check` report "no errors were found"
on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store:
35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means
not-configured, therefore the default; a malformed value falls back to the DEFAULT,
never to structure. A completed check over 5 minutes logs a WARN naming the
duration, the depth and R-401 — operator log only, no hub event, no depth change.
The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT
RECORDED, never "structure").
R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on
page LOAD, so those panels were permanently blank. backup/crossdrive implemented;
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both
storage/simulate-* deleted with their panels and JavaScript.
scripts/debug_route_gate.py fails in both directions and is registered after the
seven were resolved. 18 referenced, 18 dispatched, none orphaned.
Corrections: the dead-field warning in report/types.go said the controller runs no
integrity check and the notifiers are called from nowhere — both false since
v0.227.0. controller.yaml.example gains its missing integrity: block.
integrityCheckTimeout's "ships OFF" comment rewritten.
Nothing ever verified that the off-site copies are still readable. The
whole-guest tier has verify jobs; the tier holding the customer's documents and
photos had none -- the complete set of restic verbs this controller used
contained no `check`. We would have found out at restore time, with a customer
waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain
`restic check`; we had never run it.
R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub
allowlists both event types and carries the Hungarian text for both, the
settings checkbox exists, and the debug button posts to /api/debug/backup/
integrity. Everything was built except the part that runs. SIXTH instance of
that shape in this project.
THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by
running `unlock --remove-all` and retrying, and its own comment records why that
is safe: every caller holds the in-process single-flight mutex, so any lock it
meets is stale. A check that did not take that flag could meet a LIVE prune's
lock from this same box, remove it, and retry over the top of it. So the check
TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly
backup behind it, and a skip costs nothing because due-ness makes tomorrow try
again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never
invoked, `unlock` never in any argv. Its red-proof prints the real thing --
restic running `check` while the flag was held.
DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful
check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape,
a dated check quietly missed and never caught up. No Weekly primitive added.
THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts.
"I could not look" is not "I looked and it is broken" -- R-339 already owns
reachability, and a second alarm for the same fact trains the operator to
discount the one alarm that means the backups are damaged. A timeout is
unreachable, never damage. A failure advances due-ness (a broken store must not
be re-checked nightly); a skip and an unreachable store do not.
Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by
design. A weekly success e-mail is how people stop reading their alerts.
The customer gets a SENTENCE; restic's words go to the log, truncated (R-379:
615 bytes of raw database text reached a customer once). read-data-subset ships
OFF and a malformed value is refused at read time rather than handed to restic,
where one typo would fail the whole check.
Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK --
those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero
still passes unmodified.
Also: the monitoring page stopped promising a Sunday job that never existed, and
the debug button got its dispatch case.
PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for
already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable
since the off-site tier shipped, and other tests drive restic-backed paths
through it. A resticStepFn seam would have been WORSE here -- it would replace
the `unlock --remove-all` escalation and hide it from the assertions that must
see it. R-358's AST ordering test is converted to a real execution test instead,
which immediately surfaced something the AST walk could not: unlockStale
legitimately runs before the restore.
Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages,
rc 0. All 12 controller gates OK.
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.
The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.
The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.
It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
every failure path, so the app alarms on the next scan;
2. Begin REPLACES the set (one marker file = one operation);
3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.
Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
- drop markStopped from Begin -> "suppressed at stop = map[]"
- drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
- pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.
Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
R-329. NotifyAppStartFailures emitted severity "warn". The hub accepts exactly
{info, warning, error, critical} and silently coerces anything else to "info",
which severityNotifies then drops BEFORE both legs. Banner shown, event stored,
POST 200, no mail sent. One word.
This is the second time: DiskAlertKind.Severity emitted "warn" until v0.215.0
and its own comment records that every warning-level disk alert went to nobody.
A comment recorded the lesson and nothing enforced it. The guard is now an AST
walk over the whole controller - grep cannot work here, since "warn" appears
legitimately nine times as a healthcheck status vocabulary.
The sweep found exactly one bad severity. Its limits are stated: the walk cannot
follow a variable, so all six dynamic call sites are registered by name with the
values each can take, and a new one fails the test. Two of the six were found by
the guard, not by the hand sweep before it.
Also pinned: fillwatch.Band.Severity() returns "" for BandOK, which would vanish
the same way. It is unreachable because Check() notifies only on escalation -
but that safety lives in a different function from the one that looks unsafe, so
the test asserts the consequence rather than the mapping.
app_start_failed gains a customer toggle, DEFAULT OFF, per operator ruling. The
operator is mailed either way: processOperator never consults customer prefs.
It is deliberately NOT in operatorOnlyEvents, which would make the toggle a lie.
R-386. classifyRunStates decided "the customer stopped this" from the STATE, so
every stopped stack was assumed deliberate. Measured on demo-hp: privatebin
stopped out of band, nine scans, zero events, zero banner - while the comment
beside it claimed an out-of-band stop still alerts.
DesiredState already records the answer and has exactly one writer. Stopped ->
no alarm; Running -> alarm; absent -> UNKNOWN, keep today's behaviour AND say
so. Absent stays silent deliberately: reading it as "nobody asked" would email
about every app anyone ever stopped, fleet-wide, on the first cycle after
upgrade. The gap is bounded not silent - IntentUnknown is set and the names are
logged at INFO on the heartbeat cadence. failedRestart still lifts a Stopped
intent, or F-CRIT-1 re-opens. No new DesiredState writer.
Two settings toggles each governed two alarms. "Lemez figyelmeztetes (90%+)"
also wrote disk_critical, the drive-is-FAILING alarm. Now four honest toggles;
12 became 15. A no-op save stores the existing slice verbatim, so byte identity
is by construction - without that guard the defaults case reorders, which the
red-proof caught.
Test count 1504 -> 1522. Five red-proofs, five seen failing; one passed first
time and is reported - that mutation was inert, not the test weak.
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the
R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A
two-container app whose database exits goes unhealthy BECAUSE it cannot reach
that database - so the symptom the dead database causes was what suppressed the
alarm for it. unhealthy is not a down state, so classifyRunStates never marked
the app down and app_start_failed never fired.
Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the
F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich
failure, back through a different door.
Two things moved, and either alone leaves the defect standing: the supervised
test is hoisted above the unhealthy/starting/restarting returns, and "some
members are up" now counts ANY member not in the down bucket. The old guard was
running > 0, which made the R-51 block unreachable in exactly the case it was
written for.
IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy
container is running and folding it in reintroduces the flapping that exclusion
exists to stop. No new state was minted. Only the ORDER changed. The priority
comment was rewritten because it asserted an ordering the code no longer has.
Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they
asserted an unhealthy/starting/restarting member beat an exited peer on
unless-stopped, which pinned the defect as settled behaviour. They keep their
intent with the down member given a benign policy.
R-383. The double-failure message said the previous state's backup EXISTS,
built from the returned path without asking the filesystem - and a missing file
is one of the two ways that rollback fails. undoCopyPhrase now describes the
copy from disk: present, partial, missing (still naming where it should be), or
never written. Zero-length counts as missing.
Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two
halves of R-384 convict independently.
writeSafetyDump called DumpOne into the app's OWN unit dir and renamed the result
to pre-restore-* afterwards. DumpOne writes <stack>-<dbtype>.sql - the app's
canonical dump - so every safety dump overwrote the app's real backup and then
moved it away, leaving the app with no database backup until the next nightly
run. A local restore-from-unit in that window tells the customer the app never
had a database.
The comment beside it asserted the rename meant it 'can never overwrite the app's
real dump'. False as written, and believed for four months. Measured live before
the fix: docmost and bookstack each held only pre-restore-* files and no
canonical dump.
DumpOneTo takes the final path and derives its own .tmp from it. DumpOne keeps
its signature and calls it with the canonical name. writeSafetyDump asks for its
own name directly; the rename is gone; the comment now states the invariant and
how it is enforced.
db_dumps no longer lists the undo copies. All three consumers of Manifest.DBDumps
were grepped and named - all inside recovery_unit.go, none reads it for recovery.
The files are neither deleted nor hidden.
Tests 1485 -> 1493. FIVE red-proofs, TWO PASSED first time and both are reported:
the behavioural tests inject the dump seam so a mutation inside DumpOneTo was
invisible, and 1.3 had no test at all. Guards added at the layer each defect
lives in; both mutations then convicted.
It runs as a second process: it clears settings.json but the running controller
keeps its in-memory copy and goes on refusing. Measured on demo-hp - clear
succeeded, file correct, start button still refused until a restart.
Also records the lost-update window between the two processes, and why clearing
through the running controller (the right shape) needs an operator tier the
controller's HTTP surface does not have.
R-379 and R-380 were one failure. Both ended with a half-restored database; the
only difference was whether it looked broken. Postgres emptied and crash-looped;
MariaDB applied part of the dump and reported health=healthy with a zero-row
schema-version table. Measured live on demo-hp 2026-08-22.
The undo copy was already taken and already good - proven by hand that day on
both engines. Nothing in the product could apply it. Now it does, with the same
ImportDump call, before any restart and inside the DB-only window.
The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one
path for an app with two databases; a rollback on that would restore one and
leave the other half-written.
When the rollback also fails the app is HELD STOPPED (operator ruling): a running
app on a half-written database lets the customer make the damage permanent. Every
start path refuses it - customer button, appstop Recover, boot sweep - via the
shared driveStartGate, checked ABOVE its driveless early return because these
apps have no drive. The marker is ended so nothing auto-restarts it. The row goes
red. Cleared with --clear-restore-hold, an operator CLI route.
--single-transaction is a belt on Postgres only; MariaDB DDL is not transactional
and that is why the rollback is the fix.
R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its
middle rows out of their own database) and starts reaching the operator log,
which never had it.
R-382: the summary log prints the volume count it already held.
Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per
app, pruned from the capture side. The reported render-as-an-app symptom did NOT
reproduce - the live page was read first and had zero occurrences.
Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381
behavioural test injected below ImportDump. A guard at that layer now convicts.