ff68b0b433db6d9ac223b54357a13f0bb117836e
412 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ff68b0b433 |
v0.248.0: i18n slice 1 release A — apps and settings in English, Hungarian byte-identical (R-556)
gates / gates (push) Successful in 20s
Ten pages converted against fixtures captured from unconverted templates (
|
||
|
|
612c417024 |
v0.247.0: i18n spike — the dashboard can speak English, Hungarian byte-identical
gates / gates (push) Successful in 19s
Message bundles (internal/i18n) expanded into templates before parsing, one template set per language. Launcher, /backups, /apps/<slug> and the layout converted; household language setting, POST /settings/language, ?lang= override, report field. Parity test against fixtures captured from unconverted templates; copy gates read templates expanded; new i18n_missing_gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
0fe315b759 |
v0.246.0: an interrupted restore is told; the recovery-code reminder waits until the box can take it
gates / gates (push) Successful in 15s
MinAgent: 0.131.0 (unchanged). Requires hub v0.117.0 for restore_interrupted. R-550 (operator ruling: fix). A design reversed and recorded: the restore op-status was in memory by choice. Now restore-status.json in DataDir, written atomically at both ends of an op. At startup a record still marked running becomes a failed, interrupted result kept per app until that app's next restore, shown on /backups/restore and the off-site wizard, and raised once as restore_interrupted. Cooldowns stay in memory. R-546. The R-543 reminder bar consults the agent's own preflight ok (every blocking item, not a copy of pbs_storage_id), cached 60 s, probed only while paused. /backup/escrow shows a waiting card that polls and reloads instead of red crosses and English diagnostics. POST /api/escrow/start refuses 409 before staging or starting - the direct path chaos night used. Unknown readiness keeps the bar. Red-proofs (each seen failing): restore record across restart; main() calls both startup functions; startup helper with loading skipped; restore page card; bar held back; waiting card; start refusal. go build/vet/test ./... green, 28 packages; controller_gates --fast all OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ad398b60d9 |
v0.245.0 — R-543: the household is asked for the recovery code, the page says "szunetel" until then
gates / gates (push) Successful in 14s
Off-site backup is ON by default and does not RUN until the household creates its recovery code. The pause is the zero-knowledge escrow design and is untouched here; what was missing is that nothing ASKED, while the app-backup page promised the very copy that had never run. - a reminder bar on every authenticated page while the off-site tier is configured and its escrow is not complete, linking /backup/escrow. It is the R-241 bar, second instance: same session-cookie dismissal, back next visit, gone for good when escrowed. No second banner system. It hangs off executeTemplate, the single render choke point, so it cannot reach only the pages someone remembered. - the tier-1 file sentence renders by tier3State's own vocabulary instead of the app's shape: active -> "vedi", escrow_pending -> "vedene ... szunetel" + the route, no copy at all -> says so and names both ways out. - both fixes red-proofed: the bar test fails on BOTH pages with the hook removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2f8ff2414c |
v0.244.0: the backup page stops promising what it does not hold (R-537/R-538/R-536)
gates / gates (push) Successful in 17s
R-537 — the contents label is now PER TIER. One string computed from the app's shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so for the four class-A apps it was claiming „Adatok" for files it does not hold. R-538 — a unit restore REFUSES before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. It runs before the stack is stopped because the measured harm included the app's own wastebasket going unreachable, which still held every byte. R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion, with app_deploy_started and app_deploy_failed as the honest pair. Each fix red-proofed: seen failing with its own sentence, passing when restored. Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d3eacbb7cc |
backup tile: unknown size shows a dash, not 0 B (R-517 follow-up, measured on 9201)
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
843b319f35 |
v0.243.0: FileBrowser generated admin password (R-513); per-tier whole-guest backup truth (R-517); skip absent-storage tiers (R-518); OOM-killed worker visible (R-514)
gates / gates (push) Successful in 14s
MinAgent: 0.131.0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
406755fa8f |
docs: v0.242.0 report, context; R-489 measured limit recorded (row kept open)
gates / gates (push) Successful in 14s
A volume recreated by a unit restore carries no compose label, so the before/after difference misses it; proven on the scratch guest. Docs only, no release. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d698ce343b |
controller v0.242.0: a removed app is listed with its kept backup; five small ones (R-487 R-491 R-490 R-489 R-476 R-456)
gates / gates (push) Successful in 14s
R-487: the local backup lists are keyed on the drives, not on what is deployed — a removed app whose unit was kept is listed with the restore that reinstalls it, the picker answers for it, and the restore opens the unit where it sits. R-491: a removal clears the app's update hold. R-490: /api/system/info reaches the API router and reads the default storage path. R-489: volumes_removed is the real before/after difference, [] when none. R-476: a Tier-2 copy is dated by its data, not its manifest. R-456: the boot-orphan rule is pinned. Every fix red-proofed. |
||
|
|
3e813307cc |
controller v0.241.0: a bind-data app leans on off-site before its own unit; the hold names what the copy holds (R-479)
gates / gates (push) Successful in 13s
Operator ruling 2026-09-13. An app with classified binds walks second drive -> off-site -> own unit (its unit holds no files); volume apps keep 2 -> 1 -> 3. RestoreHold.CopyHolds records what the chosen copy holds and the sentence ends with it; older holds keep their tier-only sentence. Tests on both halves; red-proof: a layout-blind order fails the bind case. |
||
|
|
bdcbd50b42 |
controller v0.240.0: seven defects from the any-tier proof and the first nightly rotation
gates / gates (push) Successful in 13s
R-486 (P1): removing an app with its backups KEPT keeps its Tier-2 record, so the second-drive restore is no longer refused over an intact mirror. R-484: postgis/pgvector/timescaledb images are Postgres (logical dumps). R-485: the backup card sizes the recovery unit and the mirror(s). R-480: a held update's sentence leaves the card once the hold is lifted. R-477: the update's off-site lookup is one snapshots call, no stats. R-478: a copy older than this install's deploy does not count. R-474: "delete backups" deletes the unit, the mirror(s) and the prefs. Tests and red-proofs per row; evidence in felhom.eu documentation/audits/v0240-2026-09-13/ and nightly-2026-09-13-adventurelog/. |
||
|
|
b93c1543da |
controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1 (own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable counts as absent with a WARN) and leans on the first FRESH copy; the backup_max_age rule applies to whichever tier is chosen. No copy anywhere: back up first. Refused only when nothing exists and no backup can be taken. RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured unit proven current. The hold names the tier (második meghajtó / saját meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text. A successful off-site restore now lifts an update hold. The backups page still uses Tier2UnitRestorePoint unchanged. Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/. |
||
|
|
f946b0d0ca |
gates: the newest release header must state its MinAgent; backfill v0.233.0-v0.236.0 (R-470, R-472)
gates / gates (push) Successful in 12s
From hub v0.112.0 a floor above the golden is served only with a declared
MinAgent, read from this CHANGELOG's newest `## vX.Y.Z` header. New fast,
blocking gate minagent_header_gate.py: that block must contain a line
starting `**MinAgent: X.Y.Z**`; prose, a code span or a non-bold mention
does not count (decoy test; red-proof F: relaxing the match to a body
mention fails test_decoy_body_mention_is_not_the_line).
Backfill: v0.233.0, v0.234.0, v0.235.0 and v0.236.0 carried no MinAgent
line. Each now reads `**MinAgent: 0.129.0** (unchanged)`, the value of the
nearest earlier header (v0.232.0). Proof it did not move: zero commits
under controller/internal/agentapi since 2026-09-01 (last:
|
||
|
|
cbcca03061 |
v0.238.1: the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (slice 4 follow-up)
gates / gates (push) Successful in 13s
Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition (alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored the broken definition over the copy the customer is told to restore from. backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned. Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
129201abab |
v0.238.0: the page follows the update, and a held app offers no way to start it (update arc slice 4 Part 4)
gates / gates (push) Successful in 13s
No behaviour change on the box — the surface only.
- Frissítés follows the job: the button shows the phase label (polling GET /api/stacks/{name}
every 3 s) and the page reloads when updating goes false.
- An updating card offers no lifecycle button; a held card (failed update OR failed restore) shows
the hold sentence with a Mentések link and nothing that would start it; a failed update that held
nothing shows its sentence above the buttons. app_info shows the same three notices.
- The updating/held checks run BEFORE isOperational, which counts `restarting` as operational — how
the 2026-09-01 spike saw a green Frissítés beside a crash loop. Pinned with StateRestarting
fixtures; red-proofed by moving the checks after it (both tests fail).
- No new CSS, no version number.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
0d402f711d |
v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.
The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.
Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.
Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.
Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
42a73e667a |
v0.236.0: "delete my data too" deletes the data, or says that it could not (R-442)
gates / gates (push) Successful in 13s
Removal resolves the drive from the app's own app.yaml HDD_PATH (the 07 ~L437 rule), never the global cfg.Paths.HDDPath which no box sets. A data removal that cannot be resolved, or whose drive is absent, is refused with a typed RemoveRefusedError -> 409 + exact Hungarian sentence, before compose down, and the app is kept. SSD app -> hdd_paths_removed: [] never null; missing folders stated; backup-path refusals reach the response. 15 tests, two red-proofs run (pre-fix fallback -> C fails with err=nil and the handler 200s; "no drive refuses" -> D fails). Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2a56f557d0 |
v0.235.0: a delivered fix must also refresh the stored definition
gates / gates (push) Successful in 14s
Found by the LIVE validation on demo-hp, not by review. Scenario A passed - a non-image catalog change reached the pinned app on the real 15-minute cycle - and that is exactly what exposed the gap: the stored applied-compose.yml is written when the PIN is written, so the fix landed in the live compose file and not in the store. The first time the catalog then moved a version, the freeze would have rendered the pre-fix definition and reverted every fix delivered since - silently undoing the half of the operator's ruling that says fixes keep flowing. The equal-images branch now refreshes the store as it delivers. The images cannot move in that branch by construction, so no version moves and no intent is rewritten. RenderPlan gains StackDir so the syncer can write it. TestFixRefreshesTheStoredDefinition asserts both halves: the fix reaches the store, and it survives the freeze that follows. |
||
|
|
8a0e0a59ad |
v0.235.0: freeze the version, keep the fixes flowing (operator ruling 2026-09-06)
gates / gates (push) Successful in 12s
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of up -d to pick up template changes was CHOSEN and written down in its own comment. The operator ruled Option 1, and this implements it. The rule: while the catalog offers the same version you run, its fixes flow to you; the moment it moves to a newer version you are frozen until you update. NOTHING was added to any of the thirteen compose up -d call sites. Most of them are repairs - the boot reconciler, the drive-return gate, the app-stop guard - and a repair path that refuses to repair leaves a customer's app down, which is worse than the problem. They are made safe by removing the reason. app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT installed_images, which is an observation; letting a reading become a deployment is the R-166 category error one field over. Four writers, each also storing the exact definition as applied-compose.yml. UpdateStack advances the pin and re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a pin set afterwards would pull the frozen version and report success. The syncer renders instead of copying, through one nil-safe seam. Catalog images equal the pin -> verbatim, so fixes and self-healing both survive; they differ -> the WHOLE stored definition, never a substitution of refs into a newer template (wger 2.6 needs a DB config the older template cannot supply). This is deliberately not 'skip deployed apps', which was option B and was rejected. AdoptPins runs once at boot after the backfill, files only, and skips loudly rather than inventing a pin. syncer.Start() moved to after it: the initial sync would otherwise run while every app was unpinned and overwrite a deployed app's version once per boot. THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages reads the LIVE compose file, which is now the frozen one, so the comparison would have answered Naprakesz on exactly the apps that are behind - with every test green, because the new field has the same type. It now reads CatalogImages. +16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted. A test also caught the syncer writing an empty compose file over a live app. |
||
|
|
38d28b5b62 |
v0.234.0: seed installed_images at startup, so the label appears on an app nobody touched
gates / gates (push) Successful in 13s
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist - up 15 hours, running exactly the catalog pin - showing no badge at all. v0.233.0 wrote the record only from the four bring-up paths, so an app nobody restarts carried no record indefinitely. On a quiet box that is every app, which is the box we most want to see. The known limitation WAS the feature not working. BackfillInstalledImages runs once at startup, beside BackfillDesiredState and before the boot reconciler. It READS containers: starts nothing, restarts nothing, writes no compose file. It never overwrites an existing record. And it REFUSES to seed a partial observation, which is why this is not a three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a degraded app from what is visible would render 'Frissites elerheto' over an app that is perfectly current. The bring-up paths may write a partial because they follow a successful up -d where a gap is real news; a backfill meets any state. Same data, two writers, two admission rules - deliberately. Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and the string '46 napja', but the render path reads time.Now(), so it was green on the day it was written and red the next morning. Now derived. Filed as R-457 with six other candidate files named as unchecked, not accused. +5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run and reverted; the wiring and its ORDER pinned by an AST walk. |
||
|
|
8025304acc |
v0.233.0: record what each compose service actually installed, and badge whether it is current
gates / gates (push) Successful in 12s
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no auto-update, the three lifecycle buttons byte-identical. Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each entry carrying ref + repo digest + first-seen timestamp. Written by Manager.recordInstalledImages after a successful compose up from StartStack, RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never from docker-compose.yml: the syncer overwrites a deployed app's compose on a 15-minute cycle and the two disagreed for 25 minutes in the spike's own measurement. A failed write NEVER refuses the action - the deliberate opposite of SetDesiredState, because this is an observation and that is an intent. Not called from StartStackServices (the R-47 DB-only window). Its own docker seam with a context and a 30s timeout, which neither existing exec helper has. Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the recorded ref per service against what the current template pins and returns a *MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS. NO RECORD RENDERS NOTHING: absent means unknown and never means current. No version number reaches the customer and no registry is queried. Known limitation, filed not hidden: 23 catalog pins float, so those apps can read Naprakesz when the image behind the tag has moved. +17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real RestartStack plus an AST walk of the four call sites. Three companion red-proofs run and reverted. |
||
|
|
22983885f1 |
re-run CI against a register that now carries R-421
gates / gates (push) Successful in 13s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register lives in felhom.eu, so a repo citing a new row must be pushed after it. |
||
|
|
681cc663ef |
decoy sweep: eight holes in this repo's gates, all measured, all fixed (R-421)
gates / gates (push) Failing after 13s
Every gate was DECOYED - the label constructed without the fact, the gate run, the verdict recorded.
No verdict here was reached by reading, because reading is exactly how the five prior instances hid.
SCOPE IS A FACT TOO, and it was the big one. Six gates decided what to look at with os.listdir - one
directory level. Every one was green AND CORRECT, because no template subdirectory exists today; every
one would have gone blind the moment anyone added templates/partials/, which is an ordinary act. A
single planted file carrying an emoji, a native confirm(), hand-rolled row markup, a dangling JS id
reference, a templated secret and an unregistered retrieval promise passed all six.
THE CONTROL IS WHAT MAKES THAT A MEASUREMENT: mojibake and docker-v already used os.walk, saw the
identical planted file, and convicted. So the cause was the listing, not the decoy.
COMMENTS ARE NOT CODE, AND COMMENTS ARE NOT CONTROLS. debug-routes matched `case subpath == "x"` in
raw text, so a case left in a commented-out block counted as a live handler - which is R-400's
original defect (seven dead controls on the page an operator opens when something is already wrong)
reached through the one door its own gate could not see. app-row-dedup's MUST_USE check had the same
shape: a commented-out {{template "app_list_row"}} satisfied it.
Stripping is deliberately crude in debug_route_gate, and that is correct there: its own docstring
insists on ten lines that cannot rot. A // inside a string literal truncates that line, which can
only ever HIDE a reference, never invent one - it fails in the safe direction.
NOT FIXED, and left open with its decoy rather than quietly patched: R-425, offbox-rename scans a
fixed three-entry FILES list, so banned NAS branding in a NEW offbox template passes. The scope was
correct when written and silently narrows every time the feature grows a file.
test_gate_decoys.py holds 10 decoys and declares COVERS, which felhom.eu's new decoy-coverage gate
AST-parses - a substring search for coverage would be the very shape this sweep exists to find.
No Go code. No version bump. No image. No golden owed.
Survey: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
|
||
|
|
a4444088ad |
R-404: the golden NOTICE, in the repo where the debt is created - NOT A RELEASE
gates / gates (push) Successful in 12s
No version heading on purpose. No Go code, no image, no version bump; giving this one would create the exact golden debt the change is about. Until today this repo - where a release actually happens - had NO golden-currency check at all, while felhom.eu ran one on every push including documents-only ones that can neither create the debt nor clear it. The person who could act heard nothing; the person who could not act was blocked, thirteen --no-verify uses' worth. golden_notice.py is ADVISORY IN EVERY CASE, and that is the only correct behaviour rather than timidity: at the moment a release is committed the golden legitimately does not exist yet, so blocking there would refuse the commit that STARTS the process - and blocking later is the mistake being undone. NO SECOND IMPLEMENTATION: it IMPORTS felhom.eu/scripts/golden_currency_gate.py and calls that gate's own released_versions()/newest_baked(), so it is the same comparison read in the other direction. Cross-repo shape copied from instructions_gate.py; never a copy of the script, because a copy recreates the drift these gates exist to detect. An absent sibling clone is INCONCLUSIVE and silent about currency - it never guesses. controller_gates.py GAINED A FIFTH `blocking` FIELD. It could not express a reporting-only gate at all before: every registered gate's non-zero exit failed the run, so the only way to add a notice was to give it the power to refuse a push. The capability was added rather than the notice compromised (R-420). False for exactly one gate, and test_golden_notice.py asserts it stays one. Tests N1-N4 with a positive control that every other gate is still blocking. RED-PROOF RUN: making the debt branch return 1 fails N1 - in production that would refuse the commit that starts a release. |
||
|
|
62c6a8a98a |
docs(v0.232.0): CHANGELOG, three CONTEXT rulings, README (R-411/408/407, R-414, R-412a)
gates / gates (push) Successful in 14s
|
||
|
|
e43b5ec07d |
v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged. THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does, nightly, on one app. IT DOES NOT prove a restore puts data back into a running app. That stays drill work and 07 section 8 matrix row 4 is NOT moved. THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So: (1) everything declared is present, AND (2) the manifest declares what the app is supposed to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test read verdict "pass". THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate the app's shape, and GetDockerVolumes describes the running app. Database half is DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose file's parent dir, which inside a unit is the literal string "compose". Measured on all eight real units on demo-hp the counts match exactly and the naming held every time - but "held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented. THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes that test read verdict "fail". IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect: --no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test fail on "unlock" appearing in the argv. IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch does not take it (R-408) while offbox_integrity.go states that invariant as universal. DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's app. ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app> would mean a nightly background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to placement, so a proof copy can never be pushed into a live app. SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's behaviour is unchanged. NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is sound and the content is absent: different cause, different action. The hub half shipped FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted type is 400'd and vanishes. 33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller gates OK. Five red-proofs run and recorded in REPORT.md. A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242). |
||
|
|
1cfdde968f |
docs(v0.230.0): R-403 — CHANGELOG, CONTEXT rulings, README, REPORT
gates / gates (push) Failing after 13s
CHANGELOG v0.230.0, leading with the measurement rather than the fix: 120 082 104 B -> 7 036 B on the shipped v0.229.0, reproduced before anything was built. CONTEXT records three rulings: hollowness is a MANIFEST question and never a size question; the guard fences one shape and NOT shrinking, because the derived-copy rebuild is a design decision; and the rehydrate happens inside the restore because a follow-up job races the 5-minute capture. Plus the shape the live run taught: a warning that fires on everything costs the same as the comforting lie it replaces. README documents the refusal, what each surface says, and why the capture job is deliberately not guarded. REPORT leads with Part 1's result, carries the six red-proofs, the per-row Scenario D table with its seven-app control, and eight observations including R-404 filed-not-acted-on and three mistakes of mine recorded rather than tidied away. |
||
|
|
8aa95b5831 |
docs(v0.229.0): R-102 + R-103 — CHANGELOG, CONTEXT rulings, README, REUSE, REPORT
gates / gates (push) Failing after 13s
CHANGELOG v0.229.0. CONTEXT records three rulings: the source moves and the destination does not; two predicates and not one wider one (R-356's cost restated); and a destructive operation reached from a non-destructive surface carries the difference in the CONFIRM, not the label. README documents the new action and route and corrects the coverage note to the measured count. REUSE maps the four unit-directory-relative primitives and the new manager methods, with the traps. REPORT covers the live drill on demo-hp (docmost, class B, primary unit moved aside — 3 volumes of 3 and 1 database of 1 in 28.65 s, accented filename byte-identical verified as hex, the app reading its own row over TCP; Scenario D with the guest app.yaml also aside, secrets recovered=2/2), the settled count (A=7 B=45 C=1, and why the earlier 9/43/1 was wrong), the five named red-proofs, and seven observations including R-403 and a process error of mine that changed the box and is now in memory. |
||
|
|
3c49dc8ea4 |
v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged without changing its size made plain `restic check` report "no errors were found" on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store: 35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means not-configured, therefore the default; a malformed value falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs a WARN naming the duration, the depth and R-401 — operator log only, no hub event, no depth change. The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT RECORDED, never "structure"). R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on page LOAD, so those panels were permanently blank. backup/crossdrive implemented; backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both storage/simulate-* deleted with their panels and JavaScript. scripts/debug_route_gate.py fails in both directions and is registered after the seven were resolved. 18 referenced, 18 dispatched, none orphaned. Corrections: the dead-field warning in report/types.go said the controller runs no integrity check and the notifiers are called from nowhere — both false since v0.227.0. controller.yaml.example gains its missing integrity: block. integrityCheckTimeout's "ships OFF" comment rewritten. |
||
|
|
45770f2282 |
v0.227.1: the damage classifier matched restic's ordinary progress output
gates / gates (push) Successful in 11s
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and re-pushing changed bytes under a live tag is the :latest hazard with extra steps. looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A HEALTHY restic check prints "check all packs" and "check snapshots, trees and blobs" -- so any check that failed for a NON-damage reason, a connection dropped mid-run for instance, would have been classified as a corrupted repository and told the customer their backups may be damaged. That is the false alarm that teaches an operator to ignore the true one. Caught by the NEGATIVE control, built from the real bytes of a real passing check on demo-hp. The spec made the negative control mandatory and this is what it was for: a control that has only ever seen the failing case proves nothing. Signatures are now phrases from restic's own error wording. Also in this commit: CONTEXT.md records the three rulings (take the flag and skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead fields) plus the measurement a future session would otherwise assume wrongly -- THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job, the route and the config, and corrects a line that listed four debug backup routes when only two exist. REUSE gains three rows, including one that records R-398 was my own mistake so nobody re-files it. |
||
|
|
36c13cfa0a |
v0.226.1: ship the CountsUnknown fix under a NEW tag, not a rebuilt 0.226.0
gates / gates (push) Successful in 11s
0.226.0 was already deployed to demo-hp by hand when the fallback defect was found. Re-pushing a changed image under a tag that is already running somewhere is the :latest hazard with extra steps -- two different images, one name, and no way for a box to tell which it has. So the fix ships as a patch. |
||
|
|
c0c8fe67bf |
An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero result, which is honest" exposed that the sentence was FALSE. A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's fallback to RestoreApp -- which returns only an error, and whose signature is deliberately out of scope -- would have printed "ez a mentes csak a beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure direction this whole change exists to remove, re-introduced by the change. UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a fourth sentence claiming only what is known -- the restore ran, the app is back, and we cannot say what came back. RestoreApp's signature is untouched. Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam test was corrected too: its fixture has no recovery unit, so it exercises exactly this path and had been asserting the wrong sentence -- it now asserts the unknown, which is what pins the fallback to it. IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate written to stop findings dying in an overwritten REPORT.md caught a live defect instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the monitoring page advertises a weekly integrity check that does not exist) and R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST test) rather than leaving them in a file that is overwritten every session. REPORT.md is the full run record: baselines re-confirmed, per-test results, the five red-proofs with their observed output, the live validation with verbatim Hungarian messages, what was NOT validated and why, teardown across three layers, and the register 165 -> 167 -> 161. Green gate clean: 28 packages, rc 0. All 12 controller gates OK. |
||
|
|
b8af72764d |
R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21 backup-truth drill, all still in shipped code. They share one acceptance idea: a restore surface must state what it actually did, and must refuse what it cannot do. VERSION NOTE. The task specifying this targeted v0.224.0 against baseline |
||
|
|
e5eee501b5 |
R-331 (controller half): forward stats_known so the hub can tell empty from unmeasured (v0.225.0)
gates / gates (push) Successful in 12s
The hub's operator Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` for EVERY customer, because it rendered the report's `backup` object -- whose snapshot/size/integrity fields have had NO producer since disk-tier restic moved to the host agent (slice 8C). buildBackupReport leaves them zero deliberately and says so. Measured on demo-hp 2026-08-30 while that night's log said `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`. The live numbers were always in the report's `offsite` object, which the hub already reads for its Offsite page and its fill/staleness alarms. The hub fix is to render that -- and that made exactly ONE field mandatory that was not being forwarded. snapshot_count:0 means two opposite things: "holds nothing" and "never measured". R-225 measured that confusion inside this repo (a rebuilt box rendered 0 pillanatkep over a store really holding snapshot f3d9cd67), and settings.OffboxTarget.StatsKnown fixed it for the controller's own UI. It was never put on the wire, so the hub was free to make the identical mistake one layer up -- and did. OffboxReportStatus.StatsKnown now carries it, omitempty, so an older controller sends no key and a reader degrades to UNKNOWN, never to EMPTY. Absence is ignorance, not emptiness. The four dead BackupReport fields stay on the wire (historical reports in the hub store must keep parsing) but now carry a warning naming R-331 and pointing at Offsite. TestBackupReport_DeadFieldsStayZero fails the moment a producer appears for one -- the prompt to update the hub card in the SAME change rather than ship a field nothing renders. RED-PROOF: drop `StatsKnown: t.StatsKnown` -> "a MEASURED empty repository reported stats_known=<nil>". Tests assert the JSON the hub sees, not the Go struct: measured-empty and never-measured must differ ON THE WIRE, which is the entire point of the field. Green gate clean: 28 packages, rc 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
45b52b6ed5 |
R-330: live validation on demo-hp — 3 scans inside the window, 0 events
gates / gates (push) Successful in 12s
v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks stopped and restarted over 87s, three dead-app scans ran INSIDE that window (16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps that alarmed the night before on 0.223.0), zero app_start_failed pushed. The scan count is the positive control, not decoration: an absent alarm is equally consistent with "suppressed correctly" and "the scanner stopped". A first run is discarded IN THE REPORT rather than quietly dropped -- it fired 52s after a controller restart, inside deadAppBootGrace (90s), where the scan returns early and could not have alarmed whatever the code did. demo-felhom is deployed but NOT independently proven and says so: its single app cycles in ~1s, too fast for any 30s scan to land inside. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
92cebb8c95 |
R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.
The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.
The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.
It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
every failure path, so the app alarms on the next scan;
2. Begin REPLACES the set (one marker file = one operation);
3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.
Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
- drop markStopped from Begin -> "suppressed at stop = map[]"
- drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
- pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.
Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
9832760027 |
v0.223.0: the app-down alarm reached nobody (R-329), and the stop nobody heard (R-386)
gates / gates (push) Successful in 11s
R-329. NotifyAppStartFailures emitted severity "warn". The hub accepts exactly
{info, warning, error, critical} and silently coerces anything else to "info",
which severityNotifies then drops BEFORE both legs. Banner shown, event stored,
POST 200, no mail sent. One word.
This is the second time: DiskAlertKind.Severity emitted "warn" until v0.215.0
and its own comment records that every warning-level disk alert went to nobody.
A comment recorded the lesson and nothing enforced it. The guard is now an AST
walk over the whole controller - grep cannot work here, since "warn" appears
legitimately nine times as a healthcheck status vocabulary.
The sweep found exactly one bad severity. Its limits are stated: the walk cannot
follow a variable, so all six dynamic call sites are registered by name with the
values each can take, and a new one fails the test. Two of the six were found by
the guard, not by the hand sweep before it.
Also pinned: fillwatch.Band.Severity() returns "" for BandOK, which would vanish
the same way. It is unreachable because Check() notifies only on escalation -
but that safety lives in a different function from the one that looks unsafe, so
the test asserts the consequence rather than the mapping.
app_start_failed gains a customer toggle, DEFAULT OFF, per operator ruling. The
operator is mailed either way: processOperator never consults customer prefs.
It is deliberately NOT in operatorOnlyEvents, which would make the toggle a lie.
R-386. classifyRunStates decided "the customer stopped this" from the STATE, so
every stopped stack was assumed deliberate. Measured on demo-hp: privatebin
stopped out of band, nine scans, zero events, zero banner - while the comment
beside it claimed an out-of-band stop still alerts.
DesiredState already records the answer and has exactly one writer. Stopped ->
no alarm; Running -> alarm; absent -> UNKNOWN, keep today's behaviour AND say
so. Absent stays silent deliberately: reading it as "nobody asked" would email
about every app anyone ever stopped, fleet-wide, on the first cycle after
upgrade. The gap is bounded not silent - IntentUnknown is set and the names are
logged at INFO on the heartbeat cadence. failedRestart still lifts a Stopped
intent, or F-CRIT-1 re-opens. No new DesiredState writer.
Two settings toggles each governed two alarms. "Lemez figyelmeztetes (90%+)"
also wrote disk_critical, the drive-is-FAILING alarm. Now four honest toggles;
12 became 15. A no-op save stores the existing slice verbatim, so byte identity
is by construction - without that guard the defaults case reorders, which the
red-proof caught.
Test count 1504 -> 1522. Five red-proofs, five seen failing; one passed first
time and is reported - that mutation was inert, not the test weak.
|
||
|
|
5da11c4480 |
v0.222.0: ask whether a supervised member is DEAD before whether one is UNHEALTHY (R-384), and stop promising an undo copy nobody looked for (R-383)
gates / gates (push) Successful in 11s
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A two-container app whose database exits goes unhealthy BECAUSE it cannot reach that database - so the symptom the dead database causes was what suppressed the alarm for it. unhealthy is not a down state, so classifyRunStates never marked the app down and app_start_failed never fired. Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich failure, back through a different door. Two things moved, and either alone leaves the defect standing: the supervised test is hoisted above the unhealthy/starting/restarting returns, and "some members are up" now counts ANY member not in the down bucket. The old guard was running > 0, which made the R-51 block unreachable in exactly the case it was written for. IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy container is running and folding it in reintroduces the flapping that exclusion exists to stop. No new state was minted. Only the ORDER changed. The priority comment was rewritten because it asserted an ordering the code no longer has. Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they asserted an unhealthy/starting/restarting member beat an exited peer on unless-stopped, which pinned the defect as settled behaviour. They keep their intent with the down member given a benign policy. R-383. The double-failure message said the previous state's backup EXISTS, built from the returned path without asking the filesystem - and a missing file is one of the two ways that rollback fails. undoCopyPhrase now describes the copy from disk: present, partial, missing (still naming where it should be), or never written. Zero-length counts as missing. Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two halves of R-384 convict independently. |
||
|
|
da75603553 |
R-385 record: v0.221.1 gets its own CHANGELOG heading
gates / gates (push) Successful in 11s
0.221.1 was built, baked and vouched on 2026-08-23 with no entry of its own.
The prune-ordering fix (commit
|
||
|
|
810b18ab8e |
R-361 follow-on: the undo-copy prune must run above the already-current check
gates / gates (push) Successful in 11s
Excluding pre-restore-* from db_dumps made that list stable across restores, so CaptureRecoveryUnit's already-current early return began firing where it never had - and the prune, which sat after it, stopped running in exactly the case it exists for. Measured on demo-hp minutes after the change: four undo copies on disk against a cap of three. The prune is housekeeping on the dump directory and is independent of whether the manifest needs rewriting, so it belongs above the check. Pruning cannot disturb dbDumps, which no longer contains those names. Pinned by TestR361_UndoCapHoldsWhenTheUnitIsAlreadyCurrent; its red-proof moves the call back below the return and the cap fails at 5. |
||
|
|
968c968559 |
R-361: the safety dump destroyed the app's own database backup
gates / gates (push) Successful in 11s
writeSafetyDump called DumpOne into the app's OWN unit dir and renamed the result to pre-restore-* afterwards. DumpOne writes <stack>-<dbtype>.sql - the app's canonical dump - so every safety dump overwrote the app's real backup and then moved it away, leaving the app with no database backup until the next nightly run. A local restore-from-unit in that window tells the customer the app never had a database. The comment beside it asserted the rename meant it 'can never overwrite the app's real dump'. False as written, and believed for four months. Measured live before the fix: docmost and bookstack each held only pre-restore-* files and no canonical dump. DumpOneTo takes the final path and derives its own .tmp from it. DumpOne keeps its signature and calls it with the canonical name. writeSafetyDump asks for its own name directly; the rename is gone; the comment now states the invariant and how it is enforced. db_dumps no longer lists the undo copies. All three consumers of Manifest.DBDumps were grepped and named - all inside recovery_unit.go, none reads it for recovery. The files are neither deleted nor hidden. Tests 1485 -> 1493. FIVE red-proofs, TWO PASSED first time and both are reported: the behavioural tests inject the dump seam so a mutation inside DumpOneTo was invisible, and 1.3 had no test at all. Guards added at the layer each defect lives in; both mutations then convicted. |
||
|
|
1a2405e86f |
R-379: --clear-restore-hold now states the required restart
gates / gates (push) Successful in 11s
It runs as a second process: it clears settings.json but the running controller keeps its in-memory copy and goes on refusing. Measured on demo-hp - clear succeeded, file correct, start button still refused until a restart. Also records the lost-update window between the two processes, and why clearing through the running controller (the right shape) needs an operator tier the controller's HTTP surface does not have. |
||
|
|
5b52a5964d |
R-379 fix: the rollback must re-discover the DB container
gates / gates (push) Successful in 12s
Found by v0.220.0's own live walk on its first real run. writeSafetyDump
captures its DiscoveredDB before the stop; the DB-only start then re-creates the
container with a new id, so the rollback's docker exec hit a dead container and
sat in waitDBReady for 30s. The app was held for an infrastructure reason while
its data was recoverable.
Re-discover and match by {stack, engine} - what reimportDBDumpsFrom already did.
Fail closed when the container cannot be found.
No unit test caught it because they all inject the import seam and never look at
container identity. The new test asserts the identity handed to the import.
|
||
|
|
2c724c9283 |
R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts. |
||
|
|
08eb1a6e3a |
R-356: the off-site restore refused every app that has no data drive
gates / gates (push) Failing after 12s
ReconstituteFromOffsite and PlaceOffsiteRestore both resolved the restore destination with the RAW HDD_PATH and read an empty answer as "the app is not installed". For 40 of the 53 catalogue apps that answer is correctly empty and permanent, so both actions refused forever for a running, healthy app — and told the customer to reinstall it "in the same place", which those apps never offer. Separate the two questions. "Installed?" is asked of ListDeployedStacks via a new Manager.isStackDeployed that fails CLOSED on a nil provider. "Where?" is answered by GetAppDrivePath — the same resolver CaptureRecoveryUnit wrote the snapshot with, so the restore aims at the place the backup came from. The 13 drive apps are unchanged: own drive, mismatch check, ack still required. A third refusal, with its own sentence, covers installed-but-no-resolvable-root. Fixtures that marked an app "installed" by giving it an HDD path now state deployment as its own fact. No assertion weakened. |
||
|
|
5ce3a44645 |
v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s
R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label, which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no -p). The old derivation stays as the fallback and an unresolvable attribution is now loud. Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller, not the catalogue. R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside the unit, so the off-site restore had no volume leg at all — proven live with planted files: calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking folder, and the restore reported success without it. For the 40 of 53 apps that declare no data drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay with an explicit directory: ONE implementation, two callers. Volumes replay before the database and inside the stopped window. VolumesReplayed reaches the message. The comment beside the skip was half false and is corrected; the half that still holds — the live unit is the local path's source — is named, and scenario D fingerprints the whole live unit across the operation. Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the code: scenario D passed with the unit guard removed because the fingerprint had been narrowed and was blind to the unit root. |
||
|
|
f94543ee5c |
v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok, go vet clean, -race clean on the changed package - all run and read BEFORE this commit. PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT installed in this case, so there is no own drive to ask) and reads manifest.json plus the captured compose/app.yaml. Local file reads only: no network, no restic, no restore. RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as "what your backup says" would be a fabricated fact. The prefill is labelled as coming from the backup and stays editable: a memory, not a lock. PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40 have none and their data goes to the system drive, which no screen said. Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for the 40-class a recorded placement is a fact to state, never a value written into a field that does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification. PART 4 - measured before theorising, on the live off-site target: snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds. The cause is the shape already on file, so the per-app size calls now run concurrently, BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted. OffsiteInventoryList had no test at all before this. TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a template doing index/eq against an undefined key errors at RENDER time: green build, green vet, green suite, 500 on the page. Four render tests, one per branch, because the existing deploy render test only renders AutoFields and never reaches these blocks. RED-PROOFS, mutation asserted applied then reverted to 0: A three template guards dropped (count asserted 3) -> the blank form returned P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory + what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first. NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit carries no db_dumps and no volume_dumps still reports a bare completion. |
||
|
|
90f2545679 |
fix(disk-health): one physical disk must be evaluated once per run (R-335)
gates / gates (push) Successful in 9s
Found on live hardware two hours after the v0.215.0 deploy, by noticing the release's own positive observable disagreed with its own persisted artefact: the check logged '3 disk(s) evaluated' while disk-health-state.json held two records. demo-hp's c11-scratch and felhom-backup are the same NVMe and share a durable id, so one disk was walked twice per run. Not cosmetic. The loop writes a disk's record before the next entry reads it, so the second copy of an aliased disk consumed the FIRST copy's write as its prior: the disk sustained against ITSELF and reached Hiba on a first sighting, defeating truth-table row 6 — the rule that separates a one-hour benign excursion from a false critical. It would also have emitted two identical events for one drive. Latent on demo-hp only because all counters are zero. Each diskKey is now evaluated once per run. Both entries stay marked seen so neither looks like a disappeared disk, and the card still renders both rows — the dedup is about state and alerts, not display. Red-proof run and reverted: deleting the guard makes the first sighting emit Kind:2 (Hiba-from-sectors) at 8 sectors. |
||
|
|
8144a70a72 |
docs(v0.215.0): CHANGELOG, CONTEXT decisions, README feature, REUSE entries
gates / gates (push) Successful in 10s
- CHANGELOG leads with the severity fix and the live warning-vs-warn proof. - CONTEXT records the settled decisions so they are not re-litigated: Hiba is the label for predicted failure (no fourth word); sustain before count and why; the provenance of 64/55/60; phase 2 owns the new SMART attributes because they are a wire change under G-1; phase 1 state is one record per disk, not a series. - README documents the 14-row ladder, the persisted state, the hourly cadence and the five message shapes. - REUSE pins the severity wire contract on PushEvent — the defect's real home, so the next typo'd severity is caught at the table rather than in production — and records priorFor vs cardPriorFor, which differ by one observation and make the chip disagree with the email if mixed up. |
||
|
|
3ed5e3e770 |
v0.214.0 — the recovery screen stops hedging about a code it can now check (R-311)
gates / gates (push) Successful in 13s
MinAgent: 0.129.0 What was already right: the screen did not bluntly accuse. R-222/R-226 hedged, naming both causes and the kept package, and saying it could not tell them apart. That was honest - and it could not tell them apart because nothing ever looked. Agent v0.129.0 looks, so the hedge becomes an answer. New class RecoveryCodeOpensRetained on HTTP 422, gated by FeatureRetainedRecoveryClass (MinAgent 0.129.0). The gate is SEPARATE from the R-224 one because the two name different agent versions and a box can sit between them, where a 422 is a shape we did not design. ClassifyRecoveryFailure therefore takes both flags; the compiler found every call site. The message says the code is correct, names the supersession date, says the earlier package is kept, and says the CURRENT backups are unaffected - the half a customer will otherwise assume wrong. It promises NO restore: there is no in-product route to a set-aside store (R-312) and the retained package may itself predate the repository-password field. It routes to support, which can do it. The claim guard grew a surface and immediately convicted something. It scanned templates only, while every recovery message is a Go string in a handler - the highest-stakes copy in the product, never scanned. It now scans recovery_handlers.go too, and found a PRE-EXISTING unregistered claim on its first run. Six handler tests asserting which SENTENCE the customer sees; red-proofs asserted applied, including: 422 unconditional makes an agent that never looked read as having looked, and routing 400 to the new class congratulates a mistype. |