2358e561b741b3f08e254b28b29d8b6906dd52a0
152 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2358e561b7 |
R-403: a poorer copy must never delete a richer one
gates / gates (push) Successful in 11s
MEASURED FIRST, then fixed. On the shipped v0.229.0, on demo-hp, an app's Tier-2 copy went from 120 082 104 B (4 database dumps + 3 named-volume tars) to 7 036 B (none of either) in ONE nightly run, and the run recorded itself a success: 'Tier 2 copied docmost -> ... (14.9 KB, 0 leg(s), 0s)'. Evidence: felhom.eu/documentation/audits/DRILL-r403-tier2-delete-2026-08-31/. The mechanism was three individually-correct lines: RunTier2 guards the unit leg with os.Stat only (does the folder exist), rsyncMirror is rsync -a --delete, and nothing between them compared source to destination. An EMPTY unit is a folder that exists. THE GUARD. One predicate, unitCarriesData/unitIsHollow (r403_hollow.go), asking the MANIFEST and never the byte size - a big compose tree with no dumps is dangerous, a tiny unit for a tiny app is fine. Fail closed on an absent or unparseable manifest. RunTier2 skips the unit leg when the source is hollow AND the destination is not; the other legs still run, the run is not failed, and the skip is recorded for the SURFACE (CrossDriveBackup.UnitLegSkipped + UnitPackageDate) as well as logged. --delete STAYS and shrinking stays legal. 07 section 8 row 5's derived-copy rule is unchanged; the fence is exactly one shape. TestR403_DataLegShrinkIsUnaffected is the guard on the guard. THE HONESTY. A preserved package is older than the run that preserved it, so the card carries a notice and the unit-restore confirm names the PACKAGE's date - read from the mirrored manifest's own created_at, not from the status record - plus a clause saying why it is older. THE CAUSE. RestoreTier2Unit now refills a hollow or absent primary unit from the mirror it just restored from, INSIDE the call before returning. The hollow manifest was written two seconds after a restore by the 5-minute capture job; any follow-up job races it. The capture itself is NOT guarded: a capture describing an empty drive as empty is correct, and with the primary refilled there is no hollow state left to describe. Never over a complete primary, never after a failed restore. recordTier2Success and tier2UnitConfirmMsg keep their old signatures as thin callers, so no existing test needed editing. New seam unitRehydrate, separate from tier2Mirror on purpose. 22 new Go tests. Red-proofs run and reverted: A6 (predicate -> size threshold), B1 (guard removed -> the copy's 3 files are DELETED and the seam is called), B6 (a general never-shrink rule -> the shrink case fails), C2 (only-when-hollow dropped -> the complete primary is overwritten). |
||
|
|
4c8f0d2919 |
R-103: the Tier-2 refusal becomes an action
gates / gates (push) Successful in 12s
An app whose Tier-2 copy holds no file legs but a full recovery-unit mirror - 45 of the 53 catalog templates - was told to press a button on a DIFFERENT page. Since R-102 the data it is asking for is restorable from the copy it is looking at. New POST /backup/tier2/unit-restore and backupTier2UnitRestoreHandler: same guards, same restoreOpBlocked() refusal (R-351b), same async shape as the file restore beside it, plus a fail-closed pre-flight so the app is never stopped for a mirror that could not be opened. The outcome reuses unitRestoreOutcomeMsg and adds which copy overwrote the live data. The row offers the action where the refusal was, in a danger style, as a SEPARATE button. The two are not merged: one adds what is missing, the other overwrites. The confirm carries that difference in words and names the copy's date - and says so differently when that date is only an ATTEMPT (R-101). It is built from named Go constants rather than assembled inside an HTML attribute, so a test can assert it verbatim; fmtTimeStr now delegates to a package-level fmtRFC3339Local so the confirm and the outcome cannot render the same date two ways. tier2NoCoverageMsg is NARROWED to the case that remains - no legs and no openable unit - and still names the route that works. tier2UnitNotCoveredMsg is NOT deleted: it is appended where the FILE restore ran and is still exactly true of it. Tests C1-C2 and D1-D6 plus four more. Red-proofs: C1 (widen CanRestore to include HasUnit -> the unit-only cases fail), D6 (drop EndRestoreOp from the handler goroutine -> 'the restore never published a result'). |
||
|
|
0f9b796615 |
R-102: the recovery unit on the second drive becomes a way back
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any action in the product (07-backup-architecture 6.3, 7.2). Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched. reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second one invented, which is what lets the acceptance test assert the volume leg's source directory. Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken inside RestoreFromRecoveryUnitAt, not beside it. Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which refused 40 running apps for months. Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'. Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the primary -> fails with the mirror never reaching the redeploy, and with permission denied once the primary tree is unreadable). |
||
|
|
c732006d26 |
R-102 phase 1.1: split the recovery-unit path helpers, zero behaviour change
Every unit path helper took (nsRoot, stackName) and joined backups/primary/<stack>/... . That hard-coded 'primary' is the mechanism of R-102: Tier-2 mirrors the whole unit directory to <dest>/backups/secondary/<stack>/recovery-unit/ every night, and because no reader could NAME a unit outside backups/primary/, that mirror has been captured for months and read by nothing. Adds four unit-directory-relative primitives - UnitComposeDir, UnitManifestFile, UnitDBDumpDir, UnitVolumeDumpDir - each taking the recovery-unit DIRECTORY itself. The four existing (nsRoot, stackName) helpers become thin wrappers over them and keep their exact signatures and their exact return values; every current caller compiles untouched. ONE implementation, two callers - the rule restoreDockerVolumesFrom already follows in this repo. TestR102_PathWrappersAreByteIdenticalToToday pins the wrappers against hand-written literals (not re-derived from the helpers under test). Red-proof: UnitComposeDir join changed to 'compose2' -> the test fails on all three fixtures. |
||
|
|
3c49dc8ea4 |
v0.228.0 — the off-site check reads the data; the debug page stops lying (R-399 + R-400)
gates / gates (push) Successful in 12s
R-399: monitoring.integrity.read_data_subset defaults to 100%. A pack damaged without changing its size made plain `restic check` report "no errors were found" on demo-hp 2026-08-30; every read-data form caught it. Cost on that 134 MB store: 35.0s structure vs 39.2s at 100%. "off" (any case) is the off token; empty means not-configured, therefore the default; a malformed value falls back to the DEFAULT, never to structure. A completed check over 5 minutes logs a WARN naming the duration, the depth and R-401 — operator log only, no hub event, no depth change. The depth is now recorded with the verdict (LastIntegrityDepth; empty = NOT RECORDED, never "structure"). R-400: 24 debug-page references, 17 dispatched, 7 dead — three of which fetched on page LOAD, so those panels were permanently blank. backup/crossdrive implemented; backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both storage/simulate-* deleted with their panels and JavaScript. scripts/debug_route_gate.py fails in both directions and is registered after the seven were resolved. 18 referenced, 18 dispatched, none orphaned. Corrections: the dead-field warning in report/types.go said the controller runs no integrity check and the notifiers are called from nowhere — both false since v0.227.0. controller.yaml.example gains its missing integrity: block. integrityCheckTimeout's "ships OFF" comment rewritten. |
||
|
|
45770f2282 |
v0.227.1: the damage classifier matched restic's ordinary progress output
gates / gates (push) Successful in 11s
A patch and not a rebuilt 0.227.0: that tag was already running on demo-hp, and re-pushing changed bytes under a live tag is the :latest hazard with extra steps. looksLikeRepositoryDamage matched bare "pack ", "tree ", "snapshot ", "blob ". A HEALTHY restic check prints "check all packs" and "check snapshots, trees and blobs" -- so any check that failed for a NON-damage reason, a connection dropped mid-run for instance, would have been classified as a corrupted repository and told the customer their backups may be damaged. That is the false alarm that teaches an operator to ignore the true one. Caught by the NEGATIVE control, built from the real bytes of a real passing check on demo-hp. The spec made the negative control mandatory and this is what it was for: a control that has only ever seen the failing case proves nothing. Signatures are now phrases from restic's own error wording. Also in this commit: CONTEXT.md records the three rulings (take the flag and skip, due-ness not a weekday, publish on OffboxReportStatus not the R-331 dead fields) plus the measurement a future session would otherwise assume wrongly -- THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. README documents the job, the route and the config, and corrects a line that listed four debug backup routes when only two exist. REUSE gains three rows, including one that records R-398 was my own mistake so nobody re-files it. |
||
|
|
0d52a42c17 |
R-359 + R-397: the off-site store gets checked, and the advertised check becomes real
gates / gates (push) Successful in 12s
Nothing ever verified that the off-site copies are still readable. The whole-guest tier has verify jobs; the tier holding the customer's documents and photos had none -- the complete set of restic verbs this controller used contained no `check`. We would have found out at restore time, with a customer waiting. On 2026-08-21 a deliberately damaged pack was caught at once by plain `restic check`; we had never run it. R-397: NotifyIntegrityOK/NotifyIntegrityFailed existed with no caller, the hub allowlists both event types and carries the Hungarian text for both, the settings checkbox exists, and the debug button posts to /api/debug/backup/ integrity. Everything was built except the part that runs. SIXTH instance of that shape in this project. THE HAZARD SHAPES THE WHOLE DESIGN. resticStep self-heals a crash lock by running `unlock --remove-all` and retrying, and its own comment records why that is safe: every caller holds the in-process single-flight mutex, so any lock it meets is stale. A check that did not take that flag could meet a LIVE prune's lock from this same box, remove it, and retry over the top of it. So the check TAKES THE FLAG and SKIPS rather than waits -- waiting would pin the nightly backup behind it, and a skip costs nothing because due-ness makes tomorrow try again. TestR359_SkipsWhenRunningFlagHeld asserts the NON-EFFECTS: restic never invoked, `unlock` never in any argv. Its red-proof prints the real thing -- restic running `check` while the flag was held. DUE-NESS, NOT A WEEKDAY. Daily job, weekly behaviour: "is the last successful check older than 7 days?" not "is it Sunday?". R-341 is exactly the other shape, a dated check quietly missed and never caught up. No Weekly primitive added. THREE OUTCOMES, NOT TWO. Skipped, Unreachable and failed are different facts. "I could not look" is not "I looked and it is broken" -- R-339 already owns reachability, and a second alarm for the same fact trains the operator to discount the one alarm that means the backups are damaged. A timeout is unreachable, never damage. A failure advances due-ness (a broken store must not be re-checked nightly); a skip and an unreachable store do not. Success is severity `info`, which severityNotifies DROPS -- it mails NOBODY, by design. A weekly success e-mail is how people stop reading their alerts. The customer gets a SENTENCE; restic's words go to the log, truncated (R-379: 615 bytes of raw database text reached a customer once). read-data-subset ships OFF and a malformed value is refused at read time rather than handed to restic, where one typo would fail the whole check. Published on OffboxReportStatus, NOT on report.BackupReport's IntegrityOK -- those were retired by R-331 YESTERDAY and TestBackupReport_DeadFieldsStayZero still passes unmodified. Also: the monitoring page stopped promising a Sunday job that never existed, and the debug button got its dispatch case. PART 0 WAS NOT BUILT, AND R-398 WAS MY OWN MISTAKE. The seam it asked for already exists: offboxRunner/SetOffboxRunner/m.runner() has been injectable since the off-site tier shipped, and other tests drive restic-backed paths through it. A resticStepFn seam would have been WORSE here -- it would replace the `unlock --remove-all` escalation and hide it from the assertions that must see it. R-358's AST ordering test is converted to a real execution test instead, which immediately surfaced something the AST walk could not: unlockStale legitimately runs before the restore. Four red-proofs, each printing the pre-fix behaviour. Green gate: 28 packages, rc 0. All 12 controller gates OK. |
||
|
|
c0c8fe67bf |
An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero result, which is honest" exposed that the sentence was FALSE. A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's fallback to RestoreApp -- which returns only an error, and whose signature is deliberately out of scope -- would have printed "ez a mentes csak a beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure direction this whole change exists to remove, re-introduced by the change. UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a fourth sentence claiming only what is known -- the restore ran, the app is back, and we cannot say what came back. RestoreApp's signature is untouched. Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam test was corrected too: its fixture has no recovery unit, so it exercises exactly this path and had been asserting the wrong sentence -- it now asserts the unknown, which is what pins the fallback to it. IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate written to stop findings dying in an overwritten REPORT.md caught a live defect instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the monitoring page advertises a weekly integrity check that does not exist) and R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST test) rather than leaving them in a file that is overwritten every session. REPORT.md is the full run record: baselines re-confirmed, per-test results, the five red-proofs with their observed output, the live validation with verbatim Hungarian messages, what was NOT validated and why, teardown across three layers, and the register 165 -> 167 -> 161. Green gate clean: 28 packages, rc 0. All 12 controller gates OK. |
||
|
|
b8af72764d |
R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21 backup-truth drill, all still in shipped code. They share one acceptance idea: a restore surface must state what it actually did, and must refuse what it cannot do. VERSION NOTE. The task specifying this targeted v0.224.0 against baseline |
||
|
|
e5eee501b5 |
R-331 (controller half): forward stats_known so the hub can tell empty from unmeasured (v0.225.0)
gates / gates (push) Successful in 12s
The hub's operator Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` for EVERY customer, because it rendered the report's `backup` object -- whose snapshot/size/integrity fields have had NO producer since disk-tier restic moved to the host agent (slice 8C). buildBackupReport leaves them zero deliberately and says so. Measured on demo-hp 2026-08-30 while that night's log said `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`. The live numbers were always in the report's `offsite` object, which the hub already reads for its Offsite page and its fill/staleness alarms. The hub fix is to render that -- and that made exactly ONE field mandatory that was not being forwarded. snapshot_count:0 means two opposite things: "holds nothing" and "never measured". R-225 measured that confusion inside this repo (a rebuilt box rendered 0 pillanatkep over a store really holding snapshot f3d9cd67), and settings.OffboxTarget.StatsKnown fixed it for the controller's own UI. It was never put on the wire, so the hub was free to make the identical mistake one layer up -- and did. OffboxReportStatus.StatsKnown now carries it, omitempty, so an older controller sends no key and a reader degrades to UNKNOWN, never to EMPTY. Absence is ignorance, not emptiness. The four dead BackupReport fields stay on the wire (historical reports in the hub store must keep parsing) but now carry a warning naming R-331 and pointing at Offsite. TestBackupReport_DeadFieldsStayZero fails the moment a producer appears for one -- the prompt to update the hub card in the SAME change rather than ship a field nothing renders. RED-PROOF: drop `StatsKnown: t.StatsKnown` -> "a MEASURED empty repository reported stats_known=<nil>". Tests assert the JSON the hub sees, not the Go struct: measured-empty and never-measured must differ ON THE WIRE, which is the entire point of the field. Green gate clean: 28 packages, rc 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
92cebb8c95 |
R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.
The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.
The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.
It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
every failure path, so the app alarms on the next scan;
2. Begin REPLACES the set (one marker file = one operation);
3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.
Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
- drop markStopped from Begin -> "suppressed at stop = map[]"
- drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
- pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.
Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
5da11c4480 |
v0.222.0: ask whether a supervised member is DEAD before whether one is UNHEALTHY (R-384), and stop promising an undo copy nobody looked for (R-383)
gates / gates (push) Successful in 11s
R-384. aggregateState returned StateUnhealthy the moment unhealthy > 0, and the R-51 mixed-case block that asks "is a supervised member dead?" sat below it. A two-container app whose database exits goes unhealthy BECAUSE it cannot reach that database - so the symptom the dead database causes was what suppressed the alarm for it. unhealthy is not a down state, so classifyRunStates never marked the app down and app_start_failed never fired. Measured live on demo-hp 2026-08-22: bookstack-db stopped at 21:27:01 and the F-OBS heartbeat printed "0 currently down" throughout. R-51's 18-hour immich failure, back through a different door. Two things moved, and either alone leaves the defect standing: the supervised test is hoisted above the unhealthy/starting/restarting returns, and "some members are up" now counts ANY member not in the down bucket. The old guard was running > 0, which made the R-51 block unreachable in exactly the case it was written for. IsDownState is byte-identical - unhealthy stays excluded, because an unhealthy container is running and folding it in reintroduces the flapping that exclusion exists to stop. No new state was minted. Only the ORDER changed. The priority comment was rewritten because it asserted an ordering the code no longer has. Three subtests in TestAggregateState_UnchangedBranches were AMENDED: they asserted an unhealthy/starting/restarting member beat an exited peer on unless-stopped, which pinned the defect as settled behaviour. They keep their intent with the down member given a benign policy. R-383. The double-failure message said the previous state's backup EXISTS, built from the returned path without asking the filesystem - and a missing file is one of the two ways that rollback fails. undoCopyPhrase now describes the copy from disk: present, partial, missing (still naming where it should be), or never written. Zero-length counts as missing. Test count 1494 -> 1504. Four red-proofs planted, four seen failing; the two halves of R-384 convict independently. |
||
|
|
810b18ab8e |
R-361 follow-on: the undo-copy prune must run above the already-current check
gates / gates (push) Successful in 11s
Excluding pre-restore-* from db_dumps made that list stable across restores, so CaptureRecoveryUnit's already-current early return began firing where it never had - and the prune, which sat after it, stopped running in exactly the case it exists for. Measured on demo-hp minutes after the change: four undo copies on disk against a cap of three. The prune is housekeeping on the dump directory and is independent of whether the manifest needs rewriting, so it belongs above the check. Pruning cannot disturb dbDumps, which no longer contains those names. Pinned by TestR361_UndoCapHoldsWhenTheUnitIsAlreadyCurrent; its red-proof moves the call back below the return and the cap fails at 5. |
||
|
|
968c968559 |
R-361: the safety dump destroyed the app's own database backup
gates / gates (push) Successful in 11s
writeSafetyDump called DumpOne into the app's OWN unit dir and renamed the result to pre-restore-* afterwards. DumpOne writes <stack>-<dbtype>.sql - the app's canonical dump - so every safety dump overwrote the app's real backup and then moved it away, leaving the app with no database backup until the next nightly run. A local restore-from-unit in that window tells the customer the app never had a database. The comment beside it asserted the rename meant it 'can never overwrite the app's real dump'. False as written, and believed for four months. Measured live before the fix: docmost and bookstack each held only pre-restore-* files and no canonical dump. DumpOneTo takes the final path and derives its own .tmp from it. DumpOne keeps its signature and calls it with the canonical name. writeSafetyDump asks for its own name directly; the rename is gone; the comment now states the invariant and how it is enforced. db_dumps no longer lists the undo copies. All three consumers of Manifest.DBDumps were grepped and named - all inside recovery_unit.go, none reads it for recovery. The files are neither deleted nor hidden. Tests 1485 -> 1493. FIVE red-proofs, TWO PASSED first time and both are reported: the behavioural tests inject the dump seam so a mutation inside DumpOneTo was invisible, and 1.3 had no test at all. Guards added at the layer each defect lives in; both mutations then convicted. |
||
|
|
5b52a5964d |
R-379 fix: the rollback must re-discover the DB container
gates / gates (push) Successful in 12s
Found by v0.220.0's own live walk on its first real run. writeSafetyDump
captures its DiscoveredDB before the stop; the DB-only start then re-creates the
container with a new id, so the rollback's docker exec hit a dead container and
sat in waitDBReady for 30s. The app was held for an infrastructure reason while
its data was recoverable.
Re-discover and match by {stack, engine} - what reimportDBDumpsFrom already did.
Fail closed when the container cannot be found.
No unit test caught it because they all inject the import seam and never look at
container identity. The new test asserts the identity handed to the import.
|
||
|
|
2c724c9283 |
R-379/R-380: put the customer's undo copy back when a database restore fails
gates / gates (push) Successful in 11s
R-379 and R-380 were one failure. Both ended with a half-restored database; the only difference was whether it looked broken. Postgres emptied and crash-looped; MariaDB applied part of the dump and reported health=healthy with a zero-row schema-version table. Measured live on demo-hp 2026-08-22. The undo copy was already taken and already good - proven by hand that day on both engines. Nothing in the product could apply it. Now it does, with the same ImportDump call, before any restart and inside the DB-only window. The WHOLE undo set, matched on this run's stamp. writeSafetyDump returned one path for an app with two databases; a rollback on that would restore one and leave the other half-written. When the rollback also fails the app is HELD STOPPED (operator ruling): a running app on a half-written database lets the customer make the damage permanent. Every start path refuses it - customer button, appstop Recover, boot sweep - via the shared driveStartGate, checked ABOVE its driveless early return because these apps have no drive. The marker is ended so nothing auto-restarts it. The row goes red. Cleared with --clear-restore-hold, an operator CLI route. --single-transaction is a belt on Postgres only; MariaDB DDL is not transactional and that is why the rollback is the fix. R-381: the engine's stderr stops reaching the customer (615 bytes on MariaDB, its middle rows out of their own database) and starts reaching the operator log, which never had it. R-382: the summary log prints the volume count it already held. Undo copies resolve to their own app, are marked IsUndo, and are capped at 3 per app, pruned from the capture side. The reported render-as-an-app symptom did NOT reproduce - the live page was read first and had zero occurrences. Tests 1468 -> 1483. Eight red-proofs; ONE PASSED and is reported: the R-381 behavioural test injected below ImportDump. A guard at that layer now convicts. |
||
|
|
08eb1a6e3a |
R-356: the off-site restore refused every app that has no data drive
gates / gates (push) Failing after 12s
ReconstituteFromOffsite and PlaceOffsiteRestore both resolved the restore destination with the RAW HDD_PATH and read an empty answer as "the app is not installed". For 40 of the 53 catalogue apps that answer is correctly empty and permanent, so both actions refused forever for a running, healthy app — and told the customer to reinstall it "in the same place", which those apps never offer. Separate the two questions. "Installed?" is asked of ListDeployedStacks via a new Manager.isStackDeployed that fails CLOSED on a nil provider. "Where?" is answered by GetAppDrivePath — the same resolver CaptureRecoveryUnit wrote the snapshot with, so the restore aims at the place the backup came from. The 13 drive apps are unchanged: own drive, mismatch check, ack still required. A third refusal, with its own sentence, covers installed-but-no-resolvable-root. Fixtures that marked an app "installed" by giving it an HDD path now state deployment as its own fact. No assertion weakened. |
||
|
|
5ce3a44645 |
v0.218.0: attribute a DB container by its compose project, and replay volumes on the off-site restore
gates / gates (push) Successful in 11s
R-355 (first, because it is the only one where data can be lost for good). paperless-ngx's PostgreSQL was dumped into backups/primary/paperless/db-dumps/ — a directory for a stack that does not exist, on the system drive — while the app's own unit recorded db_dumps: null. The same misattribution reached writeSafetyDump, so a destructive restore of that app took NO undo copy and the fail-closed refusal was never reached. Fixed by reading the compose project label, which is the stack name by construction (compose runs with cmd.Dir set to the stack dir and no -p). The old derivation stays as the fallback and an unresolvable attribution is now loud. Catalogue sweep, proven able to convict: one affected app of 53. The fix is in the controller, not the catalogue. R-354. ReconstituteFromOffsite skipped every unit placement and the volume archives live inside the unit, so the off-site restore had no volume leg at all — proven live with planted files: calibre-web's 1,422,848-byte config archive was in the unit, the snapshot and the checking folder, and the restore reported success without it. For the 40 of 53 apps that declare no data drive that archive is the whole dataset. restoreDockerVolumesFrom is the local path's own replay with an explicit directory: ONE implementation, two callers. Volumes replay before the database and inside the stopped window. VolumesReplayed reaches the message. The comment beside the skip was half false and is corrected; the half that still holds — the live unit is the local path's source — is named, and scenario D fingerprints the whole live unit across the operation. Seven red-proofs, each asserted applied and reverted. Two found defects in the tests, not the code: scenario D passed with the unit guard removed because the fingerprint had been narrowed and was blind to the unit root. |
||
|
|
f94543ee5c |
v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok, go vet clean, -race clean on the changed package - all run and read BEFORE this commit. PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT installed in this case, so there is no own drive to ask) and reads manifest.json plus the captured compose/app.yaml. Local file reads only: no network, no restic, no restore. RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as "what your backup says" would be a fabricated fact. The prefill is labelled as coming from the backup and stays editable: a memory, not a lock. PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40 have none and their data goes to the system drive, which no screen said. Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for the 40-class a recorded placement is a fact to state, never a value written into a field that does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification. PART 4 - measured before theorising, on the live off-site target: snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds. The cause is the shape already on file, so the per-app size calls now run concurrently, BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted. OffsiteInventoryList had no test at all before this. TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a template doing index/eq against an undefined key errors at RENDER time: green build, green vet, green suite, 500 on the page. Four render tests, one per branch, because the existing deploy render test only renders AutoFields and never reaches these blocks. RED-PROOFS, mutation asserted applied then reverted to 0: A three template guards dropped (count asserted 3) -> the blank form returned P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory + what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first. NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit carries no db_dumps and no volume_dumps still reports a bare completion. |
||
|
|
985388c6e9 |
R-351: the restore compares where the backup says the data lived; second press cannot start a second run
gates / gates (push) Successful in 10s
Part 3 (not droppable) and the engine half of Part 2. No version bump yet - one bump and
one bake at the end of the session.
PART 3a - a second press really did start a second run. Established with a test BEFORE any
change: both offboxReconstituteHandler and offboxPlaceHandler answered "...elindult" and
overwrote the first restore's op/stack. Cause: every restore handler gated on
backupMgr.IsRunning() - the CONCURRENCY flag, which the restore goroutine acquires AFTER the
handler returns (offbox_reconstitute.go:180, offbox_restore.go:393). Seven sites. The wizard
had read the correct flag since v0.154.0 and said so in a comment; the handlers never moved.
New Server.restoreOpBlocked() reads BOTH flags - the display flag covers the whole off-box
restore, the concurrency flag is the only one the nightly backup holds - and the refusal now
names the running app and a route.
PART 3b - the page DOES refresh; the defect was the RESULT. backups_shared.html gated the
terminal result on a page-local sawRunning flag, so a restore that finished before the page
was opened, or inside one 3s poll, was shown to nobody. The 2026-08-21 OpenGist restore took
8.666s and no screen ever said it completed - the answer existed only in docker logs.
RestoreOpStatus.LastRecent now carries the server's verdict. The 10-minute window moved to
internal/backup as RestoreResultWindow and internal/web's constant is an alias: one
expression, two surfaces. Also removed the wizard's self-contradiction, which said the state
refreshes automatically AND that you must refresh the page.
PART 2 (engine) - every recovery unit manifest has carried drive and namespace_root since
schema 1, and NO non-test code read either back. The reconstitution opened the manifest and
took only the coherence stamp, then resolved its destination from the live app. A restore
into a different destination succeeded silently under a green message. New
backup/offbox_placement.go: CheckPlacement (pure, total), PlacementMismatchMessage,
recordedPlacementFromScratch. Compared before the safety dump and before the first byte.
A mismatch is NAMED and refused; ackPlacementChange lets the customer proceed deliberately -
a separate field from confirm=1, because one click must not carry two decisions. An UNKNOWN
recording is never a mismatch: refusing on an absence would strand every pre-field unit.
The not-installed refusal (R-253) now names the drive the backup recorded.
RED-PROOFS, each mutation asserted applied and reverted to 0:
B both guards removed (count asserted 2) -> the restore WAS seen starting with no drive
attached: no error, full 3.00s run, wrote into /tmp/mutant-destination
C Mismatch forced false -> the silent divergent restore returned
E Known() forced true -> the fabricated empty prefill appeared
D Mismatch forced true -> 8 ordinary reconstitute tests broke, proving reachability both ways
Note on D: the existing fixtures write a schema-1 manifest with NO drive, so they are
scenario-E shaped. The matching case is covered in the scenario table, not by them.
Gates 11/11 OK. Suite 28 packages ok. Hungarian verified as hex, no BOM, no mojibake sentinels.
NOT in this commit, still open: Part 2's scenario-A prefill UI, Part 1's deploy-page
visibility line, Part 1's specification document, Part 4's measurement.
|
||
|
|
89712563a0 |
R-302: the abandon banner promises only what the box can still see is true
gates / gates (push) Successful in 10s
The retrieval clause rendered unconditionally on every page and is false on a reachable state - the same screen where the orphan card says we cannot tell. The condition is a fingerprint PINNED at the decision, not a comparison against the current key. The obvious proxy asks about the wrong key: the set-aside copies were written under an older key the box no longer has, so on a twice-rebuilt box the proxy promises about copies nothing can open. Demonstrated - under the proxy, the replaced-package and legacy cases both flip back to promising. The pin is a recorded assumption and says so: nothing on the box records which key wrote those copies. Empty is not a match. A countdown started before this carries no pin and takes the cautious branch, not a backfill. A sweep of all 36 templates found a fourth instance (backups page, same condition applied) and a fifth (the confirmation screen, correctly left alone - true at the moment of the decision). New retrieval_promise_gate registers each claim with a reason rather than banning a verb: a string ban failed twice, and the honest replacement copy contains the stem. |
||
|
|
8dbbc98ff2 |
v0.207.0 — R-249: the retrieval passphrase leaves the page body; R-252/R-253: two refusals learn to say what to do
gates / gates (push) Successful in 18s
R-249. settings_security.html rendered the passphrase into a display:none span behind a Megjelenit button. That toggle stops a browser DRAWING the value and nothing else — the plaintext was in the response body of every render, so a curl of the page returned it. Found by exactly that: it landed in a session transcript while driving the documented rebuild path. The codebase already stated this rule for the recovery code and this page did not follow it (escrow_handlers.go: 'reveal (claim XHR only — R is NEVER templated server-side into HTML)'). The page now carries only HasRetrievalPassword; the value comes from POST /settings/retrieval-password/reveal — CSRF-covered because POST, no-store, and LOGGED as an act, which reading it off the markup never was. The tests assert the RAW RESPONSE BODY. Every test that asked what the customer sees passed while the bytes carried the secret; that is why this survived. Census: the render-then-hide pattern appears twice more — app_info.html (a real per-install app password in a hidden span) and deploy.html. Filed as R-254, NOT fixed here. R-252. A rebuilt box keeps its drives but loses their REGISTRATION. The restore page now states that before the customer presses anything, says the backups and drives are both still there, and links to Tarhely > Meghajtok. Page and resolver ask ONE question — HasRestoreDestination() reads the same GetSchedulableStoragePaths() the scratch resolver reads. R-253. The list promised 'a visszaallitas elobb ujratelepiti' three lines above a refusal that fired BECAUSE the app was not installed. The promise was the wrong half: reconstitution writes to the app's own GetStackHDDPath, which exists only once the CUSTOMER has chosen a drive at deploy time. Auto-reinstalling would mean the product making that choice for them. Copy now says to install first and routes to /stacks/<app>/deploy. Both notices are conditional — a healthy box renders as before, pinned by a test that fails if either becomes unconditional. |
||
|
|
72368654e4 |
R-241 part 5: escalating reminders, and operator levers for a running countdown
REMINDERS (SEC 2.3). The offer epoch now stamps when it began, and the undecided reminder escalates in EMPHASIS at 1, 3, 7 and 14 days. THE READING IS STATED BECAUSE THE SPEC IS AMBIGUOUS, and it is written into the code where it can be corrected. For an ABANDONING box, 5/3/1 are unambiguously days REMAINING before a deletion. An undecided box has no deadline - nothing counts down to anything, because SEC 7.5 deliberately does NOT auto-abandon - so 14/7/3/1 cannot be "remaining" and are taken as days ELAPSED, with the wording firming up rather than the bar appearing and disappearing. If the operator meant something else, one function changes. The stamp is re-set on every entry into the offered state, so a box that settles and is later rebuilt starts its ladder again instead of inheriting an old one. OPERATOR LEVERS (SEC 7.5). --abandon-status, --abandon-extend=N and --abandon-stop on the controller CLI, beside the existing operator subcommands. They exist because the path that ACTUALLY happens is the customer telephoning, and support needs something to press. They live on the CLI and not in the customer UI deliberately: extending a deletion the customer asked for is an operator judgement, and a customer who wants it stopped already has the self-service route - they recover with their code, which cancels it. BOTH REFUSE RATHER THAN NO-OP, in two situations: when no countdown is running, and when the store has already been deleted. A silent success is the thing an operator most easily mistakes for "handled" - they would tell the customer their data was safe when it is gone. Pinned by two tests. --abandon-extend counts from NOW, not from the old due date, and a test proves the old date passes without deleting anything. Green: go build, go vet, go test ./... all pass; controller gates OK. |
||
|
|
de39e47f53 |
R-241 part 4: the three-state surface, and the copy tells the truth about the date
FULL PAGE ONCE PER ENTRY, NOT ONCE EVER. "Most nem" used to set a flag that
nothing ever cleared, so a box that abandoned its history and was rebuilt
months later - a genuinely NEW situation - would never see the page again. The
offer now carries an EPOCH, advanced on the edge into the offered state, and a
dismissal is recorded against the epoch it was made in. A fresh entry passes
the dismissal by arithmetic, with nothing to clear and nothing that can be
forgotten to clear.
That is NOT the flag the operator's ruling forbids. The forbidden thing
remembers that the customer decided so the screen can be suppressed while the
state stays wrong. This records WHICH SITUATION a dismissal was about.
A REAL BUG, caught by the test and not by review: the first draft returned
early from recoveryInterrupts when the offer was false, so the FALLING edge
was never recorded, RecoveryOfferActive stayed true through a settled period,
and the next entry counted as a continuation. The page never came back - the
exact defect the epoch exists to fix, reintroduced inside the fix. The sync is
now unconditional and the ordering is commented as load-bearing.
THREE LEVERS, THREE SCOPES, and none of them removes the route:
- clicking the bar away -> a browser SESSION cookie, cleared on login, so
the reminder is genuinely back at the next login. Nothing persisted.
- "ne emlekeztessen ujra" -> durable, epoch-scoped, silences the BANNER ONLY.
It starts no countdown, abandons nothing, and a fresh entry reminds again.
- "most nem" -> suppresses the full page only, as before.
The entry point on /backups/remote is bound to the OFFER and to nothing else,
pinned by a test that fires all three dismissals and asserts it survives.
SEC 7.3 / Q7 - THE TRAP DOES NOT SURVIVE THIS SESSION. While a recovery is
outstanding the "Helyrealitasi kod letrehozasa" button is UNAVAILABLE, not
merely captioned: creating a new code seals the current key, demotes the
package that opens the earlier history to retained custody that no shipped
path can read (R-199), and re-enables the recovery screen through the orphan
route while invalidating the code that screen accepts. A warning beside a
button is a warning people click past. The card now explains and points at
/recovery instead.
SEC 2.4 - the abandon confirmation changes with the behaviour. It used to
promise "felretesszuk - nem toroljuk". It now states the grace in days (from
the constant the countdown actually uses, never a literal in prose), that the
sealed package goes with it, that the customer can change their mind, where
the date is visible, and that the question does not come back afterwards.
The countdown is shown on /backups/remote for the WHOLE window - the bar
elsewhere is a nudge, this is the record, and a deletion date must be findable
on a quiet day too.
Tests: once-per-entry across a full settle-and-re-enter cycle; the banner
dismissal proven to be a session cookie (MaxAge 0, no Expires) and to persist
nothing; the opt-out proven to silence the banner while leaving the offer, the
route and the countdown untouched, and to remind again on a fresh entry; the
entry point surviving all three dismissals; a settled box showing nothing; and
the back-redirect refusing "//evil.example".
An existing test (TestRecovery_E) was updated: it asserted the legacy boolean,
which the epoch replaces. It now asserts the dismissal landed on the current
epoch, which is the stronger property.
Green: go build, go vet, go test ./... all pass; controller gates OK.
|
||
|
|
a5d90ff801 |
R-241 part 3: abandoning starts a 14-day countdown that ends the question
Until now "set aside" renamed the remote store and touched neither the escrow
nor the key, so the hub went on holding a sealed package for a key the box no
longer used. Shape (c) compares those two, finds them different, and offers
recovery - correctly, and for ever. A customer who had already said "I do not
want the old data" would be asked again at every login.
The operator's ruling is that the answer is NOT a "they decided" flag: fix the
state, do not remember that it is wrong. So the decision starts a countdown,
at the end of which the set-aside store and the sealed package that protects
it are removed TOGETHER. Afterwards shape (c) has nothing to compare and the
offer falls silent on its own - because the state is right, not because
something remembers it once was not.
THE GRACE IS REAL. The recovery offer stays reachable for the whole 14 days;
that is the change-of-mind path, and a grace in which recovery is impossible
would be decorative.
BOTH HALVES OR NEITHER. Removing only the store leaves a package that opens
nothing; removing only the package leaves ciphertext nobody can ever decrypt.
The two cannot be atomic across two machines, so it is a two-phase commit:
delete the store, record a durable marker, and keep DECLARING
offsite.abandon_purge_requested until the hub's ACK stops reporting a
superseded package. A crash between the halves re-declares on the next sweep;
it never leaves the pair half-removed and silent.
HUB HALF - SEC 8.2 ANSWERED: yes, the hub was needed, and only for this.
store.PurgeSupersededEscrowForCustomer is the one place R-198's retention is
ever undone, and it never touches host_escrow (the package covering the key
the box uses now). The handler acts on the DECLARATION, never an inference,
and is placed immediately BEFORE the ACK is built - so
GetEscrowStatusForCustomer reads the effect and the SAME response closes the
box's two-phase commit. No second round-trip and no window where the box
thinks it is still owed. felhom-agent was NOT touched.
The countdown starts in ResetOrphanedRepo, NOT in the shared helper: the
helper is also the unclaimed auto-reset path, where nobody decided anything,
and an as-delivered box tidying a stranger's leftover store must not get a
customer's deletion clock. Pinned by a test.
Cancellation is wired into the recovery unlock, BEFORE the tier-up and the
listing - those can fail, and a countdown surviving a successful unlock
because a later step errored would delete the history the customer just
proved they can open.
The sweep is a Daily job at 05:10, not on the backup leg: it must run on a box
whose tier is not configured for runs. Quiet by construction on every box with
no countdown, and that silence is asserted.
Tests (all clock-injected; SEC 7.4 forbids shortening a live timer):
Scenario E (aside + package kept + countdown + offer still reachable, and
NOTHING deleted), Scenario F (both halves, the declaration repeating, the
close-out), Scenario G (cancel, path still nameable, no later deletion),
plus: not closed out while the package remains, a transport failure leaves the
countdown due and retrying, the no-op sweep issues zero remote commands, and
the unclaimed auto-reset starts no countdown.
RED-PROOFS, each with the mutation confirmed present in the file first:
F1) store deletion skipped -> Scenario F FAILS (no rm issued)
F2) declaration dropped from the report -> Scenario F FAILS (the hub is
never asked; the package would outlive the store for ever)
G) CancelAbandon made a no-op -> Scenario G FAILS (uncancellable countdown)
Green: controller and hub both build, vet and test clean; controller gates OK.
NOTHING WAS DELETED ANYWHERE - the terminal step has only ever run against
in-test fakes.
|
||
|
|
a491abef6c |
R-241 part 2: the comparison the box already makes becomes the thing that offers recovery
THE FACT WAS COMPUTED EVERY CYCLE AND KEPT NOWHERE. EscrowAutoConfirmer.Reconcile
has compared the hub's restic_pw_sha256 against the local key on every ACK since
SLICE 3. On the final-walk venue it logged, at 03:28:03Z and thirty-five minutes
before the customer looked, "the hub's escrow blob does not cover the CURRENT repo
password (hub hash 30ef574f != local 9b4a9a9d)" - and dropped it. The recovery
screen, evaluating in the same process, went on asking a question that could not
see it.
Now persisted: settings.HubEscrowKeySHA256 + HubEscrowKeyCheckedAt, recorded
UNCONDITIONALLY in Reconcile beside RecordPresence and RecordSuperseded - same
place, same reason: the box that needs it most is the rebuilt one with no target,
on which every gate below returns early.
OffsiteRecoveryOffer gains SHAPE (c): the hub holds a package for a key OTHER than
the one we are using. (a) and (b) are both proxies for that question and both have
now been wrong in opposite directions - (a) goes false the moment anything mints,
(b) is unreachable while the escrow is pending.
SEC 7.2, decided deliberately and stated in the code:
- a KNOWN DIFFERENCE offers, however old the reading. Age is not gated on. Both
sides are local; only the hub's half can be stale, and what the hub holds does
not change without a ceremony THIS box runs, which refreshes the hash on the
next ACK. Gating on age would make a box offline from the hub silently stop
offering - the exact failure this session removes. CheckedAt is persisted for
diagnosis, not as a gate.
- an ABSENT hash falls back to (a)/(b) and does NOT offer. "" is the hub
positively saying its package seals no repository password (legacy hash-less
escrow). Nothing to compare, and offering would put a permanent screen in
front of every legacy box.
The write damper: CheckedAt refreshes on every ack carrying a hash, but a save is
skipped when both the hash and the UTC day are unchanged, so an idle box does not
rewrite settings.json every fifteen minutes. It records WHEN WE LAST HEARD, not
when it last changed - the R-100 distinction.
Tests: Scenario C (a differing key offers, with both proxies asserted false first),
Scenario D (a matching key offers nothing), fact 1 still required, shape (a) still
works, and both SEC 7.2 halves.
RED-PROOFS, each with the mutation confirmed present in the file first:
D) hubHash != localHash conjunct dropped -> Scenario D FAILS (a healthy box
offered recovery forever); Scenario C still passes
WIRING) RecordEscrowKeyHash removed from the EscrowAutoConfirmer literal in
main.go -> TestMainWiresRecordEscrowKeyHash FAILS. This is the ships-inert
shape: unwired, everything compiles, every test in the package passes, the
auto-confirm still works, and shape (c) reads an empty hash forever.
Green: go build, go vet, go test ./... all pass.
|
||
|
|
763de3a025 |
R-241 part 1: the box does not mint a repository key over a sealed package
THE DEFECT. WriteOffboxSecrets auto-generated on ONE input - does the file
exist. Its two neighbours in the same file, OffsiteRecoveryOffer and
needsOffsiteCredential, both consult GetHubEscrowIdentityPresent(). The same
fact was available on three paths and used on two.
Measured on the final walk: a rebuilt box's credential self-heal reached here
at 03:18:06Z and minted 9b4a9a9d over a hub package sealing 30ef574f. The
recovery screen then correctly reported nothing recoverable under the key the
box held. The screen was honest; the minting was not. And the flag was not
merely available at that moment - it was the PRECONDITION of the chain that
reached this function, logged at 02:48:03Z, six ticks earlier.
THE GUARD IS A CONJUNCTION, deliberately: a package held AND no key present.
A box the hub holds nothing for mints exactly as before.
The refusal is a HOLDING state, not a failure. ApplyOffsiteTarget catches the
sentinel and still writes the transport (ssh key, known_hosts, coordinates),
so the recovery screen can bring the tier up the instant the escrowed key is
placed (R-219). Returning the error instead would leave needsOffsiteCredential
true forever and the hub re-staging a consumed credential on every cycle.
New declared state offsite.state=awaiting_recovery_key, shown INERT to every
existing hub reader from their code rather than assumed: offsiteheal acts on
exactly one string; isStale needs Enabled && escrowed and this carries
Enabled=false; the delivery checker skips the applied shape; an unknown state
string is ignored by encoding/json. So NO hub change is needed for this part.
OffboxAwaitingRecoveryKey is DERIVED, not stored - the operator's ruling that
the state should be fixed rather than remembered, applied to this field too.
t.Enabled is load-bearing in that predicate and was MISSING in the first
draft. The existing TestOffsiteDeclare_DisabledTargetIsNotStranded caught it,
not review: a customer who switched off-site off is not awaiting anything.
Now pinned from the new predicate's own side as well.
Tests: Scenario A (no key written; transport still written; apply holds and
stages nothing), Scenario B (first-time box still mints), idempotency, the
nil-settings fail-safe, and the Scenario E carve-out.
RED-PROOFS, each with the mutation confirmed present in the file first:
A) guard block deleted -> both Scenario A tests FAIL with
"R-241 REGRESSION: apply minted a repository password over the sealed
package"; Scenario B still passes (the mutation is specific)
B) guard over-widened (hub-package conjunct dropped) -> Scenario B FAILS
with a first-time box unable to start; Scenario A still passes
Green: go build, go vet, go test ./... all pass; controller_gates all OK.
|
||
|
|
c6b69d888e |
v0.205.0 — a run that skipped an app the customer selected is not successful (R-234)
gates / gates (push) Successful in 21s
THE VERDICT. The R-203 block already said "a warning beside a success is read as a success" and applied it to ONE of the two shapes it describes: an app missing a declared mandatory FOLDER made the run incomplete, while an app skipped ENTIRELY still reported ok. Both do now. Which skips count, decided by measurement: selected+deployed with no recovery unit YES; selected but NOT deployed no (named, with what to do — a box left amber by an app somebody removed is a status nobody reads); disconnected/decommissioned drive no (own signal); nothing selected no. LastSuccess and SnapshotCount still record what WAS captured. THE FILED MECHANISM WAS NOT THE MEASURED CAUSE, and saying so is the point. §3 stated that toggling an app on leaves it without a bundle so the first run skips it. Measured on demo-hp: the run's own pre-dump phase calls captureAllRecoveryUnits for every DEPLOYED stack, through admitApp, before the push — a unit moved aside was RECREATED and the run reported ok. That state does not survive a run. What actually produced the 2026-08-06 sequence: the manual run was dropped by the single-flight while an earlier run was still going. runOffboxBackup returned nil, the handler had already answered "A tavoli mentes elindult", and the card then showed the PREVIOUS run's green verdict — read as covering the app just selected. The decision is now taken synchronously in the handler and a dropped request says so. The nightly path still returns nil on purpose: nobody asked, and it retries. §7.3 measured before deciding: CaptureRecoveryUnit writes a few KB of compose + manifest, only ENUMERATES dumps rather than creating them, is idempotent and does NOT stop the app — and already runs inside the off-site run. So there is no wait to remove for a deployed app and NOTHING was built. 28 packages ok, 9/9 gates. Four red-proofs, each asserted to have applied. Fixture note: the shared provider's ListDeployedStacks returned nil, so Scenario A first passed for the wrong reason; fixed with an opt-in deployed set that defaults to nil. |
||
|
|
53e9bf0224 |
v0.204.0 — the restore list is keyed on the store (R-237); the size gate stops refusing in silence (R-238)
gates / gates (push) Successful in 26s
R-237: /backups/restore listed apps that are CURRENTLY DEPLOYED and CURRENTLY TOGGLED ON for future off-site backups. A rebuilt box has neither, so a household that had just lost everything was shown nothing to restore while the repository held their snapshots — measured live on the R-201 re-walk. To restore an app you had to select it, to select it you had to have installed it, and to know what to install you had to see the backup you could not see. The store is now the source of the list (offsite_restore_list.go), built on the existing R-193 OffsiteInventoryList. Installed-ness became a property OF a row, never a filter on it. Every case is answered rather than hidden: a snapshot for an app that is not installed is offered and says it will reinstall first; an installed app with no snapshot is shown as having nothing; an unreadable store renders as UNKNOWN (R-225's rule, one screen over) AND keeps the action, because "we could not look" is not "there is nothing"; no-target is its own state. The felhom-offbox and _shares marker tags are excluded from the app list. R-238 classified as a HARNESS ARTIFACT: mode=full without confirm=1 is step 1 of a deliberate two-step — it starts no job by design and redirects carrying &full_prep=<app>, which deriveWizardStep requires to reveal the commit. A driver that did not carry it forward landed back on the intent step. The operator's browser run completed the same restore. The wizard's precedence rules were NOT re-keyed: a stale ?full_prep= must never resurrect a commit button mid-restore. The residue WAS real and is fixed: neither branch of that step wrote anything to the log, so a refusal — including by the headroom gate — left no trace on the box. Both branches now log, and so does the concurrent-op refusal. resolveWizardApp is removed: it was dead once the gate moved, and its test pinned the defect's behaviour (an untoggled app refused), which would have read as policy. 28 packages ok, 9/9 gates OK. Three red-proofs, each asserted to have applied. |
||
|
|
c7446f2d6a |
R-225/R-227/R-228 Parts 2-4: unknown is not zero, the gateway speaks Hungarian, the set-aside is visible
R-225 — an unread store said '0 pillanatkép / 0 / 50 GB' above a card stating it held backups under another key. An SFTP listing found snapshot f3d9cd67 and 12 535 KB really there; snapshot_count and repo_size_bytes were simply ABSENT and the zero value spoke for them. StatsKnown is now NAMED, for the same reason OffsiteInventory.Empty is: zero is what an unread store and an empty one both look like, and on the wire 'absent' and '0' are the same bytes. The fill bar renders only when the fill is known — a 0%-wide bar is a picture of emptiness, and a picture is a claim. A measured zero still says zero. R-227 — WHICH LAYER ANSWERS: traefik, and this repo generates its config. But traefik v3 serves no static files, so a branded proxy page needs a new always-up container for every 502 on the box — out of proportion, and scoped in the report rather than built. Shipped instead: the unlock posts via fetch and answers a gateway failure in Hungarian without leaving the page. Progressive enhancement — with no JS the plain POST is unchanged and still shows the proxy's error, which the report says plainly rather than implying otherwise. R-228 — the set-aside history was recorded in orphaned_renamed_to and read by nobody: a census found zero references in any template or handler, while 12 535 KB sat at that path. It is surfaced as two facts and stops. It does NOT promise the history can be reopened, because it cannot be by anyone today (R-199's inventory is unbuilt) — and the set-aside CONFIRMATION copy was corrected for the same reason: 'a helyreállítási kód nélkül többé nem lesznek megnyithatók' implied that WITH the code they could be. The field's own comment called it 'recovery-code-recoverable', which was the same over-promise in the code. Tests: scenarios F, G, H as render tests per branch of each gate. Red-proofs, each demonstrated failing then restored: remove the StatsKnown guards (F, 'R-225 RETURNED: an unread store reports a snapshot COUNT of zero'), delete the set-aside block (H). The F assertion on the fill bar is scoped to the bar's own container — a bare width:0% search matched unrelated elements and would have passed for the wrong reason. 28 packages ok, vet clean, all controller gates OK (the emoji gate caught a warning sign in a template comment). |
||
|
|
a3499d1807 |
v0.201.0 — a correct recovery code is never called wrong again (CAMPAIGN-11) — MinAgent 0.125.0
gates / gates (push) Successful in 9s
R-216: the offsite key recovery is a coupled feature and now says so. featureProbes +
featureMinAgent 0.125.0 + a Supports gate at the unlock entry point, FAILING CLOSED — an
agent that cannot answer is named as such instead of the customer's code being blamed.
Measured live: a 404 from agent 0.120.0 came back as "we did not accept your recovery
code, check that all ten words", in 0.134 s, against a perfect code.
R-218: delete the repo-password short-circuit in needsOffsiteCredential. The declaration
stops when the TIER WORKS, not when a key exists — installing a key is the recovery
screen's whole job, so succeeding at recovery was switching off the mechanism that would
have delivered the coordinates to use it.
R-219: the unlock finishes the job — place the key, bring the tier up, then list. Without
it the promised listing could never render on the shape the screen exists for.
R-217: an unreadable store no longer claims to have opened with unattributable content
(the OffsiteInventory{} zero value). Opened / empty / unreadable are three states.
R-222: a code that is right about a RETAINED earlier package is named, not blamed. States
what the hub knows and promises nothing — no read path exists.
R-215: GET /recovery is gated on the same predicate as the interception.
Five red-proofs, each demonstrated failing and restored.
|
||
|
|
636c51e542 |
R-193: the recovery screen — unlocking, and only unlocking (v0.200.0)
A customer whose machine was rebuilt had everything needed to get their data back and no way to find out: the only route was a command line. This is the screen that closes that. IT UNLOCKS, AND ONLY UNLOCKS (operator ruling). It explains, takes the recovery code, opens the repository and shows what is in there — apps, dates, sizes. It restores nothing: restore is already per-app and lives in the backups area, and a screen that unlocks and then offers to overwrite is two decisions wearing one button. ONE CORE, TWO CALLERS. RecoverInstallCore is split out of RecoverAndInstall; the CLI wrapper keeps its exit codes and printed lines byte-identical, and the handler drives the same function. Two implementations of the one operation that can permanently lose a customer's data would drift, and only one would be tested. Asserted from source on both sides by AST. THREE WAYS OUT, none a dismiss button: recover; 'most nem' (the full page stops interrupting, the backups-area entry point stays PERMANENTLY, bound to the offer and never to the postpone flag); and 'I do not want the old data' — confirmed TWICE and reaching the SHIPPED move-aside, which sets aside and never deletes. THE CODE IS HANDLED NO MORE LOOSELY THAN ON THE COMMAND LINE: POST body only, never logged, never persisted, never echoed, cleared on every path, no-store, autocomplete off. No lockout — the code is a ten-word phrase, and locking a customer out of their own data for a typo is worse than anything it prevents. TWO DEFECTS THE TESTS CAUGHT, both fixed: an UNCLAIMED (legacy-open) box would have been shown the page, because RequireAuth passes such a box through; and the inventory nil-dereferenced when no off-site target was configured, which is exactly the pristine rebuilt shape. |
||
|
|
1214bae0a2 |
R-204 item 4 (box half): a rebuilt box DECLARES that it needs a credential (v0.199.0)
An absent off-site object has four meanings — never configured, mid-restart, a transient config read failure, and rebuilt-and-stranded — and the hub cannot tell them apart. The box can, from two local facts it holds with certainty, so it says so instead of leaving the hub to deduce it from a silence (operator ruling). The ACK's identity_blob_present is now recorded on EVERY ACK, before the gates that used to discard it: on a box with no off-site target the auto-confirm returns immediately, which is exactly a rebuilt box, so the one fact distinguishing it from a box that never had off-site backups was thrown away every cycle. The declaration needs BOTH halves — a fresh data area AND a hub-held recovery package. Freshness alone is a box that never had off-site backups; dropping that condition makes the whole fleet ask for credentials, which is what the Scenario B test exists to catch. The object carries enabled:false and zero sizes, which is what makes it inert to the hub's existing fill and staleness checkers and to a pre-upgrade hub. A configured box's JSON is byte-identical to v0.198.0's. |
||
|
|
58c703bd44 |
R-203 Part 2: a run that missed a MANDATORY directory is not a successful run (v0.197.0)
gates / gates (push) Successful in 8s
The gap was already detected and warned about, in Hungarian, naming the app and the folders -- that warning is what stopped the R-201 drill. The defect was that the run still reported `ok` beside it, and a warning standing beside a success is read as a success. last_status gains "incomplete": minted, because "ok" | "error" | "running" had nothing meaning "it ran, and this app is not fully protected". NOT "error" -- the rest of the run worked and what was captured is real, so SnapshotCount and the LastSuccess anchor still record it. Half a backup is not no backup. The gaps are now recorded STRUCTURALLY (offboxRunResult.mandatoryGaps), not only as prose, so the verdict has something to act on. It reaches the operator through the EXISTING per-run digest (backup_run_failures) rather than a new event type -- a new type is a two-repo change and the hub drops anything outside allowedEventTypes. The stat-filter gains the ClassMandatory check Tier 2 already had. It is a NO-OP today (TierOffsite admits mandatory only), so no customer-visible warning disappears -- demonstrated by widening the tier filter alone and watching the check hold the line. ANTICIPATED: calibre-web on demo-hp has exactly this gap, so its off-site status becomes incomplete the moment this ships. That is correct and is the point. Red-proofs: my first Scenario-C proof PASSED because the test only reached offboxCaptureSet while the mutation lives in runOffboxInternal -- a mutation the test cannot observe is not a red-proof, and the fix was the test. The run-level test now fails under both mutations (unreachable gap recording; unconditional ok). |
||
|
|
73efb091d9 |
R-203: the app and its backup look in the same directory — one resolver, every caller
gates / gates (push) Successful in 9s
appbackup's path helpers take a NAMESPACE ROOT. Five call sites passed a bare DRIVE path.
On an enrolled drive the two coincide, so nothing showed; on the system-data fallback they
differ by exactly the felhom-data segment, and the app then bound a directory the off-site
capture set never looked at -- while the run reported ok. Measured live on demo-hp: the app
wrote to /mnt/sys_drive/userdata/media/books, the capture set looked for
/mnt/sys_drive/felhom-data/userdata/media/books.
THE RULE NOW HAS ONE EXPRESSION. appbackup.NamespaceRootFor / IsEnrolledDrive encode the
drive-kind comparison; backup.Manager.namespaceRoot and stacks.Manager.inGuest delegate to
it. There were already TWO copies and they differed -- the backup package's compared without
filepath.Clean, the stacks package's with it, so a trailing slash from config would have
flipped the mode in one and not the other.
Sites routed through it:
- stacks/deploy.go withPathVars -> ${USERDATA_PATH} (the live defect)
- appexport/fabplan.go + export.go (via a new provider method)
- web/handlers.go FileBrowser mounts (latent: the system drive is
deliberately never a registered StoragePath, so this is the identity today)
ComputeFabBuckets now receives the namespace root, which is what ComputeCaptureSet has always
received -- so the export's classified paths and the backup's capture set describe the same
directories by construction instead of by coincidence.
Tests are table-driven over BOTH drive kinds, because this survived by being invisible on the
kind that already worked. Red-proofs observed: restoring the bare-path call fails the
system-drive row with the two paths differing by /felhom-data; inverting the drive-kind
comparison fails every enrolled row.
|
||
|
|
1b1366bb6e |
controller v0.196.0: the recovered key installs itself (R-200 plumbing half) -- MinAgent 0.125.0
gates / gates (push) Successful in 8s
--recover-offsite-install is the sibling of --recover-offsite-check: same fetch/unseal path through the agent, same STDIN discipline for R, but it PLACES the recovered repository password via InjectOffboxPassword so a rebuilt box reopens the history it inherited. Doing this by hand would put the offsite DATA key through a terminal, a clipboard and shell history. In-process the value goes agent -> this process -> the 0600 file and is rendered nowhere. The confirmation is a SECOND invocation: without --confirm-install it prints both hashes and writes nothing, so the operator sees the comparison before any write is possible. Three outcomes, named distinctly: installed (no local password -- the rebuilt-box shape), unchanged (identical key already present, nothing written), refused (a DIFFERENT key present; installing would clobber the key the current repository is encrypted under, and no force option is offered). Exit 2 for the refusal, distinct from 1 for a failed step. Red-proof: removing the confirmation gate makes the dry run write, failing the test. The R-persistence test carries a positive control -- a planted copy is found, then removed and not found -- because an absence check is worth only what its sensitivity is. |
||
|
|
9640e51321 |
controller v0.195.0: prove the offsite key comes back (R-200 plumbing half) -- MinAgent 0.125.0
gates / gates (push) Successful in 10s
--recover-offsite-check is a docker exec diagnostic in the shape of --print-reset-code: it reads the customer's recovery code from STDIN, asks the agent to fetch this host's sealed bundle and open it, and reports whether the recovered key matches the one on disk BY SHA256. Two hashes and a verdict; never a password, never R, never a blob. R comes from stdin and not a flag because a flag value is visible in ps, in shell history, in a container's command line and in any transcript of the session that ran it. IT COMPARES; IT DOES NOT INSTALL. The recovered password is never written to offbox/repo_password -- installing changes a live box on a path nobody has walked, and that link is next session's, with the drill around it. A test asserts the data dir is byte-unchanged after a check; its red-proof (adding the install call) fails it. Exit codes: 0 match, 2 clean MISMATCH, 1 a step failed -- "it failed" and "it worked and disagreed" must never share a status. A box with no local password reports distinctly: that is the rebuilt-box shape, where the next step is to install rather than compare. Nothing customer-reachable ships here: no card, no form, no preview. |
||
|
|
88897a224e |
v0.194.0 — one operator email per backup run, and nothing dropped without a trace (R-182)
gates / gates (push) Successful in 8s
MEASURED, not supposed. On 2026-08-03 nine per-app recovery_unit_capture_failed events reached the hub and TWO operator emails went out. The hub's operator cooldown key is customerID:eventType(+tier) and that event carries `app` but no `tier`, so the key held no app identifier: the first refused app took the hour's slot and every other app's failure was discarded BEFORE anything was written down, leaving no row on any channel. The obvious fix — put `app` in the key — was ruled against: on a full disk it produces one email per app, the volume problem wearing the correctness problem's clothes. internal/backup/runsummary.go: a per-run collector with exactly admissionSet's lifetime, fed by all three write legs, emitting backup_run_failures ONCE at the end and only when something failed. A clean run emits nothing. The per-app event stays and becomes the RECORD — the hub routes it record-only, stored and logged every time, never competing for an email slot. The record and the notification are now different things. Deliberate skips (disconnected, decommissioned) are excluded: they have their own alert, and a nightly email about an unplugged drive is one the operator learns to ignore. A manual run always reports: the digest carries a unique run_id the cooldown cannot collapse. Someone pressing the button is actively trying to get a backup. THE PERIODIC SWEEP GETS A DIGEST TOO. With the per-app event now record-only, a capture failure found between runs would be recorded and never notified — a new silence introduced while closing one. That path emits a digest with NO run_id, so the ordinary 1-hour cooldown caps it exactly as before while the mail now lists every failing app instead of whichever was first. A refusal is recorded ONCE, where the verdict is taken, not at the three legs that consult it — R-181's contract is one verdict per app per run. Noting it per leg listed one refused app three times and produced "2 of 1 apps failed". Found by the digest's own test, not in review. Silence is safe because the hub's deadline check raises expected_backup_missed from report freshness, independently of any mail this box sends (monitor/deadline.go:396,417). Confirmed, not assumed. 7 new tests, 4 red-proofs. The main.go seam walk did NOT fail on its first attempt — the AST test walked the backup package and not main.go; the test was fixed and the mutation re-run rather than the pass recorded. |
||
|
|
6c43bf6156 |
v0.193.1 — the refusal's size estimate is rendered in bytes, not "0.00 GiB" (R-181 follow-on)
gates / gates (push) Successful in 9s
Found by v0.193.0's own live proof run. The estimate was printed fixed to two decimal GiB, so every app under ~10 MB rendered as "estimated 0.00 GiB write" — which reads as "no estimate was available" and is the opposite of what happened. Observed live on demo-hp 08:59:46: opengist's real 178 KB estimate printed as 0.00 GiB. Shipped in the same session because it is the same defect class R-181 is about: a message an operator cannot rely on is worse than no message. The arithmetic is unchanged and still in GiB — the reserve's own unit, so the comparison against FloorFreeGiB reads directly. Only the rendering moved to humanizeBytes. estimatedWriteGiB -> estimatedWriteBytes, with the GiB conversion done once at the point of comparison. |
||
|
|
fef07c3923 |
v0.193.0 — the reserve guards the write that fills the disk, and its promise is true (R-181)
gates / gates (push) Successful in 9s
B2's capture floor (v0.192.0) was consulted in exactly ONE place — captureAllRecoveryUnits, which writes a few KB. The two legs that write the BULK into the same backups/primary/<app> tree, the DB dump and the volume dump, ran FIRST and unguarded. Measured live on demo-hp 2026-08-03 06:40:03: opengist's volume dump wrote 2.0 GB with no check, free fell to 1.0 GB, and the floor then refused the cheap write it had already lost the argument to. Its refusal message claimed "the previous unit is untouched" — measured false: that app's tar had gone 182,272 B -> 2,147,666,432 B under a stale manifest. Sixth entry in CLAUDE.md's table of shipped guarantees the code did not provide. Fix: ONE admission verdict per app per run (internal/backup/admission.go), taken before that app's FIRST write and covering all three legs — they write under one per-app root, which is why one verdict can honestly cover them. - Lazy, at the app's first write, NOT once at run start: app A's dump can put app B under the reserve, so a run-start verdict reads a disk that no longer exists. - Remembered for the run, never re-decided between an app's own legs — that is the split this closes. Reset per run. - Placed ahead of DumpAppVolumesSafe, which stops the stack as its first act, so a refused app is never bounced. After the volume-less check, which has no write. - Exactly one operator alert per refused app per run. - Leg order unchanged: volume dumps still precede the capture. The floor is now SIZE-AWARE: it asks whether THIS app's write would cross the reserve, not only whether the filesystem is already below it — which is how an app was admitted at 96% and then allowed to write 2 GB. Estimate = the app's previous .sql + .tar on disk. No history -> headroom-only, deliberately, and the alert says so. A container-based du per volume was MEASURED and rejected: 66 timed runs on demo-hp guest 9201, median ~355 ms/volume (341-404) on volumes holding tens of KB — container start-up, not the walk. Decisive on top: docker run needs the writable layer, so it can fail under exactly the pressure the reserve handles. The message was NOT weakened; the behaviour was moved so the wording became true. It now also names which term bound. Every claim is checked against a sha256 fingerprint of the tree it describes, never against the log line. Still refuses and never deletes: nothing here is generational. 11 new tests through the production functions. The DB leg cannot run without Docker, so its gate is pinned by an AST walk of backup.go asserting admitApp precedes DumpOne (strings.Contains is insufficient — a commented-out call still contains the string). 4 red-proofs demonstrated failing then restored. |
||
|
|
4be6467b50 |
v0.192.0 — the capture floor replaces the bulkhead (R-165, decision B2)
gates / gates (push) Successful in 8s
Ships BEFORE the disk-layout merge it exists for, and is harmless on a box that never gets it. The mp1 partition was a BULKHEAD as well as a ceiling: it kept a runaway capture from filling the space the container runtime needs, because /var/lib/docker was a different filesystem. After the merge it is the same one, and a full Docker data-root is a stopped box. The floor sits in captureAllRecoveryUnits, checked BEFORE anything is written: below the reserve, that ONE app's capture is refused, its previous unit is left byte-identical, the R-158 alert fires with the space figures, and the loop continues. Two terms whichever binds first (97% used / 1 GiB free) in fillwatch's shape, deliberately BEYOND its critical band (95% / 2 GiB) so the customer is always warned before a refusal can happen — a floor that fires before its own warning is a silent failure wearing a threshold. Headroom, never unit size: a per-unit cap would be R-163 rebuilt inside one volume. Refuses, never deletes: nothing here is generational, so pruning could only destroy a different app's only local copy; pruneStalePrimaryDirs is an orphan sweep, not retention, and must not be repurposed. Tests 1184 -> 1191. One fixture strengthened mid-red-proof: the "old 20 G ceiling is gone" test sat at exactly 20 GB and survived a literal UsedGB > 20 cap — hollow. Now 120 GB, and the mutation fails it. |
||
|
|
cf48214f6c |
v0.191.0 — warn before the wall comes down (R-167, R-158, R-174)
gates / gates (push) Successful in 9s
R-167: new internal/fillwatch warns the CUSTOMER before a filesystem fills. It emits the PRE-EXISTING disk_warning/disk_critical pair, which was allowlisted, copy'd, default-enabled and checkbox'd with no producer in any repo — the sixth "built but never wired" instance here. Two threshold terms (85% or 5 GiB free; critical 95%/2 GiB) because a percentage alone lies at both ends of this fleet's size range. Edge-triggered on escalation only, state persisted, hysteresis dead zone at 75%/7 GiB pinned by a test. A nil usage read is never a warning and never clears one. Per filesystem, never per app. Daily 03:30, before the nightly app-data legs. R-158: new unitNotify seam fires per app when a Tier-1 recovery-unit capture fails, loop continuing, carrying the target filesystem's used/free bytes. Operator-tier (recovery_unit_capture_failed) — deliberately NOT backup_failed, which is customer-enabled and would email the customer about a failure they cannot act on. D-c overrides R-158's own proposal here. R-174: the app-stop guard no longer starts apps onto MISSING drives — a regression in v0.189.0 code, found by review and closed the same session. SetStarter got the raw stack manager, whose StartStack has no drive gate, and Recover runs at startup. R-171 one path over. bootDriveGate could not be reused whole (its holder #2 is the guard's own marker, and holders #1/#2 read vars assigned after Recover runs), so holder #3 is extracted into a shared driveStartGate with a test pinning the delegation. ErrStartRefused splits a refusal from a failure: both keep the marker, only Failed alarms, because routing a deliberate hold into NotifyBackupFailed is the same false alarm. Tests 1157 -> 1184. All red-proofs demonstrated failing and restored. |
||
|
|
582135f861 |
v0.190.0 — the boot settle window, both gates on intent, and R-171
gates / gates (push) Successful in 8s
R-171 (a regression v0.189.0 introduced, CONFIRMED on hardware before any fix was written). Replacing isBootOrphan's container-count term with recorded intent made a drive-gate-stopped app read as a boot orphan: the gate stops apps with `compose down` (zero containers) and never touches desired_state, because it is not the customer. Observed on 9201 with the drive held unmounted — the sweep found and started it, burned both attempts, and handed it to the dead-app alarm. The write hazard did not materialise (the unbound mountpoint is host-root-owned and the guest is unprivileged) but that protection is accidental and untested. New consumer-side seam bootrecon.StartGate, fail-safe (cannot determine ⇒ do not start), wired in main.go. The rule is not new: the API's startGatedByMissingDrive already refuses this; the sweep bypassed it. R-157 mechanism A. The sweep looked once at T+5s, deriving candidates from a fleet docker was still restoring — three of six hard resets. Now a settle-then- sweep window: sample every 5s, settled after 3 identical samples, sweep ONCE at the end; ends on settled or a 50s budget, and the log says which. The budget is 50s because settle+budget+one retry must stay under the 90s dead-app grace — a test rejected 60s at 95s. A window that overruns emits a LATE RECOVERY warn rather than the grace being widened to hide it. Widening the window made two more holders reachable, so the one gate covers all three: an absent drive, a quiesce, and an in-flight app-data operation — reusing quiesce.SuppressedStacks() and a new read-only AppStopGuard.HeldStacks(). R-170. shouldRecreateOnBoot now reads desired_state with the identical three-way table; absent keeps the old hasContainers behaviour exactly. Its comment argued for the container count and was rewritten. presentStable is untouched. The two gates' agreement is pinned from both sides against one fixture table. 27/27 packages green; 6 red-proofs observed FAIL then restored. |
||
|
|
dbcb306fcf |
v0.189.0 — desired state + the app-stop crash marker (R-166 / D-b)
gates / gates (push) Successful in 8s
The box stops inferring the customer's intent from a container count and reads
what they actually asked for.
Part 1 — desired state. AppConfig gains a tri-state `desired_state`
(""/running/stopped), written ONLY by the customer's own action: the API action
switch, DeployStack, UpdateOptionalConfig's redeploy branch, and the .fab
import. Intent is written BEFORE the act and a failed write REFUSES the act.
StartStack/StopStack are deliberately not writers — 14 callers, only 2 are the
customer. bootrecon.isBootOrphan now reads intent instead of len(Containers)>0,
which closes R-157 mechanism B (a power cut or interrupted deploy left an app
with zero containers, read as a deliberate stop, and stranded silently).
ABSENT MEANS UNKNOWN, NEVER "running": every pre-v0.189.0 app.yaml reads absent,
so the legacy fallback is byte-identical to the old rule. A running-only startup
backfill converges the unambiguous cases; `stopped` is never inferred.
Part 2 — backup.AppStopGuard, a persisted marker over every stop→work→start
window (volume dump, offbox reconstitute, .fab export). Its own file, never
quiesce's. Written before the stop, cleared only after a restart that succeeded,
kept when one fails. Recover() completes before the boot reconciler is launched
and returns its outcome, which main.go reports on the existing backup_failed
event once the notifier exists. A defer is not the mechanism — a SIGKILL runs
none (Campaign 8 fault 10).
Also: SaveAppConfig rebuilt AppConfig field-by-field (the R-100 shape) and would
have dropped desired_state on every save across nine call sites. Replaced with
copy-and-overlay. Measured: app.yaml does not round-trip unknown YAML keys.
No hub change, no agent coupling, no user-visible string. 27/27 packages green;
7 red-proofs observed FAIL then restored.
|
||
|
|
4ed938cce4 |
D5: an app restore works from the drive alone (v0.188.0)
The recovery unit on the customer's drive now carries the PORTABLE secret class, so Tier-1/Tier-2 restore no longer depends on the whole-guest tier. A customer needs the drive and nothing else. Part 0's rulings overturned the brief's recommendation, on evidence: - the data_key flag is untrustworthy (4+ encryption keys the catalog itself labels as such are unflagged) -> R-127 - a DB password is not resettable in practice: POSTGRES_PASSWORD is ignored once PGDATA is non-empty, so a regenerated value leaves the app unable to authenticate against its own restored rows while the dump replay still reports success (proven on a throwaway postgres:16-alpine) Ruling (operator): type:secret travels, type:password never does, minus the nonPortableSecrets code register. Plaintext -- withholding the internet- reachable class is what licenses that, and the two are coupled. Precedence: the UNIT WINS over the guest -- the unit's secrets were captured in the same run as the dumps beside them, so they match the data being restored. The fail-closed data-key gate is unchanged. Secret values are never logged; the manifest records NAMES only. |
||
|
|
fd50a73e65 |
C9-F1 + C9-F2: a restore that restored nothing, and a crash loop nobody saw (v0.183.0)
Both are the system reporting healthy while the customer is not, and both live in the same status-derivation code. Neither is fixed by making the system quieter. C9-F1 (HIGH) — Tier-2 writes recovery-unit/ on EVERY run and RestoreTier2Files has never read it (tier2_restore.go:101-104 reads hdd/ + userdata/ only). Phase 0 enumerated all 53 catalog templates against both demo boxes: 43 apps have NO readable subtree, so the button stopped the app, restored 0 files, restarted it and said "Nincs hiányzó fájl — minden fájl megvan a helyén." — at the moment the customer pressed it because files were missing, with 156 MB of BookStack's data unread in the same copy. 9 apps have file legs but never their DB or volumes, so the same sentence was also a clean bill of health over data never opened (immich: 1.3 GB Postgres unit). Honesty half shipped: a pre-flight coverage check refuses UP FRONT without stopping the app and NAMES the action that works; a run that proceeds claims only what it EXAMINED and discloses that the database and volumes are not covered. Completeness is filed as C9-F1b — routing to the Tier-1 unit restore puts a destructive operation behind a non-destructive button, so its confirm copy has to carry that difference. C9-F4 filed: nothing reads the Tier-2 recovery-unit/ mirror, so the second local copy that exists for drive loss is unreachable by any customer action. C9-F2 (HIGH) — a crash loop was counted as working. StateRestarting is deliberately NOT added to IsDownState (that alarms on every deploy fleet-wide, the over-correction F-A1 nearly cost us); a sustained run becomes down after crashLoopAfter = 5m, set above the 120s deploy timeout, Mealie's 60s start_period and R-97b's 180s grace. The dashboard counter uses the same predicate, so it no longer contradicts the alarm on the same screen. README's claim that faults "still surface as restarting" was a wish with no test — corrected in place; it is the seventh such instance. Six red-proofs observed, including the one that matters most: adding StateRestarting to IsDownState fails the brief-restart test with "every deploy and update would page the operator". go test ./... rc=0, 27 packages, run and read separately from this commit. |
||
|
|
3f048e042b |
R-101 + F-DIAG: the restore dialog names the last SUCCESSFUL copy (v0.182.0)
Tier2LastRun is the attempt clock and was rendered as 'Legutóbbi másolat' in the restore confirm dialog. New LastSuccess + SuccessTracked anchor; tier2Update makes the three rebuild sites safe by construction. F-DIAG: six distinct causes, target-aware redaction. |
||
|
|
e000e201af |
R-100: record the offsite last-SUCCESS anchor (v0.181.0)
LastRun records an attempt, not a result. New OffboxTarget.LastSuccess, set only on the success branch via the pure offboxAnchorAfterRun rule, carried to the hub as last_success. Closes two silent-wipe sites (settings save, hub re-apply). |
||
|
|
2958946517 |
v0.172.0 — R-75: canonical import root, catalog-derived skeleton, import surfaces
${IMPORT_PATH} = <system namespace root>/userdata/import — ONE drop-zone per box,
on the system drive, injected at BOTH compose-env builders with NO per-drive
fallback (unresolvable leaves it unset so compose fails loudly rather than
quietly building a second, dead drop-zone).
Third BindRoot (RootImport) + Import list in BackupSpec, extended through
ValidateBackupSpec/ClassifyBinds. Load-bearing: a stale `userdata: import/<app>`
entry against the moved bind would be a WHOLE-BLOCK reject, taking the app's
mandatory hdd classification with it.
Exhaustive-root audit: resolveAbs/structuralGuard/ComputeCaptureSet/
ComputeFabBuckets now take importRoot explicitly (an import bind resolved
against hddPath would name a directory on the wrong drive); unresolvable is
refused loudly into Skipped. GetImportRoot added to both provider interfaces.
Catalog-derived skeleton: UserdataSkeleton() -> UserdataSkeletonCarry() +
BuildUserdataSkeleton(), SORTED. The carry-list makes zero-removals true by
construction (`documents` is in no catalog app but on both boxes) and is the
fresh-box floor. The sort is not tidiness: the naive map-order derivation
measured 20 distinct outputs from 20 identical runs, which with fbNeedsRecreate
is a fleet-wide FileBrowser restart loop.
One authoritative compose parser: ParseComposeUserdataMounts now delegates to
ParseComposeClassifiableBinds. Import root excluded from per-app migration.
Surfaces: FileBrowser /srv/beolvasas source; app-page "Hova tegyem a fajlokat?"
with PathEscape deep links (never QueryEscape) and class-driven copy;
data_paths: annotation with the Fork-3 asymmetry; system-owned beolvasas SMB
share refused server-side at handler AND store, button omitted in template.
Caught on the way: the sharing template's row struct was function-local, so
adding {{if .System}} would have 500'd every share row. ShareRow is now
package-level and the render test uses the handler's own type.
Tests 915 -> 949, all green. MinAgent unchanged.
|
||
|
|
2487681396 |
style: gofmt normalization — no logic changes
gofmt -w across the controller tree (46 files) so gofmt -l is empty — disarms the
formatting landmine where a targeted edit + accidental gofmt -w swept ~46 unrelated
files. Pure formatting: whitespace + gofmt's optional-semicolon removal in reflowed
inline closures. One doc comment reworded ('' -> 'the empty string') to avoid gofmt's
Go-1.19 doc-comment typographic substitition ('' -> curly quote) muddying its meaning.
No build/vet/test behavior change.
|