one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:
RestoreOffboxScratch - the reported one
OffboxRestorePrepareFull - the SECOND request in the customer's own two-step full-restore
flow, and the one that actually shells `restic stats`. The UI
reaches it FIRST, so flagging only the restore would have left
the collision reachable by the ordinary path.
RestoreSharesScratch - R-411's exact shape on the shares tier: unlockStale + resticStep,
a live web caller, and its sibling PlaceSharesRestore has always
taken the flag.
RestoreOffbox - no production caller today, but the same dangerous pattern.
Flagged rather than left for a future caller to inherit.
OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.
THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.
R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.
R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.
The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.
And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.
R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.
16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
This commit is contained in:
@@ -1343,7 +1343,21 @@ func (m *Manager) runOffboxInternal(ctx context.Context, apps, base, env []strin
|
||||
continue
|
||||
}
|
||||
res.backedUp++
|
||||
m.logger.Printf("[INFO] [offbox] backed up %s (%s, %d mandatory path(s))", stack, src, len(extra))
|
||||
// R-412 leg 1 — A PUSH THAT CARRIED NOTHING MUST NOT READ AS A PLAIN SUCCESS.
|
||||
//
|
||||
// Measured on demo-hp 2026-08-31: a recovery unit was destroyed mid-run, the capture rebuilt it
|
||||
// empty, and this line printed `backed up opengist (…, 0 mandatory path(s))` over a snapshot
|
||||
// holding none of the app's data. Every word of it was true and the whole of it was misleading.
|
||||
//
|
||||
// WORDING ONLY. This adds no guard and does NOT touch the capture — `07` §8.2 records why
|
||||
// guarding the capture would make the manifest lie. The reader of a log gets the fact; whether
|
||||
// the push should refuse is R-412 leg 2 and is still open.
|
||||
if unitIsHollow(src) {
|
||||
m.logger.Printf("[WARN] [offbox] backed up %s (%s, %d mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's data; the next run with a dump leg will replace it",
|
||||
stack, src, len(extra))
|
||||
} else {
|
||||
m.logger.Printf("[INFO] [offbox] backed up %s (%s, %d mandatory path(s))", stack, src, len(extra))
|
||||
}
|
||||
}
|
||||
// R-7b: the SHARES leg runs AFTER the per-app loop and BEFORE retention, so `forget --group-by
|
||||
// host,tags` covers the `_shares` group for free. It is placed BEFORE the firstErr return on
|
||||
@@ -1788,6 +1802,14 @@ func (m *Manager) offboxRecordStats(ctx context.Context, base, env []string) int
|
||||
// RestoreOffbox restores an app's latest off-box snapshot to destDir (a scratch/verify location — it does
|
||||
// NOT overwrite live data). Returns an error on failure (checks restic's own exit code).
|
||||
func (m *Manager) RestoreOffbox(ctx context.Context, stackName, destDir string) error {
|
||||
// R-411 — also found by the R-408 walk. This one has NO production caller today (only two tests
|
||||
// reach it), but it runs `unlockStale` + `resticStep restore`, which is precisely the pattern that
|
||||
// deleted a live restore's lock on demo-hp. Dead code gets resurrected; flagging it now costs
|
||||
// nothing and means a future caller inherits the guard rather than the defect.
|
||||
if err := m.acquireRunning(); err != nil {
|
||||
return err
|
||||
}
|
||||
defer m.releaseRunning()
|
||||
if !m.OffboxConfigured() {
|
||||
return fmt.Errorf("off-box backup not configured")
|
||||
}
|
||||
|
||||
@@ -27,6 +27,18 @@ import (
|
||||
// *"the in-process single-flight mutex (held by every caller of this method) proves no sibling
|
||||
// operation is live"*. Every off-site operation takes `acquireRunning` for exactly that reason.
|
||||
//
|
||||
// **THAT SENTENCE WAS FALSE FROM v0.227.0 UNTIL v0.232.0, AND NOTHING CHECKED IT.** `RestoreOffboxScratch`,
|
||||
// `OffboxRestorePrepareFull`, `RestoreSharesScratch` and `RestoreOffbox` all issued restic commands
|
||||
// without the flag. R-411 measured the consequence on demo-hp 2026-08-31: a customer full-restore's
|
||||
// `restic stats` held a repository lock, this check was therefore not blocked, met that lock, and
|
||||
// `resticStep` removed it with `unlock --remove-all` while logging *"a stale exclusive lock left by a
|
||||
// previous crash"*. There was no crash.
|
||||
//
|
||||
// The sentence is TRUE again as of v0.232.0, and it is now **pinned by
|
||||
// `TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered`** — an AST walk over this package that
|
||||
// fails when a new off-site entry point takes neither the flag nor a registered exemption. Do not
|
||||
// trust this paragraph; the test is what holds it.
|
||||
//
|
||||
// **An integrity check that did not take the flag would break that invariant.** It could meet the
|
||||
// lock of a `forget --prune` running from this same box, remove it, and retry over the top of a live
|
||||
// prune. So this check TAKES THE FLAG, and **skips rather than waits** when it cannot get it: waiting
|
||||
@@ -252,8 +264,21 @@ func (m *Manager) recordIntegrityOutcome(at time.Time, ok bool, depth string) {
|
||||
}
|
||||
}
|
||||
|
||||
// CheckOffboxIntegrity runs one off-site integrity check. It NEVER writes to the repository: `check`
|
||||
// is a read verb, and nothing here prunes, forgets, unlocks or backs up.
|
||||
// CheckOffboxIntegrity runs one off-site integrity check. It prunes nothing, forgets nothing, unlocks
|
||||
// nothing and backs up nothing.
|
||||
//
|
||||
// **IT DOES, HOWEVER, WRITE A LOCK FILE, and the earlier wording here — "It NEVER writes to the
|
||||
// repository" — was wrong (R-407).** Observed on demo-hp 2026-08-31 with a lock sampler that was
|
||||
// positively controlled first: across a real check the repository went `locks=0` -> `locks=1` for nine
|
||||
// consecutive samples -> `locks=0`. `restic check` takes a lock like most restic verbs; only `--no-lock`
|
||||
// avoids it, and this check deliberately does not pass it, because the lock is what makes the check
|
||||
// exclusive against a prune.
|
||||
//
|
||||
// **AND THE FACT NOBODY HAD, which is what made R-411 possible: `restic stats` ALSO takes a lock.**
|
||||
// Measured in a clean room the same night — nothing else running, four invocations, `locks=1`. That is
|
||||
// why a customer full-restore, whose size probe shells `stats`, could collide with this check at all.
|
||||
// The sentence is corrected in place rather than deleted, because R-360's rule is that a comment
|
||||
// claiming a guard is exactly why nobody looks for the missing one.
|
||||
//
|
||||
// Due-ness is NOT consulted here — the caller decides. The scheduled job asks IntegrityDue first; the
|
||||
// operator's debug button deliberately does not, because "run it now" is the whole point of a button.
|
||||
@@ -355,14 +380,14 @@ func looksLikeRepositoryDamage(out []byte) bool {
|
||||
// Every phrase below is restic's own error wording, taken from the 2026-08-30 damaged-pack run on
|
||||
// demo-hp or from restic's check source — never paraphrased.
|
||||
for _, sig := range []string{
|
||||
"does not match", // "Pack ID does not match, want <id>, got <id>" — the measured one
|
||||
"not found in index", // a blob or pack the index promises and the store lacks
|
||||
"blob not found", //
|
||||
"size mismatch", //
|
||||
"does not match", // "Pack ID does not match, want <id>, got <id>" — the measured one
|
||||
"not found in index", // a blob or pack the index promises and the store lacks
|
||||
"blob not found", //
|
||||
"size mismatch", //
|
||||
"ciphertext verification failed",
|
||||
"integrity error",
|
||||
"repository contains errors", // restic's own summary verdict
|
||||
"failed to load index", //
|
||||
"repository contains errors", // restic's own summary verdict
|
||||
"failed to load index", //
|
||||
} {
|
||||
if strings.Contains(s, sig) {
|
||||
return true
|
||||
|
||||
@@ -99,17 +99,37 @@ type ProofResult struct {
|
||||
// Err is a restore/plumbing failure — NOT a verdict about the backup. It never becomes the
|
||||
// "intact but empty" alarm, because a download that did not finish has seen nothing.
|
||||
Err error
|
||||
// CannotRun (R-414) is set when the proof could not START on this box at all — today the only
|
||||
// cause is that there is nowhere to put a scratch.
|
||||
//
|
||||
// IT IS SEPARATE FROM Err ON PURPOSE. An Err is transient: a download that failed says nothing
|
||||
// about the backup and nothing about the box, so it is right that it records nothing and retries.
|
||||
// This is a STANDING PROPERTY of the machine — it will be just as true tomorrow — and leaving it
|
||||
// unrecorded is what R-414 measured: `demo-felhom` failed this way every night with only a WARN,
|
||||
// and because nothing was written, `last_proof_result` stayed ABSENT — which is also what a
|
||||
// controller older than v0.231.0 sends. The hub could not tell "cannot run here" from "too old to
|
||||
// have the feature". That is the StatsKnown trap one level up, and this field is what closes it.
|
||||
CannotRun bool
|
||||
CannotWhy string
|
||||
}
|
||||
|
||||
// Verdict renders the outcome for the persisted record and the wire. "" is reserved for "no verdict
|
||||
// was reached", which is what a skip, a missing snapshot and a restore error all are.
|
||||
func (r ProofResult) Verdict() string {
|
||||
if r.CannotRun {
|
||||
// A recorded, distinguishable state — never "" , which the hub reads as "not recorded".
|
||||
return ProofResultCannotRun
|
||||
}
|
||||
if r.Skipped || r.NoSnapshot || r.Err != nil {
|
||||
return ""
|
||||
}
|
||||
return string(r.Judgement.Verdict)
|
||||
}
|
||||
|
||||
// ProofResultCannotRun is the persisted value for R-414's outcome. It is deliberately NOT one of the
|
||||
// UnitProofVerdict values: those judge a UNIT, and here no unit was ever looked at.
|
||||
const ProofResultCannotRun = "cannot_run"
|
||||
|
||||
// ProveOffboxUnit runs ONE nightly proof: pick the due app, restore its newest snapshot read-only to
|
||||
// a throwaway directory, judge it, delete the copy, and record WHICH SNAPSHOT was proved.
|
||||
//
|
||||
@@ -141,8 +161,10 @@ func (m *Manager) ProveOffboxUnit(ctx context.Context) ProofResult {
|
||||
|
||||
scratch, nsRoot, derr := m.offboxProofScratchDir(stack)
|
||||
if derr != nil {
|
||||
res.Err = derr
|
||||
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to: %v", stack, derr)
|
||||
// R-414: a standing property of this box, not a transient failure — so it is RECORDED, not
|
||||
// merely warned about. See ProofResult.CannotRun.
|
||||
res.CannotRun, res.CannotWhy = true, derr.Error()
|
||||
m.logger.Printf("[WARN] [offbox] proof: %s has nowhere to restore to, and this box will fail the same way every night until a data drive is registered: %v", stack, derr)
|
||||
return res
|
||||
}
|
||||
// The headroom gate runs BEFORE the download, through the same helper and the same Hungarian
|
||||
@@ -205,6 +227,17 @@ func (m *Manager) RecordProofVerdict(res ProofResult) {
|
||||
return
|
||||
}
|
||||
if err := m.settings.UpdateOffboxStatus(func(o *settings.OffboxTarget) {
|
||||
if v == ProofResultCannotRun {
|
||||
// RECORDED so the hub can see the state, but due-ness is deliberately NOT advanced: no
|
||||
// snapshot was proved, and marking one proved would stop the app ever being retried once a
|
||||
// drive is finally registered. A status report, not a proof.
|
||||
o.LastProofRun = res.RanAt.UTC().Format(time.RFC3339)
|
||||
o.LastProofStack = res.Stack
|
||||
o.LastProofSnapshot = ""
|
||||
o.LastProofResult = v
|
||||
o.LastProofReason = res.CannotWhy
|
||||
return
|
||||
}
|
||||
if o.ProvedSnapshots == nil {
|
||||
o.ProvedSnapshots = map[string]string{}
|
||||
}
|
||||
|
||||
@@ -177,7 +177,9 @@ func (m *Manager) offboxSnapshotSize(ctx context.Context, id string) (int64, err
|
||||
func (m *Manager) offboxRestoreScratchDir(stack string) (scratch, nsRoot string, err error) {
|
||||
// offsiteRestoreRootFor is THE place `backups/offsite-restore` is spelled (offbox_verify_copies.go)
|
||||
// — the listing/delete surface must resolve byte-identical paths to the ones written here.
|
||||
return m.offboxScratchDirIn(stack, m.offsiteRestoreRootFor)
|
||||
// unitOnly=false: the CUSTOMER's scratch can hold a full restore (bulk userdata), so it must NOT
|
||||
// fall back to the state-only system disk — see offboxScratchDirIn.
|
||||
return m.offboxScratchDirIn(stack, m.offsiteRestoreRootFor, false)
|
||||
}
|
||||
|
||||
// offboxProofScratchDir is the R-87 nightly proof's scratch, resolved by the SAME drive-preference
|
||||
@@ -190,14 +192,49 @@ func (m *Manager) offboxRestoreScratchDir(stack string) (scratch, nsRoot string,
|
||||
// copy invisible to `DeleteOffsiteRestoreCopy`, the copy listing and `OffboxFullScratchReady`, so it
|
||||
// can never be offered for placement into a live app.
|
||||
func (m *Manager) offboxProofScratchDir(stack string) (scratch, nsRoot string, err error) {
|
||||
return m.offboxScratchDirIn(stack, m.offsiteProofRootFor)
|
||||
// unitOnly=true: the proof restores ONE recovery unit (`--include <unit>`) and deletes it. For a
|
||||
// driveless app that unit already lives permanently on the system data path, so a scratch there is
|
||||
// at most a second copy of something already present — see offboxScratchDirIn.
|
||||
return m.offboxScratchDirIn(stack, m.offsiteProofRootFor, true)
|
||||
}
|
||||
|
||||
// offboxScratchDirIn holds the drive-preference rules once. `rootFor` chooses WHICH root under the
|
||||
// namespace the scratch lands in; everything else — the network-storage refusal, the ordering, the
|
||||
// R-252 wording — is shared, so the proof path can never drift from the customer path on the parts
|
||||
// that must not differ.
|
||||
func (m *Manager) offboxScratchDirIn(stack string, rootFor func(string) string) (scratch, nsRoot string, err error) {
|
||||
// R-414 — THE SYSTEM-DATA FALLBACK, AND WHY IT IS SCOPED BY WHAT IS BEING RESTORED.
|
||||
//
|
||||
// THE GAP. `demo-felhom` has ZERO registered storage paths, so steps (1)-(3) all miss and this
|
||||
// refused. The nightly proof therefore could not run AT ALL on that box — every night, with only a
|
||||
// WARN — and because its error path reaches no verdict, `last_proof_result` stayed ABSENT, which is
|
||||
// also what a controller too old to have the feature sends. The hub could not tell them apart.
|
||||
//
|
||||
// WAS IT MISSED OR DELIBERATE? Established from R-356's own commit (`08eb1a6`, 2026-08-22), whose test
|
||||
// comments say the scratch resolver *"still resolves to the registered storage path … only the
|
||||
// DESTINATION moves"* — i.e. it was OUT OF SCOPE for that change, which was about where restored data
|
||||
// LANDS. It was never ruled out on state-only grounds: the one comment about a `systemDataPath`
|
||||
// fallback belonged to `PlaceOffsiteRestore` and concerned merging bulk USERDATA onto the SSD, and
|
||||
// R-356 deliberately overruled even that. This function's own documented exclusion is
|
||||
// `cfg.Paths.DataDir` — the ROOTFS — which is a different filesystem entirely.
|
||||
//
|
||||
// SO §6.3's [DESIGN] RULE APPLIES, AND IT NOW HAS A FOURTH CONSUMER: "the restore destination is
|
||||
// resolved by the same rule as the capture destination — the drive if the app declares one, the system
|
||||
// data path otherwise."
|
||||
//
|
||||
// BUT THE TWO CALLERS ASK DIFFERENT QUESTIONS, and answering both with one predicate is the R-356
|
||||
// defect itself. So the fallback is scoped:
|
||||
//
|
||||
// - unitOnly=true (the R-87 proof): may fall back. `07` §7 records as [FACT] that a driveless app's
|
||||
// recovery unit ALREADY sits on `systemDataPath` indefinitely — "the SSD-only system-data
|
||||
// fallback" — and that the same-device placement is "intended, not a defect". The scratch is
|
||||
// bounded by that unit's own size and is deleted on every path.
|
||||
// - unitOnly=false (the customer's scratch): must NOT. A full restore pulls the app's bulk userdata,
|
||||
// and `07` §2.2 makes the internal SSD a STATE-ONLY tier. This is exactly the case the deleted
|
||||
// `PlaceOffsiteRestore` comment worried about, and the R-252 refusal below stays correct for it.
|
||||
//
|
||||
// The headroom gate still applies on the fallback path — it is the caller's `unitOnlyHeadroom`, which
|
||||
// refuses when the floor is not met, so a small system disk is protected by the same floor as a drive.
|
||||
func (m *Manager) offboxScratchDirIn(stack string, rootFor func(string) string, unitOnly bool) (scratch, nsRoot string, err error) {
|
||||
scratchFor := func(root string) (string, string) {
|
||||
return filepath.Join(rootFor(root), stack), m.namespaceRoot(root)
|
||||
}
|
||||
@@ -229,6 +266,16 @@ func (m *Manager) offboxScratchDirIn(stack string, rootFor func(string) string)
|
||||
}
|
||||
}
|
||||
}
|
||||
// (4) R-414: a UNIT-ONLY restore falls back to the system data path, which is where a driveless
|
||||
// app's unit already lives. Deliberately AFTER the network last-resort: a registered drive,
|
||||
// even a network one, is still a better scratch for ownership fidelity than the system disk.
|
||||
if unitOnly {
|
||||
if sysPath := strings.TrimSpace(m.cfg.Paths.SystemDataPath); sysPath != "" {
|
||||
s, nr := scratchFor(sysPath)
|
||||
m.logger.Printf("[INFO] [offbox] %s: no registered data drive — unit-only scratch falls back to the system data path %s (R-414; the unit already lives there)", stack, sysPath)
|
||||
return s, nr, nil
|
||||
}
|
||||
}
|
||||
// R-252: name the reason AND the way to act on it. This refusal is what a rebuilt box hits — the
|
||||
// drives are physically fine and still mounted, it is their REGISTRATION that the destroyed guest
|
||||
// took with it — and until v0.207.0 it said only that a drive was missing, which reads like data
|
||||
@@ -270,6 +317,27 @@ func (m *Manager) RestoreOffboxScratch(ctx context.Context, stack string, full b
|
||||
if !m.OffboxConfigured() {
|
||||
return fmt.Errorf("off-box backup not configured")
|
||||
}
|
||||
// R-411/R-408 — THE SINGLE-WRITER FLAG, and it must be taken HERE, before anything touches the
|
||||
// repository.
|
||||
//
|
||||
// WHAT IT COSTS TO OMIT IT, measured on demo-hp 2026-08-31 and not reasoned about: this function
|
||||
// runs `offboxSnapshotSize` for a full restore, which shells `restic stats` — and **`stats` TAKES
|
||||
// A REPOSITORY LOCK** (clean-room test: nothing else running, four invocations, the sampler reads
|
||||
// `locks=1`). Without this flag the integrity check is not blocked, starts, meets that lock, and
|
||||
// `resticStep` escalates to `unlock --remove-all` — the argv sampler caught `restore …` and
|
||||
// `unlock --remove-all` in the SAME sample at 20:50:51 — while logging *"a stale exclusive lock
|
||||
// left by a previous crash"*. There was no crash. `resticStep`'s own safety argument is that the
|
||||
// in-process mutex proves no sibling is live; this is the caller that made that false.
|
||||
//
|
||||
// BEFORE the snapshot lookup and the size probe, deliberately: a flag taken after the probe
|
||||
// protects nothing, because the probe is what takes the lock.
|
||||
//
|
||||
// The refusal shape matches the five siblings, so the handler's Hungarian wording is unchanged and
|
||||
// `restoreOpBlocked()` still refuses a second press exactly as it does today.
|
||||
if err := m.acquireRunning(); err != nil {
|
||||
return err
|
||||
}
|
||||
defer m.releaseRunning()
|
||||
if !isSafeStackName(stack) {
|
||||
return fmt.Errorf("invalid stack name")
|
||||
}
|
||||
@@ -425,6 +493,19 @@ func (m *Manager) OffboxRestorePrepareFull(ctx context.Context, stack string) (s
|
||||
if !m.OffboxConfigured() {
|
||||
return "", fmt.Errorf("off-box backup not configured")
|
||||
}
|
||||
// R-411 — THE SECOND ENTRY POINT, and the one the customer's UI actually reaches FIRST.
|
||||
//
|
||||
// The full restore is TWO HTTP requests: this one computes the size for the confirm screen, and a
|
||||
// later one does the restore. They are separate calls, so the flag taken in RestoreOffboxScratch
|
||||
// does not cover this, and NOTHING nests. `offboxSnapshotSize` below shells `restic stats`, which
|
||||
// takes a repository lock — so without this, the collision R-411 records is still reachable
|
||||
// through the ordinary two-step flow even after the restore itself is flagged.
|
||||
//
|
||||
// Found by re-reading the call graph while fixing the other one, not by the original report.
|
||||
if err := m.acquireRunning(); err != nil {
|
||||
return "", err
|
||||
}
|
||||
defer m.releaseRunning()
|
||||
if !isSafeStackName(stack) {
|
||||
return "", fmt.Errorf("invalid stack name")
|
||||
}
|
||||
|
||||
@@ -0,0 +1,224 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"go/ast"
|
||||
"go/parser"
|
||||
"go/token"
|
||||
"io/fs"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-408 — THE INVARIANT IS PINNED BY A WALK, NOT ASSERTED IN A COMMENT.
|
||||
//
|
||||
// THE DEFECT THIS EXISTS FOR IS NOT THE MISSING ACQUIRE. `offbox_integrity.go`'s header has said
|
||||
// *"Every off-site operation takes `acquireRunning` for exactly that reason"* since v0.227.0, nothing
|
||||
// ever checked it, and it was **false for months** — `RestoreOffboxScratch` and
|
||||
// `OffboxRestorePrepareFull` both issued restic commands without it. `resticStep`'s licence to run
|
||||
// `unlock --remove-all` rests entirely on that sentence being true, so a false sentence there is a
|
||||
// licence to delete a live operation's lock. Measured on demo-hp 2026-08-31, R-411.
|
||||
//
|
||||
// That is the ninth instance of this project's most-repeated class: a comment stating a guarantee the
|
||||
// code does not provide. The workspace rule is *"a comment asserting an invariant needs a test pinning
|
||||
// it, or it is a wish"*. **This is that test.**
|
||||
//
|
||||
// WHAT IT ASSERTS. Every function in `internal/backup` that reaches the off-site repository — by
|
||||
// calling `resticStep`, `resticBackupStep` or the raw exec seam `m.runner()` — must EITHER call
|
||||
// `acquireRunning` itself, OR be reachable only from a function that does, OR be registered below by
|
||||
// name with the reason it is exempt.
|
||||
//
|
||||
// IT IS AN AST WALK AND NOT `strings.Contains`, deliberately: a commented-out call still contains the
|
||||
// string. That is exactly how an earlier version of a test in this repo passed its own red-proof.
|
||||
|
||||
// offsiteExempt registers the functions that reach restic WITHOUT the flag, each with the reason.
|
||||
// Adding a line here is a deliberate act and should be argued in the commit that adds it.
|
||||
var offsiteExempt = map[string]string{
|
||||
// R-87's proof restore. It is READ-ONLY BY CONSTRUCTION: `--no-lock`, no `unlockStale`, and it goes
|
||||
// through `m.runner()` rather than `resticStep`, so the `unlock --remove-all` escalation is
|
||||
// unreachable rather than merely unlikely. It cannot remove anyone's lock and it cannot take one.
|
||||
// Its CALLER, `ProveOffboxUnit`, does take the flag — this is the inner helper.
|
||||
"restoreUnitReadOnly": "R-87: --no-lock, no unlockStale, via m.runner() — cannot take or remove a lock; its caller ProveOffboxUnit holds the flag",
|
||||
|
||||
// The lock-hygiene helpers themselves. They are only ever called from inside a function that
|
||||
// already holds the flag; flagging them would deadlock, since acquireRunning is not reentrant.
|
||||
"unlockStale": "lock hygiene, called only from inside a flag-holding caller; acquireRunning is not reentrant",
|
||||
"resticStep": "the shared step runner — its own doc comment records that every CALLER holds the flag, which is what this test pins",
|
||||
|
||||
// Probes and readers that take no lock. `restic snapshots` and `restic list` were measured on
|
||||
// demo-hp 2026-08-31 NOT to lock (6 back-to-back invocations, sampler read locks=0 throughout).
|
||||
"offboxLatestSnapshot": "restic snapshots — measured 2026-08-31 not to take a lock; always called from a flag-holding caller anyway",
|
||||
"ensureOffboxRepo": "restic cat config / init probe, called from inside flag-holding callers only",
|
||||
"offboxInventory": "restic snapshots --json, a read; no lock taken",
|
||||
// Read-only listing behind two web pages (the recovery page and the restore list). It issues only
|
||||
// `restic snapshots --json`, and `snapshots` was MEASURED on demo-hp 2026-08-31 not to take a lock
|
||||
// — six back-to-back invocations, the sampler read locks=0 throughout. Flagging it would make
|
||||
// browsing a page refuse while a backup runs, for no safety gain: it can neither take a lock nor
|
||||
// remove one.
|
||||
"OffsiteInventoryList": "restic snapshots --json only; snapshots measured 2026-08-31 not to lock, and it never routes through resticStep",
|
||||
}
|
||||
|
||||
// offsiteReachers are the calls that mean "this function talks to the off-site repository".
|
||||
var offsiteReachers = map[string]bool{
|
||||
"resticStep": true, "resticBackupStep": true, "runner": true,
|
||||
}
|
||||
|
||||
func r408WalkBackupPackage(t *testing.T) (map[string]*ast.FuncDecl, *token.FileSet) {
|
||||
t.Helper()
|
||||
root, err := filepath.Abs(".")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
fset := token.NewFileSet()
|
||||
fns := map[string]*ast.FuncDecl{}
|
||||
err = filepath.WalkDir(root, func(path string, d fs.DirEntry, werr error) error {
|
||||
if werr != nil {
|
||||
return werr
|
||||
}
|
||||
if d.IsDir() || !strings.HasSuffix(path, ".go") || strings.HasSuffix(path, "_test.go") {
|
||||
return nil
|
||||
}
|
||||
f, perr := parser.ParseFile(fset, path, nil, 0) // comments DROPPED on purpose
|
||||
if perr != nil {
|
||||
t.Fatalf("parse %s: %v — the R-408 invariant is now unasserted", path, perr)
|
||||
}
|
||||
for _, decl := range f.Decls {
|
||||
if fd, ok := decl.(*ast.FuncDecl); ok && fd.Body != nil {
|
||||
fns[fd.Name.Name] = fd
|
||||
}
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return fns, fset
|
||||
}
|
||||
|
||||
func r408Calls(fd *ast.FuncDecl) map[string]bool {
|
||||
out := map[string]bool{}
|
||||
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||
call, ok := n.(*ast.CallExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
switch fn := call.Fun.(type) {
|
||||
case *ast.Ident:
|
||||
out[fn.Name] = true
|
||||
case *ast.SelectorExpr:
|
||||
out[fn.Sel.Name] = true
|
||||
// `m.runner()(ctx, env, args...)` — the seam is invoked through the value it returns, so
|
||||
// the outer CallExpr's Fun is itself a CallExpr. Catch that shape explicitly.
|
||||
}
|
||||
if inner, ok := call.Fun.(*ast.CallExpr); ok {
|
||||
if sel, ok := inner.Fun.(*ast.SelectorExpr); ok {
|
||||
out[sel.Sel.Name] = true
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
return out
|
||||
}
|
||||
|
||||
// TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered — B1.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): removing the `acquireRunning` from
|
||||
// `RestoreOffboxScratch` makes this fail naming that function; adding an unregistered fake entry point
|
||||
// that calls `m.resticStep(...)` makes it fail naming the fake.
|
||||
func TestR408_EveryOffsiteEntryPointTakesTheFlagOrIsRegistered(t *testing.T) {
|
||||
fns, _ := r408WalkBackupPackage(t)
|
||||
|
||||
calls := map[string]map[string]bool{}
|
||||
for name, fd := range fns {
|
||||
calls[name] = r408Calls(fd)
|
||||
}
|
||||
|
||||
// A function is COVERED if it acquires the flag itself, or every path to it is through a function
|
||||
// that does. Fixed-point: start from the self-acquirers and propagate to their callees.
|
||||
covered := map[string]bool{}
|
||||
for name, c := range calls {
|
||||
if c["acquireRunning"] {
|
||||
covered[name] = true
|
||||
}
|
||||
}
|
||||
if len(covered) == 0 {
|
||||
t.Fatal("no function in internal/backup calls acquireRunning — the walk is looking in the wrong place, and a green here would be meaningless")
|
||||
}
|
||||
for i := 0; i < 12; i++ { // depth cap; the call graph here is shallow
|
||||
grew := false
|
||||
for name := range covered {
|
||||
for callee := range calls[name] {
|
||||
if _, ours := fns[callee]; ours && !covered[callee] {
|
||||
covered[callee] = true
|
||||
grew = true
|
||||
}
|
||||
}
|
||||
}
|
||||
if !grew {
|
||||
break
|
||||
}
|
||||
}
|
||||
|
||||
var offenders []string
|
||||
for name, c := range calls {
|
||||
reaches := false
|
||||
for r := range offsiteReachers {
|
||||
if c[r] {
|
||||
reaches = true
|
||||
break
|
||||
}
|
||||
}
|
||||
if !reaches || covered[name] || offsiteExempt[name] != "" {
|
||||
continue
|
||||
}
|
||||
offenders = append(offenders, name)
|
||||
}
|
||||
sort.Strings(offenders)
|
||||
|
||||
if len(offenders) > 0 {
|
||||
t.Fatalf("these functions reach the off-site repository without the single-writer flag and are not registered exempt: %v\n\n"+
|
||||
"`resticStep` escalates to `unlock --remove-all` on a lock error, and its licence to do that is\n"+
|
||||
"that every caller holds the flag. R-411 measured what happens when one does not: a live\n"+
|
||||
"customer restore's lock was deleted and logged as a crash that never happened.\n\n"+
|
||||
"Take acquireRunning, or add the function to offsiteExempt WITH the reason it cannot collide.", offenders)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR408_TheProofsReadOnlyPathIsARegisteredException — B2.
|
||||
//
|
||||
// The exemption must stay HONEST: `restoreUnitReadOnly` is exempt only because it is read-only. If it
|
||||
// ever gains `unlockStale`, or routes through `resticStep`, or loses `--no-lock`, the exemption is a
|
||||
// lie and this fails.
|
||||
func TestR408_TheProofsReadOnlyPathIsARegisteredException(t *testing.T) {
|
||||
if offsiteExempt["restoreUnitReadOnly"] == "" {
|
||||
t.Fatal("restoreUnitReadOnly must be a REGISTERED exception, with its reason, not silently absent")
|
||||
}
|
||||
fns, _ := r408WalkBackupPackage(t)
|
||||
fd := fns["restoreUnitReadOnly"]
|
||||
if fd == nil {
|
||||
t.Fatal("restoreUnitReadOnly is gone — the exemption now covers nothing and must be removed")
|
||||
}
|
||||
c := r408Calls(fd)
|
||||
if c["resticStep"] || c["resticBackupStep"] {
|
||||
t.Fatal("restoreUnitReadOnly now routes through resticStep — the unlock --remove-all escalation is reachable and the exemption is no longer true")
|
||||
}
|
||||
if c["unlockStale"] {
|
||||
t.Fatal("restoreUnitReadOnly now calls unlockStale — it issues a delete verb and the exemption is no longer true")
|
||||
}
|
||||
// And it must still pass --no-lock.
|
||||
var sawNoLock bool
|
||||
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||
if lit, ok := n.(*ast.BasicLit); ok && strings.Trim(lit.Value, `"`) == "--no-lock" {
|
||||
sawNoLock = true
|
||||
}
|
||||
return true
|
||||
})
|
||||
if !sawNoLock {
|
||||
t.Fatal("restoreUnitReadOnly no longer passes --no-lock — it can now take a lock and the exemption is no longer true")
|
||||
}
|
||||
// Its caller must hold the flag, or the exemption rests on nothing.
|
||||
if !r408Calls(fns["ProveOffboxUnit"])["acquireRunning"] {
|
||||
t.Fatal("ProveOffboxUnit no longer takes the flag — restoreUnitReadOnly's exemption depends on its caller holding it")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,179 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||
)
|
||||
|
||||
// R-411 / R-408 — the single-writer flag on the customer restore paths.
|
||||
//
|
||||
// These drive the real functions through the restic exec seam (`SetOffboxRunner`), which REUSE.md
|
||||
// names as the way to see the argv — replacing `resticStep` would hide the `unlock --remove-all`
|
||||
// escalation from exactly the assertion that must see it.
|
||||
|
||||
// lockHarness records every restic argv and lets a test hold the flag.
|
||||
type lockHarness struct {
|
||||
m *Manager
|
||||
mu sync.Mutex
|
||||
argv [][]string
|
||||
}
|
||||
|
||||
func newLockHarness(t *testing.T, stacks ...string) *lockHarness {
|
||||
t.Helper()
|
||||
m, sett := newOffboxManager(t)
|
||||
drive := t.TempDir()
|
||||
if err := sett.AddStoragePath(settings.StoragePath{Path: drive, Label: "data", Schedulable: true}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
deployed := map[string]bool{}
|
||||
for _, s := range stacks {
|
||||
deployed[s] = true
|
||||
}
|
||||
m.SetStackProvider(&offbox3aProvider{
|
||||
hdd: map[string]string{}, binds: map[string][]ClassifiedBind{},
|
||||
has: map[string]bool{}, deployed: deployed,
|
||||
})
|
||||
h := &lockHarness{m: m}
|
||||
m.SetOffboxLatestSnapshotFn(func(_ context.Context, stack string) (string, []string, error) {
|
||||
return "snap-" + stack, []string{"/mnt/sys_drive/felhom-data/backups/primary/" + stack}, nil
|
||||
})
|
||||
m.SetOffboxRunner(func(_ context.Context, _ []string, args ...string) ([]byte, error) {
|
||||
h.mu.Lock()
|
||||
h.argv = append(h.argv, append([]string{}, args...))
|
||||
h.mu.Unlock()
|
||||
return []byte("{}"), nil
|
||||
})
|
||||
m.SetOffboxFreeFn(func(string) int64 { return 100 << 30 })
|
||||
return h
|
||||
}
|
||||
|
||||
func (h *lockHarness) allArgs() [][]string {
|
||||
h.mu.Lock()
|
||||
defer h.mu.Unlock()
|
||||
return h.argv
|
||||
}
|
||||
|
||||
// TestR411_ScratchRestoreTakesTheSingleWriterFlag — A1.
|
||||
func TestR411_ScratchRestoreTakesTheSingleWriterFlag(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
// Hold the flag as a sibling operation would.
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatalf("could not take the flag: %v", err)
|
||||
}
|
||||
err := h.m.RestoreOffboxScratch(context.Background(), "kimai", false)
|
||||
if err == nil {
|
||||
t.Fatal("a scratch restore must REFUSE while the single-writer flag is held — before this fix it ran anyway, which is R-411")
|
||||
}
|
||||
if len(h.allArgs()) != 0 {
|
||||
t.Fatalf("a refused restore must invoke restic ZERO times; got %v", h.allArgs())
|
||||
}
|
||||
}
|
||||
|
||||
// TestR411_PrepareFullAlsoTakesTheFlag — the SECOND entry point, found while fixing the first.
|
||||
// The customer's real UI flow reaches this one first, and it is the one that shells `restic stats`.
|
||||
func TestR411_PrepareFullAlsoTakesTheFlag(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := h.m.OffboxRestorePrepareFull(context.Background(), "kimai"); err == nil {
|
||||
t.Fatal("the full-restore PREPARE must refuse while the flag is held — it shells `restic stats`, which takes a repository lock")
|
||||
}
|
||||
if len(h.allArgs()) != 0 {
|
||||
t.Fatalf("a refused prepare must invoke restic ZERO times; got %v", h.allArgs())
|
||||
}
|
||||
}
|
||||
|
||||
// TestR411_SharesScratchRestoreAlsoTakesTheFlag — the third, found by the R-408 walk.
|
||||
func TestR411_SharesScratchRestoreAlsoTakesTheFlag(t *testing.T) {
|
||||
h := newLockHarness(t)
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := h.m.RestoreSharesScratch(context.Background()); err == nil {
|
||||
t.Fatal("the shares scratch restore must refuse while the flag is held — it runs unlockStale + resticStep, R-411's exact shape")
|
||||
}
|
||||
if len(h.allArgs()) != 0 {
|
||||
t.Fatalf("a refused shares restore must invoke restic ZERO times; got %v", h.allArgs())
|
||||
}
|
||||
}
|
||||
|
||||
// TestR411_NoUnlockRemoveAllInAnyArgv — A3, the non-effect that matters.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01): removing the acquire from `RestoreOffboxScratch` lets the restore run
|
||||
// while the flag is held, so an integrity check could collide with it — the state this asserts is
|
||||
// impossible. The direct form of the red-proof is `TestR408_…`, which names the function.
|
||||
func TestR411_NoUnlockRemoveAllInAnyArgv(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
if err := h.m.RestoreOffboxScratch(context.Background(), "kimai", false); err != nil {
|
||||
t.Fatalf("an idle-box restore must succeed: %v", err)
|
||||
}
|
||||
for _, args := range h.allArgs() {
|
||||
joined := strings.Join(args, " ")
|
||||
if strings.Contains(joined, "--remove-all") {
|
||||
t.Fatalf("the escalation fired during an ordinary restore: %s", joined)
|
||||
}
|
||||
}
|
||||
if len(h.allArgs()) == 0 {
|
||||
t.Fatal("no restic invocation at all — the assertion above ran over nothing")
|
||||
}
|
||||
}
|
||||
|
||||
// TestR411_RestoreStillWorksWhenIdle — Scenario C. The flag must not make the common case refuse.
|
||||
func TestR411_RestoreStillWorksWhenIdle(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
if err := h.m.RestoreOffboxScratch(context.Background(), "kimai", false); err != nil {
|
||||
t.Fatalf("a restore on an idle box must still work: %v", err)
|
||||
}
|
||||
var sawRestore bool
|
||||
for _, args := range h.allArgs() {
|
||||
if containsArg(args, "restore") {
|
||||
sawRestore = true
|
||||
}
|
||||
}
|
||||
if !sawRestore {
|
||||
t.Fatal("no restore was issued")
|
||||
}
|
||||
}
|
||||
|
||||
// TestR411_FlagIsReleasedOnEveryPath — A5. A flag that leaks would wedge every nightly job.
|
||||
func TestR411_FlagIsReleasedOnEveryPath(t *testing.T) {
|
||||
t.Run("success", func(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
if err := h.m.RestoreOffboxScratch(context.Background(), "kimai", false); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal("the flag was NOT released after a successful restore")
|
||||
}
|
||||
})
|
||||
t.Run("restic error", func(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
h.m.SetOffboxRunner(func(_ context.Context, _ []string, _ ...string) ([]byte, error) {
|
||||
return []byte("boom"), os.ErrPermission
|
||||
})
|
||||
_ = h.m.RestoreOffboxScratch(context.Background(), "kimai", false)
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal("the flag was NOT released after a failed restore")
|
||||
}
|
||||
})
|
||||
t.Run("early refusal", func(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
_ = h.m.RestoreOffboxScratch(context.Background(), "no such stack!!", false)
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal("the flag was NOT released after an early refusal")
|
||||
}
|
||||
})
|
||||
t.Run("prepare", func(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
_, _ = h.m.OffboxRestorePrepareFull(context.Background(), "kimai")
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal("the flag was NOT released after the prepare step")
|
||||
}
|
||||
})
|
||||
}
|
||||
@@ -0,0 +1,115 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"log"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-412 leg 1 — a push that carried nothing must not read as a plain success.
|
||||
//
|
||||
// RUN-LEVEL on purpose, and the sibling R-234 file records why: its first version asserted the
|
||||
// capture helper alone, and its red-proof passed while the defect was untouched. The line under test
|
||||
// is emitted inside the per-app loop of the real run, so the test drives the real run and reads the
|
||||
// real logger.
|
||||
//
|
||||
// WORDING ONLY. This asserts what the run SAYS, not that it refuses — whether the push should re-read
|
||||
// the unit before sending is R-412 leg 2 and is deliberately still open.
|
||||
|
||||
// writeUnitManifest gives a unit a manifest declaring exactly what is asked for.
|
||||
func writeUnitManifest(t *testing.T, unitDir string, dbDumps, volDumps []string) {
|
||||
t.Helper()
|
||||
b, err := json.Marshal(RecoveryManifest{DBDumps: dbDumps, VolumeDumps: volDumps})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(unitDir, "manifest.json"), b, 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR412a_EmptyPushDoesNotReadAsAPlainSuccess — D1.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): restoring the single unconditional
|
||||
// `[INFO] backed up %s (%s, %d mandatory path(s))` makes this fail — the run logs a plain success over
|
||||
// a snapshot holding none of the app's data, which is exactly what was measured on demo-hp on
|
||||
// 2026-08-31.
|
||||
func TestR412a_EmptyPushDoesNotReadAsAPlainSuccess(t *testing.T) {
|
||||
drive := t.TempDir()
|
||||
m, sett, prov := classifiedOffboxManager(t, drive)
|
||||
|
||||
var buf bytes.Buffer
|
||||
m.logger = log.New(&buf, "", 0)
|
||||
|
||||
// `hollow` has a unit whose manifest declares NOTHING — the R-403 shape.
|
||||
hollow := mkUnit(t, drive, "hollow")
|
||||
writeUnitManifest(t, hollow, nil, nil)
|
||||
prov.hdd["hollow"] = drive
|
||||
prov.has["hollow"] = true
|
||||
|
||||
// `sound` has a unit that declares real data, so the contrast is in the same run.
|
||||
sound := mkUnit(t, drive, "sound")
|
||||
writeUnitManifest(t, sound, []string{"sound-postgres.sql"}, []string{"sound_data.tar"})
|
||||
prov.hdd["sound"] = drive
|
||||
prov.has["sound"] = true
|
||||
|
||||
prov.deployed = map[string]bool{"hollow": true, "sound": true}
|
||||
_ = sett.SetAppOffbox("hollow", true)
|
||||
_ = sett.SetAppOffbox("sound", true)
|
||||
|
||||
cap := &backupCapture{}
|
||||
m.SetOffboxRunner(cap.runner())
|
||||
if err := m.RunOffboxBackup(context.Background()); err != nil {
|
||||
t.Fatalf("the run itself must still succeed — this is a wording change, not a guard: %v", err)
|
||||
}
|
||||
|
||||
out := buf.String()
|
||||
|
||||
// POSITIVE CONTROL FIRST: the run must actually have logged about both apps, or every assertion
|
||||
// below is over an empty string.
|
||||
if !strings.Contains(out, "hollow") || !strings.Contains(out, "sound") {
|
||||
t.Fatalf("the run logged about neither app — the assertions below would prove nothing.\n%s", out)
|
||||
}
|
||||
// NEGATIVE CONTROL: a string that cannot be there.
|
||||
if strings.Contains(out, "ZZZ-NOT-IN-THE-LOG") {
|
||||
t.Fatal("negative control matched — the search is not discriminating")
|
||||
}
|
||||
|
||||
// The hollow app's line must say what it did NOT carry. ASCII fragments (R-364).
|
||||
var hollowLine string
|
||||
for _, l := range strings.Split(out, "\n") {
|
||||
if strings.Contains(l, "backed up hollow") {
|
||||
hollowLine = l
|
||||
}
|
||||
}
|
||||
if hollowLine == "" {
|
||||
t.Fatalf("no per-app push line for the hollow app at all.\n%s", out)
|
||||
}
|
||||
for _, frag := range []string{"NO database dump", "NO volume tar", "none of the app"} {
|
||||
if !strings.Contains(hollowLine, frag) {
|
||||
t.Fatalf("the hollow push line must say what it did not carry; %q missing from:\n %s", frag, hollowLine)
|
||||
}
|
||||
}
|
||||
if strings.Contains(hollowLine, "[INFO]") {
|
||||
t.Fatalf("a push carrying no data must not be logged at INFO like an ordinary success:\n %s", hollowLine)
|
||||
}
|
||||
|
||||
// And the SOUND app's line must be untouched — the change must not relabel healthy runs.
|
||||
var soundLine string
|
||||
for _, l := range strings.Split(out, "\n") {
|
||||
if strings.Contains(l, "backed up sound") {
|
||||
soundLine = l
|
||||
}
|
||||
}
|
||||
if soundLine == "" {
|
||||
t.Fatalf("no per-app push line for the sound app.\n%s", out)
|
||||
}
|
||||
if strings.Contains(soundLine, "NO database dump") {
|
||||
t.Fatalf("a healthy push must NOT carry the empty-push wording — a warning that fires on everything costs the same as the comforting lie it replaces:\n %s", soundLine)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,199 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
|
||||
)
|
||||
|
||||
// R-411 A2 + R-414 — the collision that can no longer happen, and the box that can now be proved.
|
||||
|
||||
// TestR411_IntegrityCheckSkipsWhileARestoreHoldsIt — A2, the whole point of Part 1.
|
||||
//
|
||||
// This is last night's collision, from the other side: with the restore now holding the flag, the
|
||||
// integrity check must SKIP rather than run, meet the lock and delete it.
|
||||
func TestR411_IntegrityCheckSkipsWhileARestoreHoldsIt(t *testing.T) {
|
||||
h := newLockHarness(t, "kimai")
|
||||
|
||||
// Stand in for a restore in flight: it now takes the flag, so hold it.
|
||||
if err := h.m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
before := h.m.settings.GetOffboxTarget().LastIntegrityCheck
|
||||
|
||||
res := h.m.CheckOffboxIntegrity(context.Background())
|
||||
|
||||
if !res.Skipped {
|
||||
t.Fatalf("the check must SKIP while a restore holds the flag; got %+v", res)
|
||||
}
|
||||
if len(h.allArgs()) != 0 {
|
||||
t.Fatalf("a skipped check must invoke restic ZERO times — this is the non-effect R-411 is about; got %v", h.allArgs())
|
||||
}
|
||||
for _, args := range h.allArgs() {
|
||||
if containsArg(args, "unlock") || containsArg(args, "--remove-all") {
|
||||
t.Fatalf("the unlock escalation fired: %v", args)
|
||||
}
|
||||
}
|
||||
if got := h.m.settings.GetOffboxTarget().LastIntegrityCheck; got != before {
|
||||
t.Fatalf("a skip must NOT advance due-ness; %q -> %q", before, got)
|
||||
}
|
||||
}
|
||||
|
||||
// ── R-414 ───────────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
// drivelessHarness is `demo-felhom`'s real shape: deployed apps, and ZERO registered storage paths.
|
||||
func drivelessHarness(t *testing.T, stacks ...string) (*Manager, *settings.Settings, *[][]string) {
|
||||
t.Helper()
|
||||
m, sett := newOffboxManager(t)
|
||||
// deliberately NO AddStoragePath — that is the whole point
|
||||
deployed := map[string]bool{}
|
||||
for _, s := range stacks {
|
||||
deployed[s] = true
|
||||
}
|
||||
m.SetStackProvider(&offbox3aProvider{
|
||||
hdd: map[string]string{}, binds: map[string][]ClassifiedBind{},
|
||||
has: map[string]bool{}, deployed: deployed,
|
||||
})
|
||||
var argv [][]string
|
||||
m.SetOffboxLatestSnapshotFn(func(_ context.Context, stack string) (string, []string, error) {
|
||||
return "snap-" + stack, []string{"/mnt/sys_drive/felhom-data/backups/primary/" + stack}, nil
|
||||
})
|
||||
m.SetOffboxRunner(func(_ context.Context, _ []string, args ...string) ([]byte, error) {
|
||||
argv = append(argv, append([]string{}, args...))
|
||||
if containsArg(args, "restore") {
|
||||
target, include := argValue(args, "--target"), argValue(args, "--include")
|
||||
if target != "" && include != "" {
|
||||
dest := filepath.Join(target, strings.TrimPrefix(include, string(filepath.Separator)))
|
||||
materialiseUnit(t, dest, unitFixture{})
|
||||
}
|
||||
}
|
||||
return []byte("{}"), nil
|
||||
})
|
||||
m.SetOffboxFreeFn(func(string) int64 { return 100 << 30 })
|
||||
return m, sett, &argv
|
||||
}
|
||||
|
||||
// TestR414_DrivelessBoxReachesAVerdict — C1.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01): reverting the resolver to refuse (dropping step 4) AND restoring the
|
||||
// Err path makes this fail with `last_proof_result` absent — which is the exact state `demo-felhom`
|
||||
// was in every night.
|
||||
func TestR414_DrivelessBoxReachesAVerdict(t *testing.T) {
|
||||
m, sett, _ := drivelessHarness(t, "opengist")
|
||||
|
||||
res := m.ProveOffboxUnit(context.Background())
|
||||
if res.Verdict() == "" {
|
||||
t.Fatalf("a driveless box must reach a RECORDED verdict, never an empty one — empty is what the hub reads as 'controller too old'; got %+v", res)
|
||||
}
|
||||
m.RecordProofVerdict(res)
|
||||
|
||||
got := sett.GetOffboxTarget().LastProofResult
|
||||
if got == "" {
|
||||
t.Fatal("last_proof_result is ABSENT after a run on a driveless box — this is R-414 exactly: the hub cannot tell 'cannot run here' from 'too old'")
|
||||
}
|
||||
// With the unit-only fallback in place the box can actually be proved.
|
||||
if got != string(UnitProofPass) && got != ProofResultCannotRun {
|
||||
t.Fatalf("unexpected verdict %q — expected a real judgement (the fallback worked) or %q", got, ProofResultCannotRun)
|
||||
}
|
||||
t.Logf("driveless box reached verdict %q", got)
|
||||
}
|
||||
|
||||
// TestR414_UnitOnlyRestoreFallsBackToSystemData — C3.
|
||||
func TestR414_UnitOnlyRestoreFallsBackToSystemData(t *testing.T) {
|
||||
m, _, _ := drivelessHarness(t, "opengist")
|
||||
scratch, nsRoot, err := m.offboxProofScratchDir("opengist")
|
||||
if err != nil {
|
||||
t.Fatalf("a UNIT-ONLY scratch must resolve on a driveless box (R-414); got %v", err)
|
||||
}
|
||||
sys := m.cfg.Paths.SystemDataPath
|
||||
if !strings.HasPrefix(filepath.Clean(scratch), filepath.Clean(sys)) {
|
||||
t.Fatalf("the unit-only scratch must fall back to the system data path %q; got %q", sys, scratch)
|
||||
}
|
||||
if nsRoot == "" {
|
||||
t.Fatal("the namespace root must be returned for the free-space probe")
|
||||
}
|
||||
if !strings.Contains(scratch, "offsite-proof") {
|
||||
t.Fatalf("it must still land in the PROOF root, not the customer's; got %q", scratch)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR414_FullRestoreDoesNotFallBack — C4. The state-only tier is protected.
|
||||
func TestR414_FullRestoreDoesNotFallBack(t *testing.T) {
|
||||
m, _, _ := drivelessHarness(t, "opengist")
|
||||
_, _, err := m.offboxRestoreScratchDir("opengist")
|
||||
if err == nil {
|
||||
t.Fatal("the CUSTOMER's scratch must still refuse on a driveless box — a full restore pulls bulk userdata and the internal SSD is a state-only tier (07 §2.2)")
|
||||
}
|
||||
// ASCII fragment, with a negative control, because an accented grep has returned 0 for strings
|
||||
// that were there (R-364).
|
||||
if !strings.Contains(err.Error(), "adatmeghajt") {
|
||||
t.Fatalf("the R-252 refusal wording must be preserved — it tells the customer what to do; got %q", err.Error())
|
||||
}
|
||||
if strings.Contains(err.Error(), "ZZZ-NOT-IN-THE-MESSAGE") {
|
||||
t.Fatal("negative control matched — the fragment search is not discriminating")
|
||||
}
|
||||
}
|
||||
|
||||
// TestR414_AbsentStillMeansNotRecorded — C2. The two meanings must stay distinct.
|
||||
func TestR414_AbsentStillMeansNotRecorded(t *testing.T) {
|
||||
m, sett, _ := drivelessHarness(t, "opengist")
|
||||
if got := sett.GetOffboxTarget().LastProofResult; got != "" {
|
||||
t.Fatalf("a box that has never proved must report ABSENT; got %q", got)
|
||||
}
|
||||
// A skip must still record nothing — only a reached outcome writes.
|
||||
if err := m.AcquireRunningForTest(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
res := m.ProveOffboxUnit(context.Background())
|
||||
m.RecordProofVerdict(res)
|
||||
if got := sett.GetOffboxTarget().LastProofResult; got != "" {
|
||||
t.Fatalf("a SKIP must leave the field absent — it looked at nothing; got %q", got)
|
||||
}
|
||||
m.ReleaseRunningForTest()
|
||||
|
||||
// And the wire must not carry the key at all when absent.
|
||||
b, err := json.Marshal(&OffboxReportStatus{})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if strings.Contains(string(b), "last_proof_result") {
|
||||
t.Fatalf("an empty status must not emit last_proof_result; got %s", b)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR414_CannotRunIsRecordedButDoesNotAdvanceDueness — the half that keeps the app retryable.
|
||||
func TestR414_CannotRunIsRecordedButDoesNotAdvanceDueness(t *testing.T) {
|
||||
m, sett, _ := drivelessHarness(t, "opengist")
|
||||
res := ProofResult{Stack: "opengist", SnapshotID: "snap-opengist",
|
||||
CannotRun: true, CannotWhy: "nincs regisztralt adatmeghajto"}
|
||||
if res.Verdict() != ProofResultCannotRun {
|
||||
t.Fatalf("cannot-run must render as %q, never as empty; got %q", ProofResultCannotRun, res.Verdict())
|
||||
}
|
||||
m.RecordProofVerdict(res)
|
||||
tt := sett.GetOffboxTarget()
|
||||
if tt.LastProofResult != ProofResultCannotRun {
|
||||
t.Fatalf("cannot-run must be RECORDED so the hub can see it; got %q", tt.LastProofResult)
|
||||
}
|
||||
if len(tt.ProvedSnapshots) != 0 {
|
||||
t.Fatalf("cannot-run must NOT advance per-snapshot due-ness — nothing was proved, and marking it proved would stop the app ever being retried; got %v", tt.ProvedSnapshots)
|
||||
}
|
||||
if tt.LastProofSnapshot != "" {
|
||||
t.Fatalf("cannot-run must not claim a proved snapshot; got %q", tt.LastProofSnapshot)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR414_NoCustomerAlarm — C5. The customer's backups are fine and there is nothing for them to do.
|
||||
func TestR414_NoCustomerAlarm(t *testing.T) {
|
||||
m, _, _ := drivelessHarness(t, "opengist")
|
||||
var pushed int
|
||||
m.SetOffboxOrphanEvent(func(string, string) { pushed++ })
|
||||
res := m.ProveOffboxUnit(context.Background())
|
||||
m.RecordProofVerdict(res)
|
||||
if pushed != 0 {
|
||||
t.Fatalf("a driveless box must raise no event from the backup layer; got %d", pushed)
|
||||
}
|
||||
}
|
||||
@@ -85,6 +85,17 @@ func (m *Manager) sharesRestoreScratchDir() (scratch, nsRoot string, err error)
|
||||
// it never touches a live share folder, the registry, or the credential — PlaceSharesRestore is the
|
||||
// deliberate second action that does.
|
||||
func (m *Manager) RestoreSharesScratch(ctx context.Context) error {
|
||||
// R-411 — FOUND BY THE R-408 WALK, not by the report that prompted it. This is the SAME shape as
|
||||
// the off-site scratch restore: it runs `unlockStale` (a delete verb against `locks/`) and then
|
||||
// `resticStep`, whose `unlock --remove-all` escalation is licensed only by every caller holding
|
||||
// this flag. Its sibling `PlaceSharesRestore` below has always taken it; this one never did.
|
||||
//
|
||||
// The handler's `restoreOpBlocked()` + `BeginRestoreOp` set the DISPLAY flag (`opRunning`), not
|
||||
// this one, so nothing nests and `acquireRunning` is safe to take here.
|
||||
if err := m.acquireRunning(); err != nil {
|
||||
return err
|
||||
}
|
||||
defer m.releaseRunning()
|
||||
if !m.OffboxConfigured() {
|
||||
return fmt.Errorf("a távoli mentés nincs beállítva")
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user