e43b5ec07d
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged. THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does, nightly, on one app. IT DOES NOT prove a restore puts data back into a running app. That stays drill work and 07 section 8 matrix row 4 is NOT moved. THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So: (1) everything declared is present, AND (2) the manifest declares what the app is supposed to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test read verdict "pass". THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate the app's shape, and GetDockerVolumes describes the running app. Database half is DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are <project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose file's parent dir, which inside a unit is the literal string "compose". Measured on all eight real units on demo-hp the counts match exactly and the naming held every time - but "held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule that is invented. THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes that test read verdict "fail". IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect: --no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test fail on "unlock" appearing in the argv. IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch does not take it (R-408) while offbox_integrity.go states that invariant as universal. DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's app. ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app> would mean a nightly background job deleting the verification copy a CUSTOMER is looking at. It is also invisible to placement, so a proof copy can never be pushed into a live app. SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's behaviour is unchanged. NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is sound and the content is absent: different cause, different action. The hub half shipped FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted type is 400'd and vanishes. 33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller gates OK. Five red-proofs run and recorded in REPORT.md. A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
154 lines
6.2 KiB
Go
154 lines
6.2 KiB
Go
package backup
|
|
|
|
import (
|
|
"fmt"
|
|
"os"
|
|
"path/filepath"
|
|
"sort"
|
|
"strings"
|
|
"time"
|
|
)
|
|
|
|
// Verification copies — the listing/delete surface for `<nsRoot>/backups/offsite-restore/<app>`.
|
|
//
|
|
// WHY THIS EXISTS (v0.147.0, feedback slice 4a): an offsite verification restore wrote its result to
|
|
// a path the customer was never told, and nothing anywhere listed what had accumulated. Pressing
|
|
// „Ellenőrző visszaállítás" produced a flash saying it had been restored "to a verification folder
|
|
// on the drive" — which folder, on which drive, and how much space it was now using were all
|
|
// invisible. So copies piled up and the only way to find them was SSH.
|
|
//
|
|
// The path segments were already open-coded in three places; offsiteRestoreRootFor() is now the one
|
|
// place `backups/offsite-restore` is spelled, and offboxRestoreScratchDir() builds on it.
|
|
|
|
// OffsiteRestoreCopy is one verification copy on disk.
|
|
type OffsiteRestoreCopy struct {
|
|
Stack string `json:"stack"` // app slug, or SharesPseudoStack for the shares copy
|
|
Path string `json:"path"` // absolute path — the thing the customer could not see
|
|
Size int64 `json:"size"` // bytes
|
|
SizeHuman string `json:"size_human"` // pre-humanized for the template
|
|
Created time.Time `json:"created"` // dir mtime; restic writes the tree once, so this is the restore time
|
|
}
|
|
|
|
// offsiteRestoreRootFor returns `<nsRoot>/backups/offsite-restore` for a drive path. THE single place
|
|
// these segments are written.
|
|
func (m *Manager) offsiteRestoreRootFor(drivePath string) string {
|
|
return filepath.Join(m.namespaceRoot(drivePath), "backups", "offsite-restore")
|
|
}
|
|
|
|
// offsiteProofRootFor returns `<nsRoot>/backups/offsite-proof` for a drive path. THE single place
|
|
// these segments are written (R-87), and deliberately a SIBLING of offsite-restore rather than a
|
|
// subdirectory of it: nothing that lists, offers or deletes a customer verification copy walks this
|
|
// root, which is the whole point — see offboxProofScratchDir for the delete it prevents.
|
|
func (m *Manager) offsiteProofRootFor(drivePath string) string {
|
|
return filepath.Join(m.namespaceRoot(drivePath), "backups", "offsite-proof")
|
|
}
|
|
|
|
// offsiteRestoreDriveRoots returns every drive path a verification copy could live under, in the same
|
|
// preference order offboxRestoreScratchDir uses to CHOOSE one — so listing can never miss a copy the
|
|
// restore path was capable of creating. Deduplicated, order preserved.
|
|
func (m *Manager) offsiteRestoreDriveRoots() []string {
|
|
seen := map[string]bool{}
|
|
var roots []string
|
|
add := func(p string) {
|
|
p = strings.TrimSpace(p)
|
|
if p == "" || seen[p] {
|
|
return
|
|
}
|
|
seen[p] = true
|
|
roots = append(roots, p)
|
|
}
|
|
// App HDDs first (offboxRestoreScratchDir's rule 1), then every schedulable path (rules 2 and 3).
|
|
if m.stackProvider != nil {
|
|
for _, s := range m.stackProvider.ListDeployedStacks() {
|
|
add(m.stackProvider.GetStackHDDPath(s.Name))
|
|
}
|
|
}
|
|
if m.settings != nil {
|
|
for _, sp := range m.settings.GetSchedulableStoragePaths() {
|
|
add(sp.Path)
|
|
}
|
|
}
|
|
return roots
|
|
}
|
|
|
|
// ListOffsiteRestoreCopies enumerates every verification copy across every candidate drive, newest
|
|
// first. Missing directories are not an error — "none yet" is the normal state.
|
|
func (m *Manager) ListOffsiteRestoreCopies() []OffsiteRestoreCopy {
|
|
sizer := m.offboxSize()
|
|
var out []OffsiteRestoreCopy
|
|
seen := map[string]bool{}
|
|
for _, drive := range m.offsiteRestoreDriveRoots() {
|
|
root := m.offsiteRestoreRootFor(drive)
|
|
entries, err := os.ReadDir(root)
|
|
if err != nil {
|
|
continue // no copies on this drive (or the drive is not mounted) — not an error
|
|
}
|
|
for _, e := range entries {
|
|
if !e.IsDir() {
|
|
continue
|
|
}
|
|
p := filepath.Join(root, e.Name())
|
|
if seen[p] {
|
|
continue // two stacks can resolve to the same drive; list each path once
|
|
}
|
|
seen[p] = true
|
|
c := OffsiteRestoreCopy{Stack: e.Name(), Path: p}
|
|
if fi, err := e.Info(); err == nil {
|
|
c.Created = fi.ModTime()
|
|
}
|
|
c.Size = sizer(p)
|
|
c.SizeHuman = humanizeBytes(c.Size)
|
|
out = append(out, c)
|
|
}
|
|
}
|
|
sort.Slice(out, func(i, j int) bool { return out[i].Created.After(out[j].Created) })
|
|
return out
|
|
}
|
|
|
|
// DeleteOffsiteRestoreCopy removes ONE verification copy.
|
|
//
|
|
// This is the only delete path v0.147.0 adds, so it is guarded twice over. The stack name must pass
|
|
// isSafeStackName (no separators, no traversal), and the resolved path must sit STRICTLY INSIDE a
|
|
// `backups/offsite-restore` root that this Manager itself computed — a path that merely looks right
|
|
// is refused. Both checks are on the RESOLVED path, not the input, so a symlinked scratch cannot
|
|
// walk the delete out of the sandbox.
|
|
func (m *Manager) DeleteOffsiteRestoreCopy(stack string) error {
|
|
if !isSafeStackName(stack) {
|
|
return fmt.Errorf("érvénytelen alkalmazásnév")
|
|
}
|
|
for _, drive := range m.offsiteRestoreDriveRoots() {
|
|
root := m.offsiteRestoreRootFor(drive)
|
|
target := filepath.Join(root, stack)
|
|
|
|
fi, err := os.Stat(target)
|
|
if err != nil || !fi.IsDir() {
|
|
continue
|
|
}
|
|
// Prefix safety: only ever remove strictly inside `backups/offsite-restore/`. Same shape as
|
|
// the F5 stale-primary prune (backup.go) — refuse loudly rather than best-effort skip, since
|
|
// reaching here with an out-of-sandbox path means a helper above is wrong.
|
|
cleanTarget := filepath.Clean(target)
|
|
cleanRoot := filepath.Clean(root) + string(filepath.Separator)
|
|
if !strings.HasPrefix(cleanTarget+string(filepath.Separator), cleanRoot) {
|
|
m.logger.Printf("[WARN] [offbox] refusing to delete verification copy outside %s: %s", root, cleanTarget)
|
|
return fmt.Errorf("a törlés útvonala kívül esik az ellenőrző mappán")
|
|
}
|
|
if err := os.RemoveAll(cleanTarget); err != nil {
|
|
return fmt.Errorf("a másolat törlése nem sikerült: %w", err)
|
|
}
|
|
m.logger.Printf("[INFO] [offbox] deleted verification copy: %s", cleanTarget)
|
|
return nil
|
|
}
|
|
return fmt.Errorf("nincs ilyen ellenőrző másolat")
|
|
}
|
|
|
|
// OffsiteRestoreScratchPath exposes WHERE a verification restore for stack would land, so the UI can
|
|
// name the full path in the completion message instead of saying "a verification folder somewhere".
|
|
func (m *Manager) OffsiteRestoreScratchPath(stack string) string {
|
|
scratch, _, err := m.offboxRestoreScratchDir(stack)
|
|
if err != nil {
|
|
return ""
|
|
}
|
|
return scratch
|
|
}
|