f94543ee5c
gates / gates (push) Successful in 10s
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok, go vet clean, -race clean on the changed package - all run and read BEFORE this commit. PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT installed in this case, so there is no own drive to ask) and reads manifest.json plus the captured compose/app.yaml. Local file reads only: no network, no restic, no restore. RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as "what your backup says" would be a fabricated fact. The prefill is labelled as coming from the backup and stays editable: a memory, not a lock. PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage field; the other 40 have none and their data goes to the system drive, which no screen said. Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for the 40-class a recorded placement is a fact to state, never a value written into a field that does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification. PART 4 - measured before theorising, on the live off-site target: snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds. The cause is the shape already on file, so the per-app size calls now run concurrently, BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted. OffsiteInventoryList had no test at all before this. TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a template doing index/eq against an undefined key errors at RENDER time: green build, green vet, green suite, 500 on the page. Four render tests, one per branch, because the existing deploy render test only renders AutoFields and never reaches these blocks. RED-PROOFS, mutation asserted applied then reverted to 0: A three template guards dropped (count asserted 3) -> the blank form returned P4 inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory + what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first. NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit carries no db_dumps and no volume_dumps still reports a bare completion.
178 lines
7.2 KiB
Go
178 lines
7.2 KiB
Go
package backup
|
|
|
|
import (
|
|
"context"
|
|
"encoding/json"
|
|
"errors"
|
|
"sort"
|
|
"sync"
|
|
"time"
|
|
)
|
|
|
|
// R-193 Part 3 — WHAT IS IN THERE. After a successful unlock the customer is shown the contents of the
|
|
// repository they just opened: which apps, from when, how big.
|
|
//
|
|
// READ-ONLY, AND THAT IS THE POINT. This restores nothing, puts nothing back, and compares nothing
|
|
// against live data. Unlocking and restoring are separate (operator ruling, 2026-08-05): restore is
|
|
// already per-app and already lives in the backups area, and a screen that unlocks and then offers to
|
|
// overwrite is two decisions wearing one button.
|
|
//
|
|
// WHY A LISTING AT ALL, rather than a success message: "unlocked" with nothing shown is
|
|
// indistinguishable from having unlocked an EMPTY store, and the customer has no way to tell whether
|
|
// what came back is the right thing. Seeing their own app names and dates is how they know.
|
|
|
|
// errNoOffsiteTarget is returned when the repository cannot even be addressed — no off-site target is
|
|
// configured on this box yet. Distinguished from a read failure because the remedy differs: this one
|
|
// resolves by itself once the tier is re-applied.
|
|
var errNoOffsiteTarget = errors.New("no off-site target is configured on this box yet")
|
|
|
|
// ErrNoOffsiteTarget reports whether err is the not-yet-configured case, so a caller can say the right
|
|
// thing rather than showing a generic failure.
|
|
func ErrNoOffsiteTarget(err error) bool { return errors.Is(err, errNoOffsiteTarget) }
|
|
|
|
// ErrNoOffsiteTargetSentinel exposes the sentinel itself so other packages — and their tests — can
|
|
// construct the not-yet-configured case. Added for R-237, whose restore list must distinguish
|
|
// "no target yet" (resolves by itself) from "could not read" (does not), and must be able to pin
|
|
// both in a table test.
|
|
func ErrNoOffsiteTargetSentinel() error { return errNoOffsiteTarget }
|
|
|
|
// OffsiteInventoryApp is one app's presence in the opened repository. Non-secret throughout.
|
|
type OffsiteInventoryApp struct {
|
|
App string // the restic tag == the stack name
|
|
LatestAt time.Time // the newest snapshot's time for this app
|
|
SizeBytes int64 // restore size of that newest snapshot (0 = could not be determined)
|
|
}
|
|
|
|
// OffsiteInventory is the whole answer, including the EMPTY case stated explicitly.
|
|
type OffsiteInventory struct {
|
|
Apps []OffsiteInventoryApp
|
|
// Empty is true when the repository opened cleanly and holds no snapshots. It is a real and
|
|
// confusing outcome — a bare list there reads as a broken page — so it is named rather than
|
|
// inferred from len(Apps)==0, which is also what a failed read looks like.
|
|
Empty bool
|
|
}
|
|
|
|
// OffsiteInventoryList opens the repository and reports what is in it, grouped per app. One
|
|
// `snapshots --json` call for the whole repo, then one `stats` per app for the newest snapshot's size.
|
|
//
|
|
// A per-app size failure is NOT fatal: the app is still listed, with SizeBytes 0, because knowing an
|
|
// app is in there matters more than knowing how big it is, and dropping it would under-report the
|
|
// customer's own data.
|
|
func (m *Manager) OffsiteInventoryList(ctx context.Context) (OffsiteInventory, error) {
|
|
var inv OffsiteInventory
|
|
// A box can hold a recovered key and still have no off-site COORDINATES — the pristine rebuilt
|
|
// shape, before its target is re-applied. Reading the repository is impossible then, and saying so
|
|
// is the honest answer; without this guard offboxBaseArgs nil-derefs on the missing target.
|
|
if !m.OffboxConfigured() {
|
|
return inv, errNoOffsiteTarget
|
|
}
|
|
t := m.settings.GetOffboxTarget()
|
|
base, env := m.offboxBaseArgs(t)
|
|
sctx, cancel := context.WithTimeout(ctx, offboxProbeTimeout)
|
|
defer cancel()
|
|
out, err := m.runner()(sctx, env, append(append([]string{}, base...), "snapshots", "--json")...)
|
|
if err != nil {
|
|
return inv, err
|
|
}
|
|
var snaps []struct {
|
|
ShortID string `json:"short_id"`
|
|
ID string `json:"id"`
|
|
Time time.Time `json:"time"`
|
|
Tags []string `json:"tags"`
|
|
}
|
|
if uerr := json.Unmarshal(out, &snaps); uerr != nil {
|
|
return inv, uerr
|
|
}
|
|
if len(snaps) == 0 {
|
|
inv.Empty = true
|
|
return inv, nil
|
|
}
|
|
// Newest snapshot per tag. A snapshot may carry several tags; each names an app it belongs to.
|
|
newest := map[string]struct {
|
|
id string
|
|
at time.Time
|
|
}{}
|
|
for _, s := range snaps {
|
|
id := s.ShortID
|
|
if id == "" {
|
|
id = s.ID
|
|
}
|
|
for _, tag := range s.Tags {
|
|
if tag == "" {
|
|
continue
|
|
}
|
|
if cur, ok := newest[tag]; !ok || s.Time.After(cur.at) {
|
|
newest[tag] = struct {
|
|
id string
|
|
at time.Time
|
|
}{id: id, at: s.Time}
|
|
}
|
|
}
|
|
}
|
|
if len(newest) == 0 {
|
|
// Snapshots exist but carry no tags — not "empty", and saying so would be a lie. Report an
|
|
// empty app list without the Empty flag; the page renders the honest in-between wording.
|
|
return inv, nil
|
|
}
|
|
// R-351 Part 4 — THE SIZE CALLS RUN CONCURRENTLY, BOUNDED.
|
|
//
|
|
// MEASURED before changing anything, on demo-hp against the live off-site target
|
|
// (u629488-sub3.your-storagebox.de:23), 2026-08-21:
|
|
//
|
|
// restic snapshots --json (once, whole repo) 2605 ms
|
|
// restic stats --mode restore-size (per app) 2697 ms each, 5 app tags, SEQUENTIAL
|
|
// => 2605 + 5*2697 = ~16.1 s
|
|
//
|
|
// which is the ten-to-fifteen seconds the page was reported to take. The cause is the shape
|
|
// already on file — one network call per app, one after another — so the fix is the same one:
|
|
// run them at once. Each call is an independent SSH round-trip to the repository and `stats` is
|
|
// a READ (restic takes a shared lock), so they do not contend.
|
|
//
|
|
// WHY BOUNDED, and why the bound is small: the target is a Hetzner Storage Box, which caps
|
|
// concurrent SSH sessions. Unbounded fan-out over a large app list would trade a slow page for
|
|
// refused connections — and a refused size call degrades to SizeBytes 0, i.e. it would quietly
|
|
// UNDER-REPORT the customer's own data rather than fail loudly. Four keeps well clear of the cap
|
|
// and still collapses the common case to a single wave.
|
|
const inventorySizeConcurrency = 4
|
|
|
|
type sized struct {
|
|
app OffsiteInventoryApp
|
|
err error
|
|
}
|
|
results := make([]sized, 0, len(newest))
|
|
var mu sync.Mutex
|
|
var wg sync.WaitGroup
|
|
sem := make(chan struct{}, inventorySizeConcurrency)
|
|
for tag, n := range newest {
|
|
wg.Add(1)
|
|
go func(tag, id string, at time.Time) {
|
|
defer wg.Done()
|
|
sem <- struct{}{}
|
|
defer func() { <-sem }()
|
|
app := OffsiteInventoryApp{App: tag, LatestAt: at}
|
|
size, serr := m.offboxSnapshotSize(ctx, id)
|
|
if serr == nil {
|
|
app.SizeBytes = size
|
|
}
|
|
mu.Lock()
|
|
results = append(results, sized{app: app, err: serr})
|
|
mu.Unlock()
|
|
}(tag, n.id, n.at)
|
|
}
|
|
wg.Wait()
|
|
// Logging happens on the caller's goroutine, after the fan-out: m.logger is shared and the
|
|
// per-app WARN is the only thing that tells an operator a size is missing rather than zero.
|
|
for _, r := range results {
|
|
if r.err != nil {
|
|
m.logger.Printf("[WARN] [offbox] inventory: size of %s's newest snapshot unknown: %v (listing it anyway)", r.app.App, r.err)
|
|
}
|
|
inv.Apps = append(inv.Apps, r.app)
|
|
}
|
|
sort.Slice(inv.Apps, func(i, j int) bool { return inv.Apps[i].App < inv.Apps[j].App })
|
|
return inv, nil
|
|
}
|
|
|
|
// HumanizeBytes exposes the shared byte formatter to the web layer so the recovery page renders sizes
|
|
// the same way every other surface does.
|
|
func HumanizeBytes(n int64) string { return humanizeBytes(n) }
|