Files
admin f94543ee5c
gates / gates (push) Successful in 10s
v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.

PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.

PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.

PART 4 - measured before theorising, on the live off-site target:
  snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
  => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.

TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.

RED-PROOFS, mutation asserted applied then reverted to 0:
  A   three template guards dropped (count asserted 3) -> the blank form returned
  P4  inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential

DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.

NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
2026-08-21 21:29:01 +02:00

179 lines
6.6 KiB
Go

package backup
import (
"bytes"
"context"
"fmt"
"log"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-controller/internal/settings"
)
// inventoryFixture wires a Manager whose repository answers a fixed snapshot list, and whose per-app
// `stats` calls are observable: how many ran at once, how many in total, and which ones fail.
//
// The delay is what makes the concurrency assertion meaningful — with instant calls a sequential
// implementation could pass a peak-of-1 check by finishing before the next one starts.
func inventoryFixture(t *testing.T, tags []string, statDelay time.Duration, failFor map[string]bool) (*Manager, *int32, *int32, *bytes.Buffer) {
t.Helper()
m, sett := newOffboxManager(t)
if err := sett.SetOffboxTarget(&settings.OffboxTarget{
Enabled: true, Host: "nas.local", Port: 22, User: "felhom", RepoPath: "/srv/repo",
Schedule: "daily", EscrowState: "escrowed",
}); err != nil {
t.Fatal(err)
}
if err := m.WriteOffboxSecrets("KEYMATERIAL", "nas.local ssh-ed25519 HOSTKEY"); err != nil {
t.Fatal(err)
}
if !m.OffboxConfigured() {
t.Fatal("fixture: the target must be configured or the listing exits early")
}
var logbuf bytes.Buffer
m.logger = log.New(&logbuf, "", 0)
var snaps []string
for i, tag := range tags {
snaps = append(snaps, fmt.Sprintf(`{"short_id":"snap%d","time":"2026-08-0%dT06:00:00Z","tags":["%s"]}`, i, i+1, tag))
}
snapshotsJSON := "[" + strings.Join(snaps, ",") + "]"
var inFlight, peak, total int32
var mu sync.Mutex
m.SetOffboxRunner(func(_ context.Context, _ []string, args ...string) ([]byte, error) {
if contains(args, "snapshots") {
return []byte(snapshotsJSON), nil
}
if !contains(args, "stats") {
return nil, nil
}
atomic.AddInt32(&total, 1)
cur := atomic.AddInt32(&inFlight, 1)
mu.Lock()
if cur > peak {
peak = cur
}
mu.Unlock()
time.Sleep(statDelay)
atomic.AddInt32(&inFlight, -1)
// args carries the snapshot id; map it back to its tag position.
for i, tag := range tags {
if contains(args, fmt.Sprintf("snap%d", i)) {
if failFor[tag] {
return nil, fmt.Errorf("injected failure for %s", tag)
}
return []byte(`{"total_size":1048576}`), nil
}
}
return []byte(`{"total_size":0}`), nil
})
return m, &peak, &total, &logbuf
}
// R-351 PART 4 — the per-app size calls run AT ONCE, and never more than the bound at once.
//
// MEASURED on demo-hp against the live off-site target before any change (2026-08-21):
// snapshots --json 2605 ms once, then stats --mode restore-size 2697 ms PER APP, sequential, over
// 5 app tags — 2605 + 5*2697 = ~16.1 s, which is the reported ten-to-fifteen seconds.
//
// THE BOUND IS THE SAFETY PROPERTY, not the speed one. The repository is a Hetzner Storage Box with
// a concurrent-SSH-session cap; a size call that is refused returns SizeBytes 0, which is a SILENT
// UNDER-REPORT of the customer's own data rather than a visible failure. So the peak is asserted,
// not just the total.
func TestOffsiteInventoryList_SizeCallsAreConcurrentButBounded(t *testing.T) {
tags := []string{"immich", "nextcloud", "opengist", "paperless-ngx", "privatebin", "calibre-web", "vaultwarden"}
m, peak, total, _ := inventoryFixture(t, tags, 40*time.Millisecond, nil)
start := time.Now()
_, err := m.OffsiteInventoryList(context.Background())
elapsed := time.Since(start)
if err != nil {
t.Fatalf("inventory: %v", err)
}
if int(*total) != len(tags) {
t.Errorf("every app needs its own size call: got %d, want %d", *total, len(tags))
}
if *peak < 2 {
t.Errorf("the size calls must run concurrently; peak in flight was %d — that is the sequential shape that cost 16s", *peak)
}
if *peak > 4 {
t.Errorf("peak in flight was %d, above the bound of 4 — an unbounded fan-out risks refused connections, "+
"and a refused size call silently under-reports the customer's data", *peak)
}
// 7 apps at 40ms: sequential would be >=280ms, four-at-a-time is two waves (~80ms).
if elapsed >= time.Duration(len(tags))*40*time.Millisecond {
t.Errorf("elapsed %v is the sequential cost — the fan-out is not taking effect", elapsed)
}
}
// Nothing may be LOST or REORDERED by the fan-out. A missing app reads as "you have no backup of
// this", which is the worst possible thing for this page to say by accident.
func TestOffsiteInventoryList_FanOutLosesNothingAndStaysSorted(t *testing.T) {
tags := []string{"vaultwarden", "immich", "calibre-web", "opengist", "nextcloud"}
m, _, _, _ := inventoryFixture(t, tags, time.Millisecond, nil)
inv, err := m.OffsiteInventoryList(context.Background())
if err != nil {
t.Fatalf("inventory: %v", err)
}
if len(inv.Apps) != len(tags) {
t.Fatalf("apps listed = %d, want %d — the fan-out dropped one, which reads as a missing backup", len(inv.Apps), len(tags))
}
seen := map[string]bool{}
for _, a := range inv.Apps {
seen[a.App] = true
if a.SizeBytes != 1048576 {
t.Errorf("%s: size = %d, want the injected value", a.App, a.SizeBytes)
}
}
for _, tag := range tags {
if !seen[tag] {
t.Errorf("%s is missing from the listing", tag)
}
}
for i := 1; i < len(inv.Apps); i++ {
if inv.Apps[i-1].App > inv.Apps[i].App {
t.Fatalf("the listing must stay sorted; %q came before %q", inv.Apps[i-1].App, inv.Apps[i].App)
}
}
if inv.Empty {
t.Error("a repository with snapshots is not Empty")
}
}
// A FAILED size call must still list the app, with size 0, AND still log the WARN. The warning is
// the only thing that distinguishes "0 bytes" from "we could not tell" — and under a fan-out it is
// the easiest thing to lose, because the logging moved off the goroutine that produced the error.
func TestOffsiteInventoryList_FailedSizeStillListsAndStillWarns(t *testing.T) {
tags := []string{"immich", "opengist"}
m, _, _, logbuf := inventoryFixture(t, tags, time.Millisecond, map[string]bool{"opengist": true})
inv, err := m.OffsiteInventoryList(context.Background())
if err != nil {
t.Fatalf("a per-app size failure must not fail the whole listing: %v", err)
}
if len(inv.Apps) != 2 {
t.Fatalf("both apps must be listed; got %d", len(inv.Apps))
}
for _, a := range inv.Apps {
if a.App == "opengist" && a.SizeBytes != 0 {
t.Errorf("an unknown size must be 0, got %d", a.SizeBytes)
}
if a.App == "immich" && a.SizeBytes == 0 {
t.Error("the healthy app's size must survive its neighbour's failure")
}
}
if !strings.Contains(logbuf.String(), "size of opengist's newest snapshot unknown") {
t.Errorf("the WARN is what separates \"0 bytes\" from \"could not tell\"; log was: %s", logbuf.String())
}
if strings.Contains(logbuf.String(), "size of immich's newest snapshot unknown") {
t.Error("a healthy app must not be warned about")
}
}