diff --git a/documentation/architecture/05-hub-architecture.md b/documentation/architecture/05-hub-architecture.md index 586c64ce..730395b5 100644 --- a/documentation/architecture/05-hub-architecture.md +++ b/documentation/architecture/05-hub-architecture.md @@ -444,5 +444,18 @@ the global-key API) open through `GetHostRecoveryCredential`, so the break-glass use). **And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md` -(R-173 — awaiting the operator's decision). +(R-173 — option A decided 2026-10-05, `09` decision 125; §16.3). + +### 16.3 The nightly database snapshot (R-173, hub v0.136.0) + +Every night at **02:00 Budapest** the hub writes `/snapshots/hub-.db` with SQLite's `VACUUM INTO`: +one statement, one point in time, every committed write included — the rows still in `hub.db-wal` too, which a file copy +of `hub.db` loses (measured in `internal/dbsnap` tests: a plain copy held 0 of 150 fresh rows). Written as `.tmp`, mode +`0600`, then renamed, so a reader never sees half a file. The newest **2** stay (the volume grew from 1 GiB to 2 GiB +for them, operator choice 2026-10-05; one snapshot measured 370 MB). Two runs never overlap (a second gets `ErrBusy`). +Each run logs `db snapshot written: (, )`. At start-up the hub runs one when the newest is older +than 24 h. **The hub ships nothing itself and holds no ep0 credential:** DooPlex's `felhom-hub-db-backup` unit picks the +newest snapshot up at 02:30, checks it (`PRAGMA integrity_check`), encrypts it and pushes it to ep0's `operator` +namespace (`runbooks/RUNBOOK-hub-db-offsite-backup.md`). Pinned by `internal/dbsnap/dbsnap_test.go` and +`cmd/hub/r173_wiring_test.go`. diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index b4a3ecc6..b62b5e19 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -881,6 +881,16 @@ its length, and both fixes cost something the household would notice — operato `-step1` — then the release's bundle, both by signed jobs. **Chosen (b)**, `felhom-agent/scripts/build-step-bundle.py`; delivered to demo-hp, demo-felhom and Tester 1 on 2026-10-05. `11` §5.4.2 rule unchanged. +### 2026-10-05 (14:05) — three operator rulings (recorded before the work; the hub-DB off-site brief) + +125. **The hub database's off-site copy goes to ep0's backup server** (option A of the hub-safety STATUS decision): + encrypted on DooPlex, pushed to a write-only namespace on ep0, restore-tested weekly, alarmed. Rejected: B, a separate + Hetzner Storage Box account (more new parts to maintain). *Operator ruling 2026-10-05.* (R-173) +126. **R-519's live test is approved:** one controller restart on scratch 9202 in the middle of a backup. *Operator ruling + 2026-10-05.* +127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer. + *Operator ruling 2026-10-05.* (R-861) + ### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief) 100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) — diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof-run1.txt b/documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof-run1.txt new file mode 100644 index 00000000..a34b4906 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof-run1.txt @@ -0,0 +1,38 @@ +### R1 — SnapshotInto copies hub.db as a file instead of VACUUM INTO +=== RUN TestSnapshot_ConsistentWithLiveDBIncludingWAL + dbsnap_test.go:104: table hosts: snapshot 0 rows, live 150 + dbsnap_test.go:104: table hub_settings: snapshot 0 rows, live 2 +--- FAIL: TestSnapshot_ConsistentWithLiveDBIncludingWAL (0.05s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.055s +FAIL + +### R2 — no busy check (two runs may overlap) +=== RUN TestSnapshot_NeverTwoAtOnce + /mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:141 +0x76 + /mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:153 +0x1a6 + /mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:141 +0x76 + /mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:151 +0x34 + /mnt/5_hdd/felhom.eu/git/felhom.eu/hub/internal/dbsnap/dbsnap_test.go:151 +0x178 +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 600.107s +FAIL + +### R3 — prune keeps everything +=== RUN TestSnapshot_KeepsNewestTwo + dbsnap_test.go:130: kept [hub-20261006T000000Z.db hub-20261007T000000Z.db hub-20261008T000000Z.db], want [hub-20261007T000000Z.db hub-20261008T000000Z.db] +--- FAIL: TestSnapshot_KeepsNewestTwo (0.07s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.075s +FAIL + +### R4 — a failed write leaves its .tmp +=== RUN TestSnapshot_FailureLeavesNoFile + dbsnap_test.go:191: left 1 file(s) behind +--- FAIL: TestSnapshot_FailureLeavesNoFile (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.007s +FAIL + +ok gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.135s + +[exited with code 0] diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof.txt b/documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof.txt new file mode 100644 index 00000000..ca21f7ca --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof.txt @@ -0,0 +1,26 @@ +Run 1 (R1–R4) in red-proof-run1.txt. R2 there HUNG (600 s timeout) — the old test blocked forever; it convicted by hanging. The test now fails cleanly; R2 re-run below. + +### R2 (re-run) — no busy check +=== RUN TestSnapshot_NeverTwoAtOnce + dbsnap_test.go:162: second Make did not return ErrBusy at once: it ran alongside the first +--- FAIL: TestSnapshot_NeverTwoAtOnce (2.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 2.007s +FAIL + +### R5 — main never schedules the snapshot +=== RUN TestR173_MainSchedulesTheDBSnapshot + r173_wiring_test.go:42: cmd/hub/main.go never calls scheduleDaily(ctx, "db-snapshot", "02:00", …) — no nightly snapshot +--- FAIL: TestR173_MainSchedulesTheDBSnapshot (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.025s +FAIL + +### R6 — main has no start-up catch-up +=== RUN TestR173_MainSchedulesTheDBSnapshot + r173_wiring_test.go:45: cmd/hub/main.go never calls dbsnap.NeedsCatchUp — a pod down at 02:00 leaves a stale snapshot +--- FAIL: TestR173_MainSchedulesTheDBSnapshot (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.023s +FAIL + diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index 73ffc304..c5ed9916 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,3 +1,18 @@ +## v0.136.0 — the hub writes a nightly, checked copy of its own database (R-173, decision A) (2026-10-05) + +**Operator action on deploy: none.** The hub's volume grows from 1 GiB to 2 GiB and joins Longhorn's nightly backup +group (`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: enabled`). + +- **Nightly snapshot (`05` §16.3).** At 02:00 Budapest the hub writes `/data/snapshots/hub-.db` with `VACUUM INTO` + (one point in time, WAL included), as `.tmp` then renamed, mode 0600; keeps the newest 2; never two at once; logs + name, bytes and duration. A start-up run fills in when the newest is older than 24 h. DooPlex picks the file up at + 02:30, checks, encrypts and pushes it to ep0 (`runbooks/RUNBOOK-hub-db-offsite-backup.md`). The hub holds no ep0 + credential. +- **Tests.** `internal/dbsnap`: integrity_check + equal row counts in every table vs. the live DB with rows still in + the WAL (precondition asserted: a plain file copy loses them), keep-2, never-two-at-once, a failed write leaves no + file, the catch-up age; `cmd/hub/r173_wiring_test.go` pins the 02:00 schedule and the catch-up in `main()`. + Red-proofs R1–R6, all convict: `documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof.txt`. + ## v0.135.0 — the hub's own safety: form protection for the password path, console passwords sealed at rest; boxes left behind are listed and alarmed; a waiting customer with no e-mail is flagged (R-135, R-133, R-604, R-530, R-508) (2026-10-05) **Operator action on deploy: none.** Scripts that POST to the hub with the operator password must now send diff --git a/hub/cmd/hub/main.go b/hub/cmd/hub/main.go index d9d680bf..98d6a287 100644 --- a/hub/cmd/hub/main.go +++ b/hub/cmd/hub/main.go @@ -17,6 +17,7 @@ import ( "gitea.dooplex.hu/admin/felhom-hub/internal/api" "gitea.dooplex.hu/admin/felhom-hub/internal/assets" "gitea.dooplex.hu/admin/felhom-hub/internal/claim" + "gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap" "gitea.dooplex.hu/admin/felhom-hub/internal/gitea" "gitea.dooplex.hu/admin/felhom-hub/internal/hetznerapi" "gitea.dooplex.hu/admin/felhom-hub/internal/intent" @@ -848,6 +849,22 @@ func main() { } }() + // Nightly DB snapshot (R-173, v0.136.0, `05` §16.3): VACUUM INTO /snapshots/, keep 2. DooPlex's + // felhom-hub-db-backup.service picks up the newest at 02:30, checks, encrypts and pushes it to ep0. A start-up + // catch-up runs one when the newest is older than 24 h (a pod that was down at 02:00). + snapMaker := &dbsnap.Maker{Store: dataStore, Dir: filepath.Join(cfg.Server.DataDir, "snapshots"), Logger: logger} + runSnapshot := func() { + sctx, cancel := context.WithTimeout(ctx, 10*time.Minute) + defer cancel() + if _, err := snapMaker.Make(sctx); err != nil { + logger.Printf("[ERROR] db snapshot failed: %v", err) + } + } + if dbsnap.NeedsCatchUp(snapMaker.Dir, time.Now(), 24*time.Hour) { + go runSnapshot() + } + go scheduleDaily(ctx, "db-snapshot", "02:00", runSnapshot, logger) + // Backup deadline checker — runs daily at 05:00 Budapest go scheduleDaily(ctx, "deadline-check", "05:00", func() { monitor.CheckBackupDeadlines(dataStore, stalenessChecker, dispatcher.ProcessEvent, logger) diff --git a/hub/cmd/hub/r173_wiring_test.go b/hub/cmd/hub/r173_wiring_test.go new file mode 100644 index 00000000..55d65e1f --- /dev/null +++ b/hub/cmd/hub/r173_wiring_test.go @@ -0,0 +1,47 @@ +package main + +import ( + "go/ast" + "go/parser" + "go/token" + "strconv" + "testing" +) + +// R-173 seam wiring: main() must schedule the nightly db snapshot — a scheduleDaily call named "db-snapshot" at +// "02:00", before DooPlex's 02:30 push — and must run the start-up catch-up (dbsnap.NeedsCatchUp). Without them the +// dbsnap package exists and DooPlex pushes nothing new. RED-PROOF: delete either call → this test fails. +func TestR173_MainSchedulesTheDBSnapshot(t *testing.T) { + f, err := parser.ParseFile(token.NewFileSet(), "main.go", nil, 0) + if err != nil { + t.Fatal(err) + } + scheduled, catchUp := false, false + ast.Inspect(f, func(n ast.Node) bool { + c, ok := n.(*ast.CallExpr) + if !ok { + return true + } + if id, ok := c.Fun.(*ast.Ident); ok && id.Name == "scheduleDaily" && len(c.Args) >= 3 { + name, _ := c.Args[1].(*ast.BasicLit) + at, _ := c.Args[2].(*ast.BasicLit) + if name != nil && at != nil { + n, _ := strconv.Unquote(name.Value) + a, _ := strconv.Unquote(at.Value) + if n == "db-snapshot" && a == "02:00" { + scheduled = true + } + } + } + if sel, ok := c.Fun.(*ast.SelectorExpr); ok && sel.Sel.Name == "NeedsCatchUp" { + catchUp = true + } + return true + }) + if !scheduled { + t.Error(`cmd/hub/main.go never calls scheduleDaily(ctx, "db-snapshot", "02:00", …) — no nightly snapshot`) + } + if !catchUp { + t.Error("cmd/hub/main.go never calls dbsnap.NeedsCatchUp — a pod down at 02:00 leaves a stale snapshot") + } +} diff --git a/hub/internal/dbsnap/dbsnap.go b/hub/internal/dbsnap/dbsnap.go new file mode 100644 index 00000000..238be62d --- /dev/null +++ b/hub/internal/dbsnap/dbsnap.go @@ -0,0 +1,150 @@ +// Package dbsnap makes the hub's nightly database snapshot (R-173, hub v0.136.0, `05` §16.3, +// runbooks/RUNBOOK-hub-db-offsite-backup.md Step 3). +// +// Every night the hub writes `/snapshots/hub-.db` with SQLite's VACUUM INTO — one consistent point +// in time, WAL-aware — then keeps the newest Keep (2). DooPlex picks up the newest one, checks it, encrypts it and +// pushes it to ep0 (Step 4). The hub never ships anything itself: it has no ep0 credential. +// +// Written as `.tmp` and renamed, so a reader never sees half a file; a run never overlaps another (a second call +// returns ErrBusy). Pinned by dbsnap_test.go. +package dbsnap + +import ( + "context" + "errors" + "fmt" + "log" + "os" + "path/filepath" + "sort" + "strings" + "sync" + "time" +) + +// Keep is how many snapshots stay on the volume (the newest). Decided by CC — operator may reverse (the volume was +// grown to 2 GiB for it, operator choice 2026-10-05). +const Keep = 2 + +// Prefix and Suffix make a snapshot's file name: hub-20261005T020000Z.db. +const ( + Prefix = "hub-" + Suffix = ".db" +) + +// ErrBusy is returned when a snapshot is already being written. +var ErrBusy = errors.New("dbsnap: a snapshot is already being written") + +// Snapshotter is the store's VACUUM INTO. +type Snapshotter interface { + SnapshotInto(ctx context.Context, path string) error +} + +// Maker writes and prunes snapshots in Dir. +type Maker struct { + Store Snapshotter + Dir string + Logger *log.Logger + Now func() time.Time + + mu sync.Mutex + running bool +} + +// Result is one snapshot written. +type Result struct { + Name string `json:"name"` + Bytes int64 `json:"bytes"` + Duration time.Duration `json:"duration_ns"` + Pruned []string `json:"pruned,omitempty"` +} + +// Make writes one snapshot and prunes to Keep. Safe to call from the nightly job and the operator button at once. +func (m *Maker) Make(ctx context.Context) (Result, error) { + m.mu.Lock() + if m.running { + m.mu.Unlock() + return Result{}, ErrBusy + } + m.running = true + m.mu.Unlock() + defer func() { m.mu.Lock(); m.running = false; m.mu.Unlock() }() + + now := time.Now + if m.Now != nil { + now = m.Now + } + if err := os.MkdirAll(m.Dir, 0o700); err != nil { + return Result{}, fmt.Errorf("dbsnap: %w", err) + } + name := Prefix + now().UTC().Format("20060102T150405Z") + Suffix + final := filepath.Join(m.Dir, name) + tmp := final + ".tmp" + _ = os.Remove(tmp) // a leftover from a crash mid-write + start := time.Now() + if err := m.Store.SnapshotInto(ctx, tmp); err != nil { + _ = os.Remove(tmp) + return Result{}, err + } + if err := os.Chmod(tmp, 0o600); err != nil { + _ = os.Remove(tmp) + return Result{}, fmt.Errorf("dbsnap: chmod: %w", err) + } + if err := os.Rename(tmp, final); err != nil { + _ = os.Remove(tmp) + return Result{}, fmt.Errorf("dbsnap: rename: %w", err) + } + fi, err := os.Stat(final) + if err != nil { + return Result{}, fmt.Errorf("dbsnap: stat: %w", err) + } + res := Result{Name: name, Bytes: fi.Size(), Duration: time.Since(start)} + res.Pruned = m.prune() + if m.Logger != nil { + m.Logger.Printf("[INFO] db snapshot written: %s (%d bytes, %s); pruned %d, keeping %d", + name, res.Bytes, res.Duration.Round(time.Millisecond), len(res.Pruned), Keep) + } + return res, nil +} + +// List returns the snapshot names in Dir, oldest first (the stamp sorts as text). +func List(dir string) []string { + entries, _ := os.ReadDir(dir) + var out []string + for _, e := range entries { + n := e.Name() + if !e.IsDir() && strings.HasPrefix(n, Prefix) && strings.HasSuffix(n, Suffix) { + out = append(out, n) + } + } + sort.Strings(out) + return out +} + +func (m *Maker) prune() []string { + names := List(m.Dir) + var pruned []string + for len(names) > Keep { + if err := os.Remove(filepath.Join(m.Dir, names[0])); err != nil && m.Logger != nil { + m.Logger.Printf("[WARN] db snapshot prune %s: %v", names[0], err) + } + pruned = append(pruned, names[0]) + names = names[1:] + } + return pruned +} + +// NeedsCatchUp reports whether the newest snapshot in dir is missing or older than maxAge — the hub runs one at +// start-up then, so a pod that was down at 02:00 does not leave DooPlex pushing a two-day-old copy. +func NeedsCatchUp(dir string, now time.Time, maxAge time.Duration) bool { + names := List(dir) + if len(names) == 0 { + return true + } + stamp := strings.TrimSuffix(strings.TrimPrefix(names[len(names)-1], Prefix), Suffix) + t, err := time.Parse("20060102T150405Z", stamp) + if err != nil { + return true + } + return now.Sub(t) > maxAge +} diff --git a/hub/internal/dbsnap/dbsnap_test.go b/hub/internal/dbsnap/dbsnap_test.go new file mode 100644 index 00000000..96411ec5 --- /dev/null +++ b/hub/internal/dbsnap/dbsnap_test.go @@ -0,0 +1,208 @@ +package dbsnap + +import ( + "context" + "database/sql" + "errors" + "fmt" + "io" + "log" + "os" + "path/filepath" + "strings" + "sync" + "testing" + "time" + + "gitea.dooplex.hu/admin/felhom-hub/internal/store" +) + +func newStore(t *testing.T) (*store.Store, string) { + t.Helper() + p := filepath.Join(t.TempDir(), "hub.db") + s, err := store.New(p, log.New(io.Discard, "", 0)) + if err != nil { + t.Fatalf("store.New: %v", err) + } + t.Cleanup(func() { s.Close() }) + return s, p +} + +func openRO(t *testing.T, p string) *sql.DB { + t.Helper() + db, err := sql.Open("sqlite", "file:"+p+"?mode=ro") + if err != nil { + t.Fatalf("open %s: %v", p, err) + } + t.Cleanup(func() { db.Close() }) + return db +} + +// rowCounts returns count(*) of every table. +func rowCounts(t *testing.T, db *sql.DB) map[string]int { + t.Helper() + rows, err := db.Query(`SELECT name FROM sqlite_master WHERE type='table' AND name NOT LIKE 'sqlite_%' ORDER BY name`) + if err != nil { + t.Fatalf("list tables: %v", err) + } + var names []string + for rows.Next() { + var n string + _ = rows.Scan(&n) + names = append(names, n) + } + rows.Close() + out := map[string]int{} + for _, n := range names { + var c int + if err := db.QueryRow(fmt.Sprintf(`SELECT count(*) FROM "%s"`, n)).Scan(&c); err != nil { + t.Fatalf("count %s: %v", n, err) + } + out[n] = c + } + return out +} + +// TestSnapshot_ConsistentWithLiveDBIncludingWAL is the consequence test: the snapshot passes integrity_check and has +// the live database's row count in EVERY table — including rows that so far live only in hub.db-wal. Precondition +// asserted in-test: copying hub.db alone (the naive file backup) LOSES those rows, so this test can tell the two apart. +func TestSnapshot_ConsistentWithLiveDBIncludingWAL(t *testing.T) { + s, live := newStore(t) + for i := 0; i < 150; i++ { + if err := s.UpsertHost(&store.Host{HostID: fmt.Sprintf("h%03d", i), CustomerID: "c1", APIKey: fmt.Sprintf("k%03d", i)}); err != nil { + t.Fatalf("UpsertHost: %v", err) + } + } + if fi, err := os.Stat(live + "-wal"); err != nil || fi.Size() == 0 { + t.Fatalf("precondition: want a non-empty WAL, got %v", err) + } + // Precondition: a plain copy of hub.db does not hold the WAL rows. + naive := filepath.Join(t.TempDir(), "naive.db") + b, _ := os.ReadFile(live) + _ = os.WriteFile(naive, b, 0o600) + if c := rowCounts(t, openRO(t, naive))["hosts"]; c == 150 { + t.Fatalf("precondition: a plain file copy already holds all 150 hosts — the test cannot tell a WAL-blind copy apart") + } + + dir := filepath.Join(t.TempDir(), "snapshots") + m := &Maker{Store: s, Dir: dir} + res, err := m.Make(context.Background()) + if err != nil { + t.Fatalf("Make: %v", err) + } + snap := openRO(t, filepath.Join(dir, res.Name)) + var ic string + if err := snap.QueryRow(`PRAGMA integrity_check`).Scan(&ic); err != nil || ic != "ok" { + t.Fatalf("integrity_check = %q, %v", ic, err) + } + want, got := rowCounts(t, openRO(t, live)), rowCounts(t, snap) + if len(want) < 10 || want["hosts"] != 150 { + t.Fatalf("live counts look wrong: %d tables, hosts=%d", len(want), want["hosts"]) + } + for tbl, n := range want { + if got[tbl] != n { + t.Errorf("table %s: snapshot %d rows, live %d", tbl, got[tbl], n) + } + } + if fi, _ := os.Stat(filepath.Join(dir, res.Name)); fi == nil || fi.Mode().Perm() != 0o600 || res.Bytes != fi.Size() { + t.Errorf("snapshot file mode/size wrong: %+v res.Bytes=%d", fi, res.Bytes) + } + if _, err := os.Stat(filepath.Join(dir, res.Name+".tmp")); !os.IsNotExist(err) { + t.Errorf("the .tmp file is left behind") + } +} + +// TestSnapshot_KeepsNewestTwo pins Keep: three runs leave the two newest. +func TestSnapshot_KeepsNewestTwo(t *testing.T) { + s, _ := newStore(t) + dir := filepath.Join(t.TempDir(), "snapshots") + base := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC) + i := 0 + m := &Maker{Store: s, Dir: dir, Now: func() time.Time { i++; return base.Add(time.Duration(i) * 24 * time.Hour) }} + for k := 0; k < 3; k++ { + if _, err := m.Make(context.Background()); err != nil { + t.Fatalf("Make %d: %v", k, err) + } + } + got := List(dir) + want := []string{"hub-20261007T000000Z.db", "hub-20261008T000000Z.db"} + if strings.Join(got, ",") != strings.Join(want, ",") { + t.Fatalf("kept %v, want %v", got, want) + } +} + +type blockingStore struct { + in, release chan struct{} + once *sync.Once +} + +func (b blockingStore) SnapshotInto(ctx context.Context, path string) error { + b.once.Do(func() { close(b.in) }) + <-b.release + return os.WriteFile(path, []byte("x"), 0o600) +} + +// TestSnapshot_NeverTwoAtOnce: a second Make while the first runs returns ErrBusy and writes nothing. +func TestSnapshot_NeverTwoAtOnce(t *testing.T) { + bs := blockingStore{in: make(chan struct{}), release: make(chan struct{}), once: &sync.Once{}} + dir := t.TempDir() + m := &Maker{Store: bs, Dir: dir} + done := make(chan error, 1) + go func() { _, err := m.Make(context.Background()); done <- err }() + <-bs.in + second := make(chan error, 1) + go func() { _, err := m.Make(context.Background()); second <- err }() + select { + case err := <-second: + if !errors.Is(err, ErrBusy) { + t.Fatalf("second Make = %v, want ErrBusy", err) + } + case <-time.After(2 * time.Second): + close(bs.release) + t.Fatalf("second Make did not return ErrBusy at once: it ran alongside the first") + } + close(bs.release) + if err := <-done; err != nil { + t.Fatalf("first Make: %v", err) + } + if n := len(List(dir)); n != 1 { + t.Fatalf("%d snapshots, want 1", n) + } + m.Now = func() time.Time { return time.Now().Add(time.Hour) } // a distinct name + if _, err := m.Make(context.Background()); err != nil { + t.Fatalf("Make after the first ended: %v", err) + } +} + +type failStore struct{} + +func (failStore) SnapshotInto(ctx context.Context, path string) error { + _ = os.WriteFile(path, []byte("half"), 0o600) + return errors.New("disk full") +} + +// TestSnapshot_FailureLeavesNoFile: a failed VACUUM INTO leaves neither a snapshot nor a .tmp for DooPlex to pick up. +func TestSnapshot_FailureLeavesNoFile(t *testing.T) { + dir := t.TempDir() + if _, err := (&Maker{Store: failStore{}, Dir: dir}).Make(context.Background()); err == nil { + t.Fatal("Make succeeded over a failing store") + } + if e, _ := os.ReadDir(dir); len(e) != 0 { + t.Fatalf("left %d file(s) behind", len(e)) + } +} + +func TestNeedsCatchUp(t *testing.T) { + dir := t.TempDir() + now := time.Date(2026, 10, 5, 12, 0, 0, 0, time.UTC) + if !NeedsCatchUp(dir, now, 24*time.Hour) { + t.Error("empty dir: want catch-up") + } + _ = os.WriteFile(filepath.Join(dir, "hub-20261005T000000Z.db"), nil, 0o600) + if NeedsCatchUp(dir, now, 24*time.Hour) { + t.Error("12 h old: want no catch-up") + } + if !NeedsCatchUp(dir, now.Add(13*time.Hour), 24*time.Hour) { + t.Error("25 h old: want catch-up") + } +} diff --git a/hub/internal/store/snapshot.go b/hub/internal/store/snapshot.go new file mode 100644 index 00000000..613b10cc --- /dev/null +++ b/hub/internal/store/snapshot.go @@ -0,0 +1,17 @@ +package store + +import ( + "context" + "fmt" +) + +// SnapshotInto writes a consistent, compacted copy of the whole database to path with SQLite's `VACUUM INTO` +// (R-173, hub v0.136.0). One statement inside one read transaction: it sees one point in time and includes every +// committed write still sitting in hub.db-wal — unlike copying hub.db (+ -wal) as files while the hub writes. The +// target must not exist. Pinned by internal/dbsnap (TestSnapshot_*). +func (s *Store) SnapshotInto(ctx context.Context, path string) error { + if _, err := s.db.ExecContext(ctx, `VACUUM INTO ?`, path); err != nil { + return fmt.Errorf("store: VACUUM INTO %s: %w", path, err) + } + return nil +} diff --git a/manifests/hub.yaml b/manifests/hub.yaml index dda57776..0a799d4f 100644 --- a/manifests/hub.yaml +++ b/manifests/hub.yaml @@ -44,14 +44,14 @@ metadata: namespace: felhom-system labels: app: hub - recurring-job-group.longhorn.io/default: disabled + recurring-job-group.longhorn.io/default: enabled spec: accessModes: - ReadWriteOnce storageClassName: longhorn resources: requests: - storage: 1Gi + storage: 2Gi # ============================================================================= # CONFIGURATION