Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 87977ff40a | |||
| 861d32a4b4 | |||
| 769c4c3cf2 | |||
| 64f704d0f7 | |||
| 208fac8027 |
+19
-1
@@ -1,4 +1,22 @@
|
||||
## unreleased (to be released as v0.147.0) — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
|
||||
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
|
||||
|
||||
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
|
||||
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
|
||||
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
|
||||
|
||||
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
|
||||
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
|
||||
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
|
||||
- **R-25 (agent half):** `POST /disks/format` and `GET /disks/format/status` return `fs_uuid`, the new filesystem's UUID
|
||||
read back after mkfs only when the bound durable id still resolves to the formatted device and the superblock is the
|
||||
requested type (empty = not verified); `DeviceProbe` gains `FSUUID` from blkid. Tests `TestFormat_*FSUUID*`; three
|
||||
red-proofs. The controller half (mount that UUID) is a controller change.
|
||||
|
||||
## v0.147.0 — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
|
||||
|
||||
Released by `scripts/release-agent.sh`: binary sha256 `642c4d196c48671c14ff653118303c5903af1670b7abb550edeaf73e701cd5b8`,
|
||||
config bundle sha256 `326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d` (tag `v0.147.0` = `f1b9b41`).
|
||||
Delivery order as for v0.146.1: signed `agent_update`, then signed `agent_config_update`.
|
||||
|
||||
MinAgent impact: none (the controller needs nothing new from this agent). Config bundle content unchanged from v0.146.1.
|
||||
|
||||
|
||||
@@ -1,7 +1,10 @@
|
||||
# REPORT — 2026-10-05 burn-down: two record corrections (no release)
|
||||
# REPORT — agent v0.147.0 (2026-10-05, burn-down round 2)
|
||||
|
||||
Full session report: `felhom.eu/REPORT-burndown-2026-10-05.md`. Baseline `e06ed97` (v0.146.1). No binary, bundle or
|
||||
version change; `go build`, `go vet ./internal/backup/`, `agent_gates.py --fast` green.
|
||||
Full session report: `felhom.eu/REPORT-burndown2-2026-10-05.md`. Baseline `d833163` (v0.146.1). Code commit `f1b9b41`
|
||||
(CI job 1365 success), tag `v0.147.0`, binary sha256 `642c4d19…`, bundle sha256 `326527d0…` (verified by download).
|
||||
|
||||
- R-291 — `scripts/retention-policy.json`: the source of the number (R-267 newest-10 prune, R-287) and readers corrected.
|
||||
- R-348 — `internal/backup/store.go`: the restart comment says what a restart really blanks.
|
||||
Rows: R-124 (recipe root namespace = ""), R-118 (no root size for an absent drive), R-269 (rotated-out token rejected at
|
||||
once), R-317 (dnsmasq install probed by its unit). Tests + red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/`.
|
||||
`go build/vet/test ./...` green; `agent_gates.py --fast` green after the release (release-complete needs the tag).
|
||||
|
||||
Delivery: see the session report (vouch, signed jobs per box, the hub System page afterwards).
|
||||
|
||||
@@ -66,7 +66,7 @@
|
||||
|---|---|---|---|---|
|
||||
| `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids |
|
||||
| `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore |
|
||||
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) <-chan error` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank |
|
||||
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) (*formatJob, <-chan error)` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank; on success `job.FSUUID` = the new filesystem's UUID, read back only when the durable id still resolves to the formatted device (R-25) — read it only after `done` delivers |
|
||||
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup stats the file on EVERY call and reloads BEFORE answering when the append-only log grew (R-269; was reload-on-miss only, v0.63.0 B3, which let a token rotated out by another process keep authorizing as a map hit): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence both ways without a restart; unchanged size = no re-read. Pinned by `TestTokenStore_RotatedOutTokenRejectedFirst` |
|
||||
| `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp |
|
||||
| `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased |
|
||||
|
||||
+26
-2
@@ -10,6 +10,7 @@ import (
|
||||
"log/slog"
|
||||
"os"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
|
||||
@@ -115,6 +116,7 @@ type Collector struct {
|
||||
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
|
||||
hostID string
|
||||
agentVersion string
|
||||
selfSHA func() string // R-349: sha256 of the running binary; default runningBinarySHA256
|
||||
logger *slog.Logger
|
||||
now func() time.Time
|
||||
}
|
||||
@@ -135,6 +137,7 @@ func NewCollector(px proxmoxReader, cf CloudflaredProber, storage StorageObserve
|
||||
temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake
|
||||
hostID: hostID,
|
||||
agentVersion: agentVersion,
|
||||
selfSHA: runningBinarySHA256,
|
||||
logger: logger,
|
||||
now: func() time.Time { return time.Now().UTC() },
|
||||
}
|
||||
@@ -337,6 +340,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
HostID: c.hostID,
|
||||
ReportedAt: c.now().Format(time.RFC3339),
|
||||
AgentVersion: c.agentVersion,
|
||||
AgentSHA256: c.agentSHA256(),
|
||||
Host: host,
|
||||
Guests: c.collectGuests(ctx),
|
||||
// storage_targets populated this slice (slice 5) via the observer; the rest stay
|
||||
@@ -431,8 +435,28 @@ const pbsWrapperPath = "/usr/local/sbin/felhom-pbs-apply"
|
||||
// unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that
|
||||
// legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is
|
||||
// needed to read it.
|
||||
func pbsWrapperSHA256() string {
|
||||
f, err := os.Open(pbsWrapperPath)
|
||||
func pbsWrapperSHA256() string { return fileSHA256(pbsWrapperPath) }
|
||||
|
||||
// selfExePath is the running binary as the kernel holds it. /proc/self/exe, not the installed path:
|
||||
// after an A/B flip the file at /usr/local/bin/felhom-agent may already be the NEXT binary while this
|
||||
// process still runs the old one, and the report must describe what runs (R-349). Test seam.
|
||||
var selfExePath = "/proc/self/exe"
|
||||
|
||||
// runningBinarySHA256 hashes the running binary ONCE per process — the bytes cannot change under a
|
||||
// running process, and re-hashing ~20 MB every report cycle buys nothing. A failed read is cached as
|
||||
// "" (UNKNOWN); it never fails the report.
|
||||
var runningBinarySHA256 = sync.OnceValue(func() string { return fileSHA256(selfExePath) })
|
||||
|
||||
func (c *Collector) agentSHA256() string {
|
||||
if c.selfSHA == nil {
|
||||
return ""
|
||||
}
|
||||
return c.selfSHA()
|
||||
}
|
||||
|
||||
// fileSHA256 is the hex sha256 of a file's bytes, or "" when it cannot be read.
|
||||
func fileSHA256(path string) string {
|
||||
f, err := os.Open(path)
|
||||
if err != nil {
|
||||
return ""
|
||||
}
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
package hub
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-349: the report carries the sha256 of the binary that is RUNNING, so the hub can tell a
|
||||
// hand-built proof binary from the vouched artifact of the same version string. The consequence
|
||||
// asserted: the wire field equals the hash of this very test binary's bytes (read independently via
|
||||
// os.Executable, a different channel from /proc/self/exe), and it is on the wire as agent_sha256.
|
||||
func TestCollect_AgentSHA256IsTheRunningBinary(t *testing.T) {
|
||||
exe, err := os.Executable()
|
||||
if err != nil {
|
||||
t.Skipf("os.Executable: %v", err)
|
||||
}
|
||||
raw, err := os.ReadFile(exe)
|
||||
if err != nil {
|
||||
t.Fatalf("read own binary: %v", err)
|
||||
}
|
||||
sum := sha256.Sum256(raw)
|
||||
want := hex.EncodeToString(sum[:])
|
||||
|
||||
px := &fakePx{node: "n", ns: newTestNodeStatus()}
|
||||
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
|
||||
r, err := c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect: %v", err)
|
||||
}
|
||||
if r.AgentSHA256 != want {
|
||||
t.Fatalf("agent_sha256 = %q, want the running binary's %q", r.AgentSHA256, want)
|
||||
}
|
||||
b, err := json.Marshal(r)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(string(b), `"agent_sha256":"`+want+`"`) {
|
||||
t.Fatalf("agent_sha256 not on the wire: %s", b)
|
||||
}
|
||||
}
|
||||
|
||||
// An unreadable binary is UNKNOWN (empty, omitted) — never a made-up hash, never a failed report.
|
||||
func TestFileSHA256_UnreadableIsEmpty(t *testing.T) {
|
||||
if got := fileSHA256(filepath.Join(t.TempDir(), "absent")); got != "" {
|
||||
t.Fatalf("absent file hashed to %q, want empty", got)
|
||||
}
|
||||
px := &fakePx{node: "n", ns: newTestNodeStatus()}
|
||||
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
|
||||
c.selfSHA = func() string { return "" }
|
||||
r, err := c.Collect(context.Background())
|
||||
if err != nil {
|
||||
t.Fatalf("Collect must not fail on an unreadable binary: %v", err)
|
||||
}
|
||||
b, _ := json.Marshal(r)
|
||||
if strings.Contains(string(b), "agent_sha256") {
|
||||
t.Fatalf("empty agent_sha256 must be omitted: %s", b)
|
||||
}
|
||||
}
|
||||
@@ -18,6 +18,16 @@ type HostReport struct {
|
||||
HostID string `json:"host_id"` // echoes config.Hub.HostID
|
||||
ReportedAt string `json:"reported_at"` // RFC3339, agent clock
|
||||
AgentVersion string `json:"agent_version"`
|
||||
// AgentSHA256 is the sha256 of the binary this process is RUNNING (read through /proc/self/exe,
|
||||
// once per process), R-349. The version string cannot tell a hand-built proof binary from the
|
||||
// published, vouched artifact of the same version — same source, different bytes (`-trimpath
|
||||
// -buildvcs=false` in release-agent.sh) — so self-update sees "already installed" and never
|
||||
// corrects it. Reporting the bytes lets the hub compare against the vouched agent_sha256, the
|
||||
// same mechanism host.wrapper_sha256 is for the PBS wrapper (R-50b(a)).
|
||||
//
|
||||
// Empty = unreadable, which the hub must treat as UNKNOWN, never as drift. Pinned by
|
||||
// TestCollect_AgentSHA256IsTheRunningBinary.
|
||||
AgentSHA256 string `json:"agent_sha256,omitempty"`
|
||||
|
||||
Host HostMetrics `json:"host"`
|
||||
Guests []Guest `json:"guests"`
|
||||
|
||||
@@ -768,6 +768,11 @@ type FormatResponse struct {
|
||||
// signature — the customer authorizes the wipe of their own data drive.
|
||||
NeedsConfirmation bool `json:"needs_confirmation,omitempty"`
|
||||
DurableID string `json:"durable_id,omitempty"` // the durable id to confirm against (user-data)
|
||||
// FSUUID (on Formatted) is the UUID of the filesystem the agent just made, verified against the
|
||||
// bound durable id after mkfs (R-25). The caller mounts THIS — re-resolving a UUID from the /dev
|
||||
// path later can name another disk if /dev re-enumerated. "" = not verified: the caller must not
|
||||
// substitute a path-resolved guess silently.
|
||||
FSUUID string `json:"fs_uuid,omitempty"`
|
||||
// PendingOp is set on a SYSTEM/BACKUP data-bearing refusal — the exact op the operator must sign.
|
||||
PendingOp *PendingOp `json:"pending_op,omitempty"`
|
||||
}
|
||||
@@ -815,7 +820,7 @@ func (s *Server) handleDiskFormatStatus(w http.ResponseWriter, r *http.Request,
|
||||
writeOK(w, map[string]any{
|
||||
"vmid": vmid, "phase": job.Phase, "device": job.Device, "fstype": job.FSType,
|
||||
"durable_id": job.DurableID, "error": job.Error, "started_at": job.StartedAt, "updated_at": job.UpdatedAt,
|
||||
"job_id": job.JobID,
|
||||
"job_id": job.JobID, "fs_uuid": job.FSUUID, // R-25: "" until done + verified
|
||||
})
|
||||
}
|
||||
|
||||
@@ -878,7 +883,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
"format refused (device may have changed since inspection): "+rerr.Error())
|
||||
return
|
||||
}
|
||||
done := s.startFormatDetached(device, blankDurable, req.FSType, true)
|
||||
job, done := s.startFormatDetached(device, blankDurable, req.FSType, true)
|
||||
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
|
||||
if err == errFormatClientGone {
|
||||
return // client gone; mkfs continues detached + the job record records the outcome
|
||||
@@ -887,7 +892,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
writeErr(w, http.StatusBadGateway, "format failed: "+err.Error())
|
||||
return
|
||||
}
|
||||
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, Reason: "blank device formatted " + req.FSType})
|
||||
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, FSUUID: job.FSUUID, Reason: "blank device formatted " + req.FSType})
|
||||
return
|
||||
}
|
||||
|
||||
@@ -924,7 +929,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
// F20-BUG3: run the destructive mkfs DETACHED off s.baseCtx (bound durable id recorded for
|
||||
// restart-recovery), so a request/client deadline can never SIGKILL it mid-write and corrupt the
|
||||
// disk. We still wait to return the synchronous result (backward-compatible with the controller).
|
||||
done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
|
||||
job, done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
|
||||
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
|
||||
if err == errFormatClientGone {
|
||||
return // client gone; the wipe continues detached + survives a restart via the job record
|
||||
@@ -936,7 +941,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
|
||||
s.logger.Warn("local-api: USER-DATA data-bearing format — CUSTOMER CONFIRMED (no operator signature)",
|
||||
"vmid", vmid, "device", device, "durable_id", deviceDurable, "fstype", req.FSType, "why", probe.Reason())
|
||||
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: true,
|
||||
Role: string(role), DurableID: deviceDurable, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
|
||||
Role: string(role), DurableID: deviceDurable, FSUUID: job.FSUUID, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
|
||||
return
|
||||
}
|
||||
|
||||
|
||||
@@ -28,6 +28,7 @@ type fakeDiskOps struct {
|
||||
unmountCalls []string
|
||||
candidates []storage.CandidateDisk // returned by ListCandidateDisks
|
||||
candErr error
|
||||
afterFormat *storage.DeviceProbe // R-25: when set, InspectDevice returns it once a format ran
|
||||
}
|
||||
|
||||
func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.CandidateDisk, error) {
|
||||
@@ -35,7 +36,12 @@ func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.Candidate
|
||||
}
|
||||
|
||||
func (f *fakeDiskOps) InspectDevice(_ context.Context, device string) (storage.DeviceProbe, error) {
|
||||
f.mu.Lock()
|
||||
p := f.probe
|
||||
if f.afterFormat != nil && len(f.formatCalls) > 0 {
|
||||
p = *f.afterFormat
|
||||
}
|
||||
f.mu.Unlock()
|
||||
p.Device = device
|
||||
return p, f.inspectErr
|
||||
}
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"net/http"
|
||||
"sync"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/storage"
|
||||
)
|
||||
|
||||
// R-25: the format answer carries the UUID of the filesystem the agent JUST made, verified against the
|
||||
// bound durable id after mkfs, so the controller mounts that filesystem rather than whatever the /dev
|
||||
// path resolves to a few requests later. The consequence asserted: the UUID on the wire (and in the
|
||||
// polled job record) is the new superblock's — and is EMPTY whenever the binding cannot be re-proved.
|
||||
|
||||
const newFSUUID = "0fc63daf-8483-4772-8e79-3d69d8477de4"
|
||||
|
||||
func confirmedFormat(t *testing.T, d *fakeDiskOps, srv *Server, fj *FormatJobStore) (string, *formatJob) {
|
||||
t.Helper()
|
||||
w := do(t, srv.Handler(), "POST", "/disks/format", "A", `{"device":"/dev/sdb1","fstype":"ext4","confirmed":true,"durable_id":"byid:wwn-/dev/sdb1"}`)
|
||||
if w.Code != http.StatusOK {
|
||||
t.Fatalf("confirmed format: %d (%s)", w.Code, w.Body.String())
|
||||
}
|
||||
var resp struct {
|
||||
Data FormatResponse `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||
t.Fatalf("decode: %v (%s)", err, w.Body.String())
|
||||
}
|
||||
if !resp.Data.Formatted {
|
||||
t.Fatalf("not formatted: %s", w.Body.String())
|
||||
}
|
||||
return resp.Data.FSUUID, waitFormatPhase(t, fj, formatPhaseDone)
|
||||
}
|
||||
|
||||
func confirmedGate() *fakeGate {
|
||||
return &fakeGate{decision: WipeDecision{Allowed: true, Tier: "customer_confirmable", Reason: "customer_confirmed"}}
|
||||
}
|
||||
|
||||
func TestFormat_ReportsNewFilesystemUUID(t *testing.T) {
|
||||
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
|
||||
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
|
||||
fj := tempFormatStore(t)
|
||||
srv := formatServer(t, d, confirmedGate(), fj)
|
||||
|
||||
got, job := confirmedFormat(t, d, srv, fj)
|
||||
if got != newFSUUID {
|
||||
t.Fatalf("response fs_uuid = %q, want the new filesystem's %q", got, newFSUUID)
|
||||
}
|
||||
if job.FSUUID != newFSUUID {
|
||||
t.Fatalf("job record fs_uuid = %q, want %q (the polled status path)", job.FSUUID, newFSUUID)
|
||||
}
|
||||
}
|
||||
|
||||
// The node moved between mkfs and the read-back: the bound durable id now resolves elsewhere. The
|
||||
// UUID must NOT be reported — reading it would name the other disk's filesystem.
|
||||
func TestFormat_FSUUIDWithheldWhenDurableIDMoved(t *testing.T) {
|
||||
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
|
||||
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
|
||||
fj := tempFormatStore(t)
|
||||
srv := formatServer(t, d, confirmedGate(), fj)
|
||||
var mu sync.Mutex
|
||||
calls := 0
|
||||
srv.reresolveWipe = func(_ context.Context, _ string) (string, error) {
|
||||
mu.Lock()
|
||||
defer mu.Unlock()
|
||||
calls++
|
||||
if calls == 1 {
|
||||
return "/dev/sdb", nil // the pre-mkfs anti-retarget re-resolve
|
||||
}
|
||||
return "/dev/sdc", nil // after mkfs: the durable id now names another node
|
||||
}
|
||||
|
||||
got, job := confirmedFormat(t, d, srv, fj)
|
||||
if got != "" || job.FSUUID != "" {
|
||||
t.Fatalf("fs_uuid reported after the durable id moved (response %q, job %q) — must be empty", got, job.FSUUID)
|
||||
}
|
||||
if calls < 2 {
|
||||
t.Fatalf("the post-mkfs re-resolve never ran (calls=%d)", calls)
|
||||
}
|
||||
}
|
||||
|
||||
// The superblock did not read back as the requested filesystem → not verified → empty.
|
||||
func TestFormat_FSUUIDWithheldOnFSTypeMismatch(t *testing.T) {
|
||||
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
|
||||
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "xfs", FSUUID: newFSUUID}}
|
||||
fj := tempFormatStore(t)
|
||||
srv := formatServer(t, d, confirmedGate(), fj)
|
||||
|
||||
got, job := confirmedFormat(t, d, srv, fj)
|
||||
if got != "" || job.FSUUID != "" {
|
||||
t.Fatalf("fs_uuid reported for a superblock of the wrong type (response %q, job %q)", got, job.FSUUID)
|
||||
}
|
||||
}
|
||||
@@ -22,6 +22,10 @@ type formatJob struct {
|
||||
Blank bool `json:"blank,omitempty"` // audit D3: blank (benign) format — recovery re-checks STILL-blank, not data-bearing
|
||||
Phase string `json:"phase"` // running | done | failed
|
||||
Error string `json:"error,omitempty"`
|
||||
// FSUUID is the filesystem UUID of the NEW filesystem, read by the agent right after mkfs on the
|
||||
// device the bound durable id still resolves to (R-25). "" = not verified (the controller must not
|
||||
// read that as a UUID). Set only on phase done.
|
||||
FSUUID string `json:"fs_uuid,omitempty"`
|
||||
StartedAt string `json:"started_at"`
|
||||
UpdatedAt string `json:"updated_at"`
|
||||
}
|
||||
@@ -96,7 +100,10 @@ func (s *FormatJobStore) save(j *formatJob) error {
|
||||
// runs to completion and records the outcome. device is the ALREADY anti-retarget-resolved device; the
|
||||
// record carries durableID so a restart can re-resolve + re-run. blank marks a benign (blank-device)
|
||||
// format, so restart recovery re-checks STILL-blank rather than data-bearing (audit D3).
|
||||
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) <-chan error {
|
||||
//
|
||||
// The returned job may be read (FSUUID) only AFTER a value arrives on done — the goroutine writes it
|
||||
// before the send, which is the happens-before edge.
|
||||
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) (*formatJob, <-chan error) {
|
||||
base := s.baseCtx
|
||||
if base == nil {
|
||||
base = context.Background()
|
||||
@@ -116,10 +123,39 @@ func (s *Server) startFormatDetached(device, durableID, fstype string, blank boo
|
||||
ctx, cancel := context.WithTimeout(base, 60*time.Minute)
|
||||
defer cancel()
|
||||
err := s.disks.Format(ctx, device, fstype)
|
||||
if err == nil {
|
||||
job.FSUUID = s.formattedFSUUID(ctx, device, durableID, fstype)
|
||||
}
|
||||
s.finishFormatJob(job, err)
|
||||
done <- err
|
||||
}()
|
||||
return done
|
||||
return job, done
|
||||
}
|
||||
|
||||
// formattedFSUUID reads the UUID of the filesystem the agent has JUST made (R-25). The caller used to
|
||||
// re-resolve the UUID from the mutable /dev path afterwards, over separate requests — a re-enumeration
|
||||
// in that window could hand it ANOTHER disk's filesystem to mount. Here the bound durable id must still
|
||||
// resolve to the very device that was formatted (and re-derive to the same id), and the superblock must
|
||||
// carry the fstype that was asked for; anything else returns "" (not verified), never a guess.
|
||||
func (s *Server) formattedFSUUID(ctx context.Context, device, durableID, fstype string) string {
|
||||
if durableID == "" || s.reresolveWipe == nil {
|
||||
return ""
|
||||
}
|
||||
// The device now holds a filesystem, so the data-bearing anti-retarget re-resolve is the right one.
|
||||
now, err := s.reresolveWipe(ctx, durableID)
|
||||
if err != nil || now != device {
|
||||
s.logger.Warn("format: new filesystem UUID NOT reported — bound durable id no longer resolves to the formatted device",
|
||||
"device", device, "durable_id", durableID, "resolves_to", now, "err", err)
|
||||
return ""
|
||||
}
|
||||
probe, err := s.disks.InspectDevice(ctx, device)
|
||||
if err != nil || !probe.Probed || probe.FSType != fstype || probe.FSUUID == "" {
|
||||
s.logger.Warn("format: new filesystem UUID NOT reported — superblock did not read back as the requested filesystem",
|
||||
"device", device, "want_fstype", fstype, "got_fstype", probe.FSType, "has_uuid", probe.FSUUID != "", "err", err)
|
||||
return ""
|
||||
}
|
||||
s.logger.Info("format: new filesystem bound to its durable id", "device", device, "durable_id", durableID, "fs_uuid", probe.FSUUID)
|
||||
return probe.FSUUID
|
||||
}
|
||||
|
||||
// finishFormatJob updates the persisted record to done/failed.
|
||||
@@ -172,7 +208,7 @@ func (s *Server) RecoverFormatJob(ctx context.Context) {
|
||||
return
|
||||
}
|
||||
s.logger.Warn("format-job recover: re-running interrupted format detached", "durable_id", job.DurableID, "device", device, "fstype", job.FSType, "blank", job.Blank)
|
||||
_ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
|
||||
_, _ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
|
||||
}
|
||||
|
||||
// nowFn returns the server clock (testable), defaulting to time.Now.
|
||||
|
||||
@@ -57,6 +57,10 @@ type DeviceProbe struct {
|
||||
HasPartitions bool `json:"has_partitions"` // child partitions present (lsblk)
|
||||
Mounted bool `json:"mounted"` // currently mounted somewhere
|
||||
FSType string `json:"fstype,omitempty"`
|
||||
// FSUUID is the filesystem UUID blkid read from the on-disk superblock (`blkid -p`, no cache),
|
||||
// "" when there is none. R-25: the format path reports it so the caller mounts the filesystem the
|
||||
// agent just made, not whatever a /dev path resolves to later.
|
||||
FSUUID string `json:"fs_uuid,omitempty"`
|
||||
}
|
||||
|
||||
// DataBearing is the conservative verdict: any signature / partition table / partition / mount —
|
||||
@@ -423,6 +427,8 @@ func (h *SudoHostOps) InspectDevice(ctx context.Context, device string) (DeviceP
|
||||
probe.FSType = v
|
||||
case "PTTYPE":
|
||||
probe.HasPartitionTable = true
|
||||
case "UUID":
|
||||
probe.FSUUID = v
|
||||
case "USAGE":
|
||||
if v != "" {
|
||||
probe.HasFilesystem = true // filesystem/raid/crypto member = data-bearing
|
||||
|
||||
@@ -208,3 +208,17 @@ func TestFormat_RejectsBadArgs(t *testing.T) {
|
||||
t.Fatalf("mkfs ran despite invalid input: %v", r.calls)
|
||||
}
|
||||
}
|
||||
|
||||
// R-25: the probe carries the superblock's filesystem UUID so the format path can report the new one.
|
||||
func TestInspect_ReadsFilesystemUUID(t *testing.T) {
|
||||
r := &scriptedRunner{
|
||||
outputs: map[string][]byte{
|
||||
"blkid": []byte("DEVNAME=/dev/sdb\nUUID=0fc63daf-8483-4772-8e79-3d69d8477de4\nTYPE=ext4\nUSAGE=filesystem\n"),
|
||||
"lsblk": []byte(`{"blockdevices":[{"name":"sdb","fstype":"ext4","pttype":null,"mountpoint":null}]}`),
|
||||
},
|
||||
}
|
||||
p, _ := newSudo(r).InspectDevice(context.Background(), "/dev/sdb")
|
||||
if p.FSUUID != "0fc63daf-8483-4772-8e79-3d69d8477de4" {
|
||||
t.Fatalf("FSUUID = %q, want the blkid UUID", p.FSUUID)
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user