Compare commits

..

5 Commits

Author SHA1 Message Date
admin 87977ff40a CHANGELOG/v0.148.0 released (shas)
gates / gates (push) Successful in 42s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-06 00:37:07 +02:00
admin 861d32a4b4 CHANGELOG: unreleased — R-349 agent_sha256, R-25 agent half (burn-down night)
gates / gates (push) Successful in 43s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:51 +02:00
admin 769c4c3cf2 R-25 (agent half): the format answer carries the NEW filesystem's UUID, bound to the durable id
After mkfs the agent re-resolves the bound durable id, requires it to name the
device it just formatted, reads the superblock back (blkid -p, requested fstype)
and returns fs_uuid in POST /disks/format and GET /disks/format/status. Anything
unverified returns "" — never a path-resolved guess. The controller half
(mount fs_uuid instead of re-resolving the UUID from the /dev path) is owed
in felhom-controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 64f704d0f7 R-349: the host report carries the sha256 of the RUNNING agent binary
A hand-built proof binary and the published artifact share a version string
but not their bytes, so no version check could see the divergence. The agent
now reports agent_sha256 (hash of /proc/self/exe, once per process; empty =
unknown) beside agent_version, the same mechanism as host.wrapper_sha256.
The hub half (compare against the vouched agent_sha256, surface drift) is
owed in the hub repo.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 21:16:11 +02:00
admin 208fac8027 CHANGELOG/REPORT: v0.147.0 released (shas)
gates / gates (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-05 18:58:37 +02:00
12 changed files with 299 additions and 17 deletions
+19 -1
View File
@@ -1,4 +1,22 @@
## unreleased (to be released as v0.147.0) — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
## v0.148.0 — the host report names the running binary's sha; the format answer carries the new filesystem's UUID (burn-down night: R-349, R-25 agent halves) (2026-10-06)
Released by `scripts/release-agent.sh`: binary sha256 `3e68a0870e0e2ce262cb4819294edddeb0a73e8c558a31611a20329a6d9ee283`
config bundle sha256 `a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de` (tag `v0.148.0` = `861d32a`).
Delivery order as for v0.147.0: signed `agent_update`, then signed `agent_config_update`.
- **R-349:** the host report carries `agent_sha256`, the sha256 of the running agent binary (read once from
`/proc/self/exe`; empty = unknown), so a hand-built binary under the vouched version name becomes visible. The hub
comparison is a separate hub change. Test `TestCollect_AgentSHA256IsTheRunningBinary`; red-proved.
- **R-25 (agent half):** `POST /disks/format` and `GET /disks/format/status` return `fs_uuid`, the new filesystem's UUID
read back after mkfs only when the bound durable id still resolves to the formatted device and the superblock is the
requested type (empty = not verified); `DeviceProbe` gains `FSUUID` from blkid. Tests `TestFormat_*FSUUID*`; three
red-proofs. The controller half (mount that UUID) is a controller change.
## v0.147.0 — the recovery recipe spells the root namespace the way PBS does; a removed drive no longer shows the root disk's size; a rotated-out token stops at once; the dnsmasq check looks at the right package (burn-down round 2: R-124, R-118, R-269, R-317) (2026-10-05)
Released by `scripts/release-agent.sh`: binary sha256 `642c4d196c48671c14ff653118303c5903af1670b7abb550edeaf73e701cd5b8`,
config bundle sha256 `326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d` (tag `v0.147.0` = `f1b9b41`).
Delivery order as for v0.146.1: signed `agent_update`, then signed `agent_config_update`.
MinAgent impact: none (the controller needs nothing new from this agent). Config bundle content unchanged from v0.146.1.
+8 -5
View File
@@ -1,7 +1,10 @@
# REPORT — 2026-10-05 burn-down: two record corrections (no release)
# REPORT — agent v0.147.0 (2026-10-05, burn-down round 2)
Full session report: `felhom.eu/REPORT-burndown-2026-10-05.md`. Baseline `e06ed97` (v0.146.1). No binary, bundle or
version change; `go build`, `go vet ./internal/backup/`, `agent_gates.py --fast` green.
Full session report: `felhom.eu/REPORT-burndown2-2026-10-05.md`. Baseline `d833163` (v0.146.1). Code commit `f1b9b41`
(CI job 1365 success), tag `v0.147.0`, binary sha256 `642c4d19…`, bundle sha256 `326527d0…` (verified by download).
- R-291 — `scripts/retention-policy.json`: the source of the number (R-267 newest-10 prune, R-287) and readers corrected.
- R-348 — `internal/backup/store.go`: the restart comment says what a restart really blanks.
Rows: R-124 (recipe root namespace = ""), R-118 (no root size for an absent drive), R-269 (rotated-out token rejected at
once), R-317 (dnsmasq install probed by its unit). Tests + red-proofs: `felhom.eu/documentation/audits/burndown2-2026-10-05/`.
`go build/vet/test ./...` green; `agent_gates.py --fast` green after the release (release-complete needs the tag).
Delivery: see the session report (vouch, signed jobs per box, the hub System page afterwards).
+1 -1
View File
@@ -66,7 +66,7 @@
|---|---|---|---|---|
| `IntentStore` (`Get/SetEnrolled/SetEjected/SetDecommissioned/OnAbsent`) | internal/storage/intent.go | `OpenIntentStore(path)` | drive intent (4-state self-heal) | Keyed by durable-id only; `OnAbsent` is the ONLY ejected→enrolled path; refuses empty ids |
| `GuestBindStore` (`Record/Remove/Guests`) | internal/localapi/guestbindstore.go | `OpenGuestBindStore(path)` | per-guest enrolled binds (F9 re-assert) | Same tmp+rename 0600 pattern as IntentStore |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) <-chan error` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank |
| `FormatJobStore` + `startFormatDetached` + `RecoverFormatJob` | internal/localapi/formatjob.go | `startFormatDetached(device, durableID, fstype, blank) (*formatJob, <-chan error)` | detached, restart-surviving mkfs (F20-BUG3) | Runs off `s.baseCtx` (60-min bound) so a request deadline can't SIGKILL mkfs; recovery re-resolves by durable id; blank jobs re-check STILL-blank; on success `job.FSUUID` = the new filesystem's UUID, read back only when the durable id still resolves to the formatted device (R-25) — read it only after `done` delivers |
| `TokenStore.Mint` / `Lookup` | internal/localapi/tokenstore.go | `Mint(vmid) (plaintext, error)` | per-guest local-API tokens | Only the SHA-256 hash persists (fsync'd append log); constant-time compare on lookup; plaintext returned exactly once. Lookup stats the file on EVERY call and reloads BEFORE answering when the append-only log grew (R-269; was reload-on-miss only, v0.63.0 B3, which let a token rotated out by another process keep authorizing as a map hit): the one-shot provisioner mints into the same file the daemon indexes — cross-process coherence both ways without a restart; unchanged size = no re-read. Pinned by `TestTokenStore_RotatedOutTokenRejectedFirst` |
| `FileNonceStore.SeenOrRecord` | internal/authz/noncestore.go | `SeenOrRecord(nonce, exp) bool` | durable anti-replay | fsync'd before returning false; prune only after exp |
| `Journal` (`Append/Latest/InFlight/AlreadyApplied`) | internal/reconcile/journal.go | `OpenJournal(path)` | op journal + idempotency + crash recovery | `Recover` consumes `InFlight()`; scratch entries special-cased |
+26 -2
View File
@@ -10,6 +10,7 @@ import (
"log/slog"
"os"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
@@ -115,6 +116,7 @@ type Collector struct {
backupTarget func() ConfiguredBackupTarget // R-109: primary backup tier id (nil → recipe records unknown)
hostID string
agentVersion string
selfSHA func() string // R-349: sha256 of the running binary; default runningBinarySHA256
logger *slog.Logger
now func() time.Time
}
@@ -135,6 +137,7 @@ func NewCollector(px proxmoxReader, cf CloudflaredProber, storage StorageObserve
temp: SysfsTempReader{}, // slice 9: real sysfs reader by default; tests inject a fake
hostID: hostID,
agentVersion: agentVersion,
selfSHA: runningBinarySHA256,
logger: logger,
now: func() time.Time { return time.Now().UTC() },
}
@@ -337,6 +340,7 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
HostID: c.hostID,
ReportedAt: c.now().Format(time.RFC3339),
AgentVersion: c.agentVersion,
AgentSHA256: c.agentSHA256(),
Host: host,
Guests: c.collectGuests(ctx),
// storage_targets populated this slice (slice 5) via the observer; the rest stay
@@ -431,8 +435,28 @@ const pbsWrapperPath = "/usr/local/sbin/felhom-pbs-apply"
// unreadable file yields "", which the hub reads as UNKNOWN rather than as drift — a host that
// legitimately has no DR wrapper must not light up amber. The file is 0755, so no privilege is
// needed to read it.
func pbsWrapperSHA256() string {
f, err := os.Open(pbsWrapperPath)
func pbsWrapperSHA256() string { return fileSHA256(pbsWrapperPath) }
// selfExePath is the running binary as the kernel holds it. /proc/self/exe, not the installed path:
// after an A/B flip the file at /usr/local/bin/felhom-agent may already be the NEXT binary while this
// process still runs the old one, and the report must describe what runs (R-349). Test seam.
var selfExePath = "/proc/self/exe"
// runningBinarySHA256 hashes the running binary ONCE per process — the bytes cannot change under a
// running process, and re-hashing ~20 MB every report cycle buys nothing. A failed read is cached as
// "" (UNKNOWN); it never fails the report.
var runningBinarySHA256 = sync.OnceValue(func() string { return fileSHA256(selfExePath) })
func (c *Collector) agentSHA256() string {
if c.selfSHA == nil {
return ""
}
return c.selfSHA()
}
// fileSHA256 is the hex sha256 of a file's bytes, or "" when it cannot be read.
func fileSHA256(path string) string {
f, err := os.Open(path)
if err != nil {
return ""
}
+64
View File
@@ -0,0 +1,64 @@
package hub
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
// R-349: the report carries the sha256 of the binary that is RUNNING, so the hub can tell a
// hand-built proof binary from the vouched artifact of the same version string. The consequence
// asserted: the wire field equals the hash of this very test binary's bytes (read independently via
// os.Executable, a different channel from /proc/self/exe), and it is on the wire as agent_sha256.
func TestCollect_AgentSHA256IsTheRunningBinary(t *testing.T) {
exe, err := os.Executable()
if err != nil {
t.Skipf("os.Executable: %v", err)
}
raw, err := os.ReadFile(exe)
if err != nil {
t.Fatalf("read own binary: %v", err)
}
sum := sha256.Sum256(raw)
want := hex.EncodeToString(sum[:])
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect: %v", err)
}
if r.AgentSHA256 != want {
t.Fatalf("agent_sha256 = %q, want the running binary's %q", r.AgentSHA256, want)
}
b, err := json.Marshal(r)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(b), `"agent_sha256":"`+want+`"`) {
t.Fatalf("agent_sha256 not on the wire: %s", b)
}
}
// An unreadable binary is UNKNOWN (empty, omitted) — never a made-up hash, never a failed report.
func TestFileSHA256_UnreadableIsEmpty(t *testing.T) {
if got := fileSHA256(filepath.Join(t.TempDir(), "absent")); got != "" {
t.Fatalf("absent file hashed to %q, want empty", got)
}
px := &fakePx{node: "n", ns: newTestNodeStatus()}
c := NewCollector(px, fakeProber{status: "running"}, nil, nil, nil, nil, "h", "0.3.0", quietLogger())
c.selfSHA = func() string { return "" }
r, err := c.Collect(context.Background())
if err != nil {
t.Fatalf("Collect must not fail on an unreadable binary: %v", err)
}
b, _ := json.Marshal(r)
if strings.Contains(string(b), "agent_sha256") {
t.Fatalf("empty agent_sha256 must be omitted: %s", b)
}
}
+10
View File
@@ -18,6 +18,16 @@ type HostReport struct {
HostID string `json:"host_id"` // echoes config.Hub.HostID
ReportedAt string `json:"reported_at"` // RFC3339, agent clock
AgentVersion string `json:"agent_version"`
// AgentSHA256 is the sha256 of the binary this process is RUNNING (read through /proc/self/exe,
// once per process), R-349. The version string cannot tell a hand-built proof binary from the
// published, vouched artifact of the same version — same source, different bytes (`-trimpath
// -buildvcs=false` in release-agent.sh) — so self-update sees "already installed" and never
// corrects it. Reporting the bytes lets the hub compare against the vouched agent_sha256, the
// same mechanism host.wrapper_sha256 is for the PBS wrapper (R-50b(a)).
//
// Empty = unreadable, which the hub must treat as UNKNOWN, never as drift. Pinned by
// TestCollect_AgentSHA256IsTheRunningBinary.
AgentSHA256 string `json:"agent_sha256,omitempty"`
Host HostMetrics `json:"host"`
Guests []Guest `json:"guests"`
+10 -5
View File
@@ -768,6 +768,11 @@ type FormatResponse struct {
// signature — the customer authorizes the wipe of their own data drive.
NeedsConfirmation bool `json:"needs_confirmation,omitempty"`
DurableID string `json:"durable_id,omitempty"` // the durable id to confirm against (user-data)
// FSUUID (on Formatted) is the UUID of the filesystem the agent just made, verified against the
// bound durable id after mkfs (R-25). The caller mounts THIS — re-resolving a UUID from the /dev
// path later can name another disk if /dev re-enumerated. "" = not verified: the caller must not
// substitute a path-resolved guess silently.
FSUUID string `json:"fs_uuid,omitempty"`
// PendingOp is set on a SYSTEM/BACKUP data-bearing refusal — the exact op the operator must sign.
PendingOp *PendingOp `json:"pending_op,omitempty"`
}
@@ -815,7 +820,7 @@ func (s *Server) handleDiskFormatStatus(w http.ResponseWriter, r *http.Request,
writeOK(w, map[string]any{
"vmid": vmid, "phase": job.Phase, "device": job.Device, "fstype": job.FSType,
"durable_id": job.DurableID, "error": job.Error, "started_at": job.StartedAt, "updated_at": job.UpdatedAt,
"job_id": job.JobID,
"job_id": job.JobID, "fs_uuid": job.FSUUID, // R-25: "" until done + verified
})
}
@@ -878,7 +883,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
"format refused (device may have changed since inspection): "+rerr.Error())
return
}
done := s.startFormatDetached(device, blankDurable, req.FSType, true)
job, done := s.startFormatDetached(device, blankDurable, req.FSType, true)
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
if err == errFormatClientGone {
return // client gone; mkfs continues detached + the job record records the outcome
@@ -887,7 +892,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
writeErr(w, http.StatusBadGateway, "format failed: "+err.Error())
return
}
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, Reason: "blank device formatted " + req.FSType})
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: false, DurableID: blankDurable, FSUUID: job.FSUUID, Reason: "blank device formatted " + req.FSType})
return
}
@@ -924,7 +929,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
// F20-BUG3: run the destructive mkfs DETACHED off s.baseCtx (bound durable id recorded for
// restart-recovery), so a request/client deadline can never SIGKILL it mid-write and corrupt the
// disk. We still wait to return the synchronous result (backward-compatible with the controller).
done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
job, done := s.startFormatDetached(device, deviceDurable, req.FSType, false)
if err := s.awaitFormat(r.Context(), done, vmid, device); err != nil {
if err == errFormatClientGone {
return // client gone; the wipe continues detached + survives a restart via the job record
@@ -936,7 +941,7 @@ func (s *Server) handleDiskFormat(w http.ResponseWriter, r *http.Request, vmid i
s.logger.Warn("local-api: USER-DATA data-bearing format — CUSTOMER CONFIRMED (no operator signature)",
"vmid", vmid, "device", device, "durable_id", deviceDurable, "fstype", req.FSType, "why", probe.Reason())
writeOK(w, FormatResponse{VMID: vmid, Device: device, Formatted: true, DataBearing: true,
Role: string(role), DurableID: deviceDurable, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
Role: string(role), DurableID: deviceDurable, FSUUID: job.FSUUID, Reason: "customer-confirmed wipe (" + probe.Reason() + ")"})
return
}
+6
View File
@@ -28,6 +28,7 @@ type fakeDiskOps struct {
unmountCalls []string
candidates []storage.CandidateDisk // returned by ListCandidateDisks
candErr error
afterFormat *storage.DeviceProbe // R-25: when set, InspectDevice returns it once a format ran
}
func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.CandidateDisk, error) {
@@ -35,7 +36,12 @@ func (f *fakeDiskOps) ListCandidateDisks(_ context.Context) ([]storage.Candidate
}
func (f *fakeDiskOps) InspectDevice(_ context.Context, device string) (storage.DeviceProbe, error) {
f.mu.Lock()
p := f.probe
if f.afterFormat != nil && len(f.formatCalls) > 0 {
p = *f.afterFormat
}
f.mu.Unlock()
p.Device = device
return p, f.inspectErr
}
+96
View File
@@ -0,0 +1,96 @@
package localapi
import (
"context"
"encoding/json"
"net/http"
"sync"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/storage"
)
// R-25: the format answer carries the UUID of the filesystem the agent JUST made, verified against the
// bound durable id after mkfs, so the controller mounts that filesystem rather than whatever the /dev
// path resolves to a few requests later. The consequence asserted: the UUID on the wire (and in the
// polled job record) is the new superblock's — and is EMPTY whenever the binding cannot be re-proved.
const newFSUUID = "0fc63daf-8483-4772-8e79-3d69d8477de4"
func confirmedFormat(t *testing.T, d *fakeDiskOps, srv *Server, fj *FormatJobStore) (string, *formatJob) {
t.Helper()
w := do(t, srv.Handler(), "POST", "/disks/format", "A", `{"device":"/dev/sdb1","fstype":"ext4","confirmed":true,"durable_id":"byid:wwn-/dev/sdb1"}`)
if w.Code != http.StatusOK {
t.Fatalf("confirmed format: %d (%s)", w.Code, w.Body.String())
}
var resp struct {
Data FormatResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
if !resp.Data.Formatted {
t.Fatalf("not formatted: %s", w.Body.String())
}
return resp.Data.FSUUID, waitFormatPhase(t, fj, formatPhaseDone)
}
func confirmedGate() *fakeGate {
return &fakeGate{decision: WipeDecision{Allowed: true, Tier: "customer_confirmable", Reason: "customer_confirmed"}}
}
func TestFormat_ReportsNewFilesystemUUID(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
got, job := confirmedFormat(t, d, srv, fj)
if got != newFSUUID {
t.Fatalf("response fs_uuid = %q, want the new filesystem's %q", got, newFSUUID)
}
if job.FSUUID != newFSUUID {
t.Fatalf("job record fs_uuid = %q, want %q (the polled status path)", job.FSUUID, newFSUUID)
}
}
// The node moved between mkfs and the read-back: the bound durable id now resolves elsewhere. The
// UUID must NOT be reported — reading it would name the other disk's filesystem.
func TestFormat_FSUUIDWithheldWhenDurableIDMoved(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "ext4", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
var mu sync.Mutex
calls := 0
srv.reresolveWipe = func(_ context.Context, _ string) (string, error) {
mu.Lock()
defer mu.Unlock()
calls++
if calls == 1 {
return "/dev/sdb", nil // the pre-mkfs anti-retarget re-resolve
}
return "/dev/sdc", nil // after mkfs: the durable id now names another node
}
got, job := confirmedFormat(t, d, srv, fj)
if got != "" || job.FSUUID != "" {
t.Fatalf("fs_uuid reported after the durable id moved (response %q, job %q) — must be empty", got, job.FSUUID)
}
if calls < 2 {
t.Fatalf("the post-mkfs re-resolve never ran (calls=%d)", calls)
}
}
// The superblock did not read back as the requested filesystem → not verified → empty.
func TestFormat_FSUUIDWithheldOnFSTypeMismatch(t *testing.T) {
d := &fakeDiskOps{probe: deviceProbeDataBearing(),
afterFormat: &storage.DeviceProbe{Probed: true, HasFilesystem: true, FSType: "xfs", FSUUID: newFSUUID}}
fj := tempFormatStore(t)
srv := formatServer(t, d, confirmedGate(), fj)
got, job := confirmedFormat(t, d, srv, fj)
if got != "" || job.FSUUID != "" {
t.Fatalf("fs_uuid reported for a superblock of the wrong type (response %q, job %q)", got, job.FSUUID)
}
}
+39 -3
View File
@@ -22,6 +22,10 @@ type formatJob struct {
Blank bool `json:"blank,omitempty"` // audit D3: blank (benign) format — recovery re-checks STILL-blank, not data-bearing
Phase string `json:"phase"` // running | done | failed
Error string `json:"error,omitempty"`
// FSUUID is the filesystem UUID of the NEW filesystem, read by the agent right after mkfs on the
// device the bound durable id still resolves to (R-25). "" = not verified (the controller must not
// read that as a UUID). Set only on phase done.
FSUUID string `json:"fs_uuid,omitempty"`
StartedAt string `json:"started_at"`
UpdatedAt string `json:"updated_at"`
}
@@ -96,7 +100,10 @@ func (s *FormatJobStore) save(j *formatJob) error {
// runs to completion and records the outcome. device is the ALREADY anti-retarget-resolved device; the
// record carries durableID so a restart can re-resolve + re-run. blank marks a benign (blank-device)
// format, so restart recovery re-checks STILL-blank rather than data-bearing (audit D3).
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) <-chan error {
//
// The returned job may be read (FSUUID) only AFTER a value arrives on done — the goroutine writes it
// before the send, which is the happens-before edge.
func (s *Server) startFormatDetached(device, durableID, fstype string, blank bool) (*formatJob, <-chan error) {
base := s.baseCtx
if base == nil {
base = context.Background()
@@ -116,10 +123,39 @@ func (s *Server) startFormatDetached(device, durableID, fstype string, blank boo
ctx, cancel := context.WithTimeout(base, 60*time.Minute)
defer cancel()
err := s.disks.Format(ctx, device, fstype)
if err == nil {
job.FSUUID = s.formattedFSUUID(ctx, device, durableID, fstype)
}
s.finishFormatJob(job, err)
done <- err
}()
return done
return job, done
}
// formattedFSUUID reads the UUID of the filesystem the agent has JUST made (R-25). The caller used to
// re-resolve the UUID from the mutable /dev path afterwards, over separate requests — a re-enumeration
// in that window could hand it ANOTHER disk's filesystem to mount. Here the bound durable id must still
// resolve to the very device that was formatted (and re-derive to the same id), and the superblock must
// carry the fstype that was asked for; anything else returns "" (not verified), never a guess.
func (s *Server) formattedFSUUID(ctx context.Context, device, durableID, fstype string) string {
if durableID == "" || s.reresolveWipe == nil {
return ""
}
// The device now holds a filesystem, so the data-bearing anti-retarget re-resolve is the right one.
now, err := s.reresolveWipe(ctx, durableID)
if err != nil || now != device {
s.logger.Warn("format: new filesystem UUID NOT reported — bound durable id no longer resolves to the formatted device",
"device", device, "durable_id", durableID, "resolves_to", now, "err", err)
return ""
}
probe, err := s.disks.InspectDevice(ctx, device)
if err != nil || !probe.Probed || probe.FSType != fstype || probe.FSUUID == "" {
s.logger.Warn("format: new filesystem UUID NOT reported — superblock did not read back as the requested filesystem",
"device", device, "want_fstype", fstype, "got_fstype", probe.FSType, "has_uuid", probe.FSUUID != "", "err", err)
return ""
}
s.logger.Info("format: new filesystem bound to its durable id", "device", device, "durable_id", durableID, "fs_uuid", probe.FSUUID)
return probe.FSUUID
}
// finishFormatJob updates the persisted record to done/failed.
@@ -172,7 +208,7 @@ func (s *Server) RecoverFormatJob(ctx context.Context) {
return
}
s.logger.Warn("format-job recover: re-running interrupted format detached", "durable_id", job.DurableID, "device", device, "fstype", job.FSType, "blank", job.Blank)
_ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
_, _ = s.startFormatDetached(device, job.DurableID, job.FSType, job.Blank) // detached; updates the record on completion
}
// nowFn returns the server clock (testable), defaulting to time.Now.
+6
View File
@@ -57,6 +57,10 @@ type DeviceProbe struct {
HasPartitions bool `json:"has_partitions"` // child partitions present (lsblk)
Mounted bool `json:"mounted"` // currently mounted somewhere
FSType string `json:"fstype,omitempty"`
// FSUUID is the filesystem UUID blkid read from the on-disk superblock (`blkid -p`, no cache),
// "" when there is none. R-25: the format path reports it so the caller mounts the filesystem the
// agent just made, not whatever a /dev path resolves to later.
FSUUID string `json:"fs_uuid,omitempty"`
}
// DataBearing is the conservative verdict: any signature / partition table / partition / mount —
@@ -423,6 +427,8 @@ func (h *SudoHostOps) InspectDevice(ctx context.Context, device string) (DeviceP
probe.FSType = v
case "PTTYPE":
probe.HasPartitionTable = true
case "UUID":
probe.FSUUID = v
case "USAGE":
if v != "" {
probe.HasFilesystem = true // filesystem/raid/crypto member = data-bearing
+14
View File
@@ -208,3 +208,17 @@ func TestFormat_RejectsBadArgs(t *testing.T) {
t.Fatalf("mkfs ran despite invalid input: %v", r.calls)
}
}
// R-25: the probe carries the superblock's filesystem UUID so the format path can report the new one.
func TestInspect_ReadsFilesystemUUID(t *testing.T) {
r := &scriptedRunner{
outputs: map[string][]byte{
"blkid": []byte("DEVNAME=/dev/sdb\nUUID=0fc63daf-8483-4772-8e79-3d69d8477de4\nTYPE=ext4\nUSAGE=filesystem\n"),
"lsblk": []byte(`{"blockdevices":[{"name":"sdb","fstype":"ext4","pttype":null,"mountpoint":null}]}`),
},
}
p, _ := newSudo(r).InspectDevice(context.Background(), "/dev/sdb")
if p.FSUUID != "0fc63daf-8483-4772-8e79-3d69d8477de4" {
t.Fatalf("FSUUID = %q, want the blkid UUID", p.FSUUID)
}
}