v0.131.0: controller supervisor (R-523); per-tier backup status + tier storage presence (R-517/R-518)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,3 +1,35 @@
|
||||
## Unreleased (to become v0.131.0) — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
|
||||
|
||||
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
|
||||
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
|
||||
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
|
||||
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
|
||||
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
|
||||
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
|
||||
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
|
||||
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
|
||||
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
|
||||
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
|
||||
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
|
||||
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
|
||||
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
|
||||
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
|
||||
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
|
||||
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
|
||||
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
|
||||
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
|
||||
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
|
||||
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
|
||||
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
|
||||
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
|
||||
after an agent restart — `last_success_source: storage`), the last attempt kept apart
|
||||
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
|
||||
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
|
||||
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
|
||||
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
|
||||
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
|
||||
last success".
|
||||
|
||||
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
|
||||
|
||||
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
|
||||
|
||||
@@ -74,6 +74,7 @@
|
||||
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
|
||||
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
|
||||
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
|
||||
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
|
||||
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
|
||||
|
||||
### Proxmox client / hub / PBS / provisioning
|
||||
|
||||
@@ -59,7 +59,7 @@ import (
|
||||
|
||||
// version is the agent version. Overridable at build time with
|
||||
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
|
||||
var version = "0.130.0"
|
||||
var version = "0.131.0"
|
||||
|
||||
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
|
||||
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
|
||||
@@ -1401,6 +1401,10 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
|
||||
// running" signal, so a deliberately stopped guest is never touched.
|
||||
go localSrv.WatchGuestPower(ctx)
|
||||
// R-523: a controller container that is simply not running (killed, stopped, a failed
|
||||
// self-update) is restarted through its bootstrap unit — nothing else watches it.
|
||||
collector.SetControllerSupervisorReporter(localSrv)
|
||||
go localSrv.WatchControllers(ctx)
|
||||
go func() { errc <- localSrv.Run(ctx) }()
|
||||
}
|
||||
if lanLoop != nil {
|
||||
@@ -1826,6 +1830,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
||||
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
|
||||
SmbCredsDir: cfg.Privileged.SmbCredsDir,
|
||||
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
|
||||
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
|
||||
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
|
||||
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
|
||||
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
|
||||
|
||||
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
|
||||
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
|
||||
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
|
||||
# view, and the docker socket. The controller reaches the agent's local API for disk management.
|
||||
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
|
||||
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
|
||||
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
|
||||
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
|
||||
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
|
||||
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
|
||||
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
|
||||
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
|
||||
-v felhom-controller-data:/opt/docker/felhom-controller \
|
||||
|
||||
@@ -103,6 +103,7 @@ type Collector struct {
|
||||
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
|
||||
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
|
||||
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
|
||||
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
|
||||
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
|
||||
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
||||
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
||||
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
|
||||
return c
|
||||
}
|
||||
|
||||
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
|
||||
type ControllerSupervisorReporter interface {
|
||||
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
|
||||
}
|
||||
|
||||
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
|
||||
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
|
||||
c.ctrlSup = r
|
||||
return c
|
||||
}
|
||||
|
||||
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
|
||||
// omitted). Returns the collector for chaining.
|
||||
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
|
||||
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
if c.pbsdr != nil {
|
||||
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
|
||||
}
|
||||
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
|
||||
if c.ctrlSup != nil {
|
||||
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
|
||||
}
|
||||
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
|
||||
if c.guestNet != nil {
|
||||
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
|
||||
|
||||
@@ -124,6 +124,32 @@ type HostReport struct {
|
||||
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
|
||||
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
|
||||
OOB *OOBStatus `json:"oob,omitempty"`
|
||||
|
||||
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
|
||||
// how many times the agent restarted a dead controller, when last and why, whether it gave up
|
||||
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
|
||||
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
|
||||
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
|
||||
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
|
||||
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
|
||||
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
|
||||
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
|
||||
}
|
||||
|
||||
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
|
||||
type ControllerSupervisorStatus struct {
|
||||
Guests []ControllerSupervisorGuest `json:"guests"`
|
||||
}
|
||||
|
||||
// ControllerSupervisorGuest is one supervised guest.
|
||||
type ControllerSupervisorGuest struct {
|
||||
VMID int `json:"vmid"`
|
||||
RestartsTotal int `json:"restarts_total"`
|
||||
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
|
||||
LastReason string `json:"last_reason,omitempty"`
|
||||
Crashloop bool `json:"crashloop"`
|
||||
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
|
||||
Parked bool `json:"parked"`
|
||||
}
|
||||
|
||||
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
|
||||
|
||||
@@ -0,0 +1,133 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"log/slog"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
|
||||
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
|
||||
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
|
||||
|
||||
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
|
||||
t.Helper()
|
||||
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
|
||||
var resp struct {
|
||||
Data BackupStatusResponse `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||
t.Fatalf("decode: %v (%s)", err, w.Body.String())
|
||||
}
|
||||
return resp.Data.Tiers
|
||||
}
|
||||
|
||||
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
|
||||
t.Helper()
|
||||
srv, err := NewServer(Options{
|
||||
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
|
||||
Storage: fakeStorage{targets: targets},
|
||||
Tokens: staticTokens{"A": 8200},
|
||||
BackupTiers: []BackupTier{
|
||||
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
|
||||
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
|
||||
},
|
||||
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
srv.baseCtx = context.Background()
|
||||
srv.now = func() time.Time { return testNow }
|
||||
return srv
|
||||
}
|
||||
|
||||
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
|
||||
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
|
||||
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
|
||||
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
|
||||
st := &fakeStore{backups: []hub.Backup{
|
||||
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
|
||||
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
|
||||
}}
|
||||
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
|
||||
tiers := tierStatesOf(t, srv)
|
||||
if len(tiers) != 2 {
|
||||
t.Fatalf("want 2 tiers, got %+v", tiers)
|
||||
}
|
||||
local, pbs := tiers[0], tiers[1]
|
||||
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
|
||||
t.Fatalf("local tier lost its successful backup: %+v", local)
|
||||
}
|
||||
if pbs.LastSuccess != nil {
|
||||
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
|
||||
}
|
||||
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
|
||||
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
|
||||
}
|
||||
if pbs.Storage != StoragePresenceAbsent {
|
||||
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
|
||||
}
|
||||
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
|
||||
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
|
||||
var resp struct {
|
||||
Data BackupStatusResponse `json:"data"`
|
||||
}
|
||||
_ = json.Unmarshal(w.Body.Bytes(), &resp)
|
||||
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
|
||||
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
|
||||
}
|
||||
}
|
||||
|
||||
// /backup/tiers advertises storage presence, tri-state.
|
||||
func TestBackupTiers_StoragePresence(t *testing.T) {
|
||||
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
|
||||
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
|
||||
var resp struct {
|
||||
Data BackupTiersResponse `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
got := map[string]string{}
|
||||
for _, ti := range resp.Data.Tiers {
|
||||
got[ti.Target] = ti.Storage
|
||||
}
|
||||
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
|
||||
t.Fatalf("storage presence wrong: %v", got)
|
||||
}
|
||||
// An unreadable storage view is "unknown", never "absent".
|
||||
srv.storage = tierErrStorage{}
|
||||
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
|
||||
t.Fatalf("unreadable storage view must be unknown, got %q", p)
|
||||
}
|
||||
}
|
||||
|
||||
type tierErrStorage struct{}
|
||||
|
||||
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
|
||||
return nil, context.DeadlineExceeded
|
||||
}
|
||||
|
||||
// A targeted request keeps the pre-R-517 bytes (no tiers array).
|
||||
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
|
||||
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
|
||||
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
|
||||
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
|
||||
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
|
||||
}
|
||||
}
|
||||
|
||||
func containsKey(b []byte, key string) bool {
|
||||
var m struct {
|
||||
Data map[string]json.RawMessage `json:"data"`
|
||||
}
|
||||
_ = json.Unmarshal(b, &m)
|
||||
_, ok := m.Data[key]
|
||||
return ok
|
||||
}
|
||||
@@ -0,0 +1,359 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// R-523 — the in-guest controller supervisor.
|
||||
//
|
||||
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
|
||||
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
|
||||
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
|
||||
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
|
||||
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
|
||||
// dashboard answered 502 for 33 minutes until the box was power-cycled.
|
||||
//
|
||||
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
|
||||
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
|
||||
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
|
||||
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
|
||||
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
|
||||
// No new privilege.
|
||||
//
|
||||
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
|
||||
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
|
||||
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
|
||||
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
|
||||
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
|
||||
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
|
||||
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
|
||||
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
|
||||
// again after 30 minutes.
|
||||
//
|
||||
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
|
||||
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
|
||||
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
|
||||
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
|
||||
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
|
||||
// record) can never read as a new restart.
|
||||
|
||||
const (
|
||||
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
|
||||
// so a killed controller is restarted 30–60 s after it died.
|
||||
controllerSupervisorInterval = 30 * time.Second
|
||||
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
|
||||
controllerSupervisorConfirm = 2
|
||||
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
|
||||
// controllerCrashloopPause.
|
||||
controllerCrashloopMax = 3
|
||||
controllerCrashloopWindow = 15 * time.Minute
|
||||
controllerCrashloopPause = 30 * time.Minute
|
||||
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
|
||||
// watchdog is indistinguishable from a dead one (standing rule 3).
|
||||
controllerSupervisorHeartbeatEvery = 20
|
||||
|
||||
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
|
||||
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
|
||||
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
|
||||
// the controller container, and a customer inside the guest cannot park the supervisor.
|
||||
ControllerParkedMarker = "controller-parked"
|
||||
|
||||
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
|
||||
)
|
||||
|
||||
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
|
||||
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
|
||||
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
|
||||
type controllerSupState struct {
|
||||
notRunningSeen int
|
||||
restarts []time.Time // restart times inside the crash-loop window (pruned)
|
||||
restartsTotal int
|
||||
lastRestartAt time.Time
|
||||
lastReason string
|
||||
crashloopSince time.Time // zero = not in a crash-loop pause
|
||||
parked bool
|
||||
}
|
||||
|
||||
type controllerSupervisor struct {
|
||||
mu sync.Mutex
|
||||
guests map[int]*controllerSupState
|
||||
sweeps int
|
||||
}
|
||||
|
||||
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
|
||||
// (staleLock) or the guest executor is not wired.
|
||||
func (s *Server) WatchControllers(ctx context.Context) {
|
||||
if s.staleLock == nil || s.guestExec == nil {
|
||||
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
|
||||
return
|
||||
}
|
||||
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
|
||||
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
|
||||
"crashloop_window", controllerCrashloopWindow.String(), "guests_dir", s.guestsStateDir())
|
||||
t := time.NewTicker(controllerSupervisorInterval)
|
||||
defer t.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-t.C:
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (s *Server) guestsStateDir() string {
|
||||
if s.guestsDir != "" {
|
||||
return s.guestsDir
|
||||
}
|
||||
return defaultGuestsStateDir
|
||||
}
|
||||
|
||||
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
|
||||
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
|
||||
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
|
||||
// not stat the file within.
|
||||
func (s *Server) provisionedGuest(vmid int) bool {
|
||||
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
|
||||
return err == nil && fi.IsDir()
|
||||
}
|
||||
|
||||
func (s *Server) controllerParked(vmid int) bool {
|
||||
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
|
||||
return err == nil
|
||||
}
|
||||
|
||||
func (s *Server) supState(vmid int) *controllerSupState {
|
||||
if s.ctrlSup.guests == nil {
|
||||
s.ctrlSup.guests = map[int]*controllerSupState{}
|
||||
}
|
||||
st := s.ctrlSup.guests[vmid]
|
||||
if st == nil {
|
||||
st = &controllerSupState{}
|
||||
s.ctrlSup.guests[vmid] = st
|
||||
}
|
||||
return st
|
||||
}
|
||||
|
||||
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
|
||||
// cycle without waiting on the ticker.
|
||||
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
|
||||
if s.staleLock == nil || s.guestExec == nil {
|
||||
return
|
||||
}
|
||||
guests, err := s.staleLock.Guests(ctx)
|
||||
if err != nil {
|
||||
// Ownership unproven ⇒ touch nothing (the guest-power rule).
|
||||
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
|
||||
return
|
||||
}
|
||||
var evaluated, down int
|
||||
for _, g := range guests {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
if !s.provisionedGuest(g.VMID) {
|
||||
continue
|
||||
}
|
||||
evaluated++
|
||||
if !s.superviseOneController(ctx, g.VMID, g.Status) {
|
||||
down++
|
||||
}
|
||||
}
|
||||
s.ctrlSup.mu.Lock()
|
||||
s.ctrlSup.sweeps++
|
||||
sweeps := s.ctrlSup.sweeps
|
||||
s.ctrlSup.mu.Unlock()
|
||||
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
|
||||
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
|
||||
"guests_evaluated", evaluated, "controllers_not_running", down)
|
||||
}
|
||||
}
|
||||
|
||||
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
|
||||
// known=false means the question could not be answered (pct exec failed for a reason other than a
|
||||
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
|
||||
// running": `docker rm` of the controller is the same outage as a kill.
|
||||
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
|
||||
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
|
||||
if err != nil {
|
||||
msg := strings.ToLower(err.Error() + " " + out)
|
||||
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
|
||||
return false, true, "absent"
|
||||
}
|
||||
return false, false, ""
|
||||
}
|
||||
status = strings.TrimSpace(out)
|
||||
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
|
||||
return status == "running" || status == "restarting", true, status
|
||||
}
|
||||
|
||||
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
|
||||
// allows. Returns false when the controller was observed not running.
|
||||
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
|
||||
now := s.clock()
|
||||
if guestStatus != "running" {
|
||||
s.resetNotRunning(vmid)
|
||||
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
|
||||
}
|
||||
|
||||
running, known, status := s.controllerRunning(ctx, vmid)
|
||||
if !known {
|
||||
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
|
||||
s.resetNotRunning(vmid)
|
||||
return true
|
||||
}
|
||||
parked := s.controllerParked(vmid)
|
||||
s.ctrlSup.mu.Lock()
|
||||
st := s.supState(vmid)
|
||||
st.parked = parked
|
||||
if running {
|
||||
st.notRunningSeen = 0
|
||||
s.ctrlSup.mu.Unlock()
|
||||
return true
|
||||
}
|
||||
st.notRunningSeen++
|
||||
seen := st.notRunningSeen
|
||||
s.ctrlSup.mu.Unlock()
|
||||
|
||||
if parked {
|
||||
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
|
||||
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
|
||||
return false
|
||||
}
|
||||
s.swapMu.Lock()
|
||||
swapping := s.swapInFlight[vmid]
|
||||
s.swapMu.Unlock()
|
||||
if swapping {
|
||||
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
|
||||
"vmid", vmid, "status", status)
|
||||
s.resetNotRunning(vmid)
|
||||
return false
|
||||
}
|
||||
if seen < controllerSupervisorConfirm {
|
||||
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
|
||||
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
|
||||
return false
|
||||
}
|
||||
lock, _, err := s.staleLock.Lock(ctx, vmid)
|
||||
if err != nil {
|
||||
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
|
||||
return false
|
||||
}
|
||||
if lock != "" {
|
||||
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
|
||||
return false
|
||||
}
|
||||
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
|
||||
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
|
||||
"vmid", vmid, "backup_running", busy, "err", berr)
|
||||
return false
|
||||
}
|
||||
|
||||
// Backoff.
|
||||
s.ctrlSup.mu.Lock()
|
||||
st = s.supState(vmid)
|
||||
if !st.crashloopSince.IsZero() {
|
||||
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
|
||||
s.ctrlSup.mu.Unlock()
|
||||
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
|
||||
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
|
||||
return false
|
||||
}
|
||||
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
|
||||
// crash-loop (the hub keys on it moving, not on it clearing).
|
||||
st.restarts = nil
|
||||
st.crashloopSince = time.Time{}
|
||||
}
|
||||
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
|
||||
if len(st.restarts) >= controllerCrashloopMax {
|
||||
st.crashloopSince = now
|
||||
n := len(st.restarts)
|
||||
s.ctrlSup.mu.Unlock()
|
||||
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
|
||||
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
|
||||
return false
|
||||
}
|
||||
s.ctrlSup.mu.Unlock()
|
||||
|
||||
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
|
||||
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
|
||||
"vmid", vmid, "status", status, "unit", bootstrapUnit)
|
||||
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
|
||||
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
|
||||
reason += "; restart FAILED: " + err.Error()
|
||||
}
|
||||
s.ctrlSup.mu.Lock()
|
||||
st = s.supState(vmid)
|
||||
st.restarts = append(st.restarts, now)
|
||||
st.restartsTotal++
|
||||
st.lastRestartAt = now
|
||||
st.lastReason = reason
|
||||
st.notRunningSeen = 0
|
||||
s.ctrlSup.mu.Unlock()
|
||||
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason)
|
||||
return false
|
||||
}
|
||||
|
||||
func (s *Server) resetNotRunning(vmid int) {
|
||||
s.ctrlSup.mu.Lock()
|
||||
defer s.ctrlSup.mu.Unlock()
|
||||
if st := s.ctrlSup.guests[vmid]; st != nil {
|
||||
st.notRunningSeen = 0
|
||||
}
|
||||
}
|
||||
|
||||
func (s *Server) clock() time.Time {
|
||||
if s.now != nil {
|
||||
return s.now()
|
||||
}
|
||||
return time.Now().UTC()
|
||||
}
|
||||
|
||||
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
|
||||
out := ts[:0]
|
||||
for _, t := range ts {
|
||||
if !t.Before(cutoff) {
|
||||
out = append(out, t)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
|
||||
// Nil when the supervisor is not wired, so the stanza is omitted.
|
||||
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
|
||||
if s.staleLock == nil || s.guestExec == nil {
|
||||
return nil
|
||||
}
|
||||
s.ctrlSup.mu.Lock()
|
||||
defer s.ctrlSup.mu.Unlock()
|
||||
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
|
||||
for vmid, st := range s.ctrlSup.guests {
|
||||
g := hub.ControllerSupervisorGuest{
|
||||
VMID: vmid,
|
||||
RestartsTotal: st.restartsTotal,
|
||||
LastReason: st.lastReason,
|
||||
Parked: st.parked,
|
||||
Crashloop: !st.crashloopSince.IsZero(),
|
||||
}
|
||||
if !st.lastRestartAt.IsZero() {
|
||||
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
|
||||
}
|
||||
if !st.crashloopSince.IsZero() {
|
||||
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
|
||||
}
|
||||
out.Guests = append(out.Guests, g)
|
||||
}
|
||||
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,251 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"io"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
|
||||
// again", and each guard is pinned by the case where acting would be wrong.
|
||||
|
||||
type supExec struct {
|
||||
mu sync.Mutex
|
||||
status map[int]string // docker .State.Status per vmid; "" = container absent
|
||||
inspectErr error // non-nil = pct exec itself failed (unknown)
|
||||
restarts map[int]int
|
||||
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
|
||||
// stays "exited").
|
||||
onRestart string
|
||||
}
|
||||
|
||||
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
switch {
|
||||
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
|
||||
if f.inspectErr != nil {
|
||||
return "", f.inspectErr
|
||||
}
|
||||
st, ok := f.status[vmid]
|
||||
if !ok || st == "" {
|
||||
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
|
||||
}
|
||||
return st + "\n", nil
|
||||
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
|
||||
if f.restarts == nil {
|
||||
f.restarts = map[int]int{}
|
||||
}
|
||||
f.restarts[vmid]++
|
||||
if f.onRestart != "" {
|
||||
f.status[vmid] = f.onRestart
|
||||
}
|
||||
return "", nil
|
||||
}
|
||||
return "", errors.New("supExec: unexpected args")
|
||||
}
|
||||
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
|
||||
return "", errors.New("supExec: no stdin exec expected")
|
||||
}
|
||||
func (f *supExec) count(vmid int) int {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
return f.restarts[vmid]
|
||||
}
|
||||
|
||||
type supClock struct{ t time.Time }
|
||||
|
||||
func (c *supClock) now() time.Time { return c.t }
|
||||
|
||||
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
for _, v := range provisioned {
|
||||
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
|
||||
s := &Server{
|
||||
staleLock: ctl,
|
||||
guestExec: ex,
|
||||
guestsDir: dir,
|
||||
swapInFlight: map[int]bool{},
|
||||
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
|
||||
now: clk.now,
|
||||
}
|
||||
return s, clk, dir
|
||||
}
|
||||
|
||||
func runningGuest(vmid int) *fakeGuestPowerCtl {
|
||||
return &fakeGuestPowerCtl{
|
||||
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
|
||||
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
|
||||
}
|
||||
}
|
||||
|
||||
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
|
||||
// the first (the bootstrap's own rm-f/run window must never be raced).
|
||||
//
|
||||
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
|
||||
// stays 0 → "the killed controller was NOT restarted — this is R-523".
|
||||
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 0 {
|
||||
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
|
||||
}
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 1 {
|
||||
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
|
||||
}
|
||||
st := s.ControllerSupervisorStatus(context.Background())
|
||||
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
|
||||
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
|
||||
}
|
||||
// Healthy again → no further restart.
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 1 {
|
||||
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
|
||||
}
|
||||
}
|
||||
|
||||
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{}, onRestart: "running"}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 1 {
|
||||
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
|
||||
}
|
||||
}
|
||||
|
||||
func TestControllerSupervisor_Guards(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
|
||||
}{
|
||||
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
|
||||
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
}},
|
||||
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
|
||||
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
|
||||
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.backupRun = map[int]bool{9201: true}
|
||||
}},
|
||||
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.backupErr = errors.New("tasks unreadable")
|
||||
}},
|
||||
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.guests[0].Status = "stopped"
|
||||
}},
|
||||
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
|
||||
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
|
||||
}},
|
||||
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.guestsErr = errors.New("api down")
|
||||
}},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||
ctl := runningGuest(9201)
|
||||
s, _, dir := supServer(t, ex, ctl, 9201)
|
||||
tc.setup(s, ex, ctl, dir)
|
||||
for i := 0; i < 4; i++ {
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
}
|
||||
if n := ex.count(9201); n != 0 {
|
||||
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
|
||||
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9202: "exited"}}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
|
||||
for i := 0; i < 3; i++ {
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
}
|
||||
if n := ex.count(9202); n != 0 {
|
||||
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
|
||||
}
|
||||
}
|
||||
|
||||
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
|
||||
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
|
||||
// again.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
|
||||
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
|
||||
// failed at "crash-looping controller restarted 10 times".
|
||||
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
|
||||
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
ctx := context.Background()
|
||||
|
||||
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
clk.t = clk.t.Add(controllerSupervisorInterval)
|
||||
}
|
||||
if n := ex.count(9201); n != controllerCrashloopMax {
|
||||
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
|
||||
}
|
||||
st := s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||
if !st.Crashloop || st.CrashloopSince == "" {
|
||||
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
|
||||
}
|
||||
// Still paused 25 minutes later.
|
||||
clk.t = clk.t.Add(15 * time.Minute)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
if n := ex.count(9201); n != controllerCrashloopMax {
|
||||
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
|
||||
}
|
||||
// After the pause: tries again.
|
||||
clk.t = clk.t.Add(controllerCrashloopPause)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
if n := ex.count(9201); n != controllerCrashloopMax+1 {
|
||||
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
|
||||
}
|
||||
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
|
||||
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
|
||||
}
|
||||
}
|
||||
|
||||
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
|
||||
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
|
||||
var m map[string][]map[string]any
|
||||
if err := json.Unmarshal(b, &m); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
g := m["guests"][0]
|
||||
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked"} {
|
||||
if _, ok := g[k]; !ok {
|
||||
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -185,6 +185,9 @@ type Options struct {
|
||||
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
|
||||
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
|
||||
StaleLock StaleLockController
|
||||
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
|
||||
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
|
||||
GuestsStateDir string
|
||||
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
|
||||
ControllerSwapStateDir string
|
||||
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
|
||||
@@ -362,6 +365,12 @@ type Server struct {
|
||||
swapMu sync.Mutex
|
||||
swapInFlight map[int]bool
|
||||
|
||||
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
|
||||
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
|
||||
guestExec GuestExecutor
|
||||
guestsDir string
|
||||
ctrlSup controllerSupervisor
|
||||
|
||||
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
|
||||
// detached pipeline runs through (tests inject; production defaults set in NewServer).
|
||||
netVerifyMu sync.Mutex
|
||||
@@ -484,7 +493,9 @@ func NewServer(o Options) (*Server, error) {
|
||||
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
|
||||
if o.ControllerSwap != nil {
|
||||
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
|
||||
s.guestExec = o.ControllerSwap
|
||||
}
|
||||
s.guestsDir = o.GuestsStateDir
|
||||
return s, nil
|
||||
}
|
||||
|
||||
@@ -1084,6 +1095,10 @@ type BackupTierInfo struct {
|
||||
Target string `json:"target"`
|
||||
CadenceSeconds int64 `json:"cadence_seconds"`
|
||||
Primary bool `json:"primary"`
|
||||
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
|
||||
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
|
||||
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
|
||||
Storage string `json:"storage,omitempty"`
|
||||
}
|
||||
|
||||
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||
@@ -1093,11 +1108,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
|
||||
Target: t.TargetID,
|
||||
CadenceSeconds: int64(t.Cadence.Seconds()),
|
||||
Primary: t.Primary,
|
||||
Storage: s.storagePresence(r.Context(), t.TargetID),
|
||||
})
|
||||
}
|
||||
writeOK(w, resp)
|
||||
}
|
||||
|
||||
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
|
||||
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
|
||||
// successful read that does not list the storage is "absent".
|
||||
func (s *Server) storagePresence(ctx context.Context, target string) string {
|
||||
if s.storage == nil || target == "" {
|
||||
return StoragePresenceUnknown
|
||||
}
|
||||
targets, err := s.storage.Observe(ctx)
|
||||
if err != nil {
|
||||
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
|
||||
return StoragePresenceUnknown
|
||||
}
|
||||
for _, t := range targets {
|
||||
if t.Name == target {
|
||||
return StoragePresencePresent
|
||||
}
|
||||
}
|
||||
return StoragePresenceAbsent
|
||||
}
|
||||
|
||||
const (
|
||||
StoragePresencePresent = "present"
|
||||
StoragePresenceAbsent = "absent"
|
||||
StoragePresenceUnknown = "unknown"
|
||||
)
|
||||
|
||||
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
|
||||
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
|
||||
// result ("presence is not success").
|
||||
type TierBackupState struct {
|
||||
Target string `json:"target"`
|
||||
Primary bool `json:"primary"`
|
||||
Storage string `json:"storage"` // present | absent | unknown
|
||||
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
|
||||
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
|
||||
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
|
||||
// LastSuccessSource is "storage".
|
||||
LastSuccess *hub.Backup `json:"last_success,omitempty"`
|
||||
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
|
||||
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
|
||||
}
|
||||
|
||||
// TierAttempt is the newest recorded attempt on a tier, successful or not.
|
||||
type TierAttempt struct {
|
||||
StartedAt string `json:"started_at"`
|
||||
Success bool `json:"success"`
|
||||
Error string `json:"error,omitempty"`
|
||||
}
|
||||
|
||||
// tierBackupStates builds the per-tier view for one guest.
|
||||
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
|
||||
out := make([]TierBackupState, 0, len(s.tiers))
|
||||
for _, t := range s.tiers {
|
||||
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
|
||||
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
|
||||
st.LastSuccess, st.LastSuccessSource = b, "record"
|
||||
} else if st.Storage != StoragePresenceAbsent {
|
||||
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
|
||||
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
|
||||
StartedAt: when.UTC().Format(time.RFC3339)}
|
||||
st.LastSuccessSource = "storage"
|
||||
}
|
||||
}
|
||||
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
|
||||
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
|
||||
}
|
||||
out = append(out, st)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// tierFromRequest resolves the `?target=` query parameter to a tier.
|
||||
//
|
||||
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
|
||||
@@ -1144,6 +1231,10 @@ type BackupStatusResponse struct {
|
||||
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
|
||||
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
|
||||
Target string `json:"target,omitempty"`
|
||||
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
|
||||
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
|
||||
// response exactly as before.
|
||||
Tiers []TierBackupState `json:"tiers,omitempty"`
|
||||
}
|
||||
|
||||
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||
@@ -1155,6 +1246,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
|
||||
// across ANY target (echo == "" → pickLatestBackup's match-any path).
|
||||
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
|
||||
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
|
||||
if echo == "" {
|
||||
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
|
||||
}
|
||||
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
|
||||
resp.Phase = job.Phase
|
||||
resp.JobID = job.JobID
|
||||
|
||||
Reference in New Issue
Block a user