Compare commits

..

9 Commits

Author SHA1 Message Date
admin 18d03bd437 v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s
Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.

Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.

Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:26:42 +02:00
admin e98b857684 REPORT: v0.131.0 supervisor + per-tier status, delivery and live validation
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 11:37:07 +02:00
admin dcdeb3d16d CHANGELOG: v0.131.0 released (tag + package verified by download)
gates / gates (push) Successful in 12s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:50:47 +02:00
admin 610804b98d v0.131.0: controller supervisor (R-523); per-tier backup status + tier storage presence (R-517/R-518)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:49:44 +02:00
admin 4586f0f7f6 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 9s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:52 +02:00
admin 205e22babe decoy sweep: no gate changed here, and that is the result (R-421)
gates / gates (push) Failing after 12s
All 29 gate scripts across the four repos were read and DECOYED - the label constructed without the
fact, the gate run, the verdict recorded. 16 were fooled. None of them were in this repo.

A decoy that nobody would write proves nothing, so the attempts that turned out illegitimate were
WITHDRAWN rather than counted. Both of this repo were withdrawn, and both are named in the audit.

The gates here that could not be given a plausible decoy are listed BY NAME in
felhom.eu/scripts/decoy_coverage_gate.py EXEMPT (R-426) as UNTESTED - not as sound. A gate nobody
tried to fool is UNKNOWN, and calling it sound would be the same confident guess this sweep exists
to find.

Survey table: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:38:58 +02:00
admin 058b945064 gate 11: register the shared observations gate
gates / gates (push) Successful in 10s
felhom.eu/scripts/observations_gate.py, invoked across the workspace like
reuse_refs_check.py and instructions_gate.py. This repo's REPORT.md has no
observations section today, so the gate passes quietly - it is registered for
the session that writes one.
2026-08-23 13:53:20 +02:00
admin 40d857b527 CHANGELOG: the fleet runs the published bytes, not the proof build (R-349)
gates / gates (push) Successful in 7s
Both boxes were first given a hand build: same source, same version
string, different bytes (256e0829 vs the published a56a92a7), because
release-agent.sh builds with -trimpath -buildvcs=false and a hand build
does not.

Nothing would have corrected it. The boxes already reported 0.130.0, so
self-update saw the vouched version as installed and would have done
nothing, forever. Every version check in the system compares the STRING.

Both reinstalled from the downloaded package; both now report a56a92a7.
2026-08-20 12:53:37 +02:00
admin 7ae6990bac v0.130.0 released: tag + package published, heading now claims it
gates / gates (push) Successful in 8s
sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3
tag v0.130.0 at 7569f34, 14,141,158 bytes.

Reproducible: rebuilding with -trimpath -buildvcs=false matches the
published artifact byte for byte (R-186's property, checked not assumed).

The heading said UNRELEASED while the fix was hand-installed on demo-hp
only -- publishing then would have pushed it onto the control box through
self-update. release-complete convicted on the release heading and was
right to; the answer was to stop claiming a release, not to bypass it.

NOT VOUCHED by this commit. Vouching is the separate operator act.
2026-08-20 12:47:25 +02:00
13 changed files with 1250 additions and 51 deletions
+107 -7
View File
@@ -1,11 +1,109 @@
## UNRELEASED — v0.130.0 candidate: the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
## Unreleased (to become v0.132.0) — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
> **Deliberately not a release heading yet, and the `release-complete` gate is doing its job by
> requiring that.** This build is **hand-installed on `demo-hp` only** so the fix can be proved
> against `demo-felhom` as an untouched control. Publishing it — tag + Gitea package — would put it
> on the control box through the self-update path and destroy the experiment. **When the operator
> authorises the publish, this heading becomes `## v0.130.0` in the same commit as the tag and the
> package**, and the gate then checks it for real. See R-347.
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
`controller_slow_crashloop`; an older hub ignores them.
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
does **not** stop restarting — the fast brake remains the only brake.
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
after an agent restart — `last_success_source: storage`), the last attempt kept apart
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
last success".
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
absent). **All four found by accident.** The gates enforce everything else here and were the one part
nothing had checked.
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
proves the cause was the listing rather than the decoy.
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
**In this repo:** no gate changed, and that is the result. `release-complete` was decoyed and is
SOUND — the sweep's attempt (a non-version heading on top of `CHANGELOG.md`) was WITHDRAWN as
illegitimate, because `HEAD_RE.search` scans the whole file and still names `v0.130.0`. The three
shared gates are covered from `felhom.eu`; the remaining two are named in the decoy-coverage
exemption list (R-426) as UNTESTED, not as sound.
## v0.130.0 — the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
> **RELEASED 2026-08-20**, on the operator's word, after the fix was proved on both boxes.
> `sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, 14,141,158 bytes,
> tag `v0.130.0` at `7569f34`. Reproducible: a rebuild with `-trimpath -buildvcs=false` matches the
> published artifact byte for byte (R-186's property, checked rather than assumed).
>
> The heading read `## UNRELEASED — v0.130.0 candidate` until this point, deliberately: while the fix
> was hand-installed on `demo-hp` only, publishing would have pushed it onto `demo-felhom` through
> self-update and destroyed the control the proof rested on. **`release-complete` convicted on the
> release heading and was right to** — the answer was to stop claiming a release, not to bypass the
> gate. See R-347.
>
> **The fleet runs these exact bytes.** Both demo boxes were first given a hand build made during the
> proof — same source, same version string, **different bytes** (`256e0829…`), because
> `release-agent.sh` builds with `-trimpath -buildvcs=false` and a hand build does not. Nothing would
> have corrected that: the boxes already reported `0.130.0`, so self-update saw the vouched version as
> installed and would have done nothing, forever. Both were reinstalled from the **downloaded package**
> and now report `a56a92a7…`. Filed as **R-349**, because every prove-then-publish train hits it.
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
@@ -5465,3 +5563,5 @@ client, signing, or storage/backup orchestration yet (later slices).
read-only `--selftest` against the demo host with TLS fingerprint pinning.
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
token) is provisioned out-of-band; the agent only consumes the token.
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->
+8
View File
@@ -101,3 +101,11 @@ the mechanism are exempt.
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
assumption, not an observation.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+39 -42
View File
@@ -1,50 +1,47 @@
# REPORT — agent v0.129.0: a correct code for an earlier package (R-311, 2026-08-12)
# REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15)
## What changed and why
Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md`
§4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6.
Yesterday's drill proved a retained escrow package **works** — unsealed with the old recovery code, it
opened a set-aside store and restored planted files byte-identical — while this agent answered that
same correct code with *"the recovery code did not open the sealed bundle"*. Nothing had ever tried
the retained packages, so a correct-but-earlier code and a mistype were genuinely indistinguishable.
## Measured first (A.1)
Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and
`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix.
- `internal/hub/client.go` — `FetchRetainedIdentityEscrow` → `GET /api/v1/hosts/<id>/escrow/retained`
(hub ≥ v0.103.0). **A 404 is a clean "none"**, not a fault: an older hub must not turn into a failed
recovery.
- `internal/escrow/recover.go` — optional `FetchRetained`, `ErrCodeOpensRetained` +
`RetainedOpenedError{SupersededAt, KeyFingerprint, Index, HasResticPassword}`. Consulted **only**
after the current package refuses.
- `internal/localapi/escrow_recover.go` — a **fifth** case on the R-224 switch: **422**, with
`opens_retained`, `superseded_at`, `retained_has_restic_pw`. Added to the switch, not a restructure.
- `cmd/felhom-agent/main.go` — the retained fetcher wired on the same self-scoped hub client.
## What shipped
- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests
that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap
in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest,
3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the
events. Same GuestExecutor and sudoers grants as the swap — no new privilege.
- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage
after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`.
- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`.
## Fail-safe, in every direction
## Red-proofs (each seen failing, then restored)
- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)`
- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3`
- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success`
nil fetcher · hub without the route (404) · transport failure · malformed package → **the original
refusal stands, unchanged**. The worst outcome of this feature breaking is the behaviour before it.
Attempts bounded (`MaxRetainedTried`, default 6) — each unwrap is ~1 s of scrypt, so an unbounded loop
would turn one wrong code into a minutes-long hang.
## Release and delivery
`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`,
verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the
box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for
`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and
Peti's box stay on 0.130.0.
## Tests — 7, with REAL age crypto
## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0)
| moment | result |
|---|---|
| idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill |
| parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** |
| kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z |
| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` |
| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z |
Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200"
line during the pause that the container state contradicts.
Real crypto because the two situations are indistinguishable **at the unwrap**; a faked unwrap would
prove nothing about what was broken. Full suite green (`go build`/`vet`/`test ./...`), agent gates OK.
## Teardown
Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker
removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release).
**Red-proof, mutation asserted applied before the run:** remove the `tryRetained` block from
`RecoverOffsiteRepoPassword` →
`err = escrow: the recovery code did not unwrap the identity escrow (wrong recovery code…)` →
`TestRecover_CodeOpensRetainedPackage_IsNotAWrongCode` FAILS. **The lie returns, in those words.**
That is the layer the lie actually lives in: removing the *controller's* case yields the neutral
message instead, because R-224's safe default catches it.
## Released and deployed
`release-agent.sh 0.129.0` — tagged `v0.129.0`, published, **verified by independent download**,
sha256 `53a54f0620afbd6d…`. Installed on `felhom-pve`, `felhom-agent --version` = 0.129.0, unit active,
journal clean (normal PBS verify cycle). **NOT VOUCHED** — that stays the operator's act.
## Bypass, stated as required
`git push --no-verify` was used **once** for the code push. The `release-complete` gate refuses a
CHANGELOG entry whose tag and package do not exist, and `release-agent.sh` refuses a tree that is not
pushed — circular by construction. The bypass was immediately followed by the real release; gates were
re-run afterwards and are **green**, and the tag+package now exist.
Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`.
+1
View File
@@ -74,6 +74,7 @@
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
### Proxmox client / hub / PBS / provisioning
+6 -1
View File
@@ -59,7 +59,7 @@ import (
// version is the agent version. Overridable at build time with
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
var version = "0.130.0"
var version = "0.131.0"
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
@@ -1401,6 +1401,10 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
// running" signal, so a deliberately stopped guest is never touched.
go localSrv.WatchGuestPower(ctx)
// R-523: a controller container that is simply not running (killed, stopped, a failed
// self-update) is restarted through its bootstrap unit — nothing else watches it.
collector.SetControllerSupervisorReporter(localSrv)
go localSrv.WatchControllers(ctx)
go func() { errc <- localSrv.Run(ctx) }()
}
if lanLoop != nil {
@@ -1826,6 +1830,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
SmbCredsDir: cfg.Privileged.SmbCredsDir,
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
+5 -1
View File
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
# view, and the docker socket. The controller reaches the agent's local API for disk management.
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
-v felhom-controller-data:/opt/docker/felhom-controller \
+16
View File
@@ -103,6 +103,7 @@ type Collector struct {
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
return c
}
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
type ControllerSupervisorReporter interface {
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
}
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
c.ctrlSup = r
return c
}
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
// omitted). Returns the collector for chaining.
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.pbsdr != nil {
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
}
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
if c.ctrlSup != nil {
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
}
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
+33
View File
@@ -124,6 +124,39 @@ type HostReport struct {
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
OOB *OOBStatus `json:"oob,omitempty"`
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
// how many times the agent restarted a dead controller, when last and why, whether it gave up
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
}
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
type ControllerSupervisorStatus struct {
Guests []ControllerSupervisorGuest `json:"guests"`
}
// ControllerSupervisorGuest is one supervised guest.
type ControllerSupervisorGuest struct {
VMID int `json:"vmid"`
RestartsTotal int `json:"restarts_total"`
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
LastReason string `json:"last_reason,omitempty"`
Crashloop bool `json:"crashloop"`
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
Parked bool `json:"parked"`
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
Restarts24h int `json:"restarts_24h"`
SlowCrashloop bool `json:"slow_crashloop"`
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
}
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
@@ -0,0 +1,133 @@
package localapi
import (
"context"
"encoding/json"
"io"
"log/slog"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
t.Helper()
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
return resp.Data.Tiers
}
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
Storage: fakeStorage{targets: targets},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
st := &fakeStore{backups: []hub.Backup{
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
}}
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
tiers := tierStatesOf(t, srv)
if len(tiers) != 2 {
t.Fatalf("want 2 tiers, got %+v", tiers)
}
local, pbs := tiers[0], tiers[1]
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
t.Fatalf("local tier lost its successful backup: %+v", local)
}
if pbs.LastSuccess != nil {
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
}
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
}
if pbs.Storage != StoragePresenceAbsent {
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
}
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &resp)
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
}
}
// /backup/tiers advertises storage presence, tri-state.
func TestBackupTiers_StoragePresence(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
var resp struct {
Data BackupTiersResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatal(err)
}
got := map[string]string{}
for _, ti := range resp.Data.Tiers {
got[ti.Target] = ti.Storage
}
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
t.Fatalf("storage presence wrong: %v", got)
}
// An unreadable storage view is "unknown", never "absent".
srv.storage = tierErrStorage{}
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
t.Fatalf("unreadable storage view must be unknown, got %q", p)
}
}
type tierErrStorage struct{}
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
return nil, context.DeadlineExceeded
}
// A targeted request keeps the pre-R-517 bytes (no tiers array).
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
}
}
func containsKey(b []byte, key string) bool {
var m struct {
Data map[string]json.RawMessage `json:"data"`
}
_ = json.Unmarshal(b, &m)
_, ok := m.Data[key]
return ok
}
+453
View File
@@ -0,0 +1,453 @@
package localapi
import (
"context"
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-523 — the in-guest controller supervisor.
//
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
// dashboard answered 502 for 33 minutes until the box was power-cycled.
//
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
// No new privilege.
//
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
// again after 30 minutes.
//
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
// record) can never read as a new restart.
const (
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
// so a killed controller is restarted 30–60 s after it died.
controllerSupervisorInterval = 30 * time.Second
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
controllerSupervisorConfirm = 2
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
// controllerCrashloopPause.
controllerCrashloopMax = 3
controllerCrashloopWindow = 15 * time.Minute
controllerCrashloopPause = 30 * time.Minute
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
// watchdog is indistinguishable from a dead one (standing rule 3).
controllerSupervisorHeartbeatEvery = 20
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
// the controller container, and a customer inside the guest cannot park the supervisor.
ControllerParkedMarker = "controller-parked"
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
// R-539 (operator ruling 3 of 2026-09-16) — the SLOW crash loop. The 3-in-15-minutes brake above
// cannot see a controller that dies every 20 minutes: no two restarts share its window, so it is
// restarted for ever and the only trace is an info event that mails nobody (measured 2026-09-16,
// R-531). A second counter over 24 hours raises a WARNING at the fifth restart. It does NOT stop
// restarting — the fast brake stays the only brake, unchanged. Every restart the supervisor
// performs counts, including one that follows a deliberate operator `docker kill` (measured
// 2026-09-15: the supervisor cannot tell a kill from a crash, and a controller that is killed five
// times a day is worth a line to the operator either way).
controllerSlowCrashloopWindow = 24 * time.Hour
controllerSlowCrashloopMax = 5
// controllerSlowCounterFile holds the 24-hour restart times and the last raise, per guest, beside
// the parked marker.
controllerSlowCounterFile = "controller-restarts-24h.json"
)
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
type controllerSupState struct {
notRunningSeen int
restarts []time.Time // restart times inside the crash-loop window (pruned)
restartsTotal int
lastRestartAt time.Time
lastReason string
crashloopSince time.Time // zero = not in a crash-loop pause
parked bool
// R-539 — the slow counter. PERSISTED, unlike everything above, and the precedent's reason does not
// apply to it: persisting the fast record could carry a stale "give up" across the restart that
// fixed it, but this record never gives anything up — it only warns. Losing it on an agent restart,
// on the other hand, would hide exactly the box it exists for (one whose agent restarts too).
restarts24h []time.Time
slowCrashloopSince time.Time // the last raise; kept after it ages out, the hub keys on it MOVING
}
type controllerSupervisor struct {
mu sync.Mutex
guests map[int]*controllerSupState
sweeps int
}
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
// (staleLock) or the guest executor is not wired.
func (s *Server) WatchControllers(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
return
}
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
"crashloop_window", controllerCrashloopWindow.String(),
"slow_crashloop_max", controllerSlowCrashloopMax, "slow_crashloop_window", controllerSlowCrashloopWindow.String(),
"guests_dir", s.guestsStateDir())
t := time.NewTicker(controllerSupervisorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s.ControllerSupervisorTick(ctx)
}
}
}
func (s *Server) guestsStateDir() string {
if s.guestsDir != "" {
return s.guestsDir
}
return defaultGuestsStateDir
}
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
// not stat the file within.
func (s *Server) provisionedGuest(vmid int) bool {
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
return err == nil && fi.IsDir()
}
func (s *Server) controllerParked(vmid int) bool {
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return err == nil
}
func (s *Server) supState(vmid int) *controllerSupState {
if s.ctrlSup.guests == nil {
s.ctrlSup.guests = map[int]*controllerSupState{}
}
st := s.ctrlSup.guests[vmid]
if st == nil {
st = &controllerSupState{}
s.loadSlowCounter(vmid, st)
s.ctrlSup.guests[vmid] = st
}
return st
}
// slowCounterRecord is the on-disk shape of the R-539 counter.
type slowCounterRecord struct {
Restarts []time.Time `json:"restarts"`
SlowCrashloopSince time.Time `json:"slow_crashloop_since,omitempty"`
}
func (s *Server) slowCounterPath(vmid int) string {
return filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), controllerSlowCounterFile)
}
// loadSlowCounter restores the persisted counter into a fresh state. Absent = a clean start; unreadable
// or corrupt = a clean start with a WARN (a warning counter must never block supervision).
func (s *Server) loadSlowCounter(vmid int, st *controllerSupState) {
b, err := os.ReadFile(s.slowCounterPath(vmid))
if err != nil {
if !os.IsNotExist(err) {
s.logger.Warn("controller-supervisor: slow counter unreadable — starting it from zero", "vmid", vmid, "err", err)
}
return
}
var rec slowCounterRecord
if err := json.Unmarshal(b, &rec); err != nil {
s.logger.Warn("controller-supervisor: slow counter corrupt — starting it from zero", "vmid", vmid, "err", err)
return
}
st.restarts24h = pruneBefore(rec.Restarts, s.clock().Add(-controllerSlowCrashloopWindow))
st.slowCrashloopSince = rec.SlowCrashloopSince
if len(st.restarts24h) > 0 || !st.slowCrashloopSince.IsZero() {
s.logger.Info("controller-supervisor: slow counter restored from disk", "vmid", vmid,
"restarts_24h", len(st.restarts24h), "slow_crashloop_since", st.slowCrashloopSince.Format(time.RFC3339))
}
}
// saveSlowCounter writes the counter atomically (tmp + rename, 0600). A failure is logged and the
// in-memory counter carries on — the next restart retries the write.
func (s *Server) saveSlowCounter(vmid int, rec slowCounterRecord) {
path := s.slowCounterPath(vmid)
b, err := json.Marshal(rec)
if err == nil {
tmp := path + ".tmp"
if err = os.WriteFile(tmp, b, 0o600); err == nil {
err = os.Rename(tmp, path)
}
}
if err != nil {
s.logger.Warn("controller-supervisor: could not persist the slow counter (kept in memory)", "vmid", vmid, "path", path, "err", err)
}
}
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
// cycle without waiting on the ticker.
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
return
}
guests, err := s.staleLock.Guests(ctx)
if err != nil {
// Ownership unproven ⇒ touch nothing (the guest-power rule).
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
return
}
var evaluated, down int
for _, g := range guests {
if ctx.Err() != nil {
return
}
if !s.provisionedGuest(g.VMID) {
continue
}
evaluated++
if !s.superviseOneController(ctx, g.VMID, g.Status) {
down++
}
}
s.ctrlSup.mu.Lock()
s.ctrlSup.sweeps++
sweeps := s.ctrlSup.sweeps
s.ctrlSup.mu.Unlock()
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
"guests_evaluated", evaluated, "controllers_not_running", down)
}
}
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
// known=false means the question could not be answered (pct exec failed for a reason other than a
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
// running": `docker rm` of the controller is the same outage as a kill.
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
if err != nil {
msg := strings.ToLower(err.Error() + " " + out)
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
return false, true, "absent"
}
return false, false, ""
}
status = strings.TrimSpace(out)
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
return status == "running" || status == "restarting", true, status
}
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
// allows. Returns false when the controller was observed not running.
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
now := s.clock()
if guestStatus != "running" {
s.resetNotRunning(vmid)
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
}
running, known, status := s.controllerRunning(ctx, vmid)
if !known {
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
s.resetNotRunning(vmid)
return true
}
parked := s.controllerParked(vmid)
s.ctrlSup.mu.Lock()
st := s.supState(vmid)
st.parked = parked
if running {
st.notRunningSeen = 0
s.ctrlSup.mu.Unlock()
return true
}
st.notRunningSeen++
seen := st.notRunningSeen
s.ctrlSup.mu.Unlock()
if parked {
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return false
}
s.swapMu.Lock()
swapping := s.swapInFlight[vmid]
s.swapMu.Unlock()
if swapping {
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
"vmid", vmid, "status", status)
s.resetNotRunning(vmid)
return false
}
if seen < controllerSupervisorConfirm {
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
return false
}
lock, _, err := s.staleLock.Lock(ctx, vmid)
if err != nil {
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
return false
}
if lock != "" {
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
return false
}
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
"vmid", vmid, "backup_running", busy, "err", berr)
return false
}
// Backoff.
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
if !st.crashloopSince.IsZero() {
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
s.ctrlSup.mu.Unlock()
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
return false
}
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
// crash-loop (the hub keys on it moving, not on it clearing).
st.restarts = nil
st.crashloopSince = time.Time{}
}
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
if len(st.restarts) >= controllerCrashloopMax {
st.crashloopSince = now
n := len(st.restarts)
s.ctrlSup.mu.Unlock()
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
return false
}
s.ctrlSup.mu.Unlock()
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
"vmid", vmid, "status", status, "unit", bootstrapUnit)
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
reason += "; restart FAILED: " + err.Error()
}
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
st.restarts = append(st.restarts, now)
st.restartsTotal++
st.lastRestartAt = now
st.lastReason = reason
st.notRunningSeen = 0
// R-539: the slow counter. Raise at most once per 24 hours — the hub mails on the raise MOVING.
st.restarts24h = append(pruneBefore(st.restarts24h, now.Add(-controllerSlowCrashloopWindow)), now)
n24 := len(st.restarts24h)
raised := false
if n24 >= controllerSlowCrashloopMax && (st.slowCrashloopSince.IsZero() || now.Sub(st.slowCrashloopSince) >= controllerSlowCrashloopWindow) {
st.slowCrashloopSince = now
raised = true
}
rec := slowCounterRecord{Restarts: append([]time.Time(nil), st.restarts24h...), SlowCrashloopSince: st.slowCrashloopSince}
s.ctrlSup.mu.Unlock()
s.saveSlowCounter(vmid, rec)
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason, "restarts_24h", n24)
if raised {
s.logger.Warn("controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop",
"vmid", vmid, "restarts_24h", n24, "window", controllerSlowCrashloopWindow.String(), "threshold", controllerSlowCrashloopMax)
}
return false
}
func (s *Server) resetNotRunning(vmid int) {
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
if st := s.ctrlSup.guests[vmid]; st != nil {
st.notRunningSeen = 0
}
}
func (s *Server) clock() time.Time {
if s.now != nil {
return s.now()
}
return time.Now().UTC()
}
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
out := ts[:0]
for _, t := range ts {
if !t.Before(cutoff) {
out = append(out, t)
}
}
return out
}
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
// Nil when the supervisor is not wired, so the stanza is omitted.
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
if s.staleLock == nil || s.guestExec == nil {
return nil
}
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
for vmid, st := range s.ctrlSup.guests {
g := hub.ControllerSupervisorGuest{
VMID: vmid,
RestartsTotal: st.restartsTotal,
LastReason: st.lastReason,
Parked: st.parked,
Crashloop: !st.crashloopSince.IsZero(),
}
if !st.lastRestartAt.IsZero() {
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
}
if !st.crashloopSince.IsZero() {
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
}
now := s.clock()
g.Restarts24h = len(pruneBefore(append([]time.Time(nil), st.restarts24h...), now.Add(-controllerSlowCrashloopWindow)))
if !st.slowCrashloopSince.IsZero() {
g.SlowCrashloopSince = st.slowCrashloopSince.UTC().Format(time.RFC3339)
g.SlowCrashloop = now.Sub(st.slowCrashloopSince) < controllerSlowCrashloopWindow
}
out.Guests = append(out.Guests, g)
}
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
return out
}
@@ -0,0 +1,350 @@
package localapi
import (
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
"strconv"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
// again", and each guard is pinned by the case where acting would be wrong.
type supExec struct {
mu sync.Mutex
status map[int]string // docker .State.Status per vmid; "" = container absent
inspectErr error // non-nil = pct exec itself failed (unknown)
restarts map[int]int
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
// stays "exited").
onRestart string
}
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
f.mu.Lock()
defer f.mu.Unlock()
switch {
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
if f.inspectErr != nil {
return "", f.inspectErr
}
st, ok := f.status[vmid]
if !ok || st == "" {
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
}
return st + "\n", nil
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
if f.restarts == nil {
f.restarts = map[int]int{}
}
f.restarts[vmid]++
if f.onRestart != "" {
f.status[vmid] = f.onRestart
}
return "", nil
}
return "", errors.New("supExec: unexpected args")
}
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
return "", errors.New("supExec: no stdin exec expected")
}
func (f *supExec) count(vmid int) int {
f.mu.Lock()
defer f.mu.Unlock()
return f.restarts[vmid]
}
type supClock struct{ t time.Time }
func (c *supClock) now() time.Time { return c.t }
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
t.Helper()
dir := t.TempDir()
for _, v := range provisioned {
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
t.Fatal(err)
}
}
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
s := &Server{
staleLock: ctl,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
now: clk.now,
}
return s, clk, dir
}
func runningGuest(vmid int) *fakeGuestPowerCtl {
return &fakeGuestPowerCtl{
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
}
}
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
// the first (the bootstrap's own rm-f/run window must never be raced).
//
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
// stays 0 → "the killed controller was NOT restarted — this is R-523".
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 0 {
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
}
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
}
st := s.ControllerSupervisorStatus(context.Background())
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
}
// Healthy again → no further restart.
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
}
}
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
}
}
func TestControllerSupervisor_Guards(t *testing.T) {
cases := []struct {
name string
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
}{
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
panic(err)
}
}},
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupRun = map[int]bool{9201: true}
}},
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupErr = errors.New("tasks unreadable")
}},
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guests[0].Status = "stopped"
}},
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
}},
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guestsErr = errors.New("api down")
}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
ctl := runningGuest(9201)
s, _, dir := supServer(t, ex, ctl, 9201)
tc.setup(s, ex, ctl, dir)
for i := 0; i < 4; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9201); n != 0 {
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
}
})
}
}
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
ex := &supExec{status: map[int]string{9202: "exited"}}
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
for i := 0; i < 3; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9202); n != 0 {
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
}
}
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
// again.
//
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
// failed at "crash-looping controller restarted 10 times".
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
s.ControllerSupervisorTick(ctx)
clk.t = clk.t.Add(controllerSupervisorInterval)
}
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
}
st := s.ControllerSupervisorStatus(ctx).Guests[0]
if !st.Crashloop || st.CrashloopSince == "" {
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
}
// Still paused 25 minutes later.
clk.t = clk.t.Add(15 * time.Minute)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
}
// After the pause: tries again.
clk.t = clk.t.Add(controllerCrashloopPause)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax+1 {
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
}
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
}
}
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
var m map[string][]map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
g := m["guests"][0]
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked", "restarts_24h", "slow_crashloop"} {
if _, ok := g[k]; !ok {
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
}
}
}
// ---- R-539 (operator ruling 3 of 2026-09-16): the SLOW crash loop ----------------------------------
// supKillOnce kills the controller and lets the supervisor restart it (two confirming sweeps), then
// moves the clock on by gap. The container comes back "running", so each restart is a separate act.
func supKillOnce(t *testing.T, s *Server, ex *supExec, clk *supClock, gap time.Duration) {
t.Helper()
before := ex.count(9201)
ex.mu.Lock()
ex.status[9201] = "exited"
ex.mu.Unlock()
s.ControllerSupervisorTick(context.Background())
clk.t = clk.t.Add(controllerSupervisorInterval)
s.ControllerSupervisorTick(context.Background())
if ex.count(9201) != before+1 {
t.Fatalf("kill was not followed by exactly one restart (restarts %d → %d)", before, ex.count(9201))
}
clk.t = clk.t.Add(gap)
}
// The consequence: a controller that dies every 20 minutes — never three times inside the 15-minute
// brake — raises slow_crashloop on the FIFTH restart in 24 hours, and the raise does not move again on
// the sixth (the hub mails on movement; once per 24 hours is the ruling).
//
// RED-PROOF: without the slow counter the stanza never sets slow_crashloop → "five restarts 20 minutes
// apart did not raise slow_crashloop — this is R-539".
func TestControllerSupervisor_SlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 1; i <= 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
g := s.ControllerSupervisorStatus(ctx).Guests[0]
if g.Crashloop {
t.Fatalf("the 15-minute brake fired on restarts 20 minutes apart — the fixture is wrong: %+v", g)
}
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("slow_crashloop raised after only 4 restarts: %+v", g)
}
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if !g.SlowCrashloop || g.SlowCrashloopSince == "" || g.Restarts24h != 5 {
t.Fatalf("five restarts 20 minutes apart did not raise slow_crashloop — this is R-539: %+v", g)
}
first := g.SlowCrashloopSince
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if g.SlowCrashloopSince != first {
t.Fatalf("the raise moved again on the 6th restart (%q → %q) — the operator would be mailed per restart", first, g.SlowCrashloopSince)
}
if !g.SlowCrashloop {
t.Fatalf("slow_crashloop cleared while the loop continues: %+v", g)
}
}
// The negative control: restarts that never reach five inside any 24 hours never raise it.
func TestControllerSupervisor_SpreadRestartsNeverSlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 8; i++ { // eight restarts, 7 hours apart: at most 4 inside any 24 hours
supKillOnce(t, s, ex, clk, 7*time.Hour)
}
g := s.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("restarts 7 hours apart raised slow_crashloop: %+v", g)
}
if g.Restarts24h > 4 {
t.Fatalf("restarts_24h=%d — the 24-hour window is not pruning", g.Restarts24h)
}
}
// An agent restart must not reset the slow counter (the ruling; a box whose AGENT also restarts would
// otherwise never reach five). The same state directory, a fresh Server.
//
// RED-PROOF: keep the counter in memory only → the second Server starts at 0 → "the agent restart
// reset the slow counter".
func TestControllerSupervisor_SlowCounterSurvivesAgentRestart(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, dir := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
s2 := &Server{
staleLock: s.staleLock,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: s.logger,
now: clk.now,
}
supKillOnce(t, s2, ex, clk, 20*time.Minute)
g := s2.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.Restarts24h != 5 || !g.SlowCrashloop {
t.Fatalf("the agent restart reset the slow counter: %+v", g)
}
}
+94
View File
@@ -185,6 +185,9 @@ type Options struct {
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
StaleLock StaleLockController
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
GuestsStateDir string
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
ControllerSwapStateDir string
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
@@ -362,6 +365,12 @@ type Server struct {
swapMu sync.Mutex
swapInFlight map[int]bool
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
guestExec GuestExecutor
guestsDir string
ctrlSup controllerSupervisor
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
// detached pipeline runs through (tests inject; production defaults set in NewServer).
netVerifyMu sync.Mutex
@@ -484,7 +493,9 @@ func NewServer(o Options) (*Server, error) {
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
if o.ControllerSwap != nil {
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
s.guestExec = o.ControllerSwap
}
s.guestsDir = o.GuestsStateDir
return s, nil
}
@@ -1084,6 +1095,10 @@ type BackupTierInfo struct {
Target string `json:"target"`
CadenceSeconds int64 `json:"cadence_seconds"`
Primary bool `json:"primary"`
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
Storage string `json:"storage,omitempty"`
}
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1093,11 +1108,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
Target: t.TargetID,
CadenceSeconds: int64(t.Cadence.Seconds()),
Primary: t.Primary,
Storage: s.storagePresence(r.Context(), t.TargetID),
})
}
writeOK(w, resp)
}
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
// successful read that does not list the storage is "absent".
func (s *Server) storagePresence(ctx context.Context, target string) string {
if s.storage == nil || target == "" {
return StoragePresenceUnknown
}
targets, err := s.storage.Observe(ctx)
if err != nil {
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
return StoragePresenceUnknown
}
for _, t := range targets {
if t.Name == target {
return StoragePresencePresent
}
}
return StoragePresenceAbsent
}
const (
StoragePresencePresent = "present"
StoragePresenceAbsent = "absent"
StoragePresenceUnknown = "unknown"
)
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
// result ("presence is not success").
type TierBackupState struct {
Target string `json:"target"`
Primary bool `json:"primary"`
Storage string `json:"storage"` // present | absent | unknown
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
// LastSuccessSource is "storage".
LastSuccess *hub.Backup `json:"last_success,omitempty"`
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
}
// TierAttempt is the newest recorded attempt on a tier, successful or not.
type TierAttempt struct {
StartedAt string `json:"started_at"`
Success bool `json:"success"`
Error string `json:"error,omitempty"`
}
// tierBackupStates builds the per-tier view for one guest.
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
out := make([]TierBackupState, 0, len(s.tiers))
for _, t := range s.tiers {
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
st.LastSuccess, st.LastSuccessSource = b, "record"
} else if st.Storage != StoragePresenceAbsent {
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
StartedAt: when.UTC().Format(time.RFC3339)}
st.LastSuccessSource = "storage"
}
}
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
}
out = append(out, st)
}
return out
}
// tierFromRequest resolves the `?target=` query parameter to a tier.
//
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
@@ -1144,6 +1231,10 @@ type BackupStatusResponse struct {
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
Target string `json:"target,omitempty"`
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
// response exactly as before.
Tiers []TierBackupState `json:"tiers,omitempty"`
}
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1155,6 +1246,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
// across ANY target (echo == "" → pickLatestBackup's match-any path).
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
if echo == "" {
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
}
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
resp.Phase = job.Phase
resp.JobID = job.JobID
+5
View File
@@ -43,6 +43,9 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
SHARED_INSTRUCTIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
# R-389 — shared, like the two above: it lives in felhom.eu/scripts/ and is never copied.
SHARED_OBSERVATIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "observations_gate.py")
# (label, absolute script path, args, fast)
GATES = [
@@ -53,6 +56,8 @@ GATES = [
# missing TAG is what actually broke every install, and the pre-push hook is the earliest place
# that can catch it.
("release-complete", os.path.join(ROOT, "scripts", "check-release-complete.py"), [], True),
# R-389 — a REPORT.md observation with no register row behind it. Fast: stdlib file reads.
("observations", SHARED_OBSERVATIONS, [ROOT], True),
]
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}