Compare commits

...

34 Commits

Author SHA1 Message Date
admin 475bdce7e4 DR bring-up refuses beside a live original (R-834): source guest present, drives bind, or unreadable config
gates / gates (push) Successful in 18s
The DR route keeps onboot 1, binds the real drives and starts the guest: right on a replaced
host, a second box on the same drives beside a live original. The restore-test's no-host-bind
half is now pinned too (measured safe live on demo-hp 2026-10-04).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-04 08:52:35 +02:00
admin d766666ff8 CHANGELOG: v0.138.0 vouched with golden 0.283.1
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 11:40:20 +02:00
admin a4c09a7c11 REPORT: v0.138.0 (R-727)
gates / gates (push) Successful in 16s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 11:00:45 +02:00
admin 904dc20466 CHANGELOG: v0.138.0 released (R-727)
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:34:25 +02:00
admin e1b8269be0 restore test takes only this box's archives (R-727): an archive encrypted with another key is another box's
gates / gates (push) Successful in 16s
Red-proof RP39. Released as v0.138.0 by release-agent.sh.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-30 10:33:52 +02:00
admin 5c68c869b6 docs: v0.137.0 vouched with golden 0.276.0 (2026-09-28)
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-28 09:45:08 +02:00
admin 728d12b1a0 CHANGELOG: v0.137.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 14:01:26 +02:00
admin dd81866b16 REPORT: v0.136.0 and v0.137.0
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 14:01:02 +02:00
admin 3ef095fb71 agent: a guest outside the agent's ACL is not a known guest (R-689, v0.136.0 regression)
gates / gates (push) Successful in 14s
PVE answers 403 permission denied, not "does not exist", for a vmid outside the felhom pool;
v0.136.0 turned that into a lookup failure and the local tier read UNKNOWN every evaluation
(measured on demo-hp). Such an archive is skipped. Red-proofed; verified read-only on demo-hp
with the pre-release binary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 14:00:42 +02:00
admin 7c986915ca CHANGELOG: v0.136.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 13:46:22 +02:00
admin 16dbc83221 agent: the restore test takes only archives of a guest that still exists (R-689, second half)
gates / gates (push) Successful in 14s
Measured on demo-hp right after v0.135.0: with the golden skipped the pick fell to a leftover
archive of guest 9100, deleted in August. "does not exist" skips it; any other lookup error
makes the tier unknown. Red-proofed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 13:45:56 +02:00
admin 9555a7f93b REPORT: v0.135.0 released and signed-delivered to both demo hosts
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 13:02:10 +02:00
admin 9ff937d8fb CHANGELOG: v0.135.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 12:16:17 +02:00
admin d4be12ca95 agent: the restore test proves only backups of a guest (R-689)
gates / gates (push) Successful in 13s
The golden template in local:backup/ was picked as the newest settled archive on demo-hp and
failed every 6 h. Candidates are now vzdump-<type>-<vmid> files or PBS ct|vm/<vmid> snapshots
with a reported vmid. Red-proofed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-27 12:15:42 +02:00
admin 7403c2a838 REPORT + CONTEXT: v0.133.0 and v0.134.0 delivered to both demo boxes; restore test back on
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-25 04:55:26 +02:00
admin 309e368731 CHANGELOG: v0.134.0 released (tag + package verified by download), not vouched
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:49:19 +02:00
admin 0722b2cdb0 agent: a whole-box backup that cannot fit its local target is skipped with a reason before anything starts (R-685)
gates / gates (push) Successful in 13s
Free space is read from GET /nodes/<node>/storage — GET /storage carries no usage (found live, before release).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:48:51 +02:00
admin 4fe2f81a32 CHANGELOG + REPORT + CONTEXT: v0.133.0 released (tag + package verified by download), not delivered
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:25:47 +02:00
admin 9bdb4dae8f v0.133.0: a restore-test can never fill a box's disk; leftovers retried on a timer (R-672, R-673)
gates / gates (push) Successful in 13s
Space preflight before anything is created (uncompressed size from the vzdump log / PBS
snapshot, x1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage
is eligible, unknown refuses, reported as a non-pass result). Failed scratch teardown and
the stale-lock sweep retried every 10 min (the sweep under the heavy-op gate). A thin
pool crossing 90% requests an immediate host report. Six red-proofs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:23:15 +02:00
admin d9864a94bf CHANGELOG: v0.132.0 vouched for Day-0 installs with golden 0.246.0 (operator decision)
gates / gates (push) Successful in 16s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 12:25:49 +02:00
admin 77cd70f7c0 REPORT: agent v0.132.0 - slow crash loop, signed delivery, live proof with the production window
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 11:30:45 +02:00
admin 1030abd7d6 CHANGELOG: v0.132.0 released (tag + package verified by download)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:27:16 +02:00
admin 18d03bd437 v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s
Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.

Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.

Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:26:42 +02:00
admin e98b857684 REPORT: v0.131.0 supervisor + per-tier status, delivery and live validation
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 11:37:07 +02:00
admin dcdeb3d16d CHANGELOG: v0.131.0 released (tag + package verified by download)
gates / gates (push) Successful in 12s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:50:47 +02:00
admin 610804b98d v0.131.0: controller supervisor (R-523); per-tier backup status + tier storage presence (R-517/R-518)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:49:44 +02:00
admin 4586f0f7f6 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 9s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:52 +02:00
admin 205e22babe decoy sweep: no gate changed here, and that is the result (R-421)
gates / gates (push) Failing after 12s
All 29 gate scripts across the four repos were read and DECOYED - the label constructed without the
fact, the gate run, the verdict recorded. 16 were fooled. None of them were in this repo.

A decoy that nobody would write proves nothing, so the attempts that turned out illegitimate were
WITHDRAWN rather than counted. Both of this repo were withdrawn, and both are named in the audit.

The gates here that could not be given a plausible decoy are listed BY NAME in
felhom.eu/scripts/decoy_coverage_gate.py EXEMPT (R-426) as UNTESTED - not as sound. A gate nobody
tried to fool is UNKNOWN, and calling it sound would be the same confident guess this sweep exists
to find.

Survey table: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:38:58 +02:00
admin 058b945064 gate 11: register the shared observations gate
gates / gates (push) Successful in 10s
felhom.eu/scripts/observations_gate.py, invoked across the workspace like
reuse_refs_check.py and instructions_gate.py. This repo's REPORT.md has no
observations section today, so the gate passes quietly - it is registered for
the session that writes one.
2026-08-23 13:53:20 +02:00
admin 40d857b527 CHANGELOG: the fleet runs the published bytes, not the proof build (R-349)
gates / gates (push) Successful in 7s
Both boxes were first given a hand build: same source, same version
string, different bytes (256e0829 vs the published a56a92a7), because
release-agent.sh builds with -trimpath -buildvcs=false and a hand build
does not.

Nothing would have corrected it. The boxes already reported 0.130.0, so
self-update saw the vouched version as installed and would have done
nothing, forever. Every version check in the system compares the STRING.

Both reinstalled from the downloaded package; both now report a56a92a7.
2026-08-20 12:53:37 +02:00
admin 7ae6990bac v0.130.0 released: tag + package published, heading now claims it
gates / gates (push) Successful in 8s
sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3
tag v0.130.0 at 7569f34, 14,141,158 bytes.

Reproducible: rebuilding with -trimpath -buildvcs=false matches the
published artifact byte for byte (R-186's property, checked not assumed).

The heading said UNRELEASED while the fix was hand-installed on demo-hp
only -- publishing then would have pushed it onto the control box through
self-update. release-complete convicted on the release heading and was
right to; the answer was to stop claiming a release, not to bypass it.

NOT VOUCHED by this commit. Vouching is the separate operator act.
2026-08-20 12:47:25 +02:00
admin 7569f34aeb CHANGELOG: correct a claim this session made and then disproved
gates / gates (push) Successful in 8s
The v0.130.0 draft said the 388 descriptors already stuck on ep0 would
persist until the PBS proxy restarted. Measured within the hour: they
clear when the AGENT restarts. demo-hp released its 199 in one second
(415 -> 216 fd); demo-felhom released the remaining 203 (220 -> 17 fd in
under two seconds). 17 is precisely ep0's t0 baseline of 2026-08-18.

They were held on both sides. Closing either side ends them. ep0 was
read-only throughout and its proxy PID never changed.
2026-08-20 12:39:12 +02:00
admin ede49b610d R-344: restore the idle-connection timeout our hand-rolled transports lost
gates / gates (push) Successful in 7s
Every client here pins TLS, so none can use http.DefaultTransport and each
hand-rolls its own. A composite literal takes IdleConnTimeout ZERO, which
means retain idle connections forever -- not "use a sane default".
pbsTargetsFromPVE builds a fresh pbs.Client every cycle and drops the
previous one, and an abandoned http.Transport does not close its
connections. One stranded socket per cycle, on both sides, forever.

Measured: 388 established connections on ep0 over 46 h, 194 per box, zero
closed in a 31-minute window. pvestatd and proxmox-backup-client made
162,404 requests in the same window and leaked none.

New leaf package internal/httpx owns the default (90s, http.DefaultTransport's
own value) and NewTransport, which returns a FRESH transport per call and
treats <=0 as "use the default", never "no timeout". pbs.Config gains
IdleConnTimeout for tests only.

hub and proxmox carried the same missing default and are corrected here for
consistency. Neither contributed to the ep0 leak -- both are built once per
process and neither talks to ep0:8007.

Tests count connections SERVER-side and model the abandonment, so they pin
the consequence rather than the field. Two red-proofs, both seen failing:
removing the timeout -> "still holds 5 open connection(s), want 0";
DisableKeepAlives -> "3 sequential requests over 3 connection(s), want 1"
(the leak test PASSES under that one -- it is the worse-fix guard that
catches it).

Not released: hand-installed on demo-hp only so demo-felhom stays the
control. CHANGELOG heading stays UNRELEASED until the publish is authorised.
2026-08-20 11:09:39 +02:00
admin f17ed11599 REPORT: agent v0.129.0 — the retained-package recovery class, released and deployed
gates / gates (push) Successful in 21s
2026-08-12 18:49:51 +02:00
46 changed files with 3678 additions and 104 deletions
+296
View File
@@ -1,3 +1,297 @@
## v0.138.0 — the restore test takes only THIS box's archives (2026-09-30, R-727, `09` §3 decision 51)
> **RELEASED 2026-09-30** by `scripts/release-agent.sh` — tag `v0.138.0` (`e1b8269`), sha256
> `55916026001790a79ebf97d32c032610cfde8e09d02979b9b9d8c2cbc5d88195`, verified by download. **Not vouched** (the golden
> keeps 0.137.0 until the next bake); delivered to the demo boxes by signed `agent_update` jobs.
>
> **VOUCHED 2026-09-30** with golden 0.283.1 (`min_agent` 0.131.0), on the operator's word: the hub logged
> `Artifact manifest set: agent=0.138.0 golden=0.283.1 min_agent="0.131.0"`. Evidence:
> `felhom.eu/documentation/audits/evidence-golden-0283-2026-09-30/`.
**MinAgent impact:** none required by any controller.
- A returning customer's PBS namespace can hold archives of EARLIER boxes: same guest id (9201), same token, written
with a different key. Measured 2026-09-30: the newest SETTLED archive was an earlier box's, and the test failed
`wrong key … manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…` every evaluation. The archive
carries no host id; it carries its key fingerprint (PVE content `encrypted`), and the storage carries its own
(`GET /storage` → `encryption-key`). `PickSettledRestoreCandidateOn` now skips — and logs by name, once — an archive
whose fingerprint is not the storage's own; an unencrypted storage is not filtered; a failed storage read is an
error (tier UNKNOWN), never "nothing to prove".
- Tests: `TestR727_TheRestoreTestTakesOnlyThisBoxsArchives` (the 2026-09-30 shape: nothing picked while this box's
archive settles, then exactly it), `TestR727_UnencryptedStorageIsNotFiltered`, `TestR727_KeyLookupFailureIsUnknown`.
Red-proof RP39: the skip removed → the earlier box's `2026-09-16T21:59:54Z` is picked.
## v0.137.0 — a guest outside the agent's ACL is not a known guest (2026-09-27, R-689, v0.136.0 regression)
> **RELEASED 2026-09-27** by `scripts/release-agent.sh` — tag `v0.137.0` (`3ef095f`), sha256 `766c9166916a1bd3674b0dc69081f8a7619e770f1402d8ad7705b395937e7627`, verified by download. **NOT vouched** (the operator's act).
>
> **VOUCHED 2026-09-28** with golden 0.276.0 (`min_agent` 0.131.0), on the operator's word of 2026-09-27: the hub logged
> `Artifact manifest set: agent=0.137.0 golden=0.276.0 min_agent="0.131.0"`, and a Day-0 test install fetched this binary
> through the manifest and sha-verified it. Evidence: `felhom.eu/documentation/audits/evidence-golden-0276-2026-09-28/`.
**MinAgent impact:** none required by any controller.
- v0.136.0 asked `GuestConfig` whether an archive's guest exists and treated anything but "does not exist" as a lookup
failure. PVE answers **403 "permission denied at /vms/<id>"** for a vmid outside the token's pool — so on demo-hp the
deleted guest 9100's archive made the local tier UNKNOWN every evaluation. A guest the agent cannot read is not one it
manages; its archive is skipped. `TestR689_AGuestOutsideTheAgentsACLIsNotAKnownGuest`, red-proofed. Verified read-only
on demo-hp with the pre-release binary before the release.
## v0.136.0 — … of a guest that still EXISTS (2026-09-27, R-689 second half)
> **RELEASED 2026-09-27** by `scripts/release-agent.sh` — tag `v0.136.0` (`16dbc83`), sha256 `2eb0b5ebe253defd68b322312bbac12418c051b0a7a0d2d1831310d97fa6d755`, verified by download. **NOT vouched** (the operator's act).
**MinAgent impact:** none required by any controller.
- **R-689, second half.** Read on demo-hp right after v0.135.0 with the read-only `-selftest=restore-test-due`: the
golden was skipped, and the pick fell to `vzdump-lxc-9100-2026_08_21…` — a leftover of a guest deleted in August. A
candidate's guest must now exist on this node (`GuestConfig`): "does not exist" skips the archive (INFO once per
volid); any other lookup failure is returned, so the tier reads UNKNOWN, never "nothing to prove". Tests
`TestR689_AnArchiveOfADeletedGuestIsNeverPicked` (red-proofed: without the check it picks the 9100 leftover),
`TestR689_AGuestLookupFailureIsUnknownNotEmpty`. A second agent release in one session — the first half was found
incomplete on the box.
## v0.135.0 — the restore test proves only backups OF A GUEST (2026-09-27, R-689)
> **RELEASED 2026-09-27** by `scripts/release-agent.sh` — tag `v0.135.0` (`d4be12c`), sha256 `ad4e75f16d338552f4588d3fe64c51cbf9651220b9d223b5386b85b6c37fd4c3`, verified by download. **NOT vouched** (the operator's act).
**MinAgent impact:** none required by any controller.
- **R-689** (`backup/runner.go` `guestBackupArchive`). demo-hp keeps its golden template in `local:backup/` — content
"backup", 654 MB, plausibly complete, the newest settled entry — and the scheduled restore test picked it every 6 h and
failed `extractconfig` with a 403, while the guest's real archive went untested. A restore-test candidate is now a
`vzdump-{lxc,qemu}-<vmid>-…` file or a PBS `backup/{ct,vm}/<vmid>/…` snapshot whose vmid the storage reports; anything
else is skipped with one INFO line per volid (not the "INCOMPLETE archive" WARN). Tests
`TestR689_TheRestoreTestNeverPicksTheGolden` (red-proof: without the check it picks the golden) and
`TestR689_GuestBackupArchiveShapes`; two older picker tests' fixtures moved to real archive names.
Evidence: `felhom.eu/documentation/audits/version-travel-2026-09-26/D1/`.
## v0.134.0 — a whole-box backup that cannot fit is skipped with a reason, before anything starts (2026-09-25 night, R-685)
> **RELEASED 2026-09-24 night** by `scripts/release-agent.sh` — tag `v0.134.0` (`0722b2c`), sha256 `7593bebe03234c7d22f3ade384e7ed7787dc659aa8c8594b3ee19af2ce81c72d`, verified by download. **NOT vouched** (Day-0 stays on the previous version).
**MinAgent impact:** none required by any controller.
- **R-685 — the backup space preflight** (`backup/runner.go` `spaceFits`). Before a vzdump to a LOCAL (non-PBS)
target, the newest archive of that guest on that target × 1.25 + 1 GiB must be free — PVE prunes old archives
only AFTER a successful backup, so the kept ones still stand during the run. A shortfall is a skip: nothing
starts, the record's Error reads `skipped: not enough space: <target> has X GiB free; the last archive of guest N
was Y GiB, so a new one needs about Z GiB …` (stable prefix `BackupSkipNoSpacePrefix`), and it reaches the
controller's tier view and, through the quiesce loop's tier notifier, the operator's `whole_guest_backup_failed`.
It FAILS OPEN on what is not known (PBS target, first backup, unreadable usage).
- **Free space is read from `GET /nodes/<node>/storage`** (`NodeStorage`), never `GET /storage` — the latter is the
cluster DEFINITIONS and carries no usage. **Found live, before release:** the first build read `/storage`,
failed open, and a real vzdump of demo-hp 9201 started from the live test; it was aborted after 5 min 16 s, no
archive left (`felhom.eu/documentation/audits/night-2026-09-25/F/`). The test fake's `ListStorage` now strips
usage like production. Red-proofs: two (`…/F/redproof-r685-*.txt`).
- Live proof (demo-hp, safe builds with a hard stop before vzdump): ×10 margin → refused, "local has 14.9 GiB
free … needs about 77.2 GiB"; release margin → "space preflight passed" need 11.3 GB, avail 16.0 GB.
## v0.133.0 — a restore-test can never fill a box's disk; leftovers retried on a timer (2026-09-24, R-672, R-673)
> **RELEASED 2026-09-24** by `scripts/release-agent.sh` — tag `v0.133.0` (`9bdb4da`), sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, verified by an independent anonymous download. **NOT delivered** (needs an operator-signed `agent_update` job per box, R-530) and **NOT vouched**.
**MinAgent impact:** none required by any controller. Hub **v0.124.0** makes a thin pool CRITICAL at 90 %
(data or metadata) and keys the storage-fill alarm per pool per 6 h; an older hub still raises its generic
90/95 % storage-fill events from the same report.
- **R-672 — the space preflight** (`reconcile/restoretest_space.go`, provider `internal/restorespace`). Before
anything is journaled or created, a restore-test needs free data ≥ restored × 1.2 + 5 GiB on its target
(`backup.restore_test_space_factor` / `backup.restore_test_space_reserve_gib`) and room in a thin pool's
metadata. `restored` is the UNCOMPRESSED size — the vzdump log's "Total bytes written" (a file-backed
archive) or the PBS snapshot size; the archive FILE is never used (9201: 6.9 GB file, 22.6 GB restore). The
target moves OFF the tested guest's own pool when another storage is eligible (active, `rootdir`, and the
agent holds Datastore.AllocateSpace there) and fits. Anything unknown refuses. A refusal is reported to the
hub as the test's result (`pass=false`, `skipped=true`, "skipped: not enough space on …") — never a pass,
never dropped. The archive config is read once (it was read twice). Measured case: demo-hp 2026-09-24 —
the restore-test that filled `local-lvm` would now be refused (needs 32.1 GB, 23.2 GB free).
- **R-672 — a failed scratch teardown is retried every 10 minutes** (`Engine.RetryScratchTeardown`, the
daemon's janitor), not only by Recover at agent start; never a vmid a running test owns; after 3 failed
tries the operator is told through a failed restore-test record naming the scratch guest.
- **R-672 — a thin pool crossing 90 % requests an immediate host report** (`Observer.SetThinHighTrigger`, the
storage watchdog's read path, every few seconds; re-armed below 85 %), so the hub's alarm sees it in
seconds, not at the next 15-minute report.
- **R-673 — the stale-lock sweep runs on the same 10-minute timer** (it ran only at start: a stale
`snapshot-delete` lock blocked 9201's whole-box backups for five hours), holding the one-heavy-operation
gate so no agent backup can start between its "no vzdump running" check and its unlock.
- Red-proofs: six, each seen failing (REPORT).
## v0.132.0 — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
> **RELEASED 2026-09-17** by `scripts/release-agent.sh` — tag `v0.132.0`, sha256 `4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321`. **Vouched 2026-09-17** for Day-0 installs on the operator's word, together with golden **0.246.0** (the hub's R-120 gate required the newer golden first). Delivered to demo-hp and the N100 by an operator-signed `agent_update` job each (operator ruling 3 of 2026-09-16), not by a floor.
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
`controller_slow_crashloop`; an older hub ignores them.
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
does **not** stop restarting — the fast brake remains the only brake.
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
after an agent restart — `last_success_source: storage`), the last attempt kept apart
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
last success".
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
absent). **All four found by accident.** The gates enforce everything else here and were the one part
nothing had checked.
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
proves the cause was the listing rather than the decoy.
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
**In this repo:** no gate changed, and that is the result. `release-complete` was decoyed and is
SOUND — the sweep's attempt (a non-version heading on top of `CHANGELOG.md`) was WITHDRAWN as
illegitimate, because `HEAD_RE.search` scans the whole file and still names `v0.130.0`. The three
shared gates are covered from `felhom.eu`; the remaining two are named in the decoy-coverage
exemption list (R-426) as UNTESTED, not as sound.
## v0.130.0 — the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
> **RELEASED 2026-08-20**, on the operator's word, after the fix was proved on both boxes.
> `sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, 14,141,158 bytes,
> tag `v0.130.0` at `7569f34`. Reproducible: a rebuild with `-trimpath -buildvcs=false` matches the
> published artifact byte for byte (R-186's property, checked rather than assumed).
>
> The heading read `## UNRELEASED — v0.130.0 candidate` until this point, deliberately: while the fix
> was hand-installed on `demo-hp` only, publishing would have pushed it onto `demo-felhom` through
> self-update and destroyed the control the proof rested on. **`release-complete` convicted on the
> release heading and was right to** — the answer was to stop claiming a release, not to bypass the
> gate. See R-347.
>
> **The fleet runs these exact bytes.** Both demo boxes were first given a hand build made during the
> proof — same source, same version string, **different bytes** (`256e0829…`), because
> `release-agent.sh` builds with `-trimpath -buildvcs=false` and a hand build does not. Nothing would
> have corrected that: the boxes already reported `0.130.0`, so self-update saw the vouched version as
> installed and would have done nothing, forever. Both were reinstalled from the **downloaded package**
> and now report `a56a92a7…`. Filed as **R-349**, because every prove-then-publish train hits it.
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
proxy whose descriptor ceiling is 65536 and whose runway at that rate was ~323 days. The connections
were held open on **both** sides: ep0 showed 388 while the two boxes showed 194 + 194, at two separate
instants, with the same source ports on each side, and **not one closed in a 31-minute window**.
`ss -tnp` on the boxes named the holder: **`felhom-agent`**, 194 of 194 on each, one PID.
**`pvestatd` and `proxmox-backup-client` made 162,404 requests in that window and leaked zero.** They
are 99.5% of the traffic to that endpoint and 0% of the leak. The agent made 811 requests — of which
387 were `GET .../snapshots` — and leaked 388 sockets. One per call, within one.
**The defect, and it is two things compounding.**
- `internal/pbs/client.go` built its transport as a composite literal:
`&http.Transport{TLSClientConfig: tlsCfg}`. That takes **`IdleConnTimeout` zero, which does not mean
"use a sane default" — it means retain idle keep-alive connections FOREVER.**
`http.DefaultTransport` sets 90s; hand-rolling the transport (which every client here must do,
because they all pin TLS) silently discards it.
- `pbsTargetsFromPVE` (`cmd/felhom-agent/main.go`) builds **a fresh `pbs.Client` every cycle**, as its
own doc comment says, and drops the previous one. An abandoned `http.Transport` does **not** close
its connections — it becomes unreachable while its `persistConn` read-loop goroutine keeps the
socket alive. So each cycle stranded exactly one connection that nothing could ever close.
The cadences reconcile without fitting: a 900s live-snapshot collect (184.7 cycles in the window) plus
a 6-hour verify loop (7.7) predicts 192.4 per box against **194 observed**.
**The fix is one field, restored to the standard library's own value.** New leaf package
`internal/httpx` owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is what
`http.DefaultTransport` uses, so there is nothing invented here to justify or tune — and
`NewTransport(tlsCfg, idleConnTimeout)`, which returns a **fresh** transport (never shared: each caller
pins a different endpoint) and treats a zero or negative timeout as **use the default, never "no
timeout"**. `pbs.Config` gains an `IdleConnTimeout` field that production leaves unset; only tests set
it, to avoid a 90-second wait.
**`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and
were corrected in the same pass — but neither contributed to the ep0 leak, and this entry must not be
read as three leaks having been found.** Both are built **once per process**, so they held one idle
connection for the life of the daemon rather than accumulating, and neither talks to ep0:8007.
**Tests, and what they deliberately do not assert.** `internal/pbs/client_leak_test.go` counts
connections **server-side** and models what `pbsTargetsFromPVE` actually does — build a client, use it
once, drop it on the floor — then asserts the connections go away. It does not assert `err == nil` and
it does not assert that some field holds some value; both were true of the leaking code.
- **Red-proof 1 (the fix):** removing `IdleConnTimeout` from `NewTransport` fails the test with
*"after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is
in the message, so the failure cannot be mistaken for a timeout with another cause. Reverted.
- **Red-proof 2 (the fix that would be worse than the bug):** setting `DisableKeepAlives: true` also
makes the leak vanish — by dialling fresh for every request, which on a box polling ~40,000 times a
day is strictly worse than what we started with. **The leak test PASSES under that mutation**;
`TestPBSClient_KeepAliveStillReuses` is what catches it, failing with *"3 sequential requests over 3
connection(s), want 1"*. Reverted.
**What this release does NOT do.** It does not reduce the poll rate (**R-336 stays open, but re-scoped
— it was never the cause of this leak**), it does not refactor `pbsTargetsFromPVE` to cache or reuse
clients (a one-line default restores the standard behaviour; a lifecycle refactor adds
cache-invalidation questions for no measurable gain), and it adds no `CloseIdleConnections` call.
**One sentence in this entry was written before the deploy and was WRONG, and it is corrected here
rather than quietly edited.** It read: *"does not clear the 388 descriptors already stuck on ep0 —
those persist until that proxy restarts."* **Measured: they clear the moment the AGENT restarts.**
Replacing the binary on `demo-hp` released exactly its 199 descriptors within one second
(415 → 216 fd), and replacing it on `demo-felhom` released the remaining 203 (**220 → 17 fd in under
two seconds**). **17 is precisely ep0's `t0` baseline** of 2026-08-18 09:51:22Z. ep0 was read-only
throughout and its proxy PID never changed. The accumulated leak was never ep0's to hold on to — it
was held on both sides, and closing either side ends it.
## v0.129.0 — a correct code for an earlier package stops being called wrong (2026-08-12, R-311)
**The measurement this fixes.** On 2026-08-12 a recovery code that provably opens a RETAINED package
@@ -5389,3 +5683,5 @@ client, signing, or storage/backup orchestration yet (later slices).
read-only `--selftest` against the demo host with TLS fingerprint pinning.
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
token) is provisioned out-of-band; the agent only consumes the token.
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->
+8
View File
@@ -101,3 +101,11 @@ the mechanism are exempt.
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
assumption, not an observation.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+17
View File
@@ -1,5 +1,22 @@
# CONTEXT — felhom-agent working state
> **2026-09-25 night — v0.133.0 AND v0.134.0 DELIVERED to both demo boxes (CC-signed `agent_update`, ruling 1);
> restore test back ON (the `-1` config kept as `agent.json.night-0925-off`). v0.134.0 = R-685:** `backup/runner.go`
> `spaceFits` — a vzdump to a LOCAL target needs free ≥ newest archive of that guest × 1.25 + 1 GiB (PVE prunes
> only after success); a shortfall is a named skip (`BackupSkipNoSpacePrefix`), fail-open on PBS / first backup /
> unknown usage. **Free space comes from `NodeStorage` (`GET /nodes/<n>/storage`) — `ListStorage` (`GET /storage`)
> has NO usage**; the first build read it and a real vzdump started in its own live test (aborted, no archive).
> demo-hp: `local_backup_retention` 1 (operator option A, saved `agent.json.pre-a4-retention`). Peti's box: nothing.
> **2026-09-24 — v0.133.0 RELEASED, NOT DELIVERED (R-672, R-673).** Restore-test space preflight
> (`reconcile/restoretest_space.go`, `internal/restorespace`): uncompressed size from the vzdump log / PBS size,
> × 1.2 + 5 GiB, thin metadata, off the tested guest's pool, unknown refuses, reported as `skipped` non-pass.
> Janitor (`cmd/felhom-agent/janitor.go`) every 10 min: `Engine.RetryScratchTeardown` + stale-lock sweep under
> the heavy-op gate. Thin pool ≥ 90 % → immediate report (hub v0.124.0 alarms). **Operator rulings 2026-09-24
> (evening):** the scheduled restore-test is OFF on both demo hosts (`backup.restore_test_eval_interval_seconds:
> -1` — note: 0 means the 6 h DEFAULT, only a negative disables) until v0.133.0 is delivered there; saved configs
> `/etc/felhom-agent/agent.json.pre-r672`. demo-hp 9201 was repaired (stop, fsck, start; two Redis AOF tails cut).
> Snapshot of the current state + open threads. Authoritative history lives in `CHANGELOG.md` (top
> entry = current); the end-of-task detail lives in `REPORT.md`.
+10 -61
View File
@@ -1,63 +1,12 @@
# REPORT — felhom-agent, 2026-08-09 (gates only)
# REPORT — 2026-09-30: v0.138.0 (R-727)
**No release. No version bump. No binary published. `scripts/` only** — nothing that runs on a
customer's machine changed, and the agent stays **v0.128.0** at `28ba8593b8`.
Full session report: `felhom.eu/REPORT-fixes-first-tester-2026-09-30.md`.
## What changed
| file | why |
|---|---|
| `scripts/retention-policy.json` **(new)** | THE retention number, in one place, with its reasoning and its honesty about where the number came from |
| `scripts/check-published-versions.py` | reads that number; bounds its assertion to the newest N; **prints what it stopped covering** |
| `scripts/check-release-complete.py` **(new)** | asserts the CHANGELOG-head version is tagged, placed in this history, and published |
| `scripts/agent_gates.py` | registers the new gate; legs 1–2 are offline so it runs in `--fast` too |
## The coupling defect, and the fix
The prune keeps the newest N; the published-versions check demanded that **every** tag be
downloadable. Nothing connected them, so CI went red at `28ba8593b8` — a commit whose own run had
been **green the day before** — and would have gone red again at the next publish when `0.121.0` was
evicted. Both now read `generic_versions_kept` from one file.
**What CI no longer covers:** a released version **older than the retention window** is no longer
asserted downloadable. Its git tag and its config tree are still asserted; only the binary's presence
is dropped. The check names the dropped versions on every run.
**The number is not a located ruling.** `generic_versions_kept: 10` is what the registry demonstrably
holds; no register row records a prune, and container packages hold 19 each. The file says so in its
own header. The principled bound is the hub's vouched `min_agent` floor — nothing can install below
it — and that is recorded as the follow-up.
## Controls, all three run
| control | expected | got |
|---|---|---|
| live run, policy = 10 | green, and it names `0.120.0` as not asserted | **exit 0**, and it did |
| widen policy to 11 | `0.120.0` re-enters the window and convicts | **exit 1**, `FAIL v0.120.0` |
| policy file removed | INCONCLUSIVE, never silently unbounded | **exit 2**, naming the path it tried |
## Red-proof of the new gate
Mutation: `CHANGELOG.md` head repointed to `## v0.129.0` — never tagged, never published. Asserted
applied (`grep -c '^## v0.129.0'` → 1). Result **exit 1**, both legs convicting:
```
- TAG v0.129.0 DOES NOT EXIST. … without the tag every install 404s mid-run, as root.
Fix: git tag -a v0.129.0 <released-commit> && git push origin v0.129.0
- PACKAGE 0.129.0 IS NOT PUBLISHED (HTTP 404 …).
Fix: bash scripts/release-agent.sh 0.129.0
```
Reverted; `git status` clean on `CHANGELOG.md`.
## Green gate
`python3 scripts/agent_gates.py` — `reuse-refs OK · instructions OK · published OK ·
release-complete OK · all agent gates OK`.
## Not done here
The deleter of `0.120.0` is **still not established** and a second attempt failed — Gitea keeps no
package-deletion trail, its container log no longer reaches the window, and the activity feed carries
no package operation. Recorded in R-287, including the withdrawal of my own earlier over-claim that
the router logs showed no DELETE: they do not cover the window, so they never said anything.
- **Measured:** a PBS archive carries its key FINGERPRINT (PVE content `encrypted`), not a host id; the storage
carries its own (`GET /storage` → `encryption-key`). The restore test now skips an archive written with another
key and logs it by name; an unencrypted storage is not filtered; a failed storage read is UNKNOWN.
- Tests `TestR727_*`; red-proof RP39 (the skip removed → the 2026-09-16 archive of an earlier box is picked).
- Released by `release-agent.sh` (tag `v0.138.0`, sha256 `55916026…8195`, verified by download); **not vouched**.
Delivered by signed `agent_update` jobs to `demo-hp-bb76ea` and `demo-felhom-8363b5` (both committed within 340 s);
`-selftest=restore-test-due` on both reads each tier normally on 0.138.0.
- ep0 (decision 51): the three drill archives in `tester-1`'s namespace removed; other namespaces byte-identical.
+4
View File
@@ -74,6 +74,7 @@
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
### Proxmox client / hub / PBS / provisioning
@@ -84,8 +85,11 @@
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
| `reconcile.PreflightRestoreSpace` + `restorespace.Provider` (v0.133.0) | internal/reconcile/restoretest_space.go, internal/restorespace/restorespace.go | `PreflightRestoreSpace(ctx, space, policy, archive, rawCfg, configured) SpaceVerdict` | ANY step that restores or copies a guest onto a storage — size it first | The restored size is UNCOMPRESSED (vzdump log "Total bytes written" / PBS snapshot size) — never the archive FILE (6.9 GB file → 22.6 GB restore, R-672); an unknown refuses; eligibility needs Datastore.AllocateSpace on `/storage/<id>` specifically (the `/` grant answers every path) |
| `Engine.RetryScratchTeardown` + the daemon janitor (v0.133.0) | internal/reconcile/restoretest_retry.go, cmd/felhom-agent/janitor.go | `RetryScratchTeardown(ctx) ScratchRetryResult` | retrying a leftover on a TIMER instead of only at start | Never `Recover` on a timer — it also resolves generic in-flight ops; a periodic sweep that unlocks guests holds the one-heavy-op gate (`InFlight.TryAcquire`) |
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
| `httpx.NewTransport` | internal/httpx/transport.go | `NewTransport(tlsCfg, idleConnTimeout) *http.Transport` | **EVERY** hand-rolled `http.Transport` in this repo — pbs, hub and proxmox all pin TLS, so none can use `http.DefaultTransport` | **R-344: never inline `&http.Transport{TLSClientConfig: ...}` again.** A composite literal takes `IdleConnTimeout` **zero, which means retain idle connections FOREVER** — `http.DefaultTransport` sets 90s and a literal does not inherit it. Combined with a client rebuilt per cycle and dropped (`pbsTargetsFromPVE`), that stranded **388 sockets on ep0 in 46 h**, held open on BOTH sides. `idleConnTimeout <= 0` means **use the default**, never "no timeout". Returns a **FRESH** transport every call — a shared one would pool connections across differently pinned endpoints. Pinned by `internal/pbs/client_leak_test.go` (server-side connection counting) + `internal/httpx/transport_test.go` |
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
+81
View File
@@ -0,0 +1,81 @@
package main
import (
"context"
"fmt"
"log/slog"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// janitorInterval is how often the leftovers of an interrupted restore-test or backup are retried
// (R-672 rule 3, R-673). Both used to be resolved ONLY at agent start: on 2026-09-24 a failed scratch
// teardown kept a full thin pool full for 2.5 h, and a stale `snapshot-delete` lock blocked 9201's
// whole-box backups for five hours — each cleared within a minute of an agent restart.
const janitorInterval = 10 * time.Minute
// janitorDeps are the janitor's seams (tests drive one pass with fakes).
type janitorDeps struct {
retryScratch func(ctx context.Context) reconcile.ScratchRetryResult
staleLocks func(ctx context.Context) // localapi Server.RecoverStaleLockedGuests; nil when the local API is off
heavy *backup.InFlight
record func(hub.RestoreTest)
now func() time.Time
logger *slog.Logger
}
// janitorPass is one pass. The stale-lock sweep runs only while holding the one-heavy-operation gate, so
// no agent backup can START between its "no vzdump is running" check and its unlock (at start-up the
// sweep ran before the backup loop existed; on a timer that ordering must be made, not assumed). A busy
// gate skips the sweep this pass — the next pass retries.
func janitorPass(ctx context.Context, d janitorDeps) {
if d.retryScratch != nil {
r := d.retryScratch(ctx)
if r.Examined > 0 {
d.logger.Info("janitor: restore-test scratch retry pass", "examined", r.Examined,
"destroyed", r.Destroyed, "already_gone", r.Clean, "failed", r.Failed)
}
for _, vmid := range r.GaveUp {
// The operator is told through the existing restore-test failure path: the hub raises
// restore_test_failed (operator) once per distinct archive — this record's archive names
// the stuck scratch guest.
if d.record != nil {
d.record(hub.RestoreTest{
SourceArchive: fmt.Sprintf("scratch-teardown:%d", vmid),
ScratchVMID: vmid,
Pass: false,
Error: fmt.Sprintf("restore-test scratch guest %d could not be torn down after %d retries — it holds its disks; remove it by hand (pct destroy %d) after checking what keeps it busy",
vmid, reconcile.MaxTeardownTries, vmid),
TestedAt: d.now().UTC().Format(time.RFC3339),
})
}
}
}
if d.staleLocks != nil {
release, busy, ok := d.heavy.TryAcquire("stale-lock-sweep")
if !ok {
d.logger.Info("janitor: stale-lock sweep deferred — a heavy operation is in flight", "busy", busy)
return
}
defer release()
d.staleLocks(ctx)
}
}
// runJanitor runs janitorPass every janitorInterval until ctx ends.
func runJanitor(ctx context.Context, d janitorDeps) {
d.logger.Info("janitor: starting (restore-test scratch retry + stale-lock sweep)", "interval", janitorInterval)
t := time.NewTicker(janitorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
janitorPass(ctx, d)
}
}
}
+60
View File
@@ -0,0 +1,60 @@
package main
import (
"context"
"io"
"log/slog"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-672 / R-673 (v0.133.0): one janitor pass, driven with fakes.
func quiet() *slog.Logger { return slog.New(slog.NewTextHandler(io.Discard, nil)) }
// A scratch the engine gave up on reaches the hub as a failed restore-test record naming it — the
// existing operator path (restore_test_failed). Never a pass.
func TestJanitor_GaveUpIsReportedAsAFailure(t *testing.T) {
var got []hub.RestoreTest
janitorPass(context.Background(), janitorDeps{
retryScratch: func(context.Context) reconcile.ScratchRetryResult {
return reconcile.ScratchRetryResult{Examined: 1, Failed: 1, GaveUp: []int{990000}}
},
heavy: &backup.InFlight{}, record: func(r hub.RestoreTest) { got = append(got, r) },
now: time.Now, logger: quiet(),
})
if len(got) != 1 || got[0].Pass || got[0].ScratchVMID != 990000 || !strings.Contains(got[0].Error, "990000") {
t.Fatalf("records = %+v — want one FAILED record naming scratch 990000", got)
}
}
// R-673: the stale-lock sweep runs only while holding the one-heavy-operation gate, so no agent backup can
// start between its "no vzdump running" check and its unlock.
//
// COMPANION RED-PROOF (REPORT): drop the TryAcquire → "the sweep ran while a backup held the gate".
func TestJanitor_StaleLockSweepWaitsForTheHeavyGate(t *testing.T) {
heavy := &backup.InFlight{}
swept := 0
d := janitorDeps{staleLocks: func(context.Context) { swept++ }, heavy: heavy, now: time.Now, logger: quiet()}
release, _, ok := heavy.TryAcquire("backup")
if !ok {
t.Fatal("setup")
}
janitorPass(context.Background(), d)
if swept != 0 {
t.Fatal("the sweep ran while a backup held the gate")
}
release()
janitorPass(context.Background(), d)
if swept != 1 {
t.Fatalf("swept %d times with the gate free — want 1", swept)
}
if _, _, ok := heavy.TryAcquire("after"); !ok {
t.Fatal("the sweep did not release the gate")
}
}
+47 -1
View File
@@ -50,6 +50,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/provision"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/restorespace"
"gitea.dooplex.hu/admin/felhom-agent/internal/selfheal"
"gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
@@ -59,7 +60,7 @@ import (
// version is the agent version. Overridable at build time with
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
var version = "0.92.1"
var version = "0.131.0"
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
@@ -897,6 +898,13 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// it finds nothing to flag. The re-mount dispatch is off the poll path (a goroutine).
storageTrigger := make(chan struct{}, 1)
loop.SetTrigger(storageTrigger)
// R-672: a thin pool crossing 90 % requests a report at once (the hub's storage-fill alarm).
observer.SetThinHighTrigger(func() {
select {
case storageTrigger <- struct{}{}:
default:
}
})
remounter := &gateRemounter{gate: gate, ops: hostOps, hostID: cfg.Hub.HostID, logger: logger}
// Drive intent store (slice 10 P3 self-heal): persisted, durable-id-keyed enroll/eject/decommission
// state. Gates the watchdog's self-heal re-mount to ENROLLED drives, and the local API records
@@ -963,6 +971,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
Logger: logger,
})
rtSpace, rtPolicy := restoreSpaceFor(cfg, px, hostOps)
engine := reconcile.NewEngine(reconcile.EngineOptions{
API: px,
Queue: queue,
@@ -971,6 +980,9 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
Gate: gate,
HostID: cfg.Hub.HostID,
Logger: logger,
// R-672: the restore-test's space preflight (nil would refuse every test — fail-closed).
RestoreSpace: rtSpace,
SpacePolicy: rtPolicy,
})
// Crash recovery (doc 03 §10): resolve any op that was in flight when the agent
@@ -1401,8 +1413,22 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
// running" signal, so a deliberately stopped guest is never touched.
go localSrv.WatchGuestPower(ctx)
// R-523: a controller container that is simply not running (killed, stopped, a failed
// self-update) is restarted through its bootstrap unit — nothing else watches it.
collector.SetControllerSupervisorReporter(localSrv)
go localSrv.WatchControllers(ctx)
go func() { errc <- localSrv.Run(ctx) }()
}
// R-672 / R-673: retry a failed restore-test teardown and sweep stale backup locks on a timer, not
// only at start-up (janitor.go).
{
jd := janitorDeps{retryScratch: engine.RetryScratchTeardown, heavy: heavyOps,
record: backupStore.RecordRestoreTest, now: time.Now, logger: logger}
if localSrv != nil {
jd.staleLocks = localSrv.RecoverStaleLockedGuests
}
go runJanitor(ctx, jd)
}
if lanLoop != nil {
lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }()
@@ -1826,6 +1852,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
SmbCredsDir: cfg.Privileged.SmbCredsDir,
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
@@ -2204,6 +2231,17 @@ func formatOrDash(t time.Time) string {
return t.UTC().Format(time.RFC3339)
}
// restoreSpaceFor builds the restore-test's space preflight (R-672): the production provider over the
// Proxmox API + the privileged `lvs` metadata read, and the configured margin.
func restoreSpaceFor(cfg config.Config, px *proxmox.Client, ops *storage.SudoHostOps) (reconcile.RestoreSpace, reconcile.SpacePolicy) {
factor, reserve := cfg.Backup.RestoreTestSpace()
p := &restorespace.Provider{API: px}
if ops != nil {
p.ThinMeta = ops.ThinPoolMetadata
}
return p, reconcile.SpacePolicy{Factor: factor, ReserveBytes: reserve}
}
func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog.Logger, archive string) int {
if err := cfg.Validate(); err != nil {
fmt.Fprintln(os.Stderr, "selftest: proxmox not configured:", err)
@@ -2234,8 +2272,10 @@ func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog
}
}
gate := reconcile.NewGate(nil, cfg.Hub.HostID, reconcile.SlogAudit{Logger: logger}, logger)
rtSpace, rtPolicy := restoreSpaceFor(cfg, px, newHostOps(cfg, logger))
engine := reconcile.NewEngine(reconcile.EngineOptions{
API: px, Queue: queue, Journal: journal, Gate: gate, HostID: cfg.Hub.HostID, Logger: logger,
RestoreSpace: rtSpace, SpacePolicy: rtPolicy,
})
fmt.Printf("=== felhom-agent %s selftest=restore-test ===\n", version)
@@ -2265,10 +2305,16 @@ func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog
RestoreTaskTimeout: restoreTaskTimeout(cfg, rtTier),
})
printJSON("restore-test record", backup.ToHubRestoreTest(res, time.Now().UTC()))
if res.Skipped && res.SkipReason != "" {
fmt.Printf(" space preflight: storage=%s required=%d avail=%d\n", res.TargetStorage, res.RequiredBytes, res.AvailBytes)
fmt.Printf("=== selftest=restore-test SKIPPED — %s ===\n", res.SkipReason)
return 4
}
if res.Skipped {
fmt.Println("=== selftest=restore-test SKIPPED (no free scratch VMID in band) ===")
return 0
}
fmt.Printf(" space preflight passed: storage=%s required=%d avail=%d\n", res.TargetStorage, res.RequiredBytes, res.AvailBytes)
if res.Err != nil || !res.Pass {
fmt.Fprintf(os.Stderr, " [FAIL] restore-test (scratch %d): %v\n", res.ScratchVMID, res.Err)
return 1
+5 -1
View File
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
# view, and the docker socket. The controller reaches the agent's local API for disk management.
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
-v felhom-controller-data:/opt/docker/felhom-controller \
+35 -12
View File
@@ -2,6 +2,7 @@ package backup
import (
"context"
"fmt"
"encoding/json"
"errors"
"io"
@@ -21,14 +22,18 @@ type fakeBackupAPI struct {
vzdumpErr error
waitErr error
cfg proxmox.GuestConfig
goneGuests map[int]bool // R-689: vmids whose config lookup answers "does not exist"
aclGuests map[int]bool // R-689: vmids outside the token's ACL — PVE answers 403 "permission denied"
cfgErr error
content []proxmox.StorageContent
contentErr error
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate)
storageErr error
vzdumps []proxmox.VzdumpOptions
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate) — DEFINITIONS: like
// production's GET /storage, it never carries usage; ListStorage strips Avail/Used (R-685's live lesson)
nodeStorages []proxmox.Storage // returned by NodeStorage (GET /nodes/{node}/storage — WITH usage)
storageErr error
vzdumps []proxmox.VzdumpOptions
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
}
func (f *fakeBackupAPI) Vzdump(_ context.Context, o proxmox.VzdumpOptions) (string, error) {
@@ -41,13 +46,30 @@ func (f *fakeBackupAPI) WaitTask(_ context.Context, _ string, _ proxmox.WaitOpti
}
return proxmox.TaskStatus{Status: "stopped", ExitStatus: "OK"}, f.waitErr
}
func (f *fakeBackupAPI) GuestConfig(_ context.Context, _ int) (proxmox.GuestConfig, error) {
func (f *fakeBackupAPI) GuestConfig(_ context.Context, vmid int) (proxmox.GuestConfig, error) {
if f.aclGuests[vmid] {
return proxmox.GuestConfig{}, fmt.Errorf("proxmox: GET /nodes/n/lxc/%d/config -> HTTP 403: permission denied at /vms/%d (missing privilege VM.Audit)", vmid, vmid)
}
if f.goneGuests[vmid] { // R-689: PVE's answer for a deleted guest
return proxmox.GuestConfig{}, fmt.Errorf("proxmox: GET /nodes/n/lxc/%d/config -> HTTP 500: Configuration file 'nodes/n/lxc/%d.conf' does not exist", vmid, vmid)
}
return f.cfg, f.cfgErr
}
func (f *fakeBackupAPI) StorageContent(_ context.Context, _ string) ([]proxmox.StorageContent, error) {
return f.content, f.contentErr
}
func (f *fakeBackupAPI) ListStorage(_ context.Context) ([]proxmox.Storage, error) {
out := make([]proxmox.Storage, len(f.storages))
for i, s := range f.storages {
s.Avail, s.Used, s.Total = 0, 0, 0 // GET /storage has no usage — a fake that had it hid R-685's defect
out[i] = s
}
return out, f.storageErr
}
func (f *fakeBackupAPI) NodeStorage(_ context.Context) ([]proxmox.Storage, error) {
if f.nodeStorages != nil {
return f.nodeStorages, f.storageErr
}
return f.storages, f.storageErr
}
func (f *fakeBackupAPI) TaskLogTail(_ context.Context, _ string, _ int) ([]string, error) {
@@ -137,13 +159,14 @@ func TestBackup_VzdumpFailureReturnsFailedRecord(t *testing.T) {
func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
const big = 4 << 30 // a plausible whole-guest archive
api := &fakeBackupAPI{content: []proxmox.StorageContent{
{VolID: "a", Content: "backup", CTime: 10, Size: big},
{VolID: "b", Content: "backup", CTime: 99, Size: big},
// R-689: real vzdump names with their vmid — only a backup OF A GUEST is a candidate.
{VolID: "local:backup/vzdump-lxc-9001-a.tar.zst", VMID: 9001, Content: "backup", CTime: 10, Size: big},
{VolID: "local:backup/vzdump-lxc-9001-b.tar.zst", VMID: 9001, Content: "backup", CTime: 99, Size: big},
{VolID: "iso", Content: "iso", CTime: 999, Size: big}, // not a backup → ignored
}}
r := NewBackupRunner(api, "local", "", "", "", quiet())
vol, err := r.PickRestoreCandidate(context.Background())
if err != nil || vol != "b" {
if err != nil || vol != "local:backup/vzdump-lxc-9001-b.tar.zst" {
t.Fatalf("pick = %q,%v want newest 'b'", vol, err)
}
// no backups → "".
@@ -163,12 +186,12 @@ func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
// `pick = "phantom" want the newest COMPLETE archive 'real'`.
func TestPickRestoreCandidate_SkipsImplausibleArchives(t *testing.T) {
api := &fakeBackupAPI{content: []proxmox.StorageContent{
{VolID: "real", Content: "backup", CTime: 10, Size: 4 << 30},
{VolID: "phantom", Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
{VolID: "felhom-pbs:backup/ct/9001/real", VMID: 9001, Content: "backup", CTime: 10, Size: 4 << 30},
{VolID: "felhom-pbs:backup/ct/9001/phantom", VMID: 9001, Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
}}
r := NewBackupRunner(api, "local", "", "", "", quiet())
vol, err := r.PickRestoreCandidate(context.Background())
if err != nil || vol != "real" {
if err != nil || vol != "felhom-pbs:backup/ct/9001/real" {
t.Fatalf("pick = %q,%v want the newest COMPLETE archive 'real'", vol, err)
}
}
+36
View File
@@ -0,0 +1,36 @@
package backup
import (
"context"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-672 (v0.133.0): a restore-test the SPACE preflight refused is the test's RESULT — recorded for the
// hub as pass=false with the reason, never dropped (a band skip still is) and never a pass.
//
// COMPANION RED-PROOF (REPORT): the scheduler's pre-v0.133.0 `if res.Skipped { return }` → "a space
// refusal never reached the host report".
func TestR672_SpaceSkipIsReportedNotDropped(t *testing.T) {
store := NewStore()
rt := &fakeRTRunner{res: reconcile.RestoreTestResult{Archive: "vol", Skipped: true,
SkipReason: "skipped: not enough space on local-lvm: restoring 21.1 GiB (vzdump log) needs 30.3 GiB free, has 21.6 GiB"}}
s := NewScheduler(SchedulerOptions{
Runner: rt, Pick: func(context.Context) (string, error) { return "vol", nil }, Store: store,
Spec: func(context.Context, string) reconcile.RestoreTestSpec {
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009}
},
Cadence: time.Hour, Logger: quiet(),
})
s.tick(context.Background())
got := store.RestoreTests(context.Background())
if len(got) != 1 {
t.Fatalf("a space refusal never reached the host report: %+v", got)
}
if got[0].Pass || !got[0].Skipped || !strings.HasPrefix(got[0].Error, "skipped: not enough space") {
t.Fatalf("record = %+v — want pass=false, skipped, the reason as the error", got[0])
}
}
+83
View File
@@ -0,0 +1,83 @@
package backup
import (
"context"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-685 (v0.134.0) — a whole-box backup that cannot fit its LOCAL target is a named skip BEFORE anything
// starts, with the numbers in the record's Error, never a vzdump that fills the disk and fails.
const gib = int64(1) << 30
func spaceAPI(avail int64, lastArchive int64, typ string) *fakeBackupAPI {
api := &fakeBackupAPI{vzdumpUPID: "UPID:vzdump:1",
storages: []proxmox.Storage{{Storage: "local", Type: typ, Content: "backup", Avail: avail}}}
if lastArchive > 0 {
api.content = []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst",
Content: "backup", VMID: 9201, Size: lastArchive, CTime: 1790280000}}
}
return api
}
// TestR685_BackupThatCannotFitIsSkipped — demo-hp's shape: an 8.2 GB archive, 4 GiB free. No vzdump is
// started; the record says why, with the numbers, under a stable prefix.
//
// COMPANION RED-PROOF (REPORT.md): drop the spaceFits call from backup() — this test fails at "a vzdump
// was started on a target that cannot hold it".
func TestR685_BackupThatCannotFitIsSkipped(t *testing.T) {
api := spaceAPI(4*gib, 8182759056, "dir")
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
rec, err := r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 0 {
t.Fatalf("a vzdump was started on a target that cannot hold it: %+v", api.vzdumps)
}
if err == nil || rec.Success || !strings.HasPrefix(rec.Error, BackupSkipNoSpacePrefix) {
t.Fatalf("want a named skip, got err=%v rec=%+v", err, rec)
}
for _, want := range []string{"4.0 GiB free", "7.6 GiB", "10.5 GiB"} {
if !strings.Contains(rec.Error, want) {
t.Errorf("the reason must carry the numbers (%q missing): %s", want, rec.Error)
}
}
}
// TestR685_BackupThatFitsRuns — the same archive with 16 GiB free (demo-hp after tonight's prune) runs.
func TestR685_BackupThatFitsRuns(t *testing.T) {
api := spaceAPI(16*gib, 8182759056, "dir")
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
_, _ = r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 1 {
t.Fatalf("a backup that fits must run, vzdumps=%d", len(api.vzdumps))
}
}
// TestR685_FailsOpen — never refuse on what is not KNOWN: a PBS target, a first backup (no archive to
// size from), an unknown free figure, or a storage list that cannot be read.
func TestR685_FailsOpen(t *testing.T) {
cases := map[string]*fakeBackupAPI{
"pbs target": spaceAPI(1*gib, 8*gib, "pbs"),
"first backup": spaceAPI(1*gib, 0, "dir"),
"avail unknown": spaceAPI(0, 8*gib, "dir"),
"storage list error": func() *fakeBackupAPI {
a := spaceAPI(1*gib, 8*gib, "dir")
a.storageErr = context.DeadlineExceeded
return a
}(),
}
for name, api := range cases {
target := "local"
if name == "pbs target" {
api.storages[0].Storage = "felhom-pbs"
target = "felhom-pbs"
}
r := NewBackupRunner(api, target, proxmox.ModeSnapshot, "", "", quiet())
_, _ = r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 1 {
t.Errorf("%s: the preflight must fail OPEN, but no vzdump ran", name)
}
}
}
+112
View File
@@ -0,0 +1,112 @@
package backup
import (
"context"
"fmt"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-689 (v0.135.0) — demo-hp keeps its golden template in `local:backup/`. It is content "backup",
// 654 MB and so "plausibly complete", and it was the newest SETTLED entry: the restore test picked it
// every 6 h and failed extractconfig with a 403 (measured 2026-09-24 10:36, 09-25 04:57 and 10:57),
// while the guest's own archive — younger than the 24 h settle — went untested and nothing was proven.
//
// COMPANION RED-PROOF (REPORT.md): drop the guestBackupArchive call from PickSettledRestoreCandidateOn —
// this test then picks `local:backup/felhom-golden-0.236.0.tar.zst`.
func TestR689_TheRestoreTestNeverPicksTheGolden(t *testing.T) {
const day = int64(86400)
now := int64(1790370000) // 2026-09-25 ~19:00Z
api := &fakeBackupAPI{content: []proxmox.StorageContent{
// the guest's real archive, settled (older than the cutoff below)
{VolID: "local:backup/vzdump-lxc-9201-2026_09_22-21_59_25.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 3*day},
// the golden: newer, settled, big, and NOT a backup of a guest
{VolID: "local:backup/felhom-golden-0.236.0.tar.zst", Content: "backup", Size: 654115664, CTime: now - 2*day},
// a hand-copied tarball that PVE happens to attribute to a vmid — the name is not a vzdump's
{VolID: "local:backup/copy-of-9201.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 2*day},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
if err != nil {
t.Fatal(err)
}
if got != "local:backup/vzdump-lxc-9201-2026_09_22-21_59_25.tar.zst" {
t.Fatalf("picked %q — the restore test must prove a backup OF A GUEST", got)
}
}
func TestR689_GuestBackupArchiveShapes(t *testing.T) {
for _, c := range []struct {
e proxmox.StorageContent
ok bool
}{
{proxmox.StorageContent{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst", VMID: 9201}, true},
{proxmox.StorageContent{VolID: "local:backup/vzdump-qemu-300-2026_09_24-21_59_25.vma.zst", VMID: 300}, true},
{proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z", VMID: 9201}, true},
{proxmox.StorageContent{VolID: "felhom-pbs:backup/vm/300/2026-07-28T05:31:14Z", VMID: 300}, true},
{proxmox.StorageContent{VolID: "local:backup/felhom-golden-0.236.0.tar.zst"}, false},
{proxmox.StorageContent{VolID: "local:backup/vzdump-lxc-9201-x.tar.zst", VMID: 9202}, false}, // vmid disagrees with the name
{proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z"}, false}, // no vmid reported
} {
if ok, why := guestBackupArchive(c.e); ok != c.ok {
t.Errorf("%s vmid=%d: ok=%v (%s), want %v", c.e.VolID, c.e.VMID, ok, why, c.ok)
}
}
}
// R-689 (v0.136.0) — the measured demo-hp shape right after v0.135.0: the golden (skipped), a leftover archive
// of guest 9100 deleted in August (settled), and today's archive of 9201 (not settled yet). The pick must be
// NOTHING — never the deleted guest's archive. With 9201's archive settled, that one.
//
// COMPANION RED-PROOF (REPORT.md): drop the known-guest check — the pick is the 9100 leftover.
func TestR689_AnArchiveOfADeletedGuestIsNeverPicked(t *testing.T) {
const day = int64(86400)
now := int64(1790476000)
api := &fakeBackupAPI{goneGuests: map[int]bool{9100: true}, content: []proxmox.StorageContent{
{VolID: "local:backup/felhom-golden-0.236.0.tar.zst", Content: "backup", Size: 654115664, CTime: now - 14*day},
{VolID: "local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst", Content: "backup", VMID: 9100, Size: 656970239, CTime: now - 37*day},
{VolID: "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 7*3600},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
if err != nil || got != "" {
t.Fatalf("picked %q err=%v — a deleted guest's archive proves nothing about this box", got, err)
}
got, _, _ = r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now, 0).UTC())
if got != "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst" {
t.Fatalf("with 9201's archive settled the pick is %q", got)
}
}
// Any OTHER lookup failure is not "the guest is gone": the tier must read UNKNOWN (an error), never
// "nothing to prove".
func TestR689_AGuestLookupFailureIsUnknownNotEmpty(t *testing.T) {
api := &fakeBackupAPI{cfgErr: fmt.Errorf("proxmox: connection refused"), content: []proxmox.StorageContent{
{VolID: "local:backup/vzdump-lxc-9201-x.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: 10},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Time{}); err == nil {
t.Fatal("a failed guest lookup read as a clean answer")
}
}
// v0.137.0 — THE MEASURED ANSWER: the agent's token sees only its pool, so for the deleted guest PVE says 403
// "permission denied at /vms/9100", not "does not exist" (demo-hp, right after v0.136.0 — the local tier read
// UNKNOWN). Such a guest is not one this agent manages: its archive is skipped, the tier is not an error.
//
// COMPANION RED-PROOF (REPORT.md): drop the "permission denied" case — the pick errors.
func TestR689_AGuestOutsideTheAgentsACLIsNotAKnownGuest(t *testing.T) {
const day = int64(86400)
now := int64(1790476000)
api := &fakeBackupAPI{aclGuests: map[int]bool{9100: true}, content: []proxmox.StorageContent{
{VolID: "local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst", Content: "backup", VMID: 9100, Size: 656970239, CTime: now - 37*day},
{VolID: "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 7*3600},
}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
if err != nil || got != "" {
t.Fatalf("picked %q err=%v — want nothing and no error (the only settled archive is not ours)", got, err)
}
}
+74
View File
@@ -0,0 +1,74 @@
package backup
import (
"context"
"errors"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
const (
thisBoxKey = "de:51:7a:18:cb:39:22:30:2c:84:f5:8b:d1:91:4b:7e:81:bb:69:b8:89:0f:57:ac:d3:59:e1:1a:62:25:11:2c"
earlierBox1 = "6b:ca:5f:3f:ca:0f:e2:3f:fb:24:62:89:bf:e7:64:59:9a:41:c5:e6:e3:9f:3f:5f:e1:71:7b:a1:9d:24:67:82"
earlierBox2 = "fe:3d:db:95:d4:df:ab:e1:7d:4a:89:fa:2b:07:53:6a:e4:d2:85:95:d1:90:27:4b:d9:c6:92:20:95:04:e5:d4"
)
// R-727 (v0.138.0) — the 2026-09-30 shape, measured on a fresh box for a returning customer: the PBS
// namespace held two archives of earlier boxes (same guest 9201, same token) and this box's own, which was not
// settled yet. The old picker chose the earlier box's newest settled archive and failed `wrong key`.
// The CONSEQUENCE asserted: no archive of another box is ever picked; with this box's archive settled it is picked.
// COMPANION RED-PROOF: remove the `ownKey != "" && !EqualFold(...)` skip → the first case picks 2026-09-16T21:59:54Z.
func TestR727_TheRestoreTestTakesOnlyThisBoxsArchives(t *testing.T) {
day := int64(86400)
now := int64(1790740000) // 2026-09-30 ~04:00Z
own := proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
content: []proxmox.StorageContent{
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
own,
},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
// 1. The night of 2026-09-30: this box's own archive is ~6 h old, not settled (cutoff 24 h) — nothing to prove.
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Unix(now-day, 0).UTC())
if err != nil {
t.Fatal(err)
}
if got != "" {
t.Fatalf("picked %q — an archive of ANOTHER box is never this box's proof (R-727)", got)
}
// 2. A day later this box's own archive is settled — it is the one picked.
got, _, err = r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Unix(now+day, 0).UTC())
if err != nil {
t.Fatal(err)
}
if got != own.VolID {
t.Fatalf("picked %q, want this box's own %q", got, own.VolID)
}
}
// An unencrypted storage (a local dir) holds only this box's vzdumps — no key filter applies.
func TestR727_UnencryptedStorageIsNotFiltered(t *testing.T) {
api := &fakeBackupAPI{
storages: []proxmox.Storage{{Storage: "local", Type: "dir"}},
content: []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_29-21_27_05.tar.zst", Content: "backup", VMID: 9201, Size: 955425507, CTime: 1790710025}},
}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
if got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Time{}); err != nil || got == "" {
t.Fatalf("got %q err %v", got, err)
}
}
// A storage-list failure makes the tier UNKNOWN (an error), never "nothing to prove".
func TestR727_KeyLookupFailureIsUnknown(t *testing.T) {
api := &fakeBackupAPI{storageErr: errors.New("proxmox: GET /storage -> HTTP 500"), content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/x", Content: "backup", VMID: 9201}}}
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err == nil {
t.Fatal("a failed key lookup must surface as an error (tier UNKNOWN)")
}
}
+185
View File
@@ -5,6 +5,7 @@ import (
"fmt"
"log/slog"
"sort"
"strconv"
"strings"
"sync"
"time"
@@ -22,6 +23,10 @@ type BackupAPI interface {
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
// ListStorage enumerates storages (name+type) — used to scope local-only retention (never prune PBS).
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
// NodeStorage is GET /nodes/{node}/storage — the storages WITH live usage (avail/used). R-685's space
// preflight reads free space HERE: ListStorage (GET /storage) is the cluster DEFINITIONS and carries no
// usage at all — measured live 2026-09-24 on demo-hp, where reading it let a backup through.
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
// TaskLogTail reads trailing task-log lines — used to read the ACTUAL vzdump mode
// (PVE may downgrade a requested snapshot to stop for a stopped guest — spike B1).
TaskLogTail(ctx context.Context, upid string, limit int) ([]string, error)
@@ -170,6 +175,16 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
rec.UncoveredVolumes = []string{}
}
// R-685 (v0.134.0): will the new archive FIT on a local target? Asked before anything runs, so a
// target that cannot hold it is a named SKIP with the numbers, not a nightly "No space left on
// device" that only the vzdump log explains (demo-hp, every night from 2026-09-23 — R-684).
if ok, why := r.spaceFits(ctx, vmid); !ok {
rec.Error = BackupSkipNoSpacePrefix + why
rec.DurationSeconds = time.Since(start).Seconds()
r.logger.Warn("backup SKIPPED by the space preflight (R-685) — nothing was started", "vmid", vmid, "target", r.target, "reason", why)
return rec, fmt.Errorf("backup: %s", rec.Error)
}
upid, err := r.api.Vzdump(ctx, proxmox.VzdumpOptions{
VMID: vmid, Storage: r.target, Mode: r.mode, Notes: r.notes,
PruneBackups: r.localPruneSpec(ctx), // local target → keep-last=N; PBS/unknown → "" (no prune)
@@ -217,6 +232,56 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
return rec, nil
}
// BackupSkipNoSpacePrefix starts a backup record's Error when the space preflight refused (R-685) — a
// stable prefix the controller's page and the hub can key on.
const BackupSkipNoSpacePrefix = "skipped: not enough space: "
// Space preflight margins (R-685): the new archive is predicted as the newest archive of this guest on the
// target × backupSpaceGrowth, plus backupSpaceFloorBytes of headroom for the host. MEASURED 2026-09-24:
// demo-hp 9201's archives grew 5.8 → 6.2 → 6.9 → 7.6 GB in four nights (+10 % a night at worst), so 1.25
// covers two nights' growth. PVE prunes old archives only AFTER a successful backup, so the free space
// must hold the new archive while every kept one still exists.
const (
backupSpaceGrowth = 1.25
backupSpaceFloorBytes = int64(1) << 30
)
// spaceFits answers whether a new archive of vmid fits on a LOCAL (non-PBS) target. It FAILS OPEN — a
// backup is the thing being protected, so an unreadable storage, an unknown type or a first backup (no
// previous archive to size from) proceeds and says so; only a POSITIVE "it does not fit" refuses.
func (r *BackupRunner) spaceFits(ctx context.Context, vmid int) (bool, string) {
// NodeStorage, never ListStorage: only the node view carries avail (see BackupAPI.NodeStorage).
stores, err := r.api.NodeStorage(ctx)
if err != nil {
r.logger.Warn("backup: space preflight could not read storage usage — proceeding (fail-open)", "target", r.target, "err", err)
return true, ""
}
var st *proxmox.Storage
for i := range stores {
if stores[i].Storage == r.target {
st = &stores[i]
break
}
}
if st == nil || st.Type == "pbs" || st.Avail <= 0 {
return true, "" // PBS dedups and has its own lifecycle; an unknown avail never refuses
}
_, last, err := r.latestArchive(ctx, vmid)
if err != nil || last <= 0 {
r.logger.Info("backup: space preflight has no previous archive to size from — proceeding", "vmid", vmid, "target", r.target)
return true, ""
}
need := int64(float64(last)*backupSpaceGrowth) + backupSpaceFloorBytes
if st.Avail >= need {
r.logger.Info("backup: space preflight passed", "vmid", vmid, "target", r.target, "last_archive_bytes", last, "need_bytes", need, "avail_bytes", st.Avail)
return true, ""
}
return false, fmt.Sprintf("%s has %s free; the last archive of guest %d was %s, so a new one needs about %s (old archives are removed only after a successful backup)",
r.target, humanGiB(st.Avail), vmid, humanGiB(last), humanGiB(need))
}
func humanGiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
// watchForSnapshot polls the running backup's task log until it sees the storage-snapshot marker
// (→ onSnapshot once) or the requested mode is reported as `stop` (→ downgraded; the marker will
// never come, so stop watching) or ctx is cancelled (backup finished). Best-effort: a log-read
@@ -292,12 +357,58 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
if err != nil {
return "", time.Time{}, err
}
// R-727 (v0.138.0): on an ENCRYPTED storage, only archives written with THIS storage's key are this box's.
// Measured 2026-09-30 on a fresh box for a returning customer: the PBS namespace still held two archives
// of earlier boxes (same guest id 9201, same token), the newest settled one was an earlier box's, and the
// test failed `wrong key` every evaluation. The archive carries no host id; its key fingerprint is the
// discriminator (PVE's content `encrypted`, the storage's `encryption-key`). A lookup failure returns
// the error — the tier reads UNKNOWN, never "nothing to prove".
ownKey, err := r.storageKeyFingerprint(ctx, target)
if err != nil {
return "", time.Time{}, fmt.Errorf("reading the key fingerprint of storage %s: %w", target, err)
}
var best string
var bestCTime int64 = -1
known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick)
for _, e := range contents {
if e.Content != "backup" {
continue
}
// R-689 (v0.135.0): only a backup OF A GUEST is a restore-test candidate. demo-hp keeps its golden
// template in `local:backup/` — content "backup", 654 MB, plausibly complete — and it was picked as
// the newest settled archive every 6 h and failed extractconfig (403) each time, while the guest's
// real archive went untested.
if ok, why := guestBackupArchive(e); !ok {
r.noteNotAGuestBackupOnce(e, why)
continue
}
if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey)))
continue
}
// R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after
// v0.135.0: with the golden skipped, the pick fell to `vzdump-lxc-9100-2026_08_21…`, a leftover of a
// guest deleted in August — proving nothing about any guest this box runs. "Does not exist" skips the
// archive; any OTHER lookup failure is returned, so the tier reads UNKNOWN, never "nothing to prove".
if _, seen := known[e.VMID]; !seen {
_, err := r.api.GuestConfig(ctx, e.VMID)
switch {
case err == nil:
known[e.VMID] = true
case strings.Contains(err.Error(), "does not exist"), strings.Contains(err.Error(), "permission denied"):
// v0.137.0: PVE answers 403 "permission denied at /vms/<id>" — not "does not exist" — for a guest
// outside the agent's ACL (the `felhom` pool). Measured on demo-hp after v0.136.0: the deleted
// guest 9100's archive made the local tier UNKNOWN every evaluation. A guest the agent cannot
// read is not one it manages; its archive is not a candidate.
known[e.VMID] = false
default:
return "", time.Time{}, fmt.Errorf("checking whether guest %d still exists: %w", e.VMID, err)
}
}
if !known[e.VMID] {
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("guest %d does not exist on this node or is not one this agent manages", e.VMID))
continue
}
if !notAfter.IsZero() && e.CTime > notAfter.Unix() {
continue // not settled yet — a newer archive is not a reason to re-prove an older one
}
@@ -520,5 +631,79 @@ func ToHubRestoreTest(res reconcile.RestoreTestResult, testedAt time.Time) hub.R
if res.Err != nil {
rt.Error = res.Err.Error()
}
if res.SkipReason != "" { // R-672: the space preflight refused — reported, never a pass
rt.Pass = false
rt.Skipped = true
rt.Error = res.SkipReason
}
return rt
}
// guestBackupArchive reports whether a storage entry is a whole-guest backup of a known guest — a
// `vzdump-<type>-<vmid>-…` file on a dir storage, or a `backup/{ct,vm}/<vmid>/<time>` snapshot on a PBS
// datastore — whose vmid the storage itself reports. Anything else in a backup content type (a golden
// template, a hand-copied tarball) is not a backup of a guest and is never restore-tested (R-689).
// Pure, so the rule is unit-tested without a storage.
func guestBackupArchive(e proxmox.StorageContent) (bool, string) {
if e.VMID <= 0 {
return false, "not a backup of a guest (the storage reports no vmid)"
}
vol := e.VolID
if i := strings.Index(vol, ":"); i >= 0 {
vol = vol[i+1:]
}
vol = strings.TrimPrefix(vol, "backup/")
vmid := strconv.Itoa(e.VMID)
switch {
case strings.HasPrefix(vol, "vzdump-lxc-"+vmid+"-"), strings.HasPrefix(vol, "vzdump-qemu-"+vmid+"-"):
return true, ""
case strings.HasPrefix(vol, "ct/"+vmid+"/"), strings.HasPrefix(vol, "vm/"+vmid+"/"):
return true, ""
}
return false, "not a vzdump archive or a PBS snapshot of guest " + vmid
}
// noteNotAGuestBackupOnce logs, once per volid, that a backup-content entry is not a restore-test
// candidate because it is not a backup of a guest (R-689). INFO, not WARN: a golden template kept in
// the backup directory is the operator's, and not a fault.
func (r *BackupRunner) noteNotAGuestBackupOnce(e proxmox.StorageContent, why string) {
r.rejectedMu.Lock()
if r.rejected == nil {
r.rejected = map[string]struct{}{}
}
_, seen := r.rejected[e.VolID]
if !seen {
r.rejected[e.VolID] = struct{}{}
}
r.rejectedMu.Unlock()
if !seen {
r.logger.Info("backup: restore-test skips an entry that is not a backup of a guest",
"target", r.target, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
}
}
// storageKeyFingerprint returns the named storage's client-side encryption key fingerprint ("" when the
// storage is not encrypted — a local dir holds only this box's own vzdumps).
func (r *BackupRunner) storageKeyFingerprint(ctx context.Context, target string) (string, error) {
sts, err := r.api.ListStorage(ctx)
if err != nil {
return "", err
}
for _, st := range sts {
if st.Storage == target {
return strings.TrimSpace(st.EncryptionKey), nil
}
}
return "", nil
}
// shortFP is the first 8 bytes of a key fingerprint, for a log line.
func shortFP(fp string) string {
if fp == "" {
return "none"
}
if len(fp) > 23 {
return fp[:23] + "…"
}
return fp
}
+3 -1
View File
@@ -208,9 +208,11 @@ func (s *Scheduler) tick(ctx context.Context) {
spec := s.spec(ctx, archive)
spec.Archive = archive
res := s.runner.RunRestoreTest(ctx, spec)
if res.Skipped {
if res.Skipped && res.SkipReason == "" {
return // already logged by the engine (no free scratch VMID)
}
// R-672: a SPACE refusal is the test's result — reported (pass=false, the reason as the error),
// never dropped and never a pass. It earns no rotation credit, so the tier stays due.
rt := ToHubRestoreTest(res, s.now())
s.store.RecordRestoreTest(rt)
// Rotation credit is given ONLY on success. A failing tier must keep sorting first, or a tier
+20
View File
@@ -345,6 +345,12 @@ type BackupConfig struct {
// between an archive settling and its proof, and the retry rate of a tier whose restore-test
// keeps failing. See defaultRestoreTestEvalInterval for the measurement it was chosen from.
RestoreTestEvalIntervalSeconds int `json:"restore_test_eval_interval_seconds"`
// RestoreTestSpaceFactor / RestoreTestSpaceReserveGiB are the restore-test's space margin (R-672,
// v0.133.0): a test starts only when the target storage has free ≥ restored × factor + reserve,
// `restored` being the UNCOMPRESSED size. 0/unset → 1.2 and 5 GiB. A test config may raise them to
// watch the refusal (the brief's live case a).
RestoreTestSpaceFactor float64 `json:"restore_test_space_factor,omitempty"`
RestoreTestSpaceReserveGiB float64 `json:"restore_test_space_reserve_gib,omitempty"`
// RestoreTestSettleSeconds is how long an archive must have sat on its tier before it is a
// restore-test candidate (R-86); 0 → default (24h), negative → 0 (no settle requirement).
// Restore-testing an archive a backup is still writing proves nothing about the backup that
@@ -615,6 +621,20 @@ func (b BackupConfig) RestoreTestEvalInterval() time.Duration {
}
}
// RestoreTestSpace returns the restore-test's space margin (R-672): factor (≥ 1) and reserve bytes.
// Unset or out-of-range → 1.2 and 5 GiB.
func (b BackupConfig) RestoreTestSpace() (factor float64, reserveBytes int64) {
factor = b.RestoreTestSpaceFactor
if factor < 1 {
factor = 1.2
}
reserveBytes = int64(b.RestoreTestSpaceReserveGiB * float64(1<<30))
if reserveBytes <= 0 {
reserveBytes = 5 << 30
}
return factor, reserveBytes
}
// RestoreTestSettle returns how long an archive must have sat before it is a restore-test
// candidate (R-86): a positive value as-is, negative → 0 (no settle requirement), 0 → the default.
//
+59
View File
@@ -0,0 +1,59 @@
// Package httpx holds the one HTTP-transport default this repo may not lose.
//
// Every client here pins TLS — PBS and PVE by leaf-cert SHA-256, the hub by an optional CA file —
// so none of them can use http.DefaultTransport and each hand-rolls its own. Hand-rolling silently
// discards DefaultTransport's settings, and one of them is load-bearing:
//
// Transport: &http.Transport{TLSClientConfig: tlsCfg} // IdleConnTimeout == 0 == NO timeout
//
// A zero IdleConnTimeout means idle keep-alive connections are retained FOREVER, not "use a sane
// default". Combined with a client that is rebuilt on a schedule and dropped (pbsTargetsFromPVE
// builds a fresh pbs.Client per cycle), every cycle strands one connection that nothing will ever
// close: the abandoned Transport becomes unreachable but its persistConn read-loop goroutine keeps
// the socket alive, and an unreachable Transport does not close its connections.
//
// Measured cost, live: 388 established connections accumulated on ep0's PBS proxy between
// 2026-08-18 09:51:22Z and 2026-08-20 08:02:13Z — 194 from each of the two boxes, held open on BOTH
// sides, one per agent poll cycle, on a proxy whose descriptor ceiling is 65536. See R-344 and
// felhom.eu/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md.
package httpx
import (
"crypto/tls"
"net/http"
"time"
)
// DefaultIdleConnTimeout is how long an idle keep-alive connection is retained before it is closed.
//
// It is 90s because that is http.DefaultTransport's own value: the fix for R-344 restores a
// standard-library default rather than inventing a number, so there is nothing here to tune and
// nothing to justify. It is comfortably shorter than every cadence that drives these clients (the
// 15-minute live-snapshot collect and the 6-hour verify loop), so a connection abandoned by one
// cycle is closed long before the next.
const DefaultIdleConnTimeout = 90 * time.Second
// NewTransport builds a FRESH *http.Transport pinned to tlsCfg, with the idle-connection timeout
// applied.
//
// Fresh, never shared: each caller pins a different endpoint, and a shared transport would pool
// connections across differently pinned servers. Reusing http.DefaultTransport for the same reason
// is not an option — it would drop the pin entirely.
//
// idleConnTimeout <= 0 means USE THE DEFAULT. It deliberately does not mean "no timeout": no-timeout
// is the bug this package exists to prevent, and an unset field must never be able to reintroduce
// it. Callers pass their configured value straight through; only tests pass a short one.
//
// Only IdleConnTimeout is set. The other DefaultTransport settings this transport also lacks
// (MaxIdleConns, TLSHandshakeTimeout, ExpectContinueTimeout) are deliberately left alone: none of
// them accumulates anything, every client bounds its whole request with http.Client.Timeout, and
// widening the change would have made the R-344 measurement unattributable.
func NewTransport(tlsCfg *tls.Config, idleConnTimeout time.Duration) *http.Transport {
if idleConnTimeout <= 0 {
idleConnTimeout = DefaultIdleConnTimeout
}
return &http.Transport{
TLSClientConfig: tlsCfg,
IdleConnTimeout: idleConnTimeout,
}
}
+74
View File
@@ -0,0 +1,74 @@
package httpx
import (
"crypto/tls"
"net/http"
"testing"
"time"
)
// TestNewTransport_ZeroMeansDefaultNeverForever is the whole point of this package.
//
// http.Transport's zero IdleConnTimeout means "retain idle connections FOREVER". Any code path that
// can reach that zero reintroduces R-344, so an unset, zero or negative value must all land on the
// default. If someone later "simplifies" NewTransport by passing the argument straight through,
// this fails.
func TestNewTransport_ZeroMeansDefaultNeverForever(t *testing.T) {
for _, tc := range []struct {
name string
in time.Duration
want time.Duration
}{
{"zero", 0, DefaultIdleConnTimeout},
{"negative", -time.Hour, DefaultIdleConnTimeout},
{"explicit short value (tests)", 50 * time.Millisecond, 50 * time.Millisecond},
{"explicit long value", time.Hour, time.Hour},
} {
t.Run(tc.name, func(t *testing.T) {
got := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, tc.in).IdleConnTimeout
if got != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v", got, tc.want)
}
if got == 0 {
t.Fatal("IdleConnTimeout is 0 — that is 'never expire', which is the R-344 defect itself")
}
})
}
}
// TestDefaultIdleConnTimeout_MatchesTheStandardLibrary pins the number to its justification.
//
// 90s is not a tuned value; it is what http.DefaultTransport uses. Reading it off the standard
// library rather than hardcoding 90 means the constant cannot drift away from the reason given for
// it in the package doc.
func TestDefaultIdleConnTimeout_MatchesTheStandardLibrary(t *testing.T) {
std, ok := http.DefaultTransport.(*http.Transport)
if !ok {
t.Skip("http.DefaultTransport is not an *http.Transport in this Go build")
}
if DefaultIdleConnTimeout != std.IdleConnTimeout {
t.Fatalf("DefaultIdleConnTimeout = %v but http.DefaultTransport uses %v — the doc comment's justification no longer holds",
DefaultIdleConnTimeout, std.IdleConnTimeout)
}
}
// TestNewTransport_IsFreshEveryCall guards the pooling property the pinning relies on.
//
// Each caller pins a DIFFERENT endpoint. A shared transport would pool connections across
// differently pinned servers, so returning a package-level singleton would be a security change
// dressed as a tidy-up.
func TestNewTransport_IsFreshEveryCall(t *testing.T) {
a := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
b := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
if a == b {
t.Fatal("NewTransport returned the SAME transport twice — connections would be pooled across differently pinned endpoints")
}
}
// TestNewTransport_KeepsTheTLSConfig — the transport gains a field; it must lose nothing.
func TestNewTransport_KeepsTheTLSConfig(t *testing.T) {
cfg := &tls.Config{MinVersion: tls.VersionTLS12, InsecureSkipVerify: true} //nolint:gosec // test only
if got := NewTransport(cfg, 0).TLSClientConfig; got != cfg {
t.Fatalf("TLSClientConfig = %p, want the config passed in (%p) — the pin would be dropped", got, cfg)
}
}
+7 -2
View File
@@ -15,6 +15,8 @@ import (
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
const reportPath = "/api/v1/host-report"
@@ -49,8 +51,11 @@ func NewClient(cfg config.HubConfig, logger *slog.Logger) (*Client, error) {
tlsCfg.RootCAs = pool
}
hc := &http.Client{
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
// R-344, consistency only: this client is built ONCE per process, so it never accumulated
// and contributed nothing to the ep0 leak. It carried the same missing default, which over
// a tunnel is how one idle connection survives long enough to fail on next use.
Transport: httpx.NewTransport(tlsCfg, 0),
}
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
}
+16
View File
@@ -103,6 +103,7 @@ type Collector struct {
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
return c
}
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
type ControllerSupervisorReporter interface {
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
}
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
c.ctrlSup = r
return c
}
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
// omitted). Returns the collector for chaining.
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.pbsdr != nil {
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
}
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
if c.ctrlSup != nil {
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
}
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
+37
View File
@@ -124,6 +124,39 @@ type HostReport struct {
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
OOB *OOBStatus `json:"oob,omitempty"`
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
// how many times the agent restarted a dead controller, when last and why, whether it gave up
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
}
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
type ControllerSupervisorStatus struct {
Guests []ControllerSupervisorGuest `json:"guests"`
}
// ControllerSupervisorGuest is one supervised guest.
type ControllerSupervisorGuest struct {
VMID int `json:"vmid"`
RestartsTotal int `json:"restarts_total"`
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
LastReason string `json:"last_reason,omitempty"`
Crashloop bool `json:"crashloop"`
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
Parked bool `json:"parked"`
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
Restarts24h int `json:"restarts_24h"`
SlowCrashloop bool `json:"slow_crashloop"`
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
}
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
@@ -437,6 +470,10 @@ type RestoreTest struct {
// mount layout, not just booted. Additive — a hub that predates them ignores the unknown keys.
MountParity string `json:"mount_parity,omitempty"`
MountInventory []string `json:"mount_inventory,omitempty"`
// Skipped (R-672, v0.133.0): the test did NOT run — the space preflight refused, and Error says
// why ("skipped: not enough space on …"). Pass is false. A hub that predates the key reads a
// failed test with that error, which is the honest reading.
Skipped bool `json:"skipped,omitempty"`
}
// PBSSnapshot is one PBS (offsite) snapshot's inventory + integrity state (doc 03 §8, slice
@@ -0,0 +1,133 @@
package localapi
import (
"context"
"encoding/json"
"io"
"log/slog"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
t.Helper()
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
return resp.Data.Tiers
}
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
Storage: fakeStorage{targets: targets},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
st := &fakeStore{backups: []hub.Backup{
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
}}
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
tiers := tierStatesOf(t, srv)
if len(tiers) != 2 {
t.Fatalf("want 2 tiers, got %+v", tiers)
}
local, pbs := tiers[0], tiers[1]
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
t.Fatalf("local tier lost its successful backup: %+v", local)
}
if pbs.LastSuccess != nil {
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
}
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
}
if pbs.Storage != StoragePresenceAbsent {
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
}
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &resp)
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
}
}
// /backup/tiers advertises storage presence, tri-state.
func TestBackupTiers_StoragePresence(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
var resp struct {
Data BackupTiersResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatal(err)
}
got := map[string]string{}
for _, ti := range resp.Data.Tiers {
got[ti.Target] = ti.Storage
}
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
t.Fatalf("storage presence wrong: %v", got)
}
// An unreadable storage view is "unknown", never "absent".
srv.storage = tierErrStorage{}
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
t.Fatalf("unreadable storage view must be unknown, got %q", p)
}
}
type tierErrStorage struct{}
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
return nil, context.DeadlineExceeded
}
// A targeted request keeps the pre-R-517 bytes (no tiers array).
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
}
}
func containsKey(b []byte, key string) bool {
var m struct {
Data map[string]json.RawMessage `json:"data"`
}
_ = json.Unmarshal(b, &m)
_, ok := m.Data[key]
return ok
}
+453
View File
@@ -0,0 +1,453 @@
package localapi
import (
"context"
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-523 — the in-guest controller supervisor.
//
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
// dashboard answered 502 for 33 minutes until the box was power-cycled.
//
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
// No new privilege.
//
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
// again after 30 minutes.
//
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
// record) can never read as a new restart.
const (
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
// so a killed controller is restarted 30–60 s after it died.
controllerSupervisorInterval = 30 * time.Second
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
controllerSupervisorConfirm = 2
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
// controllerCrashloopPause.
controllerCrashloopMax = 3
controllerCrashloopWindow = 15 * time.Minute
controllerCrashloopPause = 30 * time.Minute
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
// watchdog is indistinguishable from a dead one (standing rule 3).
controllerSupervisorHeartbeatEvery = 20
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
// the controller container, and a customer inside the guest cannot park the supervisor.
ControllerParkedMarker = "controller-parked"
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
// R-539 (operator ruling 3 of 2026-09-16) — the SLOW crash loop. The 3-in-15-minutes brake above
// cannot see a controller that dies every 20 minutes: no two restarts share its window, so it is
// restarted for ever and the only trace is an info event that mails nobody (measured 2026-09-16,
// R-531). A second counter over 24 hours raises a WARNING at the fifth restart. It does NOT stop
// restarting — the fast brake stays the only brake, unchanged. Every restart the supervisor
// performs counts, including one that follows a deliberate operator `docker kill` (measured
// 2026-09-15: the supervisor cannot tell a kill from a crash, and a controller that is killed five
// times a day is worth a line to the operator either way).
controllerSlowCrashloopWindow = 24 * time.Hour
controllerSlowCrashloopMax = 5
// controllerSlowCounterFile holds the 24-hour restart times and the last raise, per guest, beside
// the parked marker.
controllerSlowCounterFile = "controller-restarts-24h.json"
)
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
type controllerSupState struct {
notRunningSeen int
restarts []time.Time // restart times inside the crash-loop window (pruned)
restartsTotal int
lastRestartAt time.Time
lastReason string
crashloopSince time.Time // zero = not in a crash-loop pause
parked bool
// R-539 — the slow counter. PERSISTED, unlike everything above, and the precedent's reason does not
// apply to it: persisting the fast record could carry a stale "give up" across the restart that
// fixed it, but this record never gives anything up — it only warns. Losing it on an agent restart,
// on the other hand, would hide exactly the box it exists for (one whose agent restarts too).
restarts24h []time.Time
slowCrashloopSince time.Time // the last raise; kept after it ages out, the hub keys on it MOVING
}
type controllerSupervisor struct {
mu sync.Mutex
guests map[int]*controllerSupState
sweeps int
}
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
// (staleLock) or the guest executor is not wired.
func (s *Server) WatchControllers(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
return
}
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
"crashloop_window", controllerCrashloopWindow.String(),
"slow_crashloop_max", controllerSlowCrashloopMax, "slow_crashloop_window", controllerSlowCrashloopWindow.String(),
"guests_dir", s.guestsStateDir())
t := time.NewTicker(controllerSupervisorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s.ControllerSupervisorTick(ctx)
}
}
}
func (s *Server) guestsStateDir() string {
if s.guestsDir != "" {
return s.guestsDir
}
return defaultGuestsStateDir
}
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
// not stat the file within.
func (s *Server) provisionedGuest(vmid int) bool {
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
return err == nil && fi.IsDir()
}
func (s *Server) controllerParked(vmid int) bool {
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return err == nil
}
func (s *Server) supState(vmid int) *controllerSupState {
if s.ctrlSup.guests == nil {
s.ctrlSup.guests = map[int]*controllerSupState{}
}
st := s.ctrlSup.guests[vmid]
if st == nil {
st = &controllerSupState{}
s.loadSlowCounter(vmid, st)
s.ctrlSup.guests[vmid] = st
}
return st
}
// slowCounterRecord is the on-disk shape of the R-539 counter.
type slowCounterRecord struct {
Restarts []time.Time `json:"restarts"`
SlowCrashloopSince time.Time `json:"slow_crashloop_since,omitempty"`
}
func (s *Server) slowCounterPath(vmid int) string {
return filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), controllerSlowCounterFile)
}
// loadSlowCounter restores the persisted counter into a fresh state. Absent = a clean start; unreadable
// or corrupt = a clean start with a WARN (a warning counter must never block supervision).
func (s *Server) loadSlowCounter(vmid int, st *controllerSupState) {
b, err := os.ReadFile(s.slowCounterPath(vmid))
if err != nil {
if !os.IsNotExist(err) {
s.logger.Warn("controller-supervisor: slow counter unreadable — starting it from zero", "vmid", vmid, "err", err)
}
return
}
var rec slowCounterRecord
if err := json.Unmarshal(b, &rec); err != nil {
s.logger.Warn("controller-supervisor: slow counter corrupt — starting it from zero", "vmid", vmid, "err", err)
return
}
st.restarts24h = pruneBefore(rec.Restarts, s.clock().Add(-controllerSlowCrashloopWindow))
st.slowCrashloopSince = rec.SlowCrashloopSince
if len(st.restarts24h) > 0 || !st.slowCrashloopSince.IsZero() {
s.logger.Info("controller-supervisor: slow counter restored from disk", "vmid", vmid,
"restarts_24h", len(st.restarts24h), "slow_crashloop_since", st.slowCrashloopSince.Format(time.RFC3339))
}
}
// saveSlowCounter writes the counter atomically (tmp + rename, 0600). A failure is logged and the
// in-memory counter carries on — the next restart retries the write.
func (s *Server) saveSlowCounter(vmid int, rec slowCounterRecord) {
path := s.slowCounterPath(vmid)
b, err := json.Marshal(rec)
if err == nil {
tmp := path + ".tmp"
if err = os.WriteFile(tmp, b, 0o600); err == nil {
err = os.Rename(tmp, path)
}
}
if err != nil {
s.logger.Warn("controller-supervisor: could not persist the slow counter (kept in memory)", "vmid", vmid, "path", path, "err", err)
}
}
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
// cycle without waiting on the ticker.
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
return
}
guests, err := s.staleLock.Guests(ctx)
if err != nil {
// Ownership unproven ⇒ touch nothing (the guest-power rule).
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
return
}
var evaluated, down int
for _, g := range guests {
if ctx.Err() != nil {
return
}
if !s.provisionedGuest(g.VMID) {
continue
}
evaluated++
if !s.superviseOneController(ctx, g.VMID, g.Status) {
down++
}
}
s.ctrlSup.mu.Lock()
s.ctrlSup.sweeps++
sweeps := s.ctrlSup.sweeps
s.ctrlSup.mu.Unlock()
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
"guests_evaluated", evaluated, "controllers_not_running", down)
}
}
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
// known=false means the question could not be answered (pct exec failed for a reason other than a
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
// running": `docker rm` of the controller is the same outage as a kill.
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
if err != nil {
msg := strings.ToLower(err.Error() + " " + out)
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
return false, true, "absent"
}
return false, false, ""
}
status = strings.TrimSpace(out)
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
return status == "running" || status == "restarting", true, status
}
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
// allows. Returns false when the controller was observed not running.
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
now := s.clock()
if guestStatus != "running" {
s.resetNotRunning(vmid)
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
}
running, known, status := s.controllerRunning(ctx, vmid)
if !known {
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
s.resetNotRunning(vmid)
return true
}
parked := s.controllerParked(vmid)
s.ctrlSup.mu.Lock()
st := s.supState(vmid)
st.parked = parked
if running {
st.notRunningSeen = 0
s.ctrlSup.mu.Unlock()
return true
}
st.notRunningSeen++
seen := st.notRunningSeen
s.ctrlSup.mu.Unlock()
if parked {
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return false
}
s.swapMu.Lock()
swapping := s.swapInFlight[vmid]
s.swapMu.Unlock()
if swapping {
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
"vmid", vmid, "status", status)
s.resetNotRunning(vmid)
return false
}
if seen < controllerSupervisorConfirm {
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
return false
}
lock, _, err := s.staleLock.Lock(ctx, vmid)
if err != nil {
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
return false
}
if lock != "" {
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
return false
}
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
"vmid", vmid, "backup_running", busy, "err", berr)
return false
}
// Backoff.
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
if !st.crashloopSince.IsZero() {
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
s.ctrlSup.mu.Unlock()
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
return false
}
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
// crash-loop (the hub keys on it moving, not on it clearing).
st.restarts = nil
st.crashloopSince = time.Time{}
}
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
if len(st.restarts) >= controllerCrashloopMax {
st.crashloopSince = now
n := len(st.restarts)
s.ctrlSup.mu.Unlock()
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
return false
}
s.ctrlSup.mu.Unlock()
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
"vmid", vmid, "status", status, "unit", bootstrapUnit)
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
reason += "; restart FAILED: " + err.Error()
}
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
st.restarts = append(st.restarts, now)
st.restartsTotal++
st.lastRestartAt = now
st.lastReason = reason
st.notRunningSeen = 0
// R-539: the slow counter. Raise at most once per 24 hours — the hub mails on the raise MOVING.
st.restarts24h = append(pruneBefore(st.restarts24h, now.Add(-controllerSlowCrashloopWindow)), now)
n24 := len(st.restarts24h)
raised := false
if n24 >= controllerSlowCrashloopMax && (st.slowCrashloopSince.IsZero() || now.Sub(st.slowCrashloopSince) >= controllerSlowCrashloopWindow) {
st.slowCrashloopSince = now
raised = true
}
rec := slowCounterRecord{Restarts: append([]time.Time(nil), st.restarts24h...), SlowCrashloopSince: st.slowCrashloopSince}
s.ctrlSup.mu.Unlock()
s.saveSlowCounter(vmid, rec)
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason, "restarts_24h", n24)
if raised {
s.logger.Warn("controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop",
"vmid", vmid, "restarts_24h", n24, "window", controllerSlowCrashloopWindow.String(), "threshold", controllerSlowCrashloopMax)
}
return false
}
func (s *Server) resetNotRunning(vmid int) {
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
if st := s.ctrlSup.guests[vmid]; st != nil {
st.notRunningSeen = 0
}
}
func (s *Server) clock() time.Time {
if s.now != nil {
return s.now()
}
return time.Now().UTC()
}
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
out := ts[:0]
for _, t := range ts {
if !t.Before(cutoff) {
out = append(out, t)
}
}
return out
}
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
// Nil when the supervisor is not wired, so the stanza is omitted.
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
if s.staleLock == nil || s.guestExec == nil {
return nil
}
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
for vmid, st := range s.ctrlSup.guests {
g := hub.ControllerSupervisorGuest{
VMID: vmid,
RestartsTotal: st.restartsTotal,
LastReason: st.lastReason,
Parked: st.parked,
Crashloop: !st.crashloopSince.IsZero(),
}
if !st.lastRestartAt.IsZero() {
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
}
if !st.crashloopSince.IsZero() {
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
}
now := s.clock()
g.Restarts24h = len(pruneBefore(append([]time.Time(nil), st.restarts24h...), now.Add(-controllerSlowCrashloopWindow)))
if !st.slowCrashloopSince.IsZero() {
g.SlowCrashloopSince = st.slowCrashloopSince.UTC().Format(time.RFC3339)
g.SlowCrashloop = now.Sub(st.slowCrashloopSince) < controllerSlowCrashloopWindow
}
out.Guests = append(out.Guests, g)
}
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
return out
}
@@ -0,0 +1,350 @@
package localapi
import (
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
"strconv"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
// again", and each guard is pinned by the case where acting would be wrong.
type supExec struct {
mu sync.Mutex
status map[int]string // docker .State.Status per vmid; "" = container absent
inspectErr error // non-nil = pct exec itself failed (unknown)
restarts map[int]int
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
// stays "exited").
onRestart string
}
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
f.mu.Lock()
defer f.mu.Unlock()
switch {
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
if f.inspectErr != nil {
return "", f.inspectErr
}
st, ok := f.status[vmid]
if !ok || st == "" {
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
}
return st + "\n", nil
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
if f.restarts == nil {
f.restarts = map[int]int{}
}
f.restarts[vmid]++
if f.onRestart != "" {
f.status[vmid] = f.onRestart
}
return "", nil
}
return "", errors.New("supExec: unexpected args")
}
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
return "", errors.New("supExec: no stdin exec expected")
}
func (f *supExec) count(vmid int) int {
f.mu.Lock()
defer f.mu.Unlock()
return f.restarts[vmid]
}
type supClock struct{ t time.Time }
func (c *supClock) now() time.Time { return c.t }
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
t.Helper()
dir := t.TempDir()
for _, v := range provisioned {
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
t.Fatal(err)
}
}
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
s := &Server{
staleLock: ctl,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
now: clk.now,
}
return s, clk, dir
}
func runningGuest(vmid int) *fakeGuestPowerCtl {
return &fakeGuestPowerCtl{
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
}
}
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
// the first (the bootstrap's own rm-f/run window must never be raced).
//
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
// stays 0 → "the killed controller was NOT restarted — this is R-523".
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 0 {
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
}
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
}
st := s.ControllerSupervisorStatus(context.Background())
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
}
// Healthy again → no further restart.
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
}
}
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
}
}
func TestControllerSupervisor_Guards(t *testing.T) {
cases := []struct {
name string
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
}{
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
panic(err)
}
}},
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupRun = map[int]bool{9201: true}
}},
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupErr = errors.New("tasks unreadable")
}},
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guests[0].Status = "stopped"
}},
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
}},
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guestsErr = errors.New("api down")
}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
ctl := runningGuest(9201)
s, _, dir := supServer(t, ex, ctl, 9201)
tc.setup(s, ex, ctl, dir)
for i := 0; i < 4; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9201); n != 0 {
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
}
})
}
}
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
ex := &supExec{status: map[int]string{9202: "exited"}}
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
for i := 0; i < 3; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9202); n != 0 {
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
}
}
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
// again.
//
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
// failed at "crash-looping controller restarted 10 times".
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
s.ControllerSupervisorTick(ctx)
clk.t = clk.t.Add(controllerSupervisorInterval)
}
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
}
st := s.ControllerSupervisorStatus(ctx).Guests[0]
if !st.Crashloop || st.CrashloopSince == "" {
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
}
// Still paused 25 minutes later.
clk.t = clk.t.Add(15 * time.Minute)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
}
// After the pause: tries again.
clk.t = clk.t.Add(controllerCrashloopPause)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax+1 {
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
}
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
}
}
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
var m map[string][]map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
g := m["guests"][0]
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked", "restarts_24h", "slow_crashloop"} {
if _, ok := g[k]; !ok {
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
}
}
}
// ---- R-539 (operator ruling 3 of 2026-09-16): the SLOW crash loop ----------------------------------
// supKillOnce kills the controller and lets the supervisor restart it (two confirming sweeps), then
// moves the clock on by gap. The container comes back "running", so each restart is a separate act.
func supKillOnce(t *testing.T, s *Server, ex *supExec, clk *supClock, gap time.Duration) {
t.Helper()
before := ex.count(9201)
ex.mu.Lock()
ex.status[9201] = "exited"
ex.mu.Unlock()
s.ControllerSupervisorTick(context.Background())
clk.t = clk.t.Add(controllerSupervisorInterval)
s.ControllerSupervisorTick(context.Background())
if ex.count(9201) != before+1 {
t.Fatalf("kill was not followed by exactly one restart (restarts %d → %d)", before, ex.count(9201))
}
clk.t = clk.t.Add(gap)
}
// The consequence: a controller that dies every 20 minutes — never three times inside the 15-minute
// brake — raises slow_crashloop on the FIFTH restart in 24 hours, and the raise does not move again on
// the sixth (the hub mails on movement; once per 24 hours is the ruling).
//
// RED-PROOF: without the slow counter the stanza never sets slow_crashloop → "five restarts 20 minutes
// apart did not raise slow_crashloop — this is R-539".
func TestControllerSupervisor_SlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 1; i <= 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
g := s.ControllerSupervisorStatus(ctx).Guests[0]
if g.Crashloop {
t.Fatalf("the 15-minute brake fired on restarts 20 minutes apart — the fixture is wrong: %+v", g)
}
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("slow_crashloop raised after only 4 restarts: %+v", g)
}
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if !g.SlowCrashloop || g.SlowCrashloopSince == "" || g.Restarts24h != 5 {
t.Fatalf("five restarts 20 minutes apart did not raise slow_crashloop — this is R-539: %+v", g)
}
first := g.SlowCrashloopSince
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if g.SlowCrashloopSince != first {
t.Fatalf("the raise moved again on the 6th restart (%q → %q) — the operator would be mailed per restart", first, g.SlowCrashloopSince)
}
if !g.SlowCrashloop {
t.Fatalf("slow_crashloop cleared while the loop continues: %+v", g)
}
}
// The negative control: restarts that never reach five inside any 24 hours never raise it.
func TestControllerSupervisor_SpreadRestartsNeverSlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 8; i++ { // eight restarts, 7 hours apart: at most 4 inside any 24 hours
supKillOnce(t, s, ex, clk, 7*time.Hour)
}
g := s.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("restarts 7 hours apart raised slow_crashloop: %+v", g)
}
if g.Restarts24h > 4 {
t.Fatalf("restarts_24h=%d — the 24-hour window is not pruning", g.Restarts24h)
}
}
// An agent restart must not reset the slow counter (the ruling; a box whose AGENT also restarts would
// otherwise never reach five). The same state directory, a fresh Server.
//
// RED-PROOF: keep the counter in memory only → the second Server starts at 0 → "the agent restart
// reset the slow counter".
func TestControllerSupervisor_SlowCounterSurvivesAgentRestart(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, dir := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
s2 := &Server{
staleLock: s.staleLock,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: s.logger,
now: clk.now,
}
supKillOnce(t, s2, ex, clk, 20*time.Minute)
g := s2.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.Restarts24h != 5 || !g.SlowCrashloop {
t.Fatalf("the agent restart reset the slow counter: %+v", g)
}
}
+94
View File
@@ -185,6 +185,9 @@ type Options struct {
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
StaleLock StaleLockController
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
GuestsStateDir string
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
ControllerSwapStateDir string
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
@@ -362,6 +365,12 @@ type Server struct {
swapMu sync.Mutex
swapInFlight map[int]bool
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
guestExec GuestExecutor
guestsDir string
ctrlSup controllerSupervisor
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
// detached pipeline runs through (tests inject; production defaults set in NewServer).
netVerifyMu sync.Mutex
@@ -484,7 +493,9 @@ func NewServer(o Options) (*Server, error) {
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
if o.ControllerSwap != nil {
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
s.guestExec = o.ControllerSwap
}
s.guestsDir = o.GuestsStateDir
return s, nil
}
@@ -1084,6 +1095,10 @@ type BackupTierInfo struct {
Target string `json:"target"`
CadenceSeconds int64 `json:"cadence_seconds"`
Primary bool `json:"primary"`
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
Storage string `json:"storage,omitempty"`
}
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1093,11 +1108,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
Target: t.TargetID,
CadenceSeconds: int64(t.Cadence.Seconds()),
Primary: t.Primary,
Storage: s.storagePresence(r.Context(), t.TargetID),
})
}
writeOK(w, resp)
}
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
// successful read that does not list the storage is "absent".
func (s *Server) storagePresence(ctx context.Context, target string) string {
if s.storage == nil || target == "" {
return StoragePresenceUnknown
}
targets, err := s.storage.Observe(ctx)
if err != nil {
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
return StoragePresenceUnknown
}
for _, t := range targets {
if t.Name == target {
return StoragePresencePresent
}
}
return StoragePresenceAbsent
}
const (
StoragePresencePresent = "present"
StoragePresenceAbsent = "absent"
StoragePresenceUnknown = "unknown"
)
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
// result ("presence is not success").
type TierBackupState struct {
Target string `json:"target"`
Primary bool `json:"primary"`
Storage string `json:"storage"` // present | absent | unknown
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
// LastSuccessSource is "storage".
LastSuccess *hub.Backup `json:"last_success,omitempty"`
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
}
// TierAttempt is the newest recorded attempt on a tier, successful or not.
type TierAttempt struct {
StartedAt string `json:"started_at"`
Success bool `json:"success"`
Error string `json:"error,omitempty"`
}
// tierBackupStates builds the per-tier view for one guest.
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
out := make([]TierBackupState, 0, len(s.tiers))
for _, t := range s.tiers {
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
st.LastSuccess, st.LastSuccessSource = b, "record"
} else if st.Storage != StoragePresenceAbsent {
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
StartedAt: when.UTC().Format(time.RFC3339)}
st.LastSuccessSource = "storage"
}
}
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
}
out = append(out, st)
}
return out
}
// tierFromRequest resolves the `?target=` query parameter to a tier.
//
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
@@ -1144,6 +1231,10 @@ type BackupStatusResponse struct {
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
Target string `json:"target,omitempty"`
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
// response exactly as before.
Tiers []TierBackupState `json:"tiers,omitempty"`
}
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1155,6 +1246,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
// across ANY target (echo == "" → pickLatestBackup's match-any path).
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
if echo == "" {
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
}
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
resp.Phase = job.Phase
resp.JobID = job.JobID
+14 -2
View File
@@ -10,6 +10,8 @@ import (
"net/url"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no
@@ -31,6 +33,11 @@ type Config struct {
Secret string // token secret (from <id>.pw)
Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy.
Timeout time.Duration
// IdleConnTimeout bounds how long this client's idle keep-alive connections are retained.
// Zero means httpx.DefaultIdleConnTimeout (90s) — it does NOT mean "no timeout", which is the
// R-344 defect. Production leaves it unset; only tests set it, to avoid a 90-second wait.
IdleConnTimeout time.Duration
}
// NewClient builds a fingerprint-pinned, token-authed PBS client.
@@ -55,8 +62,13 @@ func NewClient(cfg Config) (*Client, error) {
authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret,
namespace: cfg.Namespace,
http: &http.Client{
Timeout: timeout,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: timeout,
// R-344: this transport MUST come from httpx. pbsTargetsFromPVE (cmd/felhom-agent/
// main.go) builds a fresh Client every cycle and drops the previous one, so a
// transport with no idle timeout strands one connection per cycle, forever, on both
// sides. That leaked 388 sockets onto ep0 in 46 hours. Pinned by
// TestAbandonedClientsReleaseTheirConnections — do not inline an http.Transport here.
Transport: httpx.NewTransport(tlsCfg, cfg.IdleConnTimeout),
},
}, nil
}
+173
View File
@@ -0,0 +1,173 @@
package pbs
import (
"context"
"net"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// connCounter is the SERVER-side observer. It counts what the server actually holds, which is the
// only thing that answers the question this file exists for: a client that believes it closed a
// connection, and a server still holding the socket, is precisely the R-344 shape. Asserting on
// anything client-side would be asserting the mechanism instead of the consequence.
type connCounter struct {
mu sync.Mutex
open int
total int // every connection ever accepted — how many times the client DIALLED
}
func (c *connCounter) hook(_ net.Conn, s http.ConnState) {
c.mu.Lock()
defer c.mu.Unlock()
switch s {
case http.StateNew:
c.open++
c.total++
case http.StateClosed, http.StateHijacked:
c.open--
}
}
func (c *connCounter) counts() (open, total int) {
c.mu.Lock()
defer c.mu.Unlock()
return c.open, c.total
}
// waitForOpen polls until the server holds want connections, or fails naming what it still holds.
func (c *connCounter) waitForOpen(t *testing.T, want int, within time.Duration, what string) {
t.Helper()
deadline := time.Now().Add(within)
for {
open, total := c.counts()
if open == want {
return
}
if time.Now().After(deadline) {
t.Fatalf("%s: after %s the server still holds %d open connection(s), want %d (%d dialled in total)",
what, within, open, want, total)
}
time.Sleep(5 * time.Millisecond)
}
}
// newCountingPBSServer is newPBSTestServer plus a ConnState hook. Kept separate rather than
// changing the shared helper, so the existing tests are untouched by this file.
func newCountingPBSServer(t *testing.T, fn http.HandlerFunc) (*httptest.Server, string, *connCounter) {
t.Helper()
cc := &connCounter{}
ts := httptest.NewUnstartedServer(fn)
ts.Config.ConnState = cc.hook
ts.StartTLS()
t.Cleanup(ts.Close)
return ts, fingerprintOf(ts), cc
}
// TestAbandonedClientsReleaseTheirConnections is Scenario A, and it is the load-bearing test for
// R-344.
//
// It models what pbsTargetsFromPVE actually does — build a client, use it once, drop it on the
// floor without closing anything — and asserts the CONSEQUENCE on the server: the connections go
// away. Before the fix every one of these stayed established forever on both sides; 388 of them
// accumulated on ep0 in 46 hours.
//
// Deliberately NOT asserted: that err == nil, or that IdleConnTimeout holds some value. Both were
// true of the leaking code.
func TestAbandonedClientsReleaseTheirConnections(t *testing.T) {
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte(`{"data":[]}`))
})
host, port := hostPort(t, ts.URL)
const cycles = 5
for i := 0; i < cycles; i++ {
// One fresh client per "cycle", exactly as pbsTargetsFromPVE builds one per collect.
c, err := NewClient(Config{
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
IdleConnTimeout: 50 * time.Millisecond, // production uses the 90s default
})
if err != nil {
t.Fatal(err)
}
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
t.Fatalf("cycle %d: %v", i, err)
}
_ = c // dropped here — nothing closes it, nothing can
}
if _, total := cc.counts(); total != cycles {
t.Fatalf("setup is not modelling the leak: want %d separate dials (one per abandoned client), got %d", cycles, total)
}
cc.waitForOpen(t, 0, 5*time.Second, "abandoned pbs.Clients")
}
// TestAbandonedClientsReleaseTheirConnections_ProductionDefaultIsUsable pins the value that ships.
//
// The field being settable is exactly how it could silently become zero again — and zero used to
// mean "never expire". This asserts the production path (Config leaving it unset) lands on the
// standard-library default, so the leak cannot return through an unset field.
func TestPBSClient_UnsetIdleTimeoutUsesTheDefault(t *testing.T) {
for _, tc := range []struct {
name string
cfg time.Duration
want time.Duration
}{
{"unset — the production path", 0, httpx.DefaultIdleConnTimeout},
{"explicit zero is NOT no-timeout", 0, httpx.DefaultIdleConnTimeout},
{"negative is NOT no-timeout", -time.Second, httpx.DefaultIdleConnTimeout},
{"an explicit value is honoured", 3 * time.Second, 3 * time.Second},
} {
t.Run(tc.name, func(t *testing.T) {
c, err := NewClient(Config{
Server: "pbs.example", Fingerprint: strings.Repeat("ab", 32),
TokenID: "u@pbs!t", Secret: "s", IdleConnTimeout: tc.cfg,
})
if err != nil {
t.Fatal(err)
}
tr, ok := c.http.Transport.(*http.Transport)
if !ok {
t.Fatalf("transport is %T, not *http.Transport — the httpx wiring was replaced", c.http.Transport)
}
if tr.IdleConnTimeout != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v (zero would mean connections are retained FOREVER — that is R-344)", tr.IdleConnTimeout, tc.want)
}
})
}
}
// TestPBSClient_KeepAliveStillReuses is Scenario C, and it is the guard against a "fix" that is
// worse than the bug.
//
// Disabling keep-alive entirely would also make the leak go away — by dialling a fresh connection
// for every single request, which on a box polling ~40,000 times a day is strictly worse than what
// we started with. The fix must retire IDLE connections without stopping reuse.
func TestPBSClient_KeepAliveStillReuses(t *testing.T) {
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte(`{"data":[]}`))
})
host, port := hostPort(t, ts.URL)
c, err := NewClient(Config{
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
IdleConnTimeout: 30 * time.Second, // long enough that reuse is what is being measured
})
if err != nil {
t.Fatal(err)
}
for i := 0; i < 3; i++ {
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
t.Fatalf("request %d: %v", i, err)
}
}
if _, total := cc.counts(); total != 1 {
t.Fatalf("one client made 3 sequential requests over %d connection(s), want 1 — keep-alive reuse is broken, which would make the poll load WORSE than the leak", total)
}
}
+6 -2
View File
@@ -10,6 +10,8 @@ import (
"net/url"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// doer is the minimal HTTP surface the client needs; *http.Client satisfies it.
@@ -65,8 +67,10 @@ func NewClient(cfg Config) (*Client, error) {
timeout = 30 * time.Second
}
hc := &http.Client{
Timeout: timeout,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: timeout,
// R-344, consistency only: built ONCE per process, so it never accumulated and contributed
// nothing to the ep0 leak. Same missing default, corrected for the same reason.
Transport: httpx.NewTransport(tlsCfg, 0),
}
return &Client{
base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json",
+15 -9
View File
@@ -219,15 +219,18 @@ type Storage struct {
UsedFraction float64 `json:"used_fraction,omitempty"`
// Type-specific config (durable_id sources).
Server string `json:"server,omitempty"` // nfs/cifs/pbs server host
Export string `json:"export,omitempty"` // nfs export path
Share string `json:"share,omitempty"` // cifs share name
Datastore string `json:"datastore,omitempty"` // pbs datastore name
Fingerprint string `json:"fingerprint,omitempty"` // pbs server cert fingerprint
Username string `json:"username,omitempty"` // pbs auth id, e.g. "felhom@pbs!n100"
Namespace string `json:"namespace,omitempty"` // pbs namespace ("" = root; per-customer tenancy = S4)
VGName string `json:"vgname,omitempty"` // lvm/lvmthin volume group
ThinPool string `json:"thinpool,omitempty"` // lvmthin pool LV name
Server string `json:"server,omitempty"` // nfs/cifs/pbs server host
// EncryptionKey is the storage's own client-side key FINGERPRINT (pbs; the key itself stays in
// /etc/pve/priv). R-727: an archive encrypted with any other key was written by another box.
EncryptionKey string `json:"encryption-key,omitempty"`
Export string `json:"export,omitempty"` // nfs export path
Share string `json:"share,omitempty"` // cifs share name
Datastore string `json:"datastore,omitempty"` // pbs datastore name
Fingerprint string `json:"fingerprint,omitempty"` // pbs server cert fingerprint
Username string `json:"username,omitempty"` // pbs auth id, e.g. "felhom@pbs!n100"
Namespace string `json:"namespace,omitempty"` // pbs namespace ("" = root; per-customer tenancy = S4)
VGName string `json:"vgname,omitempty"` // lvm/lvmthin volume group
ThinPool string `json:"thinpool,omitempty"` // lvmthin pool LV name
}
// StorageContent is one entry of GET /nodes/{node}/storage/{store}/content
@@ -239,4 +242,7 @@ type StorageContent struct {
Size int64 `json:"size"`
CTime int64 `json:"ctime"`
VMID int `json:"vmid,omitempty"`
// Encrypted is the fingerprint of the key a PBS archive was encrypted with ("" = not encrypted). R-727:
// the restore test reads it to tell THIS box's archives from an earlier box's in the same namespace.
Encrypted string `json:"encrypted,omitempty"`
}
+57
View File
@@ -302,6 +302,19 @@ func (e *Engine) runBringUp(ctx context.Context, spec BringUpSpec, res *BringUpR
}
}
// R-834: DR keeps the archive's `onboot: 1`, binds the host's REAL drives (4d) and STARTS the guest
// — right on a replaced host, where the original is gone. Beside a LIVE original it would be a
// second controller for the same household on the same drives. So DR refuses when this host
// still carries the original (the archive's source VMID) or any guest that binds the drives.
// A copy beside the original is the restore-test's job (onboot=0, throwaway stand-ins, torn
// down) or the runbook's beside-restore. Pinned by TestRunBringUp_DRRefusesBesideALiveOriginal.
if spec.Mode == ModeDRGuestLoss {
if why := e.liveOriginalBeside(ctx, lxc, spec.Archive); why != "" {
res.Err = fmt.Errorf("reconcile: dr bring-up refused: %s — a DR restore beside a live original would run two boxes on the same drives (R-834)", why)
return
}
}
base := JournalEntry{OpID: e.bringUpOpID(spec.VMID), VMID: spec.VMID, Kind: bringUpKind, Rollback: true}
// OWN the rollback BEFORE any mutation. From here a crash leaves an in-flight Rollback
@@ -734,3 +747,47 @@ func net0MAC(cfg proxmox.GuestConfig) string {
}
return ""
}
// archiveSourceVMID reads the source guest's VMID from a backup volid: a vzdump file
// (`…/vzdump-lxc-<vmid>-<date>.tar.zst`) or a PBS snapshot (`…:backup/ct/<vmid>/<time>`). 0 = unknown.
func archiveSourceVMID(archive string) int {
if i := strings.Index(archive, "vzdump-lxc-"); i >= 0 {
rest := archive[i+len("vzdump-lxc-"):]
if j := strings.Index(rest, "-"); j > 0 {
if n, err := strconv.Atoi(rest[:j]); err == nil {
return n
}
}
}
if i := strings.Index(archive, "ct/"); i >= 0 {
rest := archive[i+len("ct/"):]
if j := strings.Index(rest, "/"); j > 0 {
if n, err := strconv.Atoi(rest[:j]); err == nil {
return n
}
}
}
return 0
}
// liveOriginalBeside says why a DR bring-up would land beside a live original ("" = it would not):
// the archive's source guest still exists here, or a guest binds the drives parent. Fails CLOSED: a
// guest whose config cannot be read cannot be ruled out.
func (e *Engine) liveOriginalBeside(ctx context.Context, lxc []proxmox.Guest, archive string) string {
src := archiveSourceVMID(archive)
for _, g := range lxc {
if src > 0 && g.VMID == src {
return fmt.Sprintf("the archive's source guest %d still exists on this host (status %s)", g.VMID, g.Status)
}
cfg, err := e.api.GuestConfig(ctx, g.VMID)
if err != nil {
return fmt.Sprintf("guest %d's config could not be read to rule out a live original: %v", g.VMID, err)
}
for slot, v := range cfg.MountPoints() {
if source, _, _ := strings.Cut(v, ","); source == structuralParentDir {
return fmt.Sprintf("guest %d binds the household drives (%s %s)", g.VMID, slot, structuralParentDir)
}
}
}
return ""
}
+138
View File
@@ -0,0 +1,138 @@
package reconcile
import (
"context"
"encoding/json"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-834: a whole-guest restore BESIDE a live original must never come up as a second box on the same
// drives. The DR route keeps `onboot: 1`, binds the real drives and starts the guest, so it refuses
// when the original (or any guest binding the drives) is still on this host; on a replaced host it
// proceeds and keeps its binds. Red-proof: make liveOriginalBeside return "" and the refusals pass
// the restore through (the "refused" sub-tests fail).
func TestRunBringUp_DRRefusesBesideALiveOriginal(t *testing.T) {
const target = 9299
drivesBind := proxmox.GuestConfig{Extra: map[string]json.RawMessage{
"mp8": json.RawMessage(`"/mnt/felhom-drives,mp=/mnt/felhom-drives"`),
}}
cases := []struct {
name string
archive string
lxc []proxmox.Guest
cfg map[int]proxmox.GuestConfig
refuse string // substring of the refusal; "" = must proceed
}{
{"the source guest still exists", "local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst",
[]proxmox.Guest{{VMID: 9201, Status: "running"}}, map[int]proxmox.GuestConfig{9201: scratchCfg()}, "source guest 9201"},
{"another guest binds the drives (PBS archive)", "felhom-pbs:backup/ct/9201/2026-10-04T02:34:55Z",
[]proxmox.Guest{{VMID: 9300, Status: "stopped"}}, map[int]proxmox.GuestConfig{9300: drivesBind}, "binds the household drives"},
{"a guest whose config cannot be read", "local:backup/vzdump-lxc-9201-x.tar.zst",
[]proxmox.Guest{{VMID: 9400, Status: "running"}}, map[int]proxmox.GuestConfig{}, "could not be read"},
{"replaced host: only an unrelated scratch guest", "local:backup/vzdump-lxc-9201-x.tar.zst",
[]proxmox.Guest{{VMID: 9202, Status: "running"}}, map[int]proxmox.GuestConfig{9202: scratchCfg()}, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
cfg := c.cfg
cfg[target] = scratchCfg()
api := &fakeAPI{lxc: c.lxc, cfg: cfg}
e, fr, _, q := newDREngine(t, api)
defer q.Close()
res := e.RunBringUp(context.Background(), BringUpSpec{
Mode: ModeDRGuestLoss, Archive: c.archive, VMID: target, RestoreStorage: "local-lvm", KeepMAC: true,
})
if c.refuse != "" {
if res.Err == nil || !strings.Contains(res.Err.Error(), c.refuse) || !strings.Contains(res.Err.Error(), "R-834") {
t.Fatalf("want a refusal naming %q, got %+v", c.refuse, res)
}
if len(api.restores) != 0 || len(api.starts) != 0 || len(fr.cmds) != 0 {
t.Fatalf("a refused DR touched the host: restores=%d starts=%v cmds=%v", len(api.restores), api.starts, fr.cmds)
}
return
}
if res.Err != nil || !res.Pass {
t.Fatalf("a DR on a replaced host must proceed, got %+v", res)
}
// … and there it keeps the REAL drives bind (the right binds on a replaced host).
joined := strings.Join(fr.cmds, "\n")
if !strings.Contains(joined, "-mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives") {
t.Fatalf("the DR guest lost its drives bind: %v", fr.cmds)
}
})
}
}
// Provisioning restores the GOLDEN (no drives, onboot set by the back-half on purpose): a drives-
// binding guest on the host does not block it — the rule is DR's alone.
func TestRunBringUp_ProvisionNotBlockedByADrivesBind(t *testing.T) {
api := &fakeAPI{
lxc: []proxmox.Guest{{VMID: 9201, Status: "running"}},
cfg: map[int]proxmox.GuestConfig{
9201: {Extra: map[string]json.RawMessage{"mp8": json.RawMessage(`"/mnt/felhom-drives,mp=/mnt/felhom-drives"`)}},
9203: scratchCfg(),
},
}
e, _, q := newEngine(t, api, EmptyProvider{})
defer q.Close()
res := e.RunBringUp(context.Background(), BringUpSpec{Mode: ModeProvision, Archive: "local:vztmpl/felhom-golden.tar.zst", VMID: 9203, RestoreStorage: "local-lvm"})
if res.Err != nil || !res.Pass {
t.Fatalf("provision must proceed, got %+v", res)
}
}
func TestArchiveSourceVMID(t *testing.T) {
for in, want := range map[string]int{
"local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst": 9201,
"felhom-pbs:backup/ct/9201/2026-10-04T02:34:55Z": 9201,
"tmp-dooplex-copy:backup/ct/9201/2026-10-03T19:00:00Z": 9201,
"local:vztmpl/felhom-golden.tar.zst": 0,
"vol": 0,
} {
if got := archiveSourceVMID(in); got != want {
t.Errorf("archiveSourceVMID(%q) = %d, want %d", in, got, want)
}
}
}
// R-834, the restore-test route: its scratch guest sits BESIDE the live original by design, so it must
// carry no host-path bind — the archive's mp8 (the household's drives) and mp9 (the original's
// bootstrap) are replaced by throwaway volumes AT restore time. Measured live 2026-10-04 on demo-hp
// (`audits/backup-close-2026-10-04/partA/`): onboot 0 and no host bind on every poll. Red-proof:
// make drRestoreOverrides return the archive's own mp8 value and this fails.
func TestRestoreTest_NoHostPathBindBesideTheOriginal(t *testing.T) {
api := &fakeAPI{
cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()},
extractCfg: "hostname: demo-hp\nonboot: 1\nrootfs: local-lvm:vm-9201-disk-0,size=16G\n" +
"mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G\n" +
"mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\n" +
"mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1\n",
}
e, _, q := newEngine(t, api, EmptyProvider{})
defer q.Close()
_ = e.RunRestoreTest(context.Background(), RestoreTestSpec{
Archive: "local:backup/vzdump-lxc-9201-x.tar.zst", RestoreStorage: "local-lvm",
ScratchMin: 990000, ScratchMax: 990009, SourceTier: "local",
})
if len(api.restores) != 1 {
t.Fatalf("want one restore, got %+v", api.restores)
}
r := api.restores[0]
for _, slot := range []string{"mp8", "mp9"} {
v, ok := r.MountOverrides[slot]
if !ok || strings.HasPrefix(v, "/") {
t.Fatalf("%s = %q (present=%v): the scratch beside the original must get a throwaway volume, never the host path", slot, v, ok)
}
}
for slot, v := range r.MountOverrides {
if strings.HasPrefix(v, "/") {
t.Fatalf("%s carries a host path %q into the scratch guest", slot, v)
}
}
if r.ConfigOverrides["onboot"] != "0" {
t.Fatalf("onboot = %q, want 0", r.ConfigOverrides["onboot"])
}
}
+1 -1
View File
@@ -52,7 +52,7 @@ func newDREngine(t *testing.T, api GuestAPI) (*Engine, *fakeRunner, string, *Que
t.Cleanup(q.Close)
fr := &fakeRunner{}
sd := t.TempDir()
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: EmptyProvider{}, HostRunner: fr, StateDir: sd})
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: EmptyProvider{}, HostRunner: fr, StateDir: sd, RestoreSpace: roomySpace{}})
return e, fr, sd, q
}
+30
View File
@@ -44,6 +44,17 @@ type Engine struct {
opSeq uint64 // atomic; makes each op id unique per attempt
// restoreSpace + spacePolicy are the restore-test's space preflight (R-672). nil space REFUSES
// every restore-test (fail-closed) — see restoretest_space.go.
restoreSpace RestoreSpace
spacePolicy SpacePolicy
// scratchMu guards activeScratch (the vmids a running restore-test owns — the retry timer never
// touches those) and teardownTries (failed timer retries per journal op, R-672 rule 3).
scratchMu sync.Mutex
activeScratch map[int]bool
teardownTries map[string]int
// lastRes records the most recent successful Reconcile Result (v0.90.0, R-28 fast-tick source).
// The fast-tick reads it to decide convergence: actionable drift is Planned − Pending > 0 (a
// destructive pending_signature refusal is EXPECTED state, not drift to hammer on). lastOK is
@@ -86,6 +97,10 @@ type EngineOptions struct {
HostRunner proxmox.Runner
// StateDir is the agent state dir ("" → /var/lib/felhom-agent); only the 4d swap reads it.
StateDir string
// RestoreSpace is the restore-test's space preflight (R-672). nil → every restore-test is REFUSED
// with its reason (fail-closed). SpacePolicy zero → DefaultSpacePolicy.
RestoreSpace RestoreSpace
SpacePolicy SpacePolicy
}
// NewEngine builds an Engine. The Queue is shared (the single §10 choke point); the
@@ -124,9 +139,24 @@ func NewEngine(opts EngineOptions) *Engine {
logger: logger,
hostRun: opts.HostRunner,
stateDir: stateDir,
restoreSpace: opts.RestoreSpace,
spacePolicy: policyOrDefault(opts.SpacePolicy),
activeScratch: map[int]bool{},
teardownTries: map[string]int{},
}
}
func policyOrDefault(p SpacePolicy) SpacePolicy {
if p.Factor < 1 {
p.Factor = DefaultSpacePolicy.Factor
}
if p.ReserveBytes <= 0 {
p.ReserveBytes = DefaultSpacePolicy.ReserveBytes
}
return p
}
// Result summarizes one Reconcile pass.
type Result struct {
Planned int
+1 -1
View File
@@ -213,7 +213,7 @@ func newEngine(t *testing.T, api GuestAPI, provider DesiredProvider) (*Engine, *
t.Cleanup(func() { j.Close() })
q := NewQueue()
t.Cleanup(q.Close)
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: provider})
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: provider, RestoreSpace: roomySpace{}})
return e, j, q
}
+40 -11
View File
@@ -50,10 +50,18 @@ type RestoreTestResult struct {
ScratchVMID int
Pass bool
Verified string // "boot+running" this slice
Skipped bool // no free scratch VMID in band → test not run
Err error
StartedAt time.Time
Duration time.Duration
Skipped bool // test not run: no free scratch VMID in band, or the space preflight refused (R-672)
// SkipReason is set when the SPACE PREFLIGHT refused (R-672): the test did not run, and this is
// reported to the hub as the test's result (pass=false), never as a pass. Empty for a band skip.
SkipReason string
// TargetStorage is where the restore went (rule 2 may move it off the tested guest's pool);
// RequiredBytes/AvailBytes are rule 1's figures.
TargetStorage string
RequiredBytes int64
AvailBytes int64
Err error
StartedAt time.Time
Duration time.Duration
// StartWarnings holds the warning line(s) the guest-start task emitted (e.g. the
// systemd-nesting advisory). Populated only when the start exited "WARNINGS: N";
// always surfaced, NEVER used to decide pass/fail (the verdict is liveness — waitRunning).
@@ -136,6 +144,29 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
return res
}
// R-672: the space preflight, BEFORE anything is journaled or created.
rawCfg, err := e.api.ExtractArchiveConfig(ctx, spec.Archive)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test extract archive config: %w", err)
return res
}
v := PreflightRestoreSpace(ctx, e.restoreSpace, e.spacePolicy, spec.Archive, rawCfg, spec.RestoreStorage)
res.TargetStorage, res.RequiredBytes, res.AvailBytes = v.Storage, v.Required, v.Avail
if !v.OK {
res.Skipped = true
res.SkipReason = "skipped: " + v.Reason
e.logger.Warn("restore-test SKIPPED by the space preflight (R-672) — nothing was created",
"archive", spec.Archive, "storage", v.Storage, "required_bytes", v.Required, "avail_bytes", v.Avail, "reason", v.Reason)
res.Duration = time.Since(now)
return res
}
if v.Storage != spec.RestoreStorage {
e.logger.Info("restore-test: restoring OFF the tested guest's own pool (R-672 rule 2)",
"configured", spec.RestoreStorage, "avoided", v.Avoided, "target", v.Storage)
}
e.logger.Info("restore-test: space preflight passed", "storage", v.Storage, "required_bytes", v.Required, "avail_bytes", v.Avail)
spec.RestoreStorage = v.Storage
lxc, err := e.api.ListLXC(ctx)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test list guests: %w", err)
@@ -164,7 +195,9 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
// Serialize on the scratch VMID's lane (inherits §10), and capture the result.
var vmidOccupied bool
ch := e.queue.Submit(vmid, func() error {
vmidOccupied = e.runScratchTest(ctx, vmid, spec, &res)
e.markScratch(vmid, true)
defer e.markScratch(vmid, false)
vmidOccupied = e.runScratchTest(ctx, vmid, spec, rawCfg, &res)
return res.Err
})
<-ch
@@ -182,7 +215,7 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
// runScratchTest is the journaled body (runs on vmid's queue lane). The occupied return is true
// ONLY when PVE synchronously refused the restore because the vmid already holds a guest (one
// the pool-blind band scan couldn't see) — the caller then advances to the next band vmid (F2).
func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestSpec, res *RestoreTestResult) (occupied bool) {
func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestSpec, rawCfg string, res *RestoreTestResult) (occupied bool) {
base := JournalEntry{OpID: e.scratchOpID(vmid), VMID: vmid, Kind: scratchKind, Scratch: true}
// OWN the scratch guest's cleanup BEFORE any mutation. From here, a crash is recoverable.
@@ -215,11 +248,7 @@ func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestS
// genuinely EXTRACTED — full fidelity; the added runtime IS the verification), the two
// structural binds → throwaway stand-ins. An unreadable archive config or an unknown
// topology REFUSES up front — never restore a partial guest to "verify" it.
rawCfg, err := e.api.ExtractArchiveConfig(ctx, spec.Archive)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test extract archive config: %w", err)
return false
}
// The archive's config was read ONCE, by the space preflight (R-672), and is passed in.
mountOverrides, err := drRestoreOverrides(rawCfg, spec.RestoreStorage)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test: %w", err)
+94
View File
@@ -0,0 +1,94 @@
package reconcile
import "context"
// ── A failed scratch teardown is retried on a TIMER, not only at agent start (R-672 rule 3) ────────
//
// MEASURED 2026-09-24 on demo-hp: the scheduled restore-test's teardown failed (`lvremove … contains a
// filesystem in use`, a transient hold) and logged "left for Recover" — and Recover runs ONLY at agent
// start, so the 22 GiB scratch guest sat in the full pool for 2.5 hours until an agent restart. The
// timer calls RetryScratchTeardown every 10 minutes: the SAME resolution as Recover (recoverScratch —
// the gate's benign scratch destroy, idempotent when the guest is already gone), restricted to Scratch
// entries that carry a launch-proof UPID and that no running restore-test owns. After
// MaxTeardownTries failed attempts for one entry the operator is told (the caller reports it); the
// timer keeps trying.
// MaxTeardownTries is how many failed timer retries of one scratch entry happen before the operator
// is told.
const MaxTeardownTries = 3
// ScratchRetryResult summarizes one timer pass.
type ScratchRetryResult struct {
Examined int
Destroyed int
Clean int // already gone
Failed int
// GaveUp lists the scratch vmids whose failed tries reached MaxTeardownTries IN THIS PASS — each
// is reported exactly once (the caller tells the operator).
GaveUp []int
}
func (e *Engine) markScratch(vmid int, active bool) {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
if active {
e.activeScratch[vmid] = true
} else {
delete(e.activeScratch, vmid)
}
}
func (e *Engine) scratchActive(vmid int) bool {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
return e.activeScratch[vmid]
}
// RetryScratchTeardown is the timer's pass. It never touches a non-Scratch entry (unlike Recover,
// which also resolves generic in-flight operations and must therefore run only at start), never an
// entry without a launch-proof UPID (nothing was created), and never a vmid a running test owns.
func (e *Engine) RetryScratchTeardown(ctx context.Context) ScratchRetryResult {
var out ScratchRetryResult
if e.journal == nil {
return out
}
for _, entry := range e.journal.InFlight() {
if !entry.Scratch || entry.UPID == "" || e.scratchActive(entry.VMID) {
continue
}
out.Examined++
var r RecoverResult
e.recoverScratch(ctx, entry, &r)
switch {
case r.ScratchDestroyed > 0:
out.Destroyed++
e.forgetTries(entry.OpID)
case r.ScratchClean > 0:
out.Clean++
e.forgetTries(entry.OpID)
default:
out.Failed++
n := e.addTry(entry.OpID)
e.logger.Warn("restore-test: scratch teardown retry failed (timer)", "vmid", entry.VMID, "op_id", entry.OpID, "try", n)
if n == MaxTeardownTries {
out.GaveUp = append(out.GaveUp, entry.VMID)
e.logger.Error("restore-test: scratch guest still NOT torn down after repeated retries — telling the operator",
"vmid", entry.VMID, "tries", n)
}
}
}
return out
}
func (e *Engine) addTry(op string) int {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
e.teardownTries[op]++
return e.teardownTries[op]
}
func (e *Engine) forgetTries(op string) {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
delete(e.teardownTries, op)
}
@@ -0,0 +1,73 @@
package reconcile
import (
"context"
"errors"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 rule 3 (v0.133.0): a failed scratch teardown is retried on a TIMER. The consequence asserted:
// the leaked scratch guest is destroyed by a timer pass (not only by a restart's Recover), the operator
// is told exactly once after MaxTeardownTries failures, and a vmid a running test owns is never touched.
// leakScratch runs a restore-test whose teardown fails, leaving scratch 990000 in-flight — the
// 2026-09-24 shape ("lvremove … contains a filesystem in use").
func leakScratch(t *testing.T) (*Engine, *fakeAPI, *Journal) {
t.Helper()
api := &fakeAPI{cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}, restoreUPID: "UPID:r", destroyErr: errors.New("lvremove: contains a filesystem in use")}
e, j := spaceEngine(t, api, roomySpace{})
e.RunRestoreTest(context.Background(), RestoreTestSpec{Archive: "local:backup/x.tar.zst", RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009})
if len(j.InFlight()) != 1 {
t.Fatalf("setup: want the scratch left in-flight after a failed teardown, got %+v", j.InFlight())
}
api.lxc = []proxmox.Guest{{VMID: 990000}}
return e, api, j
}
// COMPANION RED-PROOF (REPORT): make RetryScratchTeardown return without touching the journal (the
// v0.132.0 shape — only Recover at start resolved a leak) → "the leaked scratch was not destroyed by
// the timer".
func TestRetry_TheTimerDestroysALeakedScratch(t *testing.T) {
e, api, j := leakScratch(t)
api.destroyErr = nil // the transient hold is gone
before := len(api.destroys)
r := e.RetryScratchTeardown(context.Background())
if r.Destroyed != 1 || len(api.destroys) != before+1 || api.destroys[len(api.destroys)-1] != 990000 {
t.Fatalf("the leaked scratch was not destroyed by the timer: result=%+v destroys=%v", r, api.destroys)
}
if len(j.InFlight()) != 0 {
t.Fatalf("the entry is still in flight after a successful retry: %+v", j.InFlight())
}
if r2 := e.RetryScratchTeardown(context.Background()); r2.Examined != 0 {
t.Fatalf("a resolved entry was examined again: %+v", r2)
}
}
func TestRetry_OperatorToldOnceAfterThreeFailures(t *testing.T) {
e, _, _ := leakScratch(t)
var gave [][]int
for i := 0; i < MaxTeardownTries+2; i++ {
gave = append(gave, e.RetryScratchTeardown(context.Background()).GaveUp)
}
for i, g := range gave {
want := 0
if i == MaxTeardownTries-1 {
want = 1
}
if len(g) != want {
t.Fatalf("pass %d gave up on %v — want the operator told exactly once, on pass %d", i+1, g, MaxTeardownTries)
}
}
}
func TestRetry_NeverTouchesARunningTest(t *testing.T) {
e, api, _ := leakScratch(t)
api.destroyErr = nil
e.markScratch(990000, true) // a restore-test is (again) working on this vmid
before := len(api.destroys)
if r := e.RetryScratchTeardown(context.Background()); r.Examined != 0 || len(api.destroys) != before {
t.Fatalf("the timer touched a scratch a running test owns: %+v destroys=%v", r, api.destroys)
}
}
+176
View File
@@ -0,0 +1,176 @@
package reconcile
import (
"context"
"fmt"
"sort"
"strings"
)
// ── The restore-test's space preflight (R-672, agent v0.133.0) ─────────────────────────────────────
//
// MEASURED 2026-09-24 on demo-hp: the scheduled restore-test restored 9201's archive into `local-lvm`
// — the SAME thin pool that holds 9201 — with no free-space check. The pool reached 100 %
// (`out_of_data_space`, `error_if_no_space`), and 9201's rootfs and data volume remounted READ-ONLY.
// Evidence: felhom.eu `documentation/audits/night-2026-09-24/C-02…C-07`, `audits/r672-2026-09-24/`.
//
// THE THREE RULES, all decided here before ANY mutation (before the scratch entry is even journaled):
// 1. SPACE FIRST. The target storage must have free data ≥ restored × factor + reserve (defaults 1.2 and
// 5 GiB, `backup.restore_test_space_factor` / `backup.restore_test_space_reserve_gib`), and a thin
// pool's metadata must have room for the same share. `restored` is the UNCOMPRESSED size — the
// archive FILE is the wrong number: 9201's archive was 6.9 GB and its restore wrote 22.6 GB, so
// "file × 1.2 + 5 GiB" (14.3 GB) would have let the 2026-09-24 test run into a pool with 23 GB free.
// 2. KEEP OFF THE TESTED GUEST'S POOL when another eligible storage (active, takes `rootdir`, and the
// agent holds Datastore.AllocateSpace on it) passes rule 1. With only one, rule 1 decides.
// 3. UNKNOWN REFUSES. An unreadable size, an unreadable storage or an unknown thin-pool metadata fill is
// a skip with its reason, never a guess — the fail-safe direction of every guard in this project.
// A refusal is reported to the hub as the test's RESULT ("skipped: …", pass=false), never as a pass.
// Pinned by restoretest_space_test.go.
// RestoreSpace is the preflight's seam onto the host. Production: internal/restorespace.
type RestoreSpace interface {
// RestoredBytes is how many bytes restoring `archive` will write (uncompressed), and where that
// figure came from (for the log and the refusal).
RestoredBytes(ctx context.Context, archive string) (bytes int64, source string, err error)
// Free reports the storage's free data bytes and, for a thin pool, its metadata-used fraction.
Free(ctx context.Context, storage string) (StorageFree, error)
// Eligible lists the storages a restore-test may target: active, content `rootdir`, and the agent
// holds Datastore.AllocateSpace there.
Eligible(ctx context.Context) ([]string, error)
}
// StorageFree is one storage's free space as the preflight judges it.
type StorageFree struct {
AvailBytes int64
UsedBytes int64
Thin bool
// MetaUsedFraction is the thin pool's metadata use (0..1); MetaKnown false = could not be read.
MetaUsedFraction float64
MetaKnown bool
}
// SpacePolicy is rule 1's margin.
type SpacePolicy struct {
Factor float64 // ≥ 1
ReserveBytes int64
}
// DefaultSpacePolicy is 1.2 × restored + 5 GiB.
var DefaultSpacePolicy = SpacePolicy{Factor: 1.2, ReserveBytes: 5 << 30}
// SpaceVerdict is the preflight's answer.
type SpaceVerdict struct {
OK bool
Storage string // the storage the restore goes to (when OK) or was judged (when not)
Required int64
Avail int64
Reason string // empty when OK
// Avoided is the tested guest's own storage, when rule 2 moved the restore off it.
Avoided string
}
// requiredBytes is rule 1's figure.
func (p SpacePolicy) requiredBytes(restored int64) int64 {
f := p.Factor
if f < 1 {
f = DefaultSpacePolicy.Factor
}
return int64(float64(restored)*f) + p.ReserveBytes
}
// fits judges one storage against rule 1 (data AND thin metadata). An unknown metadata fill on a thin
// pool refuses (rule 3).
func fits(fr StorageFree, required int64) (bool, string) {
if fr.AvailBytes < required {
return false, fmt.Sprintf("needs %s free, has %s", gib(required), gib(fr.AvailBytes))
}
if fr.Thin {
if !fr.MetaKnown {
return false, "thin-pool metadata fill unknown"
}
// The metadata a restore of `required` bytes needs, in the pool's own proportion of metadata to
// data. A pool with no data yet has no proportion to read → only the absolute ceiling applies.
need := 0.0
if fr.UsedBytes > 0 {
need = fr.MetaUsedFraction * float64(required) / float64(fr.UsedBytes)
}
if fr.MetaUsedFraction+need > 0.9 {
return false, fmt.Sprintf("thin-pool metadata would reach %.0f%% (now %.0f%%)", 100*(fr.MetaUsedFraction+need), 100*fr.MetaUsedFraction)
}
}
return true, ""
}
func gib(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
// sourceStorages returns the storage ids that hold the ARCHIVED guest's volumes (rootfs and every mpN
// that names a `storage:volume`), read from the archive's own embedded config — the guest under test.
// Bind mounts (a leading "/") carry no storage.
func sourceStorages(rawCfg string) map[string]bool {
out := map[string]bool{}
for k, v := range archiveCurrentConfig(rawCfg) {
if k != "rootfs" && !(strings.HasPrefix(k, "mp") && len(k) > 2 && k[2] >= '0' && k[2] <= '9') {
continue
}
vol := strings.TrimSpace(strings.SplitN(strings.TrimSpace(v), ",", 2)[0])
if vol == "" || strings.HasPrefix(vol, "/") {
continue
}
if st, _, ok := strings.Cut(vol, ":"); ok && st != "" {
out[st] = true
}
}
return out
}
// PreflightRestoreSpace applies the three rules. `configured` is `backup.restore_storage`.
func PreflightRestoreSpace(ctx context.Context, space RestoreSpace, policy SpacePolicy, archive, rawCfg, configured string) SpaceVerdict {
if space == nil {
return SpaceVerdict{Storage: configured, Reason: "no space check is wired — refusing (fail-closed)"}
}
restored, src, err := space.RestoredBytes(ctx, archive)
if err != nil || restored <= 0 {
return SpaceVerdict{Storage: configured, Reason: fmt.Sprintf("cannot tell how much the restore writes (%v)", err)}
}
required := policy.requiredBytes(restored)
own := sourceStorages(rawCfg)
// Rule 2: the configured storage holds the guest under test → try the others first.
var order []string
avoided := ""
if own[configured] {
eligible, eerr := space.Eligible(ctx)
if eerr == nil {
sort.Strings(eligible)
for _, s := range eligible {
if s != configured && !own[s] {
order = append(order, s)
}
}
}
if len(order) > 0 {
avoided = configured
}
}
order = append(order, configured)
var last SpaceVerdict
for _, s := range order {
fr, ferr := space.Free(ctx, s)
if ferr != nil {
last = SpaceVerdict{Storage: s, Required: required, Reason: fmt.Sprintf("cannot read free space on %s (%v)", s, ferr)}
continue
}
ok, why := fits(fr, required)
v := SpaceVerdict{OK: ok, Storage: s, Required: required, Avail: fr.AvailBytes}
if ok {
if s != configured {
v.Avoided = avoided
}
return v
}
v.Reason = fmt.Sprintf("not enough space on %s: restoring %s (%s) %s", s, gib(restored), src, why)
last = v
}
return last
}
@@ -0,0 +1,156 @@
package reconcile
import (
"context"
"errors"
"path/filepath"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 (v0.133.0) — the restore-test's space preflight. Every test asserts the CONSEQUENCE: whether the
// Proxmox API was asked to restore anything, where to, and what the result says — never only the verdict.
const gb = int64(1000 * 1000 * 1000)
// fakeSpace is a configurable RestoreSpace.
type fakeSpace struct {
restored int64
restoredErr error
free map[string]StorageFree
freeErr map[string]error
eligible []string
}
func (f fakeSpace) RestoredBytes(context.Context, string) (int64, string, error) {
return f.restored, "fake", f.restoredErr
}
func (f fakeSpace) Free(_ context.Context, s string) (StorageFree, error) {
if err := f.freeErr[s]; err != nil {
return StorageFree{}, err
}
fr, ok := f.free[s]
if !ok {
return StorageFree{}, errors.New("unknown storage")
}
return fr, nil
}
func (f fakeSpace) Eligible(context.Context) ([]string, error) { return f.eligible, nil }
// thin9201 is demo-hp's local-lvm at 10:29 on 2026-09-24, just before the restore-test that filled it:
// 23.2 GB free, 33.3 GB used, metadata 2.65 %.
var thin9201 = StorageFree{AvailBytes: 23210892 * 1024, UsedBytes: 33277043 * 1024, Thin: true, MetaUsedFraction: 0.0265, MetaKnown: true}
// archive9201 is 9201's archive config: both volumes on local-lvm.
const archive9201 = "hostname: demo-hp\nrootfs: local-lvm:vm-9201-disk-0,size=32G\nmp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G\nmp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\n"
func spaceEngine(t *testing.T, api *fakeAPI, sp RestoreSpace) (*Engine, *Journal) {
t.Helper()
j, err := OpenJournal(filepath.Join(t.TempDir(), "journal.log"))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { j.Close() })
q := NewQueue()
t.Cleanup(q.Close)
return NewEngine(EngineOptions{API: api, Queue: q, Journal: j, RestoreSpace: sp}), j
}
func run9201(e *Engine) RestoreTestResult {
return e.RunRestoreTest(context.Background(), RestoreTestSpec{
Archive: "local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst", RestoreStorage: "local-lvm",
ScratchMin: 990000, ScratchMax: 990009, SourceTier: "local",
})
}
// TestSpace_The2026_09_24TestIsRefused replays R-672: 9201's archive restores 22.6 GB (its vzdump log),
// the pool has 23.2 GB free. The test must NOT start — no restore call, no journaled scratch — and the
// result must say why, as a non-pass.
//
// COMPANION RED-PROOFS (REPORT): (1) the preflight removed (v0.132.0's shape) → a restore into local-lvm
// is issued; (2) `restored` taken from the archive FILE (6.9 GB, the brief's "archive × 1.2 + 5 GiB") →
// 6.9×1.2+5.4 = 13.7 GB < 23.2 GB free, so the test starts — the defect the uncompressed size exists for.
func TestSpace_The2026_09_24TestIsRefused(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
e, j := spaceEngine(t, api, fakeSpace{restored: 22607360000, free: map[string]StorageFree{"local-lvm": thin9201}})
res := run9201(e)
if len(api.restores) != 0 {
t.Fatalf("a restore was issued into a pool that cannot take it: %+v", api.restores)
}
if len(j.InFlight()) != 0 {
t.Fatalf("a scratch entry was journaled for a test that must not start: %+v", j.InFlight())
}
if res.Pass || !res.Skipped || !strings.Contains(res.SkipReason, "not enough space on local-lvm") {
t.Fatalf("result = pass=%v skipped=%v reason=%q — want a non-pass skip naming the storage", res.Pass, res.Skipped, res.SkipReason)
}
if res.RequiredBytes < 32*gb || res.AvailBytes != thin9201.AvailBytes {
t.Fatalf("required=%d avail=%d — want ≥ 32 GB required (22.6 × 1.2 + 5 GiB) against 23.2 GB", res.RequiredBytes, res.AvailBytes)
}
}
// TestSpace_KeepsOffTheTestedGuestsPool — rule 2: another eligible storage that fits takes the restore.
func TestSpace_KeepsOffTheTestedGuestsPool(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
e, _ := spaceEngine(t, api, fakeSpace{restored: 22607360000, eligible: []string{"local-lvm", "big-dir"},
free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: 10 * gb, Thin: true, MetaKnown: true}, "big-dir": {AvailBytes: 500 * gb}}})
res := run9201(e)
if len(api.restores) != 1 || api.restores[0].Storage != "big-dir" {
t.Fatalf("restores = %+v — want ONE restore onto big-dir, off 9201's own pool", api.restores)
}
if res.TargetStorage != "big-dir" {
t.Fatalf("target=%q", res.TargetStorage)
}
for k, v := range api.restores[0].MountOverrides {
if strings.HasPrefix(v, "local-lvm:") {
t.Fatalf("%s still lands on the tested guest's pool: %s", k, v)
}
}
}
// TestSpace_OnlyOnePool_RuleOneDecides — no other eligible storage: the tested guest's pool is used when it
// fits (demo-hp's real shape: nvme-scratch takes rootdir but the agent holds no AllocateSpace there).
func TestSpace_OnlyOnePool_RuleOneDecides(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
e, _ := spaceEngine(t, api, fakeSpace{restored: 2 * gb, eligible: []string{"local-lvm"}, free: map[string]StorageFree{"local-lvm": thin9201}})
res := run9201(e)
if len(api.restores) != 1 || api.restores[0].Storage != "local-lvm" || res.Skipped || res.TargetStorage != "local-lvm" {
t.Fatalf("restores=%+v skipped=%v — a 2 GB restore fits 23 GB free on the only pool", api.restores, res.Skipped)
}
}
// TestSpace_UnknownRefuses — rule 3, one case per unknown. Nothing is restored in any of them.
func TestSpace_UnknownRefuses(t *testing.T) {
cases := map[string]RestoreSpace{
"no space check wired": nil,
"restore size unknown": fakeSpace{restoredErr: errors.New("no vzdump log"), free: map[string]StorageFree{"local-lvm": thin9201}},
"free space unreadable": fakeSpace{restored: gb, freeErr: map[string]error{"local-lvm": errors.New("api down")}},
"thin metadata unknown": fakeSpace{restored: gb, free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: gb, Thin: true}}},
"metadata would overrun": fakeSpace{restored: 10 * gb, free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: 10 * gb, Thin: true, MetaUsedFraction: 0.5, MetaKnown: true}}},
}
for name, sp := range cases {
t.Run(name, func(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
var e *Engine
if sp == nil {
e, _ = spaceEngine(t, api, nil)
} else {
e, _ = spaceEngine(t, api, sp)
}
res := run9201(e)
if len(api.restores) != 0 || res.Pass || !res.Skipped || res.SkipReason == "" {
t.Fatalf("restores=%d pass=%v skipped=%v reason=%q — an unknown must refuse before anything moves",
len(api.restores), res.Pass, res.Skipped, res.SkipReason)
}
})
}
}
// TestSpace_SourceStorages reads the tested guest's pools from the ARCHIVE's config, binds excluded.
func TestSpace_SourceStorages(t *testing.T) {
got := sourceStorages(archive9201 + "mp1: other:vm-9201-disk-2,mp=/x,size=1G\n[snap]\nrootfs: snapstore:x\n")
if !got["local-lvm"] || !got["other"] || got["snapstore"] || len(got) != 2 {
t.Fatalf("sourceStorages = %v — want local-lvm + other, binds and snapshot sections excluded", got)
}
}
+16
View File
@@ -0,0 +1,16 @@
package reconcile
import "context"
// roomySpace is the permissive RestoreSpace the pre-R-672 restore-test tests run with: 1 GiB restored,
// 1 TiB free, thin metadata known and low. The space rules themselves are pinned in
// restoretest_space_test.go.
type roomySpace struct{}
func (roomySpace) RestoredBytes(context.Context, string) (int64, string, error) {
return 1 << 30, "test", nil
}
func (roomySpace) Free(context.Context, string) (StorageFree, error) {
return StorageFree{AvailBytes: 1 << 40, UsedBytes: 1 << 30, Thin: true, MetaUsedFraction: 0.01, MetaKnown: true}, nil
}
func (roomySpace) Eligible(context.Context) ([]string, error) { return nil, nil }
+176
View File
@@ -0,0 +1,176 @@
// Package restorespace is the production seam behind reconcile.RestoreSpace (R-672, agent v0.133.0):
// how much a restore of an archive writes, how much a storage has free, and which storages a
// restore-test may target. Every read that cannot answer returns an error — the preflight then
// REFUSES (reconcile/restoretest_space.go rule 3); nothing here guesses.
package restorespace
import (
"context"
"fmt"
"os"
"path"
"regexp"
"strconv"
"strings"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// API is the Proxmox subset the provider reads.
type API interface {
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
Permissions(ctx context.Context, aclPath string) (map[string]int, error)
}
// Provider implements reconcile.RestoreSpace.
type Provider struct {
API API
// ThinMeta reads a thin pool's metadata-used fraction (storage.HostOps.ThinPoolMetadata).
ThinMeta func(ctx context.Context, vg, pool string) (float64, bool)
// ReadFile reads a vzdump log; nil → os.ReadFile.
ReadFile func(name string) ([]byte, error)
}
var _ reconcile.RestoreSpace = (*Provider)(nil)
// totalWrittenRe is vzdump's own count of the bytes tar wrote into the archive — the UNCOMPRESSED size,
// i.e. what a restore writes back ("INFO: Total bytes written: 22607360000 (22GiB, 49MiB/s)").
var totalWrittenRe = regexp.MustCompile(`Total bytes written:\s*(\d+)`)
// archiveExts are the vzdump archive suffixes; the log is the archive name without it + ".log".
var archiveExts = []string{".tar.zst", ".tar.gz", ".tar.lzo", ".tgz", ".tar"}
func (p *Provider) readFile(name string) ([]byte, error) {
if p.ReadFile != nil {
return p.ReadFile(name)
}
return os.ReadFile(name)
}
// RestoredBytes: a file-backed archive → its vzdump log's "Total bytes written"; a PBS archive → the
// size Proxmox reports for the snapshot (its logical, uncompressed size). The archive FILE size is never
// used: it is compressed (6.9 GB for a 22.6 GB restore, measured 2026-09-24).
func (p *Provider) RestoredBytes(ctx context.Context, archive string) (int64, string, error) {
id, vol, ok := strings.Cut(archive, ":")
if !ok || id == "" || vol == "" {
return 0, "", fmt.Errorf("not a storage volid: %q", archive)
}
st, err := p.storageConfig(ctx, id)
if err != nil {
return 0, "", err
}
switch st.Type {
case "pbs":
items, err := p.API.StorageContent(ctx, id)
if err != nil {
return 0, "", fmt.Errorf("list %s: %w", id, err)
}
for _, it := range items {
if it.VolID == archive && it.Size > 0 {
return it.Size, "pbs snapshot size", nil
}
}
return 0, "", fmt.Errorf("archive %s not listed on %s with a size", archive, id)
default:
if st.Path == "" {
return 0, "", fmt.Errorf("storage %s (%s) has no path to read a vzdump log from", id, st.Type)
}
base := path.Base(vol) // "backup/vzdump-lxc-…tar.zst" → "vzdump-lxc-…tar.zst"
stem := ""
for _, ext := range archiveExts {
if strings.HasSuffix(base, ext) {
stem = strings.TrimSuffix(base, ext)
break
}
}
if stem == "" {
return 0, "", fmt.Errorf("unknown archive suffix: %s", base)
}
logPath := path.Join(st.Path, "dump", stem+".log")
b, err := p.readFile(logPath)
if err != nil {
return 0, "", fmt.Errorf("read vzdump log %s: %w", logPath, err)
}
m := totalWrittenRe.FindSubmatch(b)
if m == nil {
return 0, "", fmt.Errorf("vzdump log %s carries no \"Total bytes written\"", logPath)
}
n, err := strconv.ParseInt(string(m[1]), 10, 64)
if err != nil || n <= 0 {
return 0, "", fmt.Errorf("vzdump log %s: bad byte count %q", logPath, m[1])
}
return n, "vzdump log: total bytes written", nil
}
}
func (p *Provider) storageConfig(ctx context.Context, id string) (proxmox.Storage, error) {
all, err := p.API.ListStorage(ctx)
if err != nil {
return proxmox.Storage{}, fmt.Errorf("list storage config: %w", err)
}
for _, s := range all {
if s.Storage == id {
return s, nil
}
}
return proxmox.Storage{}, fmt.Errorf("storage %s not configured", id)
}
// Free reads the node's live usage for `storage`; a thin pool adds its metadata fill.
func (p *Provider) Free(ctx context.Context, storage string) (reconcile.StorageFree, error) {
live, err := p.API.NodeStorage(ctx)
if err != nil {
return reconcile.StorageFree{}, fmt.Errorf("node storage: %w", err)
}
for _, s := range live {
if s.Storage != storage {
continue
}
if s.Active != 1 {
return reconcile.StorageFree{}, fmt.Errorf("storage %s is not active", storage)
}
fr := reconcile.StorageFree{AvailBytes: s.Avail, UsedBytes: s.Used, Thin: s.Type == "lvmthin"}
if fr.Thin {
cfg, cerr := p.storageConfig(ctx, storage)
if cerr == nil && p.ThinMeta != nil && cfg.VGName != "" && cfg.ThinPool != "" {
fr.MetaUsedFraction, fr.MetaKnown = p.ThinMeta(ctx, cfg.VGName, cfg.ThinPool)
}
}
return fr, nil
}
return reconcile.StorageFree{}, fmt.Errorf("storage %s not reported by the node", storage)
}
// Eligible: active, content includes `rootdir`, and the agent holds Datastore.AllocateSpace on
// /storage/<id>. The permission is read for the SPECIFIC privilege — the box-wide grant answers every
// path with inherited privileges (proxmox.Client.Permissions).
func (p *Provider) Eligible(ctx context.Context) ([]string, error) {
live, err := p.API.NodeStorage(ctx)
if err != nil {
return nil, fmt.Errorf("node storage: %w", err)
}
var out []string
for _, s := range live {
if s.Active != 1 || !hasContent(s.Content, "rootdir") {
continue
}
privs, perr := p.API.Permissions(ctx, "/storage/"+s.Storage)
if perr != nil || privs["Datastore.AllocateSpace"] != 1 {
continue
}
out = append(out, s.Storage)
}
return out, nil
}
func hasContent(list, want string) bool {
for _, c := range strings.Split(list, ",") {
if strings.TrimSpace(c) == want {
return true
}
}
return false
}
+133
View File
@@ -0,0 +1,133 @@
package restorespace
import (
"context"
"errors"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 (v0.133.0). The provider behind the restore-test's space preflight. No test reaches a real
// Proxmox or a real file: the API and ReadFile are fakes.
type fakeAPI struct {
cfg []proxmox.Storage
live []proxmox.Storage
content map[string][]proxmox.StorageContent
perms map[string]map[string]int
}
func (f fakeAPI) ListStorage(context.Context) ([]proxmox.Storage, error) { return f.cfg, nil }
func (f fakeAPI) NodeStorage(context.Context) ([]proxmox.Storage, error) { return f.live, nil }
func (f fakeAPI) StorageContent(_ context.Context, s string) ([]proxmox.StorageContent, error) {
return f.content[s], nil
}
func (f fakeAPI) Permissions(_ context.Context, p string) (map[string]int, error) {
if m, ok := f.perms[p]; ok {
return m, nil
}
return map[string]int{}, nil
}
const archive = "local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst"
// The real log's tail (demo-hp, 2026-09-23): 22.6 GB written, a 6.91 GB archive file.
const vzdumpLog = "2026-09-23 07:02:47 INFO: Total bytes written: 22607360000 (22GiB, 49MiB/s)\n2026-09-23 07:02:47 INFO: archive file size: 6.91GB\n"
// demoHP is demo-hp's storage layout: `local` (dir, backups), `local-lvm` (thin), `nvme-scratch` (dir,
// rootdir — but the agent holds NO grant there), a pbs.
func demoHP() fakeAPI {
return fakeAPI{
cfg: []proxmox.Storage{
{Storage: "local", Type: "dir", Path: "/var/lib/vz"},
{Storage: "local-lvm", Type: "lvmthin", VGName: "pve", ThinPool: "data"},
{Storage: "nvme-scratch", Type: "dir", Path: "/mnt/hdd_1"},
{Storage: "felhom-pbs", Type: "pbs"},
},
live: []proxmox.Storage{
{Storage: "local", Type: "dir", Content: "vztmpl,backup,iso,import", Active: 1, Avail: 4 << 30, Used: 34 << 30},
{Storage: "local-lvm", Type: "lvmthin", Content: "images,rootdir", Active: 1, Avail: 23210892 * 1024, Used: 33277043 * 1024},
{Storage: "nvme-scratch", Type: "dir", Content: "images,rootdir", Active: 1, Avail: 800 << 30},
{Storage: "felhom-pbs", Type: "pbs", Content: "backup", Active: 0},
},
content: map[string][]proxmox.StorageContent{
"local": {{VolID: archive, Size: 7417540996}},
"felhom-pbs": {{VolID: "felhom-pbs:backup/ct/9201/2026-09-23T02:00:00Z", Size: 21 << 30}},
},
perms: map[string]map[string]int{
"/storage/local": {"Datastore.Audit": 1, "Datastore.AllocateSpace": 1},
"/storage/local-lvm": {"Datastore.Audit": 1, "Datastore.AllocateSpace": 1},
// nvme-scratch: only the inherited box-wide Datastore.Audit — the trap proxmox.Permissions names.
"/storage/nvme-scratch": {"Datastore.Audit": 1},
},
}
}
// TestRestoredBytes_ReadsTheUncompressedSize — the vzdump log's "Total bytes written", never the archive
// FILE size (6.9 GB for a 22.6 GB restore).
//
// COMPANION RED-PROOF (REPORT): return the storage content's Size for a dir storage → 7417540996, and
// this test fails at "the compressed file size was used".
func TestRestoredBytes_ReadsTheUncompressedSize(t *testing.T) {
var asked string
p := &Provider{API: demoHP(), ReadFile: func(n string) ([]byte, error) { asked = n; return []byte(vzdumpLog), nil }}
n, src, err := p.RestoredBytes(context.Background(), archive)
if err != nil {
t.Fatal(err)
}
if n == 7417540996 {
t.Fatal("the compressed file size was used — the restore writes 3× that")
}
if n != 22607360000 || asked != "/var/lib/vz/dump/vzdump-lxc-9201-2026_09_23-06_55_25.log" {
t.Fatalf("n=%d from %q (log %q)", n, src, asked)
}
}
func TestRestoredBytes_UnknownIsAnError(t *testing.T) {
cases := map[string]*Provider{
"no log": {API: demoHP(), ReadFile: func(string) ([]byte, error) { return nil, errors.New("ENOENT") }},
"log without size": {API: demoHP(), ReadFile: func(string) ([]byte, error) { return []byte("ERROR: failed\n"), nil }},
}
for name, p := range cases {
if n, _, err := p.RestoredBytes(context.Background(), archive); err == nil {
t.Fatalf("%s: got %d, want an error (the preflight then refuses)", name, n)
}
}
if _, _, err := (&Provider{API: demoHP()}).RestoredBytes(context.Background(), "local:backup/weird.vma"); err == nil {
t.Fatal("an unknown archive suffix must be an error")
}
}
func TestRestoredBytes_PBS(t *testing.T) {
p := &Provider{API: demoHP()}
n, _, err := p.RestoredBytes(context.Background(), "felhom-pbs:backup/ct/9201/2026-09-23T02:00:00Z")
if err != nil || n != 21<<30 {
t.Fatalf("n=%d err=%v", n, err)
}
}
// TestEligible_NeedsTheSpecificGrant — nvme-scratch takes rootdir but the agent holds only the inherited
// Datastore.Audit there, so it is NOT eligible; `local` holds no rootdir.
func TestEligible_NeedsTheSpecificGrant(t *testing.T) {
got, err := (&Provider{API: demoHP()}).Eligible(context.Background())
if err != nil || len(got) != 1 || got[0] != "local-lvm" {
t.Fatalf("eligible = %v (%v) — want only local-lvm on demo-hp", got, err)
}
}
func TestFree_ThinCarriesMetadata(t *testing.T) {
p := &Provider{API: demoHP(), ThinMeta: func(_ context.Context, vg, pool string) (float64, bool) {
if vg != "pve" || pool != "data" {
t.Fatalf("metadata read for %s/%s", vg, pool)
}
return 0.0265, true
}}
fr, err := p.Free(context.Background(), "local-lvm")
if err != nil || !fr.Thin || !fr.MetaKnown || fr.MetaUsedFraction != 0.0265 || fr.AvailBytes != 23210892*1024 {
t.Fatalf("free = %+v err=%v", fr, err)
}
if _, err := p.Free(context.Background(), "felhom-pbs"); err == nil {
t.Fatal("an inactive storage must be an error")
}
}
+41
View File
@@ -6,6 +6,7 @@ import (
"log/slog"
"regexp"
"strings"
"sync"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
@@ -35,6 +36,45 @@ type Observer struct {
host HostReader
ops HostOps
logger *slog.Logger
// R-672 (v0.133.0): a thin pool crossing thinPoolAlarmFraction (data OR metadata) requests an
// out-of-band host report at once, so the hub's storage-fill alarm sees it in seconds instead of at
// the next 15-minute report. Rising edge per pool; re-armed below thinPoolRearmFraction.
highMu sync.Mutex
high map[string]bool
onThinHigh func()
}
const (
thinPoolAlarmFraction = 0.90
thinPoolRearmFraction = 0.85
)
// SetThinHighTrigger wires the out-of-band report request (main: the storage trigger channel).
func (o *Observer) SetThinHighTrigger(f func()) { o.onThinHigh = f }
// noteThinFill is the edge detector. key separates data from metadata so each has its own edge.
func (o *Observer) noteThinFill(key string, frac float64) {
o.highMu.Lock()
if o.high == nil {
o.high = map[string]bool{}
}
fire := false
switch {
case frac >= thinPoolAlarmFraction && !o.high[key]:
o.high[key] = true
fire = true
case frac < thinPoolRearmFraction && o.high[key]:
o.high[key] = false
}
o.highMu.Unlock()
if fire {
o.logger.Error("storage: thin pool crossed 90% — requesting an immediate host report (the hub alarms)",
"pool", key, "fraction", frac)
if o.onThinHigh != nil {
o.onThinHigh()
}
}
}
// NewObserver builds an Observer. host defaults to a ProcHostReader; logger to the
@@ -260,6 +300,7 @@ func (o *Observer) build(s proxmox.Storage, mounts []Mount) observed {
o.logger.Warn("storage: lvmthin pool data fill is high (a full pool corrupts every guest on it)",
"storage", s.Storage, "data_used_fraction", frac)
}
o.noteThinFill(s.Storage+"/data", frac)
}
// SMART-only device hint (v0.95.0): a dir-storage that lives INSIDE a shared filesystem (the
+34
View File
@@ -0,0 +1,34 @@
package storage
import (
"context"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 (v0.133.0): a thin pool crossing 90 % requests an out-of-band host report ONCE, through the
// watchdog's own read path (Known — every few seconds), so the hub's storage-fill alarm sees the pool in
// seconds, not at the next 15-minute report. Re-armed below 85 %.
//
// COMPANION RED-PROOF (REPORT): remove the noteThinFill call from the data path → "no report was
// requested when the pool crossed 90 %".
func TestThinHigh_RequestsOneReportPerCrossing(t *testing.T) {
pool := proxmox.Storage{Storage: "local-lvm", Type: "lvmthin", Content: "rootdir,images", Total: 1000, Active: 1}
api := &fakeStorageAPI{node: "n", cluster: []proxmox.Storage{pool}}
o := NewObserver(api, &fakeHostReader{}, nil, quietLogger())
asked := 0
o.SetThinHighTrigger(func() { asked++ })
for i, used := range []int64{800, 910, 950, 1000, 840, 920} {
p := pool
p.Used, p.Avail, p.UsedFraction = used, 1000-used, float64(used)/1000
api.nodeSt = []proxmox.Storage{p}
if _, err := o.Known(context.Background()); err != nil {
t.Fatal(err)
}
want := map[int]int{0: 0, 1: 1, 2: 1, 3: 1, 4: 1, 5: 2}[i]
if asked != want {
t.Fatalf("after %d/1000 used: %d report requests, want %d", used, asked, want)
}
}
}
+5
View File
@@ -43,6 +43,9 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
SHARED_INSTRUCTIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
# R-389 — shared, like the two above: it lives in felhom.eu/scripts/ and is never copied.
SHARED_OBSERVATIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "observations_gate.py")
# (label, absolute script path, args, fast)
GATES = [
@@ -53,6 +56,8 @@ GATES = [
# missing TAG is what actually broke every install, and the pre-push hook is the earliest place
# that can catch it.
("release-complete", os.path.join(ROOT, "scripts", "check-release-complete.py"), [], True),
# R-389 — a REPORT.md observation with no register row behind it. Fast: stdlib file reads.
("observations", SHARED_OBSERVATIONS, [ROOT], True),
]
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}