Compare commits
39 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 9cac3462bb | |||
| 1bb8608e88 | |||
| b84e0dd1bd | |||
| 23a8ef3de4 | |||
| 596238cc2e | |||
| 475bdce7e4 | |||
| d766666ff8 | |||
| a4c09a7c11 | |||
| 904dc20466 | |||
| e1b8269be0 | |||
| 5c68c869b6 | |||
| 728d12b1a0 | |||
| dd81866b16 | |||
| 3ef095fb71 | |||
| 7c986915ca | |||
| 16dbc83221 | |||
| 9555a7f93b | |||
| 9ff937d8fb | |||
| d4be12ca95 | |||
| 7403c2a838 | |||
| 309e368731 | |||
| 0722b2cdb0 | |||
| 4fe2f81a32 | |||
| 9bdb4dae8f | |||
| d9864a94bf | |||
| 77cd70f7c0 | |||
| 1030abd7d6 | |||
| 18d03bd437 | |||
| e98b857684 | |||
| dcdeb3d16d | |||
| 610804b98d | |||
| 4586f0f7f6 | |||
| 205e22babe | |||
| 058b945064 | |||
| 40d857b527 | |||
| 7ae6990bac | |||
| 7569f34aeb | |||
| ede49b610d | |||
| f17ed11599 |
@@ -9,3 +9,6 @@
|
||||
|
||||
# go
|
||||
/vendor/
|
||||
|
||||
# Python bytecode written by configs/test_felhom_os_apply.py
|
||||
configs/__pycache__/
|
||||
|
||||
+317
@@ -1,3 +1,318 @@
|
||||
## v0.139.0 — a DR restore never lands beside a live original (2026-10-04, R-834)
|
||||
|
||||
> **RELEASED 2026-10-04** by `scripts/release-agent.sh` — tag `v0.139.0` (`475bdce`), sha256
|
||||
> `8534a9be368a6d24d8065db77436e86900443c5c6554c71f6fe91c2bdb9d0b9c`, verified by download. **Not vouched.**
|
||||
|
||||
**MinAgent impact:** none required by any controller.
|
||||
|
||||
- The DR bring-up (`--selftest=bring-up -mode dr`) keeps the archive's `onboot: 1`, binds the host's REAL drives
|
||||
(`mp8 /mnt/felhom-drives`) and STARTS the guest — right on a replaced host, wrong beside a live original (a second
|
||||
controller for the same household on the same drives). It now REFUSES, before any restore, when the archive's
|
||||
source guest still exists on the host, when any guest binds the drives parent, or when a guest's config cannot be
|
||||
read (fail closed). On a replaced host it proceeds and keeps its binds, unchanged.
|
||||
- The restore-test was MEASURED safe live on demo-hp (onboot 0 and throwaway stand-ins for mp8/mp9 from the first
|
||||
config read to teardown); a test now pins its "no host path" half beside the existing onboot test.
|
||||
- No sudoers change: the restore-test sets onboot 0 through the API create call, and DR refuses rather than degrade,
|
||||
so no `-onboot 0` line is needed.
|
||||
- Tests: `TestRunBringUp_DRRefusesBesideALiveOriginal` (source guest present / drives bind on another guest / an
|
||||
unreadable config refuse; a replaced host proceeds and keeps the drives bind), `TestRunBringUp_ProvisionNotBlockedByADrivesBind`,
|
||||
`TestArchiveSourceVMID`, `TestRestoreTest_NoHostPathBindBesideTheOriginal`. Red-proofs: the DR check returning ""
|
||||
→ three refusal cases restore and START; the restore-test's mp8 override set to the host path → fails.
|
||||
|
||||
## v0.138.0 — the restore test takes only THIS box's archives (2026-09-30, R-727, `09` §3 decision 51)
|
||||
|
||||
> **RELEASED 2026-09-30** by `scripts/release-agent.sh` — tag `v0.138.0` (`e1b8269`), sha256
|
||||
> `55916026001790a79ebf97d32c032610cfde8e09d02979b9b9d8c2cbc5d88195`, verified by download. **Not vouched** (the golden
|
||||
> keeps 0.137.0 until the next bake); delivered to the demo boxes by signed `agent_update` jobs.
|
||||
>
|
||||
> **VOUCHED 2026-09-30** with golden 0.283.1 (`min_agent` 0.131.0), on the operator's word: the hub logged
|
||||
> `Artifact manifest set: agent=0.138.0 golden=0.283.1 min_agent="0.131.0"`. Evidence:
|
||||
> `felhom.eu/documentation/audits/evidence-golden-0283-2026-09-30/`.
|
||||
|
||||
**MinAgent impact:** none required by any controller.
|
||||
|
||||
- A returning customer's PBS namespace can hold archives of EARLIER boxes: same guest id (9201), same token, written
|
||||
with a different key. Measured 2026-09-30: the newest SETTLED archive was an earlier box's, and the test failed
|
||||
`wrong key … manifest's key 6b:ca:5f:3f… does not match provided key de:51:7a:18…` every evaluation. The archive
|
||||
carries no host id; it carries its key fingerprint (PVE content `encrypted`), and the storage carries its own
|
||||
(`GET /storage` → `encryption-key`). `PickSettledRestoreCandidateOn` now skips — and logs by name, once — an archive
|
||||
whose fingerprint is not the storage's own; an unencrypted storage is not filtered; a failed storage read is an
|
||||
error (tier UNKNOWN), never "nothing to prove".
|
||||
- Tests: `TestR727_TheRestoreTestTakesOnlyThisBoxsArchives` (the 2026-09-30 shape: nothing picked while this box's
|
||||
archive settles, then exactly it), `TestR727_UnencryptedStorageIsNotFiltered`, `TestR727_KeyLookupFailureIsUnknown`.
|
||||
Red-proof RP39: the skip removed → the earlier box's `2026-09-16T21:59:54Z` is picked.
|
||||
|
||||
## v0.137.0 — a guest outside the agent's ACL is not a known guest (2026-09-27, R-689, v0.136.0 regression)
|
||||
|
||||
> **RELEASED 2026-09-27** by `scripts/release-agent.sh` — tag `v0.137.0` (`3ef095f`), sha256 `766c9166916a1bd3674b0dc69081f8a7619e770f1402d8ad7705b395937e7627`, verified by download. **NOT vouched** (the operator's act).
|
||||
>
|
||||
> **VOUCHED 2026-09-28** with golden 0.276.0 (`min_agent` 0.131.0), on the operator's word of 2026-09-27: the hub logged
|
||||
> `Artifact manifest set: agent=0.137.0 golden=0.276.0 min_agent="0.131.0"`, and a Day-0 test install fetched this binary
|
||||
> through the manifest and sha-verified it. Evidence: `felhom.eu/documentation/audits/evidence-golden-0276-2026-09-28/`.
|
||||
|
||||
**MinAgent impact:** none required by any controller.
|
||||
|
||||
- v0.136.0 asked `GuestConfig` whether an archive's guest exists and treated anything but "does not exist" as a lookup
|
||||
failure. PVE answers **403 "permission denied at /vms/<id>"** for a vmid outside the token's pool — so on demo-hp the
|
||||
deleted guest 9100's archive made the local tier UNKNOWN every evaluation. A guest the agent cannot read is not one it
|
||||
manages; its archive is skipped. `TestR689_AGuestOutsideTheAgentsACLIsNotAKnownGuest`, red-proofed. Verified read-only
|
||||
on demo-hp with the pre-release binary before the release.
|
||||
|
||||
## v0.136.0 — … of a guest that still EXISTS (2026-09-27, R-689 second half)
|
||||
|
||||
> **RELEASED 2026-09-27** by `scripts/release-agent.sh` — tag `v0.136.0` (`16dbc83`), sha256 `2eb0b5ebe253defd68b322312bbac12418c051b0a7a0d2d1831310d97fa6d755`, verified by download. **NOT vouched** (the operator's act).
|
||||
|
||||
**MinAgent impact:** none required by any controller.
|
||||
|
||||
- **R-689, second half.** Read on demo-hp right after v0.135.0 with the read-only `-selftest=restore-test-due`: the
|
||||
golden was skipped, and the pick fell to `vzdump-lxc-9100-2026_08_21…` — a leftover of a guest deleted in August. A
|
||||
candidate's guest must now exist on this node (`GuestConfig`): "does not exist" skips the archive (INFO once per
|
||||
volid); any other lookup failure is returned, so the tier reads UNKNOWN, never "nothing to prove". Tests
|
||||
`TestR689_AnArchiveOfADeletedGuestIsNeverPicked` (red-proofed: without the check it picks the 9100 leftover),
|
||||
`TestR689_AGuestLookupFailureIsUnknownNotEmpty`. A second agent release in one session — the first half was found
|
||||
incomplete on the box.
|
||||
|
||||
## v0.135.0 — the restore test proves only backups OF A GUEST (2026-09-27, R-689)
|
||||
|
||||
> **RELEASED 2026-09-27** by `scripts/release-agent.sh` — tag `v0.135.0` (`d4be12c`), sha256 `ad4e75f16d338552f4588d3fe64c51cbf9651220b9d223b5386b85b6c37fd4c3`, verified by download. **NOT vouched** (the operator's act).
|
||||
|
||||
**MinAgent impact:** none required by any controller.
|
||||
|
||||
- **R-689** (`backup/runner.go` `guestBackupArchive`). demo-hp keeps its golden template in `local:backup/` — content
|
||||
"backup", 654 MB, plausibly complete, the newest settled entry — and the scheduled restore test picked it every 6 h and
|
||||
failed `extractconfig` with a 403, while the guest's real archive went untested. A restore-test candidate is now a
|
||||
`vzdump-{lxc,qemu}-<vmid>-…` file or a PBS `backup/{ct,vm}/<vmid>/…` snapshot whose vmid the storage reports; anything
|
||||
else is skipped with one INFO line per volid (not the "INCOMPLETE archive" WARN). Tests
|
||||
`TestR689_TheRestoreTestNeverPicksTheGolden` (red-proof: without the check it picks the golden) and
|
||||
`TestR689_GuestBackupArchiveShapes`; two older picker tests' fixtures moved to real archive names.
|
||||
Evidence: `felhom.eu/documentation/audits/version-travel-2026-09-26/D1/`.
|
||||
|
||||
## v0.134.0 — a whole-box backup that cannot fit is skipped with a reason, before anything starts (2026-09-25 night, R-685)
|
||||
|
||||
> **RELEASED 2026-09-24 night** by `scripts/release-agent.sh` — tag `v0.134.0` (`0722b2c`), sha256 `7593bebe03234c7d22f3ade384e7ed7787dc659aa8c8594b3ee19af2ce81c72d`, verified by download. **NOT vouched** (Day-0 stays on the previous version).
|
||||
|
||||
**MinAgent impact:** none required by any controller.
|
||||
|
||||
- **R-685 — the backup space preflight** (`backup/runner.go` `spaceFits`). Before a vzdump to a LOCAL (non-PBS)
|
||||
target, the newest archive of that guest on that target × 1.25 + 1 GiB must be free — PVE prunes old archives
|
||||
only AFTER a successful backup, so the kept ones still stand during the run. A shortfall is a skip: nothing
|
||||
starts, the record's Error reads `skipped: not enough space: <target> has X GiB free; the last archive of guest N
|
||||
was Y GiB, so a new one needs about Z GiB …` (stable prefix `BackupSkipNoSpacePrefix`), and it reaches the
|
||||
controller's tier view and, through the quiesce loop's tier notifier, the operator's `whole_guest_backup_failed`.
|
||||
It FAILS OPEN on what is not known (PBS target, first backup, unreadable usage).
|
||||
- **Free space is read from `GET /nodes/<node>/storage`** (`NodeStorage`), never `GET /storage` — the latter is the
|
||||
cluster DEFINITIONS and carries no usage. **Found live, before release:** the first build read `/storage`,
|
||||
failed open, and a real vzdump of demo-hp 9201 started from the live test; it was aborted after 5 min 16 s, no
|
||||
archive left (`felhom.eu/documentation/audits/night-2026-09-25/F/`). The test fake's `ListStorage` now strips
|
||||
usage like production. Red-proofs: two (`…/F/redproof-r685-*.txt`).
|
||||
- Live proof (demo-hp, safe builds with a hard stop before vzdump): ×10 margin → refused, "local has 14.9 GiB
|
||||
free … needs about 77.2 GiB"; release margin → "space preflight passed" need 11.3 GB, avail 16.0 GB.
|
||||
|
||||
## v0.133.0 — a restore-test can never fill a box's disk; leftovers retried on a timer (2026-09-24, R-672, R-673)
|
||||
|
||||
> **RELEASED 2026-09-24** by `scripts/release-agent.sh` — tag `v0.133.0` (`9bdb4da`), sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, verified by an independent anonymous download. **NOT delivered** (needs an operator-signed `agent_update` job per box, R-530) and **NOT vouched**.
|
||||
|
||||
**MinAgent impact:** none required by any controller. Hub **v0.124.0** makes a thin pool CRITICAL at 90 %
|
||||
(data or metadata) and keys the storage-fill alarm per pool per 6 h; an older hub still raises its generic
|
||||
90/95 % storage-fill events from the same report.
|
||||
|
||||
- **R-672 — the space preflight** (`reconcile/restoretest_space.go`, provider `internal/restorespace`). Before
|
||||
anything is journaled or created, a restore-test needs free data ≥ restored × 1.2 + 5 GiB on its target
|
||||
(`backup.restore_test_space_factor` / `backup.restore_test_space_reserve_gib`) and room in a thin pool's
|
||||
metadata. `restored` is the UNCOMPRESSED size — the vzdump log's "Total bytes written" (a file-backed
|
||||
archive) or the PBS snapshot size; the archive FILE is never used (9201: 6.9 GB file, 22.6 GB restore). The
|
||||
target moves OFF the tested guest's own pool when another storage is eligible (active, `rootdir`, and the
|
||||
agent holds Datastore.AllocateSpace there) and fits. Anything unknown refuses. A refusal is reported to the
|
||||
hub as the test's result (`pass=false`, `skipped=true`, "skipped: not enough space on …") — never a pass,
|
||||
never dropped. The archive config is read once (it was read twice). Measured case: demo-hp 2026-09-24 —
|
||||
the restore-test that filled `local-lvm` would now be refused (needs 32.1 GB, 23.2 GB free).
|
||||
- **R-672 — a failed scratch teardown is retried every 10 minutes** (`Engine.RetryScratchTeardown`, the
|
||||
daemon's janitor), not only by Recover at agent start; never a vmid a running test owns; after 3 failed
|
||||
tries the operator is told through a failed restore-test record naming the scratch guest.
|
||||
- **R-672 — a thin pool crossing 90 % requests an immediate host report** (`Observer.SetThinHighTrigger`, the
|
||||
storage watchdog's read path, every few seconds; re-armed below 85 %), so the hub's alarm sees it in
|
||||
seconds, not at the next 15-minute report.
|
||||
- **R-673 — the stale-lock sweep runs on the same 10-minute timer** (it ran only at start: a stale
|
||||
`snapshot-delete` lock blocked 9201's whole-box backups for five hours), holding the one-heavy-operation
|
||||
gate so no agent backup can start between its "no vzdump running" check and its unlock.
|
||||
- Red-proofs: six, each seen failing (REPORT).
|
||||
|
||||
## v0.132.0 — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
|
||||
|
||||
> **RELEASED 2026-09-17** by `scripts/release-agent.sh` — tag `v0.132.0`, sha256 `4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321`. **Vouched 2026-09-17** for Day-0 installs on the operator's word, together with golden **0.246.0** (the hub's R-120 gate required the newer golden first). Delivered to demo-hp and the N100 by an operator-signed `agent_update` job each (operator ruling 3 of 2026-09-16), not by a floor.
|
||||
|
||||
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
|
||||
`controller_slow_crashloop`; an older hub ignores them.
|
||||
|
||||
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
|
||||
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
|
||||
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
|
||||
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
|
||||
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
|
||||
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
|
||||
does **not** stop restarting — the fast brake remains the only brake.
|
||||
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
|
||||
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
|
||||
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
|
||||
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
|
||||
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
|
||||
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
|
||||
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
|
||||
|
||||
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
|
||||
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
|
||||
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
|
||||
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
|
||||
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
|
||||
|
||||
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
|
||||
|
||||
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
|
||||
|
||||
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
|
||||
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
|
||||
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
|
||||
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
|
||||
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
|
||||
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
|
||||
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
|
||||
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
|
||||
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
|
||||
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
|
||||
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
|
||||
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
|
||||
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
|
||||
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
|
||||
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
|
||||
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
|
||||
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
|
||||
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
|
||||
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
|
||||
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
|
||||
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
|
||||
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
|
||||
after an agent restart — `last_success_source: storage`), the last attempt kept apart
|
||||
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
|
||||
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
|
||||
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
|
||||
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
|
||||
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
|
||||
last success".
|
||||
|
||||
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
|
||||
|
||||
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
|
||||
|
||||
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
|
||||
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
|
||||
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
|
||||
absent). **All four found by accident.** The gates enforce everything else here and were the one part
|
||||
nothing had checked.
|
||||
|
||||
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
|
||||
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
|
||||
|
||||
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
|
||||
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
|
||||
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
|
||||
proves the cause was the listing rather than the decoy.
|
||||
|
||||
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
|
||||
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
|
||||
|
||||
**In this repo:** no gate changed, and that is the result. `release-complete` was decoyed and is
|
||||
SOUND — the sweep's attempt (a non-version heading on top of `CHANGELOG.md`) was WITHDRAWN as
|
||||
illegitimate, because `HEAD_RE.search` scans the whole file and still names `v0.130.0`. The three
|
||||
shared gates are covered from `felhom.eu`; the remaining two are named in the decoy-coverage
|
||||
exemption list (R-426) as UNTESTED, not as sound.
|
||||
|
||||
## v0.130.0 — the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
|
||||
|
||||
> **RELEASED 2026-08-20**, on the operator's word, after the fix was proved on both boxes.
|
||||
> `sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, 14,141,158 bytes,
|
||||
> tag `v0.130.0` at `7569f34`. Reproducible: a rebuild with `-trimpath -buildvcs=false` matches the
|
||||
> published artifact byte for byte (R-186's property, checked rather than assumed).
|
||||
>
|
||||
> The heading read `## UNRELEASED — v0.130.0 candidate` until this point, deliberately: while the fix
|
||||
> was hand-installed on `demo-hp` only, publishing would have pushed it onto `demo-felhom` through
|
||||
> self-update and destroyed the control the proof rested on. **`release-complete` convicted on the
|
||||
> release heading and was right to** — the answer was to stop claiming a release, not to bypass the
|
||||
> gate. See R-347.
|
||||
>
|
||||
> **The fleet runs these exact bytes.** Both demo boxes were first given a hand build made during the
|
||||
> proof — same source, same version string, **different bytes** (`256e0829…`), because
|
||||
> `release-agent.sh` builds with `-trimpath -buildvcs=false` and a hand build does not. Nothing would
|
||||
> have corrected that: the boxes already reported `0.130.0`, so self-update saw the vouched version as
|
||||
> installed and would have done nothing, forever. Both were reinstalled from the **downloaded package**
|
||||
> and now report `a56a92a7…`. Filed as **R-349**, because every prove-then-publish train hits it.
|
||||
|
||||
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
|
||||
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
|
||||
proxy whose descriptor ceiling is 65536 and whose runway at that rate was ~323 days. The connections
|
||||
were held open on **both** sides: ep0 showed 388 while the two boxes showed 194 + 194, at two separate
|
||||
instants, with the same source ports on each side, and **not one closed in a 31-minute window**.
|
||||
`ss -tnp` on the boxes named the holder: **`felhom-agent`**, 194 of 194 on each, one PID.
|
||||
|
||||
**`pvestatd` and `proxmox-backup-client` made 162,404 requests in that window and leaked zero.** They
|
||||
are 99.5% of the traffic to that endpoint and 0% of the leak. The agent made 811 requests — of which
|
||||
387 were `GET .../snapshots` — and leaked 388 sockets. One per call, within one.
|
||||
|
||||
**The defect, and it is two things compounding.**
|
||||
|
||||
- `internal/pbs/client.go` built its transport as a composite literal:
|
||||
`&http.Transport{TLSClientConfig: tlsCfg}`. That takes **`IdleConnTimeout` zero, which does not mean
|
||||
"use a sane default" — it means retain idle keep-alive connections FOREVER.**
|
||||
`http.DefaultTransport` sets 90s; hand-rolling the transport (which every client here must do,
|
||||
because they all pin TLS) silently discards it.
|
||||
- `pbsTargetsFromPVE` (`cmd/felhom-agent/main.go`) builds **a fresh `pbs.Client` every cycle**, as its
|
||||
own doc comment says, and drops the previous one. An abandoned `http.Transport` does **not** close
|
||||
its connections — it becomes unreachable while its `persistConn` read-loop goroutine keeps the
|
||||
socket alive. So each cycle stranded exactly one connection that nothing could ever close.
|
||||
|
||||
The cadences reconcile without fitting: a 900s live-snapshot collect (184.7 cycles in the window) plus
|
||||
a 6-hour verify loop (7.7) predicts 192.4 per box against **194 observed**.
|
||||
|
||||
**The fix is one field, restored to the standard library's own value.** New leaf package
|
||||
`internal/httpx` owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is what
|
||||
`http.DefaultTransport` uses, so there is nothing invented here to justify or tune — and
|
||||
`NewTransport(tlsCfg, idleConnTimeout)`, which returns a **fresh** transport (never shared: each caller
|
||||
pins a different endpoint) and treats a zero or negative timeout as **use the default, never "no
|
||||
timeout"**. `pbs.Config` gains an `IdleConnTimeout` field that production leaves unset; only tests set
|
||||
it, to avoid a 90-second wait.
|
||||
|
||||
**`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and
|
||||
were corrected in the same pass — but neither contributed to the ep0 leak, and this entry must not be
|
||||
read as three leaks having been found.** Both are built **once per process**, so they held one idle
|
||||
connection for the life of the daemon rather than accumulating, and neither talks to ep0:8007.
|
||||
|
||||
**Tests, and what they deliberately do not assert.** `internal/pbs/client_leak_test.go` counts
|
||||
connections **server-side** and models what `pbsTargetsFromPVE` actually does — build a client, use it
|
||||
once, drop it on the floor — then asserts the connections go away. It does not assert `err == nil` and
|
||||
it does not assert that some field holds some value; both were true of the leaking code.
|
||||
|
||||
- **Red-proof 1 (the fix):** removing `IdleConnTimeout` from `NewTransport` fails the test with
|
||||
*"after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is
|
||||
in the message, so the failure cannot be mistaken for a timeout with another cause. Reverted.
|
||||
- **Red-proof 2 (the fix that would be worse than the bug):** setting `DisableKeepAlives: true` also
|
||||
makes the leak vanish — by dialling fresh for every request, which on a box polling ~40,000 times a
|
||||
day is strictly worse than what we started with. **The leak test PASSES under that mutation**;
|
||||
`TestPBSClient_KeepAliveStillReuses` is what catches it, failing with *"3 sequential requests over 3
|
||||
connection(s), want 1"*. Reverted.
|
||||
|
||||
**What this release does NOT do.** It does not reduce the poll rate (**R-336 stays open, but re-scoped
|
||||
— it was never the cause of this leak**), it does not refactor `pbsTargetsFromPVE` to cache or reuse
|
||||
clients (a one-line default restores the standard behaviour; a lifecycle refactor adds
|
||||
cache-invalidation questions for no measurable gain), and it adds no `CloseIdleConnections` call.
|
||||
|
||||
**One sentence in this entry was written before the deploy and was WRONG, and it is corrected here
|
||||
rather than quietly edited.** It read: *"does not clear the 388 descriptors already stuck on ep0 —
|
||||
those persist until that proxy restarts."* **Measured: they clear the moment the AGENT restarts.**
|
||||
Replacing the binary on `demo-hp` released exactly its 199 descriptors within one second
|
||||
(415 → 216 fd), and replacing it on `demo-felhom` released the remaining 203 (**220 → 17 fd in under
|
||||
two seconds**). **17 is precisely ep0's `t0` baseline** of 2026-08-18 09:51:22Z. ep0 was read-only
|
||||
throughout and its proxy PID never changed. The accumulated leak was never ep0's to hold on to — it
|
||||
was held on both sides, and closing either side ends it.
|
||||
|
||||
## v0.129.0 — a correct code for an earlier package stops being called wrong (2026-08-12, R-311)
|
||||
|
||||
**The measurement this fixes.** On 2026-08-12 a recovery code that provably opens a RETAINED package
|
||||
@@ -5389,3 +5704,5 @@ client, signing, or storage/backup orchestration yet (later slices).
|
||||
read-only `--selftest` against the demo host with TLS fingerprint pinning.
|
||||
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
|
||||
token) is provisioned out-of-band; the agent only consumes the token.
|
||||
|
||||
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->
|
||||
|
||||
@@ -101,3 +101,11 @@ the mechanism are exempt.
|
||||
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
|
||||
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
|
||||
assumption, not an observation.
|
||||
|
||||
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
|
||||
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
|
||||
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
|
||||
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
|
||||
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
|
||||
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
|
||||
`os.listdir`, and a glob over a hand-maintained list.
|
||||
|
||||
+17
@@ -1,5 +1,22 @@
|
||||
# CONTEXT — felhom-agent working state
|
||||
|
||||
|
||||
> **2026-09-25 night — v0.133.0 AND v0.134.0 DELIVERED to both demo boxes (CC-signed `agent_update`, ruling 1);
|
||||
> restore test back ON (the `-1` config kept as `agent.json.night-0925-off`). v0.134.0 = R-685:** `backup/runner.go`
|
||||
> `spaceFits` — a vzdump to a LOCAL target needs free ≥ newest archive of that guest × 1.25 + 1 GiB (PVE prunes
|
||||
> only after success); a shortfall is a named skip (`BackupSkipNoSpacePrefix`), fail-open on PBS / first backup /
|
||||
> unknown usage. **Free space comes from `NodeStorage` (`GET /nodes/<n>/storage`) — `ListStorage` (`GET /storage`)
|
||||
> has NO usage**; the first build read it and a real vzdump started in its own live test (aborted, no archive).
|
||||
> demo-hp: `local_backup_retention` 1 (operator option A, saved `agent.json.pre-a4-retention`). Peti's box: nothing.
|
||||
|
||||
> **2026-09-24 — v0.133.0 RELEASED, NOT DELIVERED (R-672, R-673).** Restore-test space preflight
|
||||
> (`reconcile/restoretest_space.go`, `internal/restorespace`): uncompressed size from the vzdump log / PBS size,
|
||||
> × 1.2 + 5 GiB, thin metadata, off the tested guest's pool, unknown refuses, reported as `skipped` non-pass.
|
||||
> Janitor (`cmd/felhom-agent/janitor.go`) every 10 min: `Engine.RetryScratchTeardown` + stale-lock sweep under
|
||||
> the heavy-op gate. Thin pool ≥ 90 % → immediate report (hub v0.124.0 alarms). **Operator rulings 2026-09-24
|
||||
> (evening):** the scheduled restore-test is OFF on both demo hosts (`backup.restore_test_eval_interval_seconds:
|
||||
> -1` — note: 0 means the 6 h DEFAULT, only a negative disables) until v0.133.0 is delivered there; saved configs
|
||||
> `/etc/felhom-agent/agent.json.pre-r672`. demo-hp 9201 was repaired (stop, fsck, start; two Redis AOF tails cut).
|
||||
> Snapshot of the current state + open threads. Authoritative history lives in `CHANGELOG.md` (top
|
||||
> entry = current); the end-of-task detail lives in `REPORT.md`.
|
||||
|
||||
|
||||
@@ -1,63 +1,10 @@
|
||||
# REPORT — felhom-agent, 2026-08-09 (gates only)
|
||||
# REPORT — 2026-10-04: v0.139.0 (R-834)
|
||||
|
||||
**No release. No version bump. No binary published. `scripts/` only** — nothing that runs on a
|
||||
customer's machine changed, and the agent stays **v0.128.0** at `28ba8593b8`.
|
||||
Full session report: `felhom.eu/REPORT-backup-close-os-spike-2026-10-04.md`.
|
||||
|
||||
## What changed
|
||||
|
||||
| file | why |
|
||||
|---|---|
|
||||
| `scripts/retention-policy.json` **(new)** | THE retention number, in one place, with its reasoning and its honesty about where the number came from |
|
||||
| `scripts/check-published-versions.py` | reads that number; bounds its assertion to the newest N; **prints what it stopped covering** |
|
||||
| `scripts/check-release-complete.py` **(new)** | asserts the CHANGELOG-head version is tagged, placed in this history, and published |
|
||||
| `scripts/agent_gates.py` | registers the new gate; legs 1–2 are offline so it runs in `--fast` too |
|
||||
|
||||
## The coupling defect, and the fix
|
||||
|
||||
The prune keeps the newest N; the published-versions check demanded that **every** tag be
|
||||
downloadable. Nothing connected them, so CI went red at `28ba8593b8` — a commit whose own run had
|
||||
been **green the day before** — and would have gone red again at the next publish when `0.121.0` was
|
||||
evicted. Both now read `generic_versions_kept` from one file.
|
||||
|
||||
**What CI no longer covers:** a released version **older than the retention window** is no longer
|
||||
asserted downloadable. Its git tag and its config tree are still asserted; only the binary's presence
|
||||
is dropped. The check names the dropped versions on every run.
|
||||
|
||||
**The number is not a located ruling.** `generic_versions_kept: 10` is what the registry demonstrably
|
||||
holds; no register row records a prune, and container packages hold 19 each. The file says so in its
|
||||
own header. The principled bound is the hub's vouched `min_agent` floor — nothing can install below
|
||||
it — and that is recorded as the follow-up.
|
||||
|
||||
## Controls, all three run
|
||||
|
||||
| control | expected | got |
|
||||
|---|---|---|
|
||||
| live run, policy = 10 | green, and it names `0.120.0` as not asserted | **exit 0**, and it did |
|
||||
| widen policy to 11 | `0.120.0` re-enters the window and convicts | **exit 1**, `FAIL v0.120.0` |
|
||||
| policy file removed | INCONCLUSIVE, never silently unbounded | **exit 2**, naming the path it tried |
|
||||
|
||||
## Red-proof of the new gate
|
||||
|
||||
Mutation: `CHANGELOG.md` head repointed to `## v0.129.0` — never tagged, never published. Asserted
|
||||
applied (`grep -c '^## v0.129.0'` → 1). Result **exit 1**, both legs convicting:
|
||||
|
||||
```
|
||||
- TAG v0.129.0 DOES NOT EXIST. … without the tag every install 404s mid-run, as root.
|
||||
Fix: git tag -a v0.129.0 <released-commit> && git push origin v0.129.0
|
||||
- PACKAGE 0.129.0 IS NOT PUBLISHED (HTTP 404 …).
|
||||
Fix: bash scripts/release-agent.sh 0.129.0
|
||||
```
|
||||
|
||||
Reverted; `git status` clean on `CHANGELOG.md`.
|
||||
|
||||
## Green gate
|
||||
|
||||
`python3 scripts/agent_gates.py` — `reuse-refs OK · instructions OK · published OK ·
|
||||
release-complete OK · all agent gates OK`.
|
||||
|
||||
## Not done here
|
||||
|
||||
The deleter of `0.120.0` is **still not established** and a second attempt failed — Gitea keeps no
|
||||
package-deletion trail, its container log no longer reaches the window, and the activity feed carries
|
||||
no package operation. Recorded in R-287, including the withdrawal of my own earlier over-claim that
|
||||
the router logs showed no DELETE: they do not cover the window, so they never said anything.
|
||||
- **Measured** on demo-hp: the scheduled restore-test's scratch guest has `onboot: 0` and throwaway stand-ins for
|
||||
mp8/mp9 on every config read until teardown — it was already safe. Evidence `felhom.eu/documentation/audits/backup-close-2026-10-04/partA/`.
|
||||
- **Fixed:** the DR bring-up refuses beside a live original (source guest present, a drives bind on another guest,
|
||||
or an unreadable config). On a replaced host it is unchanged.
|
||||
- Tests `TestRunBringUp_DRRefusesBesideALiveOriginal`, `TestRestoreTest_NoHostPathBindBesideTheOriginal` and two
|
||||
more; both rules red-proved. No sudoers change (said why in the CHANGELOG).
|
||||
|
||||
@@ -74,6 +74,7 @@
|
||||
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
|
||||
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
|
||||
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
|
||||
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
|
||||
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
|
||||
|
||||
### Proxmox client / hub / PBS / provisioning
|
||||
@@ -84,8 +85,11 @@
|
||||
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
|
||||
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
|
||||
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
|
||||
| `reconcile.PreflightRestoreSpace` + `restorespace.Provider` (v0.133.0) | internal/reconcile/restoretest_space.go, internal/restorespace/restorespace.go | `PreflightRestoreSpace(ctx, space, policy, archive, rawCfg, configured) SpaceVerdict` | ANY step that restores or copies a guest onto a storage — size it first | The restored size is UNCOMPRESSED (vzdump log "Total bytes written" / PBS snapshot size) — never the archive FILE (6.9 GB file → 22.6 GB restore, R-672); an unknown refuses; eligibility needs Datastore.AllocateSpace on `/storage/<id>` specifically (the `/` grant answers every path) |
|
||||
| `Engine.RetryScratchTeardown` + the daemon janitor (v0.133.0) | internal/reconcile/restoretest_retry.go, cmd/felhom-agent/janitor.go | `RetryScratchTeardown(ctx) ScratchRetryResult` | retrying a leftover on a TIMER instead of only at start | Never `Recover` on a timer — it also resolves generic in-flight ops; a periodic sweep that unlocks guests holds the one-heavy-op gate (`InFlight.TryAcquire`) |
|
||||
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
|
||||
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
|
||||
| `httpx.NewTransport` | internal/httpx/transport.go | `NewTransport(tlsCfg, idleConnTimeout) *http.Transport` | **EVERY** hand-rolled `http.Transport` in this repo — pbs, hub and proxmox all pin TLS, so none can use `http.DefaultTransport` | **R-344: never inline `&http.Transport{TLSClientConfig: ...}` again.** A composite literal takes `IdleConnTimeout` **zero, which means retain idle connections FOREVER** — `http.DefaultTransport` sets 90s and a literal does not inherit it. Combined with a client rebuilt per cycle and dropped (`pbsTargetsFromPVE`), that stranded **388 sockets on ep0 in 46 h**, held open on BOTH sides. `idleConnTimeout <= 0` means **use the default**, never "no timeout". Returns a **FRESH** transport every call — a shared one would pool connections across differently pinned endpoints. Pinned by `internal/pbs/client_leak_test.go` (server-side connection counting) + `internal/httpx/transport_test.go` |
|
||||
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
|
||||
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
|
||||
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
)
|
||||
|
||||
// janitorInterval is how often the leftovers of an interrupted restore-test or backup are retried
|
||||
// (R-672 rule 3, R-673). Both used to be resolved ONLY at agent start: on 2026-09-24 a failed scratch
|
||||
// teardown kept a full thin pool full for 2.5 h, and a stale `snapshot-delete` lock blocked 9201's
|
||||
// whole-box backups for five hours — each cleared within a minute of an agent restart.
|
||||
const janitorInterval = 10 * time.Minute
|
||||
|
||||
// janitorDeps are the janitor's seams (tests drive one pass with fakes).
|
||||
type janitorDeps struct {
|
||||
retryScratch func(ctx context.Context) reconcile.ScratchRetryResult
|
||||
staleLocks func(ctx context.Context) // localapi Server.RecoverStaleLockedGuests; nil when the local API is off
|
||||
heavy *backup.InFlight
|
||||
record func(hub.RestoreTest)
|
||||
now func() time.Time
|
||||
logger *slog.Logger
|
||||
}
|
||||
|
||||
// janitorPass is one pass. The stale-lock sweep runs only while holding the one-heavy-operation gate, so
|
||||
// no agent backup can START between its "no vzdump is running" check and its unlock (at start-up the
|
||||
// sweep ran before the backup loop existed; on a timer that ordering must be made, not assumed). A busy
|
||||
// gate skips the sweep this pass — the next pass retries.
|
||||
func janitorPass(ctx context.Context, d janitorDeps) {
|
||||
if d.retryScratch != nil {
|
||||
r := d.retryScratch(ctx)
|
||||
if r.Examined > 0 {
|
||||
d.logger.Info("janitor: restore-test scratch retry pass", "examined", r.Examined,
|
||||
"destroyed", r.Destroyed, "already_gone", r.Clean, "failed", r.Failed)
|
||||
}
|
||||
for _, vmid := range r.GaveUp {
|
||||
// The operator is told through the existing restore-test failure path: the hub raises
|
||||
// restore_test_failed (operator) once per distinct archive — this record's archive names
|
||||
// the stuck scratch guest.
|
||||
if d.record != nil {
|
||||
d.record(hub.RestoreTest{
|
||||
SourceArchive: fmt.Sprintf("scratch-teardown:%d", vmid),
|
||||
ScratchVMID: vmid,
|
||||
Pass: false,
|
||||
Error: fmt.Sprintf("restore-test scratch guest %d could not be torn down after %d retries — it holds its disks; remove it by hand (pct destroy %d) after checking what keeps it busy",
|
||||
vmid, reconcile.MaxTeardownTries, vmid),
|
||||
TestedAt: d.now().UTC().Format(time.RFC3339),
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
if d.staleLocks != nil {
|
||||
release, busy, ok := d.heavy.TryAcquire("stale-lock-sweep")
|
||||
if !ok {
|
||||
d.logger.Info("janitor: stale-lock sweep deferred — a heavy operation is in flight", "busy", busy)
|
||||
return
|
||||
}
|
||||
defer release()
|
||||
d.staleLocks(ctx)
|
||||
}
|
||||
}
|
||||
|
||||
// runJanitor runs janitorPass every janitorInterval until ctx ends.
|
||||
func runJanitor(ctx context.Context, d janitorDeps) {
|
||||
d.logger.Info("janitor: starting (restore-test scratch retry + stale-lock sweep)", "interval", janitorInterval)
|
||||
t := time.NewTicker(janitorInterval)
|
||||
defer t.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-t.C:
|
||||
janitorPass(ctx, d)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,60 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io"
|
||||
"log/slog"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
)
|
||||
|
||||
// R-672 / R-673 (v0.133.0): one janitor pass, driven with fakes.
|
||||
|
||||
func quiet() *slog.Logger { return slog.New(slog.NewTextHandler(io.Discard, nil)) }
|
||||
|
||||
// A scratch the engine gave up on reaches the hub as a failed restore-test record naming it — the
|
||||
// existing operator path (restore_test_failed). Never a pass.
|
||||
func TestJanitor_GaveUpIsReportedAsAFailure(t *testing.T) {
|
||||
var got []hub.RestoreTest
|
||||
janitorPass(context.Background(), janitorDeps{
|
||||
retryScratch: func(context.Context) reconcile.ScratchRetryResult {
|
||||
return reconcile.ScratchRetryResult{Examined: 1, Failed: 1, GaveUp: []int{990000}}
|
||||
},
|
||||
heavy: &backup.InFlight{}, record: func(r hub.RestoreTest) { got = append(got, r) },
|
||||
now: time.Now, logger: quiet(),
|
||||
})
|
||||
if len(got) != 1 || got[0].Pass || got[0].ScratchVMID != 990000 || !strings.Contains(got[0].Error, "990000") {
|
||||
t.Fatalf("records = %+v — want one FAILED record naming scratch 990000", got)
|
||||
}
|
||||
}
|
||||
|
||||
// R-673: the stale-lock sweep runs only while holding the one-heavy-operation gate, so no agent backup can
|
||||
// start between its "no vzdump running" check and its unlock.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT): drop the TryAcquire → "the sweep ran while a backup held the gate".
|
||||
func TestJanitor_StaleLockSweepWaitsForTheHeavyGate(t *testing.T) {
|
||||
heavy := &backup.InFlight{}
|
||||
swept := 0
|
||||
d := janitorDeps{staleLocks: func(context.Context) { swept++ }, heavy: heavy, now: time.Now, logger: quiet()}
|
||||
release, _, ok := heavy.TryAcquire("backup")
|
||||
if !ok {
|
||||
t.Fatal("setup")
|
||||
}
|
||||
janitorPass(context.Background(), d)
|
||||
if swept != 0 {
|
||||
t.Fatal("the sweep ran while a backup held the gate")
|
||||
}
|
||||
release()
|
||||
janitorPass(context.Background(), d)
|
||||
if swept != 1 {
|
||||
t.Fatalf("swept %d times with the gate free — want 1", swept)
|
||||
}
|
||||
if _, _, ok := heavy.TryAcquire("after"); !ok {
|
||||
t.Fatal("the sweep did not release the gate")
|
||||
}
|
||||
}
|
||||
+124
-2
@@ -48,8 +48,10 @@ import (
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/pbsdr"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/poke"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/provision"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/osupdate"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/restorespace"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/selfheal"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
|
||||
@@ -59,7 +61,7 @@ import (
|
||||
|
||||
// version is the agent version. Overridable at build time with
|
||||
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
|
||||
var version = "0.92.1"
|
||||
var version = "0.131.0"
|
||||
|
||||
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
|
||||
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
|
||||
@@ -241,6 +243,8 @@ func main() {
|
||||
os.Exit(runSelftestRestoreTest(context.Background(), cfg, logger, archive))
|
||||
case "restore-test-due":
|
||||
os.Exit(runSelftestRestoreTestDue(context.Background(), cfg, logger))
|
||||
case "os-update":
|
||||
os.Exit(runSelftestOSUpdate(context.Background(), cfg, logger, vmid))
|
||||
case "pbs-verify":
|
||||
os.Exit(runSelftestPBSVerify(context.Background(), cfg, logger))
|
||||
case "lanresolver":
|
||||
@@ -838,6 +842,10 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// The "Down" channel sync hook: on each heartbeat, fetch desired-state when the generation
|
||||
// advances. The loop calls it via the EnvelopeObserver seam (hub does not import desired).
|
||||
desiredSyncer := desired.NewSyncer(client, desiredProvider, logger)
|
||||
// OS updates, guest fast lane (agent v0.140.0, `11-os-updates.md` §8 step 2): the leg consumes the hub's
|
||||
// os_update block and runs after each successful primary whole-guest backup (wired on the local API below).
|
||||
osLeg := newOSLeg(cfg, client, logger)
|
||||
desiredSyncer.AddConsumer(osLeg)
|
||||
// S5: consume a host_loss restore_directive into an inspectable restore PLAN (derive + surface,
|
||||
// execute nothing). The recipe is fetched on-demand (rare directive) via a fresh Collect.
|
||||
desiredSyncer.AddConsumer(dr.NewConsumer(func(ctx context.Context) *hub.DRRecipeHostHalf {
|
||||
@@ -897,6 +905,13 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// it finds nothing to flag. The re-mount dispatch is off the poll path (a goroutine).
|
||||
storageTrigger := make(chan struct{}, 1)
|
||||
loop.SetTrigger(storageTrigger)
|
||||
// R-672: a thin pool crossing 90 % requests a report at once (the hub's storage-fill alarm).
|
||||
observer.SetThinHighTrigger(func() {
|
||||
select {
|
||||
case storageTrigger <- struct{}{}:
|
||||
default:
|
||||
}
|
||||
})
|
||||
remounter := &gateRemounter{gate: gate, ops: hostOps, hostID: cfg.Hub.HostID, logger: logger}
|
||||
// Drive intent store (slice 10 P3 self-heal): persisted, durable-id-keyed enroll/eject/decommission
|
||||
// state. Gates the watchdog's self-heal re-mount to ENROLLED drives, and the local API records
|
||||
@@ -963,6 +978,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
Logger: logger,
|
||||
})
|
||||
|
||||
rtSpace, rtPolicy := restoreSpaceFor(cfg, px, hostOps)
|
||||
engine := reconcile.NewEngine(reconcile.EngineOptions{
|
||||
API: px,
|
||||
Queue: queue,
|
||||
@@ -971,6 +987,9 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
Gate: gate,
|
||||
HostID: cfg.Hub.HostID,
|
||||
Logger: logger,
|
||||
// R-672: the restore-test's space preflight (nil would refuse every test — fail-closed).
|
||||
RestoreSpace: rtSpace,
|
||||
SpacePolicy: rtPolicy,
|
||||
})
|
||||
|
||||
// Crash recovery (doc 03 §10): resolve any op that was in flight when the agent
|
||||
@@ -1100,6 +1119,17 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
},
|
||||
}
|
||||
localSrv := buildLocalAPIServer(cfg, px, backupStore, heavyOps, observer, driveKnown, hostOps, gate, collector, client, intentRec, guestBindStore, formatJobStore, logRing, escrowCeremonyCfg, logger, &localTokens)
|
||||
if localSrv != nil {
|
||||
localSrv.SetAfterPrimaryBackup(func(ctx context.Context, vmid int) {
|
||||
// Let the controller finish bringing its apps back after the backup, then run (still under the gate).
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-time.After(90 * time.Second):
|
||||
}
|
||||
osLeg.Run(ctx, vmid, "night")
|
||||
})
|
||||
}
|
||||
if localTokens != nil {
|
||||
defer localTokens.Close()
|
||||
}
|
||||
@@ -1401,8 +1431,22 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
|
||||
// running" signal, so a deliberately stopped guest is never touched.
|
||||
go localSrv.WatchGuestPower(ctx)
|
||||
// R-523: a controller container that is simply not running (killed, stopped, a failed
|
||||
// self-update) is restarted through its bootstrap unit — nothing else watches it.
|
||||
collector.SetControllerSupervisorReporter(localSrv)
|
||||
go localSrv.WatchControllers(ctx)
|
||||
go func() { errc <- localSrv.Run(ctx) }()
|
||||
}
|
||||
// R-672 / R-673: retry a failed restore-test teardown and sweep stale backup locks on a timer, not
|
||||
// only at start-up (janitor.go).
|
||||
{
|
||||
jd := janitorDeps{retryScratch: engine.RetryScratchTeardown, heavy: heavyOps,
|
||||
record: backupStore.RecordRestoreTest, now: time.Now, logger: logger}
|
||||
if localSrv != nil {
|
||||
jd.staleLocks = localSrv.RecoverStaleLockedGuests
|
||||
}
|
||||
go runJanitor(ctx, jd)
|
||||
}
|
||||
if lanLoop != nil {
|
||||
lanServers = 1
|
||||
go func() { errc <- lanLoop.Run(ctx) }()
|
||||
@@ -1826,6 +1870,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
||||
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
|
||||
SmbCredsDir: cfg.Privileged.SmbCredsDir,
|
||||
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
|
||||
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
|
||||
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
|
||||
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
|
||||
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
|
||||
@@ -2204,6 +2249,17 @@ func formatOrDash(t time.Time) string {
|
||||
return t.UTC().Format(time.RFC3339)
|
||||
}
|
||||
|
||||
// restoreSpaceFor builds the restore-test's space preflight (R-672): the production provider over the
|
||||
// Proxmox API + the privileged `lvs` metadata read, and the configured margin.
|
||||
func restoreSpaceFor(cfg config.Config, px *proxmox.Client, ops *storage.SudoHostOps) (reconcile.RestoreSpace, reconcile.SpacePolicy) {
|
||||
factor, reserve := cfg.Backup.RestoreTestSpace()
|
||||
p := &restorespace.Provider{API: px}
|
||||
if ops != nil {
|
||||
p.ThinMeta = ops.ThinPoolMetadata
|
||||
}
|
||||
return p, reconcile.SpacePolicy{Factor: factor, ReserveBytes: reserve}
|
||||
}
|
||||
|
||||
func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog.Logger, archive string) int {
|
||||
if err := cfg.Validate(); err != nil {
|
||||
fmt.Fprintln(os.Stderr, "selftest: proxmox not configured:", err)
|
||||
@@ -2234,8 +2290,10 @@ func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog
|
||||
}
|
||||
}
|
||||
gate := reconcile.NewGate(nil, cfg.Hub.HostID, reconcile.SlogAudit{Logger: logger}, logger)
|
||||
rtSpace, rtPolicy := restoreSpaceFor(cfg, px, newHostOps(cfg, logger))
|
||||
engine := reconcile.NewEngine(reconcile.EngineOptions{
|
||||
API: px, Queue: queue, Journal: journal, Gate: gate, HostID: cfg.Hub.HostID, Logger: logger,
|
||||
RestoreSpace: rtSpace, SpacePolicy: rtPolicy,
|
||||
})
|
||||
|
||||
fmt.Printf("=== felhom-agent %s selftest=restore-test ===\n", version)
|
||||
@@ -2265,10 +2323,16 @@ func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog
|
||||
RestoreTaskTimeout: restoreTaskTimeout(cfg, rtTier),
|
||||
})
|
||||
printJSON("restore-test record", backup.ToHubRestoreTest(res, time.Now().UTC()))
|
||||
if res.Skipped && res.SkipReason != "" {
|
||||
fmt.Printf(" space preflight: storage=%s required=%d avail=%d\n", res.TargetStorage, res.RequiredBytes, res.AvailBytes)
|
||||
fmt.Printf("=== selftest=restore-test SKIPPED — %s ===\n", res.SkipReason)
|
||||
return 4
|
||||
}
|
||||
if res.Skipped {
|
||||
fmt.Println("=== selftest=restore-test SKIPPED (no free scratch VMID in band) ===")
|
||||
return 0
|
||||
}
|
||||
fmt.Printf(" space preflight passed: storage=%s required=%d avail=%d\n", res.TargetStorage, res.RequiredBytes, res.AvailBytes)
|
||||
if res.Err != nil || !res.Pass {
|
||||
fmt.Fprintf(os.Stderr, " [FAIL] restore-test (scratch %d): %v\n", res.ScratchVMID, res.Err)
|
||||
return 1
|
||||
@@ -3432,8 +3496,66 @@ func (f *selftestFlag) Set(v string) error {
|
||||
f.mode = "identity-consume"
|
||||
case "controller-swap":
|
||||
f.mode = "controller-swap"
|
||||
case "os-update":
|
||||
f.mode = "os-update"
|
||||
case "wgtunnel": // dispatched since S3 but refused here until 2026-10-04 (TestSelftestFlag_AcceptsEveryDispatchedMode)
|
||||
f.mode = "wgtunnel"
|
||||
default:
|
||||
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap)", v)
|
||||
return fmt.Errorf("invalid --selftest value %q (want read|task|hub|storage|backup|restore-test|restore-test-due|pbs-verify|lanresolver|wgtunnel|bring-up|provision|escrow-create|escrow-consume|identity-consume|controller-swap|os-update)", v)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// newOSLeg builds the OS-update leg (agent v0.140.0). The wrapper runs through sudo (FELHOM_OSAPPLY); the plan and
|
||||
// the once-per-night marker live in the agent's own os/ dir.
|
||||
func newOSLeg(cfg config.Config, client *hub.Client, logger *slog.Logger) *osupdate.Leg {
|
||||
mode := proxmox.RunnerMode(cfg.Privileged.Mode)
|
||||
if mode == "" {
|
||||
mode = proxmox.RunnerSudo
|
||||
}
|
||||
l := &osupdate.Leg{
|
||||
Runner: &proxmox.ExecRunner{Mode: mode, SudoPath: cfg.Privileged.SudoPath},
|
||||
Logger: logger,
|
||||
PlanDir: osupdate.DefaultPlanDir,
|
||||
StatePath: filepath.Join(osupdate.DefaultPlanDir, "last-night-run"),
|
||||
}
|
||||
if client != nil {
|
||||
l.Hub = client
|
||||
}
|
||||
return l
|
||||
}
|
||||
|
||||
// runSelftestOSUpdate is the OS leg's DEBUG ACTION (agent v0.140.0): one pass for -vmid, now, exactly as the night
|
||||
// runs it after a backup — the hub's os_update block (fetched fresh), the wrapper via sudo, the health wait, the
|
||||
// report to the hub — with trigger "debug" (never throttled, and it does NOT count as a night run for approval).
|
||||
// Run it as the agent user: sudo -u felhom-agent felhom-agent --config … --selftest=os-update -vmid 9201
|
||||
func runSelftestOSUpdate(ctx context.Context, cfg config.Config, logger *slog.Logger, vmid int) int {
|
||||
if vmid <= 0 {
|
||||
fmt.Fprintln(os.Stderr, "selftest=os-update: -vmid is required")
|
||||
return 2
|
||||
}
|
||||
client, err := hub.NewClient(cfg.Hub, logger)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "selftest=os-update: hub client:", err)
|
||||
return 1
|
||||
}
|
||||
leg := newOSLeg(cfg, client, logger)
|
||||
resp, err := client.FetchDesiredState(ctx)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "selftest=os-update: desired state:", err)
|
||||
return 1
|
||||
}
|
||||
leg.SetBlock(resp.DesiredState.OSUpdate)
|
||||
b := leg.Block()
|
||||
fmt.Printf("=== felhom-agent %s selftest=os-update vmid=%d ring=%d enabled=%v release=%v ===\n", version, vmid, b.Ring, b.Enabled, b.Release != nil)
|
||||
rep := leg.Run(ctx, vmid, "debug")
|
||||
printJSON("os-update report", map[string]any{"run_id": rep.RunID, "ring": rep.Ring, "release_id": rep.ReleaseID,
|
||||
"mode": rep.Mode, "outcome": rep.Outcome, "healthy": rep.Healthy, "health_reason": rep.HealthReason,
|
||||
"upgraded": rep.Upgraded, "pending": len(rep.Pending), "not_covered": rep.NotCovered,
|
||||
"restart_needed": rep.RestartNeeded, "docker_restart_needed": rep.DockerRestartNeeded, "refused": rep.Refused})
|
||||
switch rep.Outcome {
|
||||
case "applied", "nothing", "inventory", "skipped":
|
||||
return 0
|
||||
}
|
||||
return 1
|
||||
}
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"regexp"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Every mode the dispatcher (`switch selftest.mode`) runs must be ACCEPTED by the --selftest flag. Found live
|
||||
// 2026-10-04: --selftest=os-update had a dispatch case and a function but the flag's allow-list refused it, so the
|
||||
// debug action could not run. Red-proof: drop the "os-update" case from selftestFlag.Set and this fails.
|
||||
func TestSelftestFlag_AcceptsEveryDispatchedMode(t *testing.T) {
|
||||
src, err := os.ReadFile("main.go")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
s := string(src)
|
||||
i := strings.Index(s, "switch selftest.mode {")
|
||||
if i < 0 {
|
||||
t.Fatal("dispatch switch not found")
|
||||
}
|
||||
block := s[i:]
|
||||
block = block[:strings.Index(block, "\n\t}\n")]
|
||||
modes := regexp.MustCompile(`(?m)^\tcase "([a-z-]+)":`).FindAllStringSubmatch(block, -1)
|
||||
if len(modes) < 5 {
|
||||
t.Fatalf("parsed only %d dispatch cases — the parser is wrong", len(modes))
|
||||
}
|
||||
var bad []string
|
||||
for _, m := range modes {
|
||||
var f selftestFlag
|
||||
if err := f.Set(m[1]); err != nil {
|
||||
bad = append(bad, m[1])
|
||||
}
|
||||
}
|
||||
if len(bad) > 0 {
|
||||
t.Fatalf("dispatched but refused by --selftest: %v", bad)
|
||||
}
|
||||
}
|
||||
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
|
||||
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
|
||||
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
|
||||
# view, and the docker socket. The controller reaches the agent's local API for disk management.
|
||||
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
|
||||
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
|
||||
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
|
||||
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
|
||||
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
|
||||
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
|
||||
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
|
||||
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
|
||||
-v felhom-controller-data:/opt/docker/felhom-controller \
|
||||
|
||||
@@ -299,10 +299,17 @@ Cmnd_Alias FELHOM_SELFHEAL = \
|
||||
# argument after the numeric vmid is a literal, so the grant cannot be widened by anything the guest or
|
||||
# the hub says. The address read is deliberately NOT duplicated here — it is already FELHOM_DNSMASQ's,
|
||||
# and the same command must not be granted twice under two names.
|
||||
# OS updates, guest fast lane (`11-os-updates.md` §5.4.1, agent v0.140.0). The ONLY entry: the root wrapper with
|
||||
# one plan file in the agent's own os/ dir. Every safety rule (no removal, no downgrade, no new or unlisted package,
|
||||
# Debian origin only, the box's own customer guest only) lives in the wrapper, red-proved per rule
|
||||
# (configs/test_felhom_os_apply.py). The agent gets NO apt grant of its own.
|
||||
Cmnd_Alias FELHOM_OSAPPLY = \
|
||||
/usr/local/sbin/felhom-os-apply --plan /var/lib/felhom-agent/os/plan-*.json
|
||||
|
||||
Cmnd_Alias FELHOM_GUESTNET = \
|
||||
/usr/sbin/pct exec [0-9]* -- ip route show default, \
|
||||
/usr/sbin/pct exec [0-9]* -- cat /etc/network/interfaces, \
|
||||
/usr/sbin/pct exec [0-9]* -- pgrep -x dhclient, \
|
||||
/usr/sbin/pct exec [0-9]* -- dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0
|
||||
|
||||
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN
|
||||
felhom-agent ALL=(root) NOPASSWD: FELHOM_MOUNT, FELHOM_DISK, FELHOM_PROVISION, FELHOM_FORMAT, FELHOM_DNSMASQ, FELHOM_GUESTHOOK, FELHOM_INTERMEDIARY, FELHOM_CONTROLLERSWAP, FELHOM_STALELOCK, FELHOM_NETMOUNT, FELHOM_WG, FELHOM_SELFUPDATE, FELHOM_SSHD, FELHOM_OOB, FELHOM_PBSDR, FELHOM_BACKUPTARGET, FELHOM_SELFHEAL, FELHOM_ESCROW, FELHOM_GUESTNET, FELHOM_SCRATCH_TEARDOWN, FELHOM_OSAPPLY
|
||||
|
||||
Executable
+486
@@ -0,0 +1,486 @@
|
||||
#!/usr/bin/python3
|
||||
# felhom-os-apply — the ROOT half of the agent's operating-system update leg (`11-os-updates.md` §5.4.1).
|
||||
#
|
||||
# Install as /usr/local/sbin/felhom-os-apply (0755 root:root). The non-root agent invokes it via `sudo -n`
|
||||
# (FELHOM_OSAPPLY alias) with EXACTLY: felhom-os-apply --plan /var/lib/felhom-agent/os/plan-<id>.json
|
||||
# Nothing else on the command line is accepted. Python 3, standard library only (a JSON plan cannot be parsed
|
||||
# safely in sh). Tests: configs/test_felhom_os_apply.py (a fake runner; nothing real is executed).
|
||||
#
|
||||
# THE TRUST MODEL. The plan is written by the agent, so a broken-into agent writes whatever plan it likes. The
|
||||
# protection is therefore what this file REFUSES, not where the plan came from: no removal, no downgrade, no new
|
||||
# package, no package outside the plan, only Debian origin in the fast lane, only the box's own customer guest.
|
||||
# Package signatures stay Debian's: apt checks every Release file against the guest's keyring, including the
|
||||
# snapshot.debian.org fallback (decision 79). Nothing here is overridable from the environment.
|
||||
#
|
||||
# THIS RELEASE: layer "guest", lane "fast" only. The host layer and the slow lane exist in the interface and are
|
||||
# REFUSED (R3, R12) until `11` §8 steps 3 and 5 enable them.
|
||||
#
|
||||
# Modes (plan field "mode"):
|
||||
# inventory read-only for packages: `apt-get update` in the guest, then report what is installed (with origin),
|
||||
# what is pending, restart-needed and health. Installs nothing.
|
||||
# apply repair first, check every refusal on an `apt-get -s` simulation of EXACTLY name=version, then
|
||||
# install, clean, and report the same as inventory.
|
||||
# health report health only (the agent polls it after a run).
|
||||
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
|
||||
# OSAPPLY-REPORT <one JSON object>
|
||||
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import stat
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
|
||||
PLAN_DIR = "/var/lib/felhom-agent/os"
|
||||
PLAN_RE = re.compile(r"^plan-[A-Za-z0-9._-]{1,80}\.json$")
|
||||
AGENT_USER = "felhom-agent"
|
||||
FAST_ORIGINS = ("Debian", "Debian-Security")
|
||||
# Debian package name and version grammar (Debian policy §5.6.1, §5.6.12).
|
||||
NAME_RE = re.compile(r"^[a-z0-9][a-z0-9+.-]+$")
|
||||
VERSION_RE = re.compile(r"^(?:[0-9]+:)?[0-9][A-Za-z0-9.+~-]*$")
|
||||
SNAP_RE = re.compile(r"^[0-9]{8}T[0-9]{6}Z$")
|
||||
RESERVED_VMIDS = set(range(990000, 990010)) | {9999}
|
||||
DRIVES_PARENT = "/mnt/felhom-drives"
|
||||
SNAPSHOT_LIST = "/etc/apt/sources.list.d/felhom-os-snapshot.list"
|
||||
APT_ENV = ["env", "DEBIAN_FRONTEND=noninteractive", "APT_LISTCHANGES_FRONTEND=none", "NEEDRESTART_MODE=l", "LC_ALL=C"]
|
||||
DPKG_OPTS = ["-o", "Dpkg::Options::=--force-confold", "-o", "Dpkg::Options::=--force-confdef"]
|
||||
MIN_FREE = 500 * 1024 * 1024
|
||||
|
||||
|
||||
class Refused(Exception):
|
||||
def __init__(self, code, reason):
|
||||
super().__init__(f"{code} {reason}")
|
||||
self.code, self.reason = code, reason
|
||||
|
||||
|
||||
class Runner:
|
||||
"""Runs commands for real. Tests replace it with a fake. `guest` runs inside the container via pct exec."""
|
||||
|
||||
def host(self, argv, timeout=600):
|
||||
p = subprocess.run(argv, capture_output=True, text=True, timeout=timeout)
|
||||
return p.returncode, p.stdout, p.stderr
|
||||
|
||||
def guest(self, vmid, argv, timeout=1800):
|
||||
return self.host(["/usr/sbin/pct", "exec", str(vmid), "--"] + argv, timeout)
|
||||
|
||||
def read_file(self, path):
|
||||
with open(path) as f:
|
||||
return f.read()
|
||||
|
||||
def stat(self, path):
|
||||
return os.lstat(path)
|
||||
|
||||
def agent_uid(self):
|
||||
import pwd
|
||||
return pwd.getpwnam(AGENT_USER).pw_uid
|
||||
|
||||
def log(self, line):
|
||||
print(line, file=sys.stderr, flush=True)
|
||||
try:
|
||||
subprocess.run(["logger", "-t", "felhom-os-apply", line], timeout=10)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
class Apply:
|
||||
def __init__(self, runner, plan_path):
|
||||
self.r = runner
|
||||
self.plan_path = plan_path
|
||||
self.report = {"refused": None, "mode": None}
|
||||
|
||||
# ---------- checks ----------
|
||||
def load_plan(self):
|
||||
p = self.plan_path
|
||||
d, base = os.path.dirname(p), os.path.basename(p)
|
||||
if d != PLAN_DIR or not PLAN_RE.match(base) or ".." in p:
|
||||
raise Refused("R1", f"the plan must be {PLAN_DIR}/plan-<id>.json, got {p!r}")
|
||||
try:
|
||||
st = self.r.stat(p)
|
||||
except OSError as e:
|
||||
raise Refused("R1", f"cannot stat the plan: {e}")
|
||||
if not stat.S_ISREG(st.st_mode):
|
||||
raise Refused("R1", "the plan is not a regular file (a symlink or a device is refused)")
|
||||
if st.st_uid != self.r.agent_uid():
|
||||
raise Refused("R1", f"the plan is not owned by {AGENT_USER}")
|
||||
if st.st_size > 2 * 1024 * 1024:
|
||||
raise Refused("R1", "the plan is larger than 2 MB")
|
||||
try:
|
||||
plan = json.loads(self.r.read_file(p))
|
||||
except (OSError, ValueError) as e:
|
||||
raise Refused("R1", f"the plan is not valid JSON: {e}")
|
||||
if not isinstance(plan, dict):
|
||||
raise Refused("R1", "the plan is not a JSON object")
|
||||
return plan
|
||||
|
||||
def check_plan(self, plan):
|
||||
mode = plan.get("mode", "apply")
|
||||
if mode not in ("apply", "inventory", "health"):
|
||||
raise Refused("R11", f"unknown mode {mode!r}")
|
||||
if plan.get("layer") != "guest":
|
||||
raise Refused("R12", f"layer {plan.get('layer')!r} is refused in this release (guest only)")
|
||||
if plan.get("lane", "fast") != "fast":
|
||||
raise Refused("R3", "the slow lane is refused in this release")
|
||||
vmid = plan.get("vmid")
|
||||
if not isinstance(vmid, int) or isinstance(vmid, bool) or vmid <= 0:
|
||||
raise Refused("R11", f"vmid must be a positive integer, got {vmid!r}")
|
||||
rid = plan.get("release_id", "")
|
||||
if not isinstance(rid, str) or not re.match(r"^[A-Za-z0-9._:-]{1,80}$", rid):
|
||||
raise Refused("R11", f"release_id {rid!r} is not a plain id")
|
||||
if plan.get("allow_new"):
|
||||
raise Refused("R6", "allow_new is a slow-lane field; the fast lane never adds a package")
|
||||
pk = plan.get("packages", [])
|
||||
if not isinstance(pk, list) or (mode == "apply" and not pk):
|
||||
raise Refused("R11", "packages must be a non-empty list in apply mode")
|
||||
seen = set()
|
||||
for e in pk:
|
||||
if not isinstance(e, dict):
|
||||
raise Refused("R11", "every package entry must be an object")
|
||||
n, v, o = e.get("name"), e.get("version"), e.get("origin")
|
||||
if not isinstance(n, str) or not NAME_RE.match(n):
|
||||
raise Refused("R11", f"package name {n!r} is not a Debian package name")
|
||||
if not isinstance(v, str) or not VERSION_RE.match(v):
|
||||
raise Refused("R11", f"version {v!r} of {n} is not a Debian version string")
|
||||
if n in seen:
|
||||
raise Refused("R11", f"package {n} is named twice")
|
||||
seen.add(n)
|
||||
if o not in FAST_ORIGINS:
|
||||
raise Refused("R2", f"{n}: origin {o!r} is not Debian / Debian-Security (the fast lane, `11` C3)")
|
||||
snap = plan.get("snapshot", "")
|
||||
if snap and not SNAP_RE.match(snap):
|
||||
raise Refused("R11", f"snapshot {snap!r} is not YYYYMMDDTHHMMSSZ")
|
||||
return mode, vmid
|
||||
|
||||
def check_guest(self, vmid):
|
||||
if vmid in RESERVED_VMIDS:
|
||||
raise Refused("R10", f"vmid {vmid} is a reserved scratch vmid")
|
||||
try:
|
||||
conf = self.r.read_file(f"/etc/pve/lxc/{vmid}.conf")
|
||||
except OSError:
|
||||
raise Refused("R10", f"vmid {vmid} is not a container on this host")
|
||||
cur = conf.split("\n[", 1)[0] # the current config, not a snapshot section
|
||||
binds = [l for l in cur.splitlines() if re.match(r"^mp[0-9]+: " + re.escape(DRIVES_PARENT) + r",", l)]
|
||||
if not binds:
|
||||
raise Refused("R10", f"vmid {vmid} does not bind {DRIVES_PARENT} — it is not this box's customer guest")
|
||||
lock = [l for l in cur.splitlines() if l.startswith("lock:")]
|
||||
if lock:
|
||||
raise Refused("R9", f"vmid {vmid} is locked ({lock[0].split(':', 1)[1].strip()}) — a backup or restore is running")
|
||||
rc, out, _ = self.r.host(["/usr/sbin/pct", "status", str(vmid)])
|
||||
if rc != 0 or "running" not in out:
|
||||
raise Refused("R10", f"vmid {vmid} is not running")
|
||||
|
||||
# ---------- guest helpers ----------
|
||||
def g(self, argv, timeout=1800):
|
||||
return self.r.guest(self.vmid, argv, timeout)
|
||||
|
||||
def installed(self):
|
||||
rc, out, _ = self.g(["dpkg-query", "-W", "-f", "${Package}\t${Version}\t${db:Status-Abbrev}\n"])
|
||||
res = {}
|
||||
for l in out.splitlines():
|
||||
parts = l.split("\t")
|
||||
if len(parts) == 3 and parts[2].startswith("ii"):
|
||||
res[parts[0]] = parts[1]
|
||||
return res
|
||||
|
||||
def dpkg_cmp(self, a, op, b):
|
||||
rc, _, _ = self.g(["dpkg", "--compare-versions", a, op, b])
|
||||
return rc == 0
|
||||
|
||||
def madison(self, name):
|
||||
rc, out, _ = self.g(["apt-cache", "madison", name])
|
||||
vs = set()
|
||||
for l in out.splitlines():
|
||||
f = [x.strip() for x in l.split("|")]
|
||||
if len(f) >= 3 and f[0] == name:
|
||||
vs.add(f[1])
|
||||
return vs
|
||||
|
||||
def simulate(self, args):
|
||||
rc, out, err = self.g(APT_ENV + ["apt-get", "-s", "-q"] + args)
|
||||
inst, remv = [], []
|
||||
for l in out.splitlines():
|
||||
m = re.match(r"^Inst (\S+) (?:\[([^]]*)\] )?\((\S+) (.*?) \[[a-z0-9]+\]\)", l)
|
||||
if m:
|
||||
inst.append({"name": m.group(1), "from": m.group(2), "to": m.group(3), "origin": m.group(4)})
|
||||
m = re.match(r"^Remv (\S+)", l)
|
||||
if m:
|
||||
remv.append(m.group(1))
|
||||
return rc, inst, remv, out + err
|
||||
|
||||
@staticmethod
|
||||
def origin_name(origin):
|
||||
# "Debian:13.7/stable, Debian-Security:13/stable-security" -> {"Debian", "Debian-Security"}
|
||||
return {o.strip().split(":")[0] for o in origin.split(",") if o.strip()}
|
||||
|
||||
def free_bytes(self):
|
||||
rc, out, _ = self.g(["df", "-B1", "--output=avail", "/"])
|
||||
try:
|
||||
return int(out.strip().splitlines()[-1])
|
||||
except (ValueError, IndexError):
|
||||
return -1
|
||||
|
||||
def apt_lock_held(self):
|
||||
rc, out, _ = self.g(["fuser", "/var/lib/dpkg/lock-frontend", "/var/lib/dpkg/lock"])
|
||||
return rc == 0 and out.strip() != ""
|
||||
|
||||
def health(self):
|
||||
"""The guest's signals: every container's state + health, the controller's own health, the network."""
|
||||
rc, out, _ = self.g(["docker", "ps", "-a", "--format", "{{.Names}}\t{{.State}}\t{{.Status}}"], timeout=60)
|
||||
cont = {}
|
||||
for l in out.splitlines():
|
||||
p = l.split("\t")
|
||||
if len(p) == 3:
|
||||
h = "healthy" if "(healthy)" in p[2] else "unhealthy" if "(unhealthy)" in p[2] else \
|
||||
"starting" if "(health: starting)" in p[2] else "none"
|
||||
cont[p[0]] = {"state": p[1], "health": h}
|
||||
nrc, _, _ = self.g(["getent", "hosts", "deb.debian.org"], timeout=30)
|
||||
return {"docker_ok": rc == 0, "containers": cont,
|
||||
"controller": cont.get("felhom-controller", {}).get("health", "absent"),
|
||||
"network_ok": nrc == 0}
|
||||
|
||||
def restart_needed(self):
|
||||
"""Processes still mapping deleted files, OUTSIDE docker containers (C11)."""
|
||||
script = ('for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; '
|
||||
'grep -q "docker" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done')
|
||||
rc, out, _ = self.g(["sh", "-c", script], timeout=120)
|
||||
procs = sorted({l.split(" ", 1)[1] for l in out.splitlines() if " " in l})
|
||||
pid1 = any(l.split(" ", 1)[0] == "1" for l in out.splitlines())
|
||||
return procs, pid1
|
||||
|
||||
def inventory(self):
|
||||
inst = self.installed()
|
||||
names = sorted(inst)
|
||||
origins = {}
|
||||
for i in range(0, len(names), 200):
|
||||
rc, out, _ = self.g(["apt-cache", "policy"] + names[i:i + 200])
|
||||
cur, star = None, False
|
||||
for l in out.splitlines():
|
||||
if not l.startswith(" "):
|
||||
cur, star = l.rstrip(":"), False
|
||||
continue
|
||||
s = l.strip()
|
||||
if s.startswith("*** "):
|
||||
star = True
|
||||
continue
|
||||
if star and cur and re.match(r"^[0-9-]+ ", s):
|
||||
if "/var/lib/dpkg/status" in s:
|
||||
origins.setdefault(cur, "local")
|
||||
else:
|
||||
origins[cur] = s
|
||||
continue
|
||||
if star and not re.match(r"^[0-9-]+ ", s):
|
||||
star = False
|
||||
# Map an index URL to an origin name the hub understands.
|
||||
def oname(src):
|
||||
if src in (None, "local"):
|
||||
return "unknown"
|
||||
if "security" in src and "debian" in src:
|
||||
return "Debian-Security"
|
||||
if "docker.com" in src:
|
||||
return "Docker"
|
||||
if "debian" in src:
|
||||
return "Debian"
|
||||
return "other"
|
||||
rc, pend, remv, _ = self.simulate(["dist-upgrade"])
|
||||
return {
|
||||
"installed": [{"name": n, "version": inst[n], "origin": oname(origins.get(n))} for n in names],
|
||||
"pending": [{"name": p["name"], "from": p["from"], "to": p["to"],
|
||||
"origin": sorted(self.origin_name(p["origin"]))} for p in pend],
|
||||
}
|
||||
|
||||
# ---------- the run ----------
|
||||
def run(self):
|
||||
plan = self.load_plan()
|
||||
self.mode, self.vmid = self.check_plan(plan)
|
||||
self.report.update(mode=self.mode, release_id=plan.get("release_id"), vmid=self.vmid)
|
||||
self.check_guest(self.vmid)
|
||||
log = self.r.log
|
||||
if self.mode == "health":
|
||||
self.report["health"] = self.health()
|
||||
return 0
|
||||
log(f"os-apply: START release={plan.get('release_id')} layer=guest:{self.vmid} lane=fast mode={self.mode} packages={len(plan.get('packages', []))}")
|
||||
if self.apt_lock_held():
|
||||
raise Refused("R9", "another apt/dpkg holds the lock in the guest")
|
||||
self.report["health_before"] = self.health()
|
||||
if self.mode == "apply":
|
||||
self.repair()
|
||||
rc, out, err = self.g(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
|
||||
if rc != 0:
|
||||
raise Refused("R7", f"apt-get update failed in the guest: {(out + err).strip().splitlines()[-1:]}")
|
||||
if self.mode == "apply":
|
||||
rc = self.apply(plan)
|
||||
if rc:
|
||||
return rc
|
||||
procs, pid1 = self.restart_needed()
|
||||
self.report.update(self.inventory())
|
||||
self.report["restart_needed"] = procs
|
||||
self.report["docker_restart_needed"] = any(p in ("dockerd", "containerd") for p in procs)
|
||||
self.report["reboot_needed"] = pid1
|
||||
self.report["health_after"] = self.health()
|
||||
return 0
|
||||
|
||||
def repair(self):
|
||||
rc, before, _ = self.g(["dpkg", "--audit"])
|
||||
self.g(APT_ENV + ["dpkg", "--configure", "-a", "--force-confold"])
|
||||
rc2, out, err = self.g(APT_ENV + ["apt-get", "-f", "install", "-y", "-q"] + DPKG_OPTS)
|
||||
_, after, _ = self.g(["dpkg", "--audit"])
|
||||
configured = len([l for l in before.splitlines() if l.startswith(" ")])
|
||||
fixed = len(re.findall(r"^Setting up ", out, re.M))
|
||||
self.report["repair"] = {"half_configured_before": configured, "fixed": fixed, "clean_after": after.strip() == ""}
|
||||
self.r.log(f"os-apply: REPAIR configured={configured} fixed={fixed}")
|
||||
if after.strip():
|
||||
raise Refused("R13", "dpkg is still broken after the repair: " + after.strip().splitlines()[0])
|
||||
|
||||
def apply(self, plan):
|
||||
log = self.r.log
|
||||
inst = self.installed()
|
||||
upgrade, already, notinst = [], 0, 0
|
||||
for e in plan["packages"]:
|
||||
n, v = e["name"], e["version"]
|
||||
if n not in inst:
|
||||
notinst += 1
|
||||
continue
|
||||
if not self.dpkg_cmp(v, "gt", inst[n]):
|
||||
already += 1
|
||||
continue
|
||||
upgrade.append((n, v))
|
||||
from_snap = 0
|
||||
missing = [(n, v) for n, v in upgrade if v not in self.madison(n)]
|
||||
if missing:
|
||||
snap = plan.get("snapshot", "")
|
||||
if not snap:
|
||||
raise Refused("R7", f"{missing[0][0]}={missing[0][1]} is not downloadable and the plan names no snapshot")
|
||||
self.add_snapshot_sources(snap)
|
||||
still = [(n, v) for n, v in missing if v not in self.madison(n)]
|
||||
if still:
|
||||
self.remove_snapshot_sources()
|
||||
raise Refused("R7", f"{still[0][0]}={still[0][1]} is not downloadable, not even from snapshot {snap}")
|
||||
from_snap = len(missing)
|
||||
try:
|
||||
log(f"os-apply: PLAN upgrade={len(upgrade)} already={already} not-installed={notinst} from-snapshot={from_snap}")
|
||||
self.report["plan"] = {"upgrade": len(upgrade), "already": already, "not_installed": notinst, "from_snapshot": from_snap}
|
||||
if not upgrade:
|
||||
self.report["upgraded"] = []
|
||||
log("os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)")
|
||||
return 0
|
||||
args = ["install", "--only-upgrade", "--no-install-recommends"] + [f"{n}={v}" for n, v in upgrade]
|
||||
rc, sim, remv, text = self.simulate(args)
|
||||
if rc != 0:
|
||||
tail = text.strip().splitlines()[-1] if text.strip() else ""
|
||||
raise Refused("R7", "the simulation failed: " + tail)
|
||||
if remv:
|
||||
raise Refused("R4", f"the plan would remove {', '.join(remv[:5])}")
|
||||
want = dict(upgrade)
|
||||
for p in sim:
|
||||
if p["from"] is None:
|
||||
raise Refused("R6", f"the plan would add a package that is not installed: {p['name']}")
|
||||
if p["name"] not in want:
|
||||
raise Refused("R6", f"the plan would touch {p['name']}, which is not in the plan")
|
||||
if p["to"] != want[p["name"]]:
|
||||
raise Refused("R6", f"{p['name']} would go to {p['to']}, not the approved {want[p['name']]}")
|
||||
if not self.dpkg_cmp(p["to"], "gt", p["from"]):
|
||||
raise Refused("R5", f"{p['name']} would be downgraded {p['from']} -> {p['to']}")
|
||||
if not self.origin_name(p["origin"]) & set(FAST_ORIGINS):
|
||||
raise Refused("R2", f"{p['name']} would come from {p['origin']}, not Debian")
|
||||
need = self.download_bytes(args)
|
||||
free = self.free_bytes()
|
||||
if free >= 0 and free < max(MIN_FREE, 3 * need):
|
||||
raise Refused("R8", f"free space {free} B is below max(500 MB, 3 x download {need} B)")
|
||||
t0 = time.time()
|
||||
rc, out, err = self.g(APT_ENV + ["apt-get", "-y", "-q"] + DPKG_OPTS + args)
|
||||
secs = time.time() - t0
|
||||
# dpkg says "Installing new version of config file X" when X was NOT changed locally (the package's new
|
||||
# version is taken), and "Configuration file 'X'" + "Keeping old config file" when it was (--force-confold
|
||||
# keeps the local one; the package's version lands as X.dpkg-dist). Measured live 2026-10-04 (debian_version).
|
||||
conflict = None
|
||||
for l in (out + err).splitlines():
|
||||
m = re.search(r"Installing new version of config file (\S+?)\s*\.\.\.", l)
|
||||
if m:
|
||||
log(f"os-apply: CONFFILE updated {m.group(1)} (it was not changed locally)")
|
||||
m = re.search(r"Configuration file '([^']+)'", l)
|
||||
if m:
|
||||
conflict = m.group(1)
|
||||
if conflict and "Keeping old config file" in l:
|
||||
log(f"os-apply: CONFFILE kept {conflict} (changed locally; the package's version is {conflict}.dpkg-dist)")
|
||||
self.report.setdefault("conffiles_kept", []).append(conflict)
|
||||
conflict = None
|
||||
self.g(["apt-get", "clean"])
|
||||
if rc != 0:
|
||||
_, aud, _ = self.g(["dpkg", "--audit"])
|
||||
first = aud.strip().splitlines()[0] if aud.strip() else "clean"
|
||||
log(f"os-apply: FAILED rc={rc} step=install — dpkg state: {first}")
|
||||
self.report["failed"] = {"rc": rc, "dpkg_audit": first, "tail": (out + err).strip().splitlines()[-3:]}
|
||||
return 3
|
||||
self.report["upgraded"] = [{"name": n, "version": v} for n, v in upgrade]
|
||||
self.report["seconds"] = round(secs, 1)
|
||||
log(f"os-apply: DONE rc=0 seconds={secs:.1f} upgraded={len(upgrade)}")
|
||||
return 0
|
||||
finally:
|
||||
if from_snap:
|
||||
self.remove_snapshot_sources()
|
||||
|
||||
def download_bytes(self, args):
|
||||
rc, out, _ = self.g(APT_ENV + ["apt-get", "-s", "-o", "Debug::NoLocking=1", "--print-uris", "-q"] + args)
|
||||
total = 0
|
||||
for l in out.splitlines():
|
||||
m = re.match(r"^'[^']+' \S+ ([0-9]+) ", l)
|
||||
if m:
|
||||
total += int(m.group(1))
|
||||
return total
|
||||
|
||||
def add_snapshot_sources(self, snap):
|
||||
rc, out, _ = self.g(["sh", "-c", ". /etc/os-release && echo $VERSION_CODENAME"])
|
||||
code = out.strip()
|
||||
if not re.match(r"^[a-z]+$", code):
|
||||
raise Refused("R7", f"cannot read the guest's Debian codename ({code!r})")
|
||||
body = (f"deb [check-valid-until=no] http://snapshot.debian.org/archive/debian/{snap} {code} main\n"
|
||||
f"deb [check-valid-until=no] http://snapshot.debian.org/archive/debian-security/{snap} {code}-security main\n")
|
||||
self.r.guest_write(self.vmid, SNAPSHOT_LIST, body)
|
||||
self.r.log(f"os-apply: SNAPSHOT using snapshot.debian.org/{snap} for versions no longer published (decision 79)")
|
||||
rc, out, err = self.g(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
|
||||
if rc != 0:
|
||||
self.remove_snapshot_sources()
|
||||
raise Refused("R7", "apt-get update against snapshot.debian.org failed")
|
||||
|
||||
def remove_snapshot_sources(self):
|
||||
self.g(["rm", "-f", SNAPSHOT_LIST])
|
||||
self.g(APT_ENV + ["apt-get", "-q", "update"], timeout=600)
|
||||
|
||||
|
||||
def guest_write(self, vmid, path, body):
|
||||
"""Write a small text file inside the guest via `pct exec … tee` (stdin), never via a shell string."""
|
||||
p = subprocess.run(["/usr/sbin/pct", "exec", str(vmid), "--", "tee", path], input=body,
|
||||
capture_output=True, text=True, timeout=60)
|
||||
if p.returncode != 0:
|
||||
raise Refused("R7", f"could not write {path} in the guest")
|
||||
|
||||
|
||||
Runner.guest_write = guest_write
|
||||
|
||||
|
||||
def main(argv, runner=None):
|
||||
r = runner or Runner()
|
||||
if len(argv) != 3 or argv[1] != "--plan":
|
||||
r.log("os-apply: REFUSED: R1 usage: felhom-os-apply --plan /var/lib/felhom-agent/os/plan-<id>.json")
|
||||
print("OSAPPLY-REPORT " + json.dumps({"refused": {"code": "R1", "reason": "usage"}}))
|
||||
return 2
|
||||
a = Apply(r, argv[2])
|
||||
try:
|
||||
rc = a.run()
|
||||
except Refused as e:
|
||||
r.log(f"os-apply: REFUSED: {e.code} {e.reason}")
|
||||
a.report["refused"] = {"code": e.code, "reason": e.reason}
|
||||
rc = 2
|
||||
except subprocess.TimeoutExpired as e:
|
||||
r.log(f"os-apply: FAILED rc=124 step=timeout — {e.cmd}")
|
||||
a.report["failed"] = {"rc": 124, "timeout": str(e.cmd)[:200]}
|
||||
rc = 3
|
||||
print("OSAPPLY-REPORT " + json.dumps(a.report, sort_keys=True))
|
||||
return rc
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if os.geteuid() != 0:
|
||||
print("felhom-os-apply: must run as root (via sudo)", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
sys.exit(main(sys.argv))
|
||||
@@ -0,0 +1,467 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Tests for configs/felhom-os-apply (`11` §5.4.1). A fake runner plays the host and the guest: nothing is executed
|
||||
for real except the local `dpkg --compare-versions` (pure, no network). Every refusal R1–R13 has a test; the red-proof
|
||||
(each test fails when its rule is removed) is `audits/os-guest-lane-2026-10-04/partB/redproof.txt`.
|
||||
|
||||
Run: python3 configs/test_felhom_os_apply.py (also run by internal/osupdate's Go test)
|
||||
"""
|
||||
import importlib.machinery
|
||||
import importlib.util
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
import re
|
||||
import stat as statmod
|
||||
import subprocess
|
||||
import unittest
|
||||
|
||||
HERE = pathlib.Path(__file__).resolve().parent
|
||||
_loader = importlib.machinery.SourceFileLoader("osapply", os.environ.get("OSAPPLY_UNDER_TEST", str(HERE / "felhom-os-apply"))) # red-proof seam
|
||||
_spec = importlib.util.spec_from_loader("osapply", _loader)
|
||||
osapply = importlib.util.module_from_spec(_spec)
|
||||
_loader.exec_module(osapply)
|
||||
|
||||
PLAN = "/var/lib/felhom-agent/os/plan-t1.json"
|
||||
CONF_OK = ("arch: amd64\nmp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G\n"
|
||||
"mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\nrootfs: local-lvm:vm-9201-disk-0,size=32G\n")
|
||||
DEB = "Debian:13.7/stable"
|
||||
SEC = "Debian-Security:13/stable-security"
|
||||
|
||||
|
||||
def dpkg_cmp(a, op, b):
|
||||
return subprocess.run(["dpkg", "--compare-versions", a, op, b]).returncode == 0
|
||||
|
||||
|
||||
class St:
|
||||
def __init__(self, mode=statmod.S_IFREG | 0o600, uid=999, size=100):
|
||||
self.st_mode, self.st_uid, self.st_size = mode, uid, size
|
||||
|
||||
|
||||
class Fake:
|
||||
"""The host + one guest. `installed` / `live` (name -> versions in the live archive) / `snapshot` (versions
|
||||
the snapshot archive adds) / `extra_sim` (lines the simulation adds) / `dpkg_audit` / `free`."""
|
||||
|
||||
def __init__(self):
|
||||
self.plan = {"release_id": "os-t1", "layer": "guest", "lane": "fast", "vmid": 9201, "mode": "apply",
|
||||
"snapshot": "20261004T080000Z",
|
||||
"packages": [{"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"},
|
||||
{"name": "openssl", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}]}
|
||||
self.files = {"/etc/pve/lxc/9201.conf": CONF_OK}
|
||||
self.stats = {PLAN: St()}
|
||||
self.installed = {"libc6": "2.41-12+deb13u3", "openssl": "3.5.6-1~deb13u1", "bash": "5.2.37-2+b9"}
|
||||
self.live = {"libc6": {"2.41-12+deb13u4"}, "openssl": {"3.5.7-1~deb13u3"}}
|
||||
self.snapshot = {}
|
||||
self.snap_active = False
|
||||
self.extra_sim = []
|
||||
self.dpkg_audit = ""
|
||||
self.free = 10 * 1024 ** 3
|
||||
self.install_rc = 0
|
||||
self.calls = []
|
||||
self.logs = []
|
||||
self.written = {}
|
||||
self.status = "status: running"
|
||||
self.lock_held = False
|
||||
|
||||
# Runner interface
|
||||
def read_file(self, p):
|
||||
if p == PLAN:
|
||||
return json.dumps(self.plan)
|
||||
if p not in self.files:
|
||||
raise OSError("no such file")
|
||||
return self.files[p]
|
||||
|
||||
def stat(self, p):
|
||||
if p not in self.stats:
|
||||
raise OSError("no such file")
|
||||
return self.stats[p]
|
||||
|
||||
def agent_uid(self):
|
||||
return 999
|
||||
|
||||
def log(self, line):
|
||||
self.logs.append(line)
|
||||
|
||||
def host(self, argv, timeout=600):
|
||||
self.calls.append(("host", argv))
|
||||
if argv[1] == "status":
|
||||
return 0, self.status + "\n", ""
|
||||
return 1, "", "unexpected host call"
|
||||
|
||||
def guest_write(self, vmid, path, body):
|
||||
self.written[path] = body
|
||||
if path == osapply.SNAPSHOT_LIST:
|
||||
self.snap_active = True
|
||||
|
||||
def avail(self, n):
|
||||
v = set(self.live.get(n, set()))
|
||||
if self.snap_active:
|
||||
v |= self.snapshot.get(n, set())
|
||||
return v
|
||||
|
||||
def guest(self, vmid, argv, timeout=1800):
|
||||
self.calls.append(("guest", vmid, argv))
|
||||
a = [x for x in argv if not re.match(r"^[A-Z_]+=", x) and x != "env"]
|
||||
cmd = a[0]
|
||||
if cmd == "dpkg-query":
|
||||
return 0, "".join(f"{n}\t{v}\tii \n" for n, v in self.installed.items()), ""
|
||||
if cmd == "dpkg" and a[1] == "--compare-versions":
|
||||
return (0 if dpkg_cmp(a[2], a[3], a[4]) else 1), "", ""
|
||||
if cmd == "dpkg" and a[1] == "--audit":
|
||||
return 0, self.dpkg_audit, ""
|
||||
if cmd == "dpkg" and a[1] == "--configure":
|
||||
return 0, "", ""
|
||||
if cmd == "fuser":
|
||||
return (0, " 123", "") if self.lock_held else (1, "", "")
|
||||
if cmd == "apt-cache" and a[1] == "madison":
|
||||
return 0, "".join(f" {a[2]} | {v} | http://deb.debian.org trixie/main amd64 Packages\n" for v in self.avail(a[2])), ""
|
||||
if cmd == "apt-cache" and a[1] == "policy":
|
||||
out = ""
|
||||
for n in a[2:]:
|
||||
out += f"{n}:\n Installed: {self.installed.get(n)}\n Version table:\n *** {self.installed.get(n)} 500\n 500 http://deb.debian.org/debian trixie/main amd64 Packages\n"
|
||||
return 0, out, ""
|
||||
if cmd == "apt-get":
|
||||
if "update" in a:
|
||||
return 0, "", ""
|
||||
if "clean" in a:
|
||||
return 0, "", ""
|
||||
if "-f" in a:
|
||||
self.dpkg_audit = ""
|
||||
return 0, "Setting up x (1) ...\n" if getattr(self, "repaired", False) else "", ""
|
||||
if "-s" in a:
|
||||
return self.sim(a)
|
||||
if "install" in a:
|
||||
if self.install_rc:
|
||||
return self.install_rc, "", "E: boom"
|
||||
for x in a:
|
||||
if "=" in x and not x.startswith("-") and "::" not in x:
|
||||
n, v = x.split("=", 1)
|
||||
self.installed[n] = v
|
||||
return 0, getattr(self, "install_out", "Setting up libc6 ...\n"), ""
|
||||
if cmd == "df":
|
||||
return 0, f"Avail\n{self.free}\n", ""
|
||||
if cmd == "docker":
|
||||
return 0, "felhom-controller\trunning\tUp 1 hour (healthy)\napp\trunning\tUp 1 hour (healthy)\n", ""
|
||||
if cmd == "getent":
|
||||
return 0, "1.2.3.4 deb.debian.org\n", ""
|
||||
if cmd == "sh":
|
||||
if "os-release" in a[2]:
|
||||
return 0, "trixie\n", ""
|
||||
return 0, "", ""
|
||||
if cmd == "rm":
|
||||
self.snap_active = False
|
||||
return 0, "", ""
|
||||
return 1, "", f"unexpected guest call {a}"
|
||||
|
||||
def sim(self, a):
|
||||
if "--print-uris" in a:
|
||||
return 0, "'http://x/libc6.deb' libc6.deb 4000000 SHA256:x\n", ""
|
||||
if "dist-upgrade" in a:
|
||||
return 0, "Inst bash [5.2.37-2+b9] (5.2.37-2+b10 Debian:13.7/stable [amd64])\n", ""
|
||||
out = ""
|
||||
for x in a:
|
||||
if "=" in x and not x.startswith("-") and "::" not in x:
|
||||
n, v = x.split("=", 1)
|
||||
if v not in self.avail(n):
|
||||
return 100, "", f"E: Version '{v}' for '{n}' was not found"
|
||||
origin = SEC if n == "openssl" else DEB
|
||||
out += f"Inst {n} [{self.installed[n]}] ({v} {origin} [amd64])\n"
|
||||
out += "".join(l + "\n" for l in self.extra_sim)
|
||||
return 0, out, ""
|
||||
|
||||
|
||||
def run(f):
|
||||
import io
|
||||
import contextlib
|
||||
buf = io.StringIO()
|
||||
with contextlib.redirect_stdout(buf):
|
||||
rc = osapply.main(["felhom-os-apply", "--plan", PLAN], runner=f)
|
||||
line = [l for l in buf.getvalue().splitlines() if l.startswith("OSAPPLY-REPORT ")][-1]
|
||||
return rc, json.loads(line[len("OSAPPLY-REPORT "):])
|
||||
|
||||
|
||||
class Happy(unittest.TestCase):
|
||||
def test_apply_installs_exactly_the_plan(self):
|
||||
f = Fake()
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u4")
|
||||
self.assertEqual(f.installed["openssl"], "3.5.7-1~deb13u3")
|
||||
self.assertEqual(f.installed["bash"], "5.2.37-2+b9", "a package outside the plan was changed")
|
||||
self.assertEqual(rep["plan"]["upgrade"], 2)
|
||||
self.assertIn("installed", rep)
|
||||
self.assertEqual(rep["pending"][0]["name"], "bash")
|
||||
self.assertTrue(any(l.startswith("os-apply: REPAIR ") for l in f.logs), "the repair line must always print")
|
||||
self.assertTrue(any(l.startswith("os-apply: DONE rc=0") for l in f.logs))
|
||||
inst = [c for c in f.calls if c[0] == "guest" and "install" in c[2] and "-s" not in c[2] and "-f" not in c[2]]
|
||||
self.assertTrue(inst and "Dpkg::Options::=--force-confold" in inst[0][2], "must keep existing config files")
|
||||
|
||||
def test_already_current_is_a_no_op(self):
|
||||
f = Fake()
|
||||
f.installed.update(libc6="2.41-12+deb13u4", openssl="3.5.7-1~deb13u3")
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0)
|
||||
self.assertEqual(rep["plan"]["upgrade"], 0)
|
||||
|
||||
def test_inventory_installs_nothing(self):
|
||||
f = Fake()
|
||||
f.plan["mode"] = "inventory"
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(f.installed["libc6"], "2.41-12+deb13u3")
|
||||
self.assertIn("installed", rep)
|
||||
self.assertIn("restart_needed", rep)
|
||||
|
||||
def test_health_mode(self):
|
||||
f = Fake()
|
||||
f.plan["mode"] = "health"
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0)
|
||||
self.assertEqual(rep["health"]["controller"], "healthy")
|
||||
|
||||
|
||||
class Repair(unittest.TestCase):
|
||||
def test_repair_runs_first_and_is_reported(self):
|
||||
f = Fake()
|
||||
f.dpkg_audit = "The following packages have been unpacked but not yet configured.\n perl Larry Wall's\n"
|
||||
f.repaired = True
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(rep["repair"]["half_configured_before"], 1)
|
||||
self.assertEqual(rep["repair"]["fixed"], 1)
|
||||
order = [i for i, c in enumerate(f.calls) if c[0] == "guest" and c[2][-1:] != ["update"]]
|
||||
first_cfg = next(i for i, c in enumerate(f.calls) if c[0] == "guest" and "--configure" in c[2])
|
||||
first_upd = next(i for i, c in enumerate(f.calls) if c[0] == "guest" and "update" in c[2])
|
||||
self.assertLess(first_cfg, first_upd, "the repair must run before anything else touches apt")
|
||||
self.assertTrue(order)
|
||||
|
||||
|
||||
class Snapshot(unittest.TestCase):
|
||||
def test_a_replaced_version_comes_from_the_snapshot(self):
|
||||
f = Fake()
|
||||
f.live["openssl"] = {"3.5.7-1~deb13u4"} # Debian moved on
|
||||
f.snapshot["openssl"] = {"3.5.7-1~deb13u3"}
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertEqual(rep["plan"]["from_snapshot"], 1)
|
||||
self.assertEqual(f.installed["openssl"], "3.5.7-1~deb13u3", "must install the APPROVED version, not the newer one")
|
||||
body = f.written[osapply.SNAPSHOT_LIST]
|
||||
self.assertIn("snapshot.debian.org/archive/debian/20261004T080000Z trixie main", body)
|
||||
self.assertIn("debian-security/20261004T080000Z trixie-security main", body)
|
||||
self.assertFalse(f.snap_active, "the temporary snapshot sources must be removed after the run")
|
||||
|
||||
def test_snapshot_does_not_have_it_either(self):
|
||||
f = Fake()
|
||||
f.live["openssl"] = set()
|
||||
rc, rep = run(f)
|
||||
self.assertEqual((rc, rep["refused"]["code"]), (2, "R7"))
|
||||
self.assertFalse(f.snap_active)
|
||||
|
||||
|
||||
class Refusals(unittest.TestCase):
|
||||
def refused(self, f, code):
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 2, rep)
|
||||
self.assertEqual(rep["refused"]["code"], code, rep)
|
||||
self.assertTrue(any(l.startswith(f"os-apply: REFUSED: {code} ") for l in f.logs), f.logs)
|
||||
inst = [c for c in f.calls if c[0] == "guest" and "install" in c[2] and "-s" not in c[2] and "-f" not in c[2]]
|
||||
self.assertEqual(inst, [], "a refusal must install nothing")
|
||||
return rep
|
||||
|
||||
def test_R1_usage(self):
|
||||
import io
|
||||
import contextlib
|
||||
f = Fake()
|
||||
with contextlib.redirect_stdout(io.StringIO()):
|
||||
self.assertEqual(osapply.main(["felhom-os-apply", "--plan", PLAN, "--extra"], runner=f), 2)
|
||||
self.assertEqual(osapply.main(["felhom-os-apply", "--plan"], runner=f), 2)
|
||||
|
||||
def test_R1_path_outside_the_plan_dir(self):
|
||||
import io
|
||||
import contextlib
|
||||
f = Fake()
|
||||
with contextlib.redirect_stdout(io.StringIO()):
|
||||
rc = osapply.main(["felhom-os-apply", "--plan", "/tmp/plan-x.json"], runner=f)
|
||||
self.assertEqual(rc, 2)
|
||||
self.assertTrue(any("R1" in l for l in f.logs))
|
||||
|
||||
def test_R1_symlink(self):
|
||||
f = Fake()
|
||||
f.stats[PLAN] = St(mode=statmod.S_IFLNK | 0o777)
|
||||
self.refused(f, "R1")
|
||||
|
||||
def test_R1_not_owned_by_the_agent(self):
|
||||
f = Fake()
|
||||
f.stats[PLAN] = St(uid=0)
|
||||
self.refused(f, "R1")
|
||||
|
||||
def test_R2_non_debian_origin_in_the_plan(self):
|
||||
f = Fake()
|
||||
f.plan["packages"][0]["origin"] = "Proxmox"
|
||||
self.refused(f, "R2")
|
||||
|
||||
def test_R2_non_debian_origin_in_the_simulation(self):
|
||||
f = Fake()
|
||||
f.installed["libc6"] = "2.41-12+deb13u3"
|
||||
orig = f.sim
|
||||
|
||||
def sim(a):
|
||||
rc, out, err = orig(a)
|
||||
return rc, out.replace("Debian:13.7/stable", "Proxmox Debian Repository:stable"), err
|
||||
f.sim = sim
|
||||
self.refused(f, "R2")
|
||||
|
||||
def test_R3_slow_lane(self):
|
||||
f = Fake()
|
||||
f.plan["lane"] = "slow"
|
||||
self.refused(f, "R3")
|
||||
|
||||
def test_R4_removal(self):
|
||||
f = Fake()
|
||||
f.extra_sim = ["Remv bash [5.2.37-2+b9]"]
|
||||
self.refused(f, "R4")
|
||||
|
||||
def test_R5_downgrade(self):
|
||||
f = Fake()
|
||||
f.installed["libc6"] = "2.41-12+deb13u4"
|
||||
f.plan["packages"] = [{"name": "openssl", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}]
|
||||
f.extra_sim = ["Inst openssl [3.5.6-1~deb13u1] (3.5.5-1 Debian:13.7/stable [amd64])"]
|
||||
orig = f.sim
|
||||
|
||||
def sim(a): # the simulation answers with a LOWER version than installed
|
||||
rc, out, err = orig(a)
|
||||
return rc, "\n".join(l for l in out.splitlines() if not l.startswith("Inst openssl [3.5.6-1~deb13u1] (3.5.7")) + "\n", err
|
||||
f.sim = sim
|
||||
f.plan["packages"][0]["version"] = "3.5.7-1~deb13u3"
|
||||
rep = run(f)[1]
|
||||
# The plan asks 3.5.7; the simulation goes to 3.5.5: that is BOTH a wrong version (R6) and a downgrade.
|
||||
self.assertIn(rep["refused"]["code"], ("R5", "R6"))
|
||||
|
||||
def test_R5_downgrade_exact(self):
|
||||
f = Fake()
|
||||
f.plan["packages"] = [{"name": "openssl", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}]
|
||||
f.installed["openssl"] = "3.5.6-1~deb13u1"
|
||||
orig = f.sim
|
||||
|
||||
def sim(a):
|
||||
rc, out, err = orig(a)
|
||||
return rc, out.replace("[3.5.6-1~deb13u1]", "[3.5.8-1]"), err
|
||||
f.sim = sim
|
||||
self.refused(f, "R5")
|
||||
|
||||
def test_R6_new_package(self):
|
||||
f = Fake()
|
||||
f.extra_sim = ["Inst newthing (1.0 Debian:13.7/stable [amd64])"]
|
||||
self.refused(f, "R6")
|
||||
|
||||
def test_R6_unlisted_package(self):
|
||||
f = Fake()
|
||||
f.extra_sim = ["Inst bash [5.2.37-2+b9] (5.2.37-2+b10 Debian:13.7/stable [amd64])"]
|
||||
self.refused(f, "R6")
|
||||
|
||||
def test_R6_allow_new_is_slow_lane(self):
|
||||
f = Fake()
|
||||
f.plan["allow_new"] = ["proxmox-kernel-x"]
|
||||
self.refused(f, "R6")
|
||||
|
||||
def test_R7_not_downloadable_and_no_snapshot(self):
|
||||
f = Fake()
|
||||
f.live["openssl"] = set()
|
||||
f.plan["snapshot"] = ""
|
||||
self.refused(f, "R7")
|
||||
|
||||
def test_R8_free_space(self):
|
||||
f = Fake()
|
||||
f.free = 100 * 1024 * 1024
|
||||
self.refused(f, "R8")
|
||||
|
||||
def test_R9_guest_locked_by_a_backup(self):
|
||||
f = Fake()
|
||||
f.files["/etc/pve/lxc/9201.conf"] = CONF_OK + "lock: backup\n"
|
||||
self.refused(f, "R9")
|
||||
|
||||
def test_R9_apt_lock_held(self):
|
||||
f = Fake()
|
||||
f.lock_held = True
|
||||
self.refused(f, "R9")
|
||||
|
||||
def test_R10_not_the_boxs_own_guest(self):
|
||||
f = Fake()
|
||||
f.files["/etc/pve/lxc/9201.conf"] = CONF_OK.replace("mp8: /mnt/felhom-drives,", "mp8: /mnt/hdd_1/scratch,")
|
||||
self.refused(f, "R10")
|
||||
|
||||
def test_R10_reserved_vmid(self):
|
||||
f = Fake()
|
||||
f.plan["vmid"] = 990003
|
||||
self.refused(f, "R10")
|
||||
|
||||
def test_R10_bind_only_in_a_snapshot_section(self):
|
||||
f = Fake()
|
||||
f.files["/etc/pve/lxc/9201.conf"] = "rootfs: x\n[snap1]\nmp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\n"
|
||||
self.refused(f, "R10")
|
||||
|
||||
def test_R10_not_running(self):
|
||||
f = Fake()
|
||||
f.status = "status: stopped"
|
||||
self.refused(f, "R10")
|
||||
|
||||
def test_R11_duplicate(self):
|
||||
f = Fake()
|
||||
f.plan["packages"].append(dict(f.plan["packages"][0]))
|
||||
self.refused(f, "R11")
|
||||
|
||||
def test_R11_bad_version_string(self):
|
||||
f = Fake()
|
||||
f.plan["packages"][0]["version"] = "1.0; rm -rf /"
|
||||
self.refused(f, "R11")
|
||||
|
||||
def test_R11_bad_name(self):
|
||||
f = Fake()
|
||||
f.plan["packages"][0]["name"] = "--purge"
|
||||
self.refused(f, "R11")
|
||||
|
||||
def test_R12_host_layer(self):
|
||||
f = Fake()
|
||||
f.plan["layer"] = "host"
|
||||
self.refused(f, "R12")
|
||||
|
||||
def test_R13_repair_does_not_fix_it(self):
|
||||
f = Fake()
|
||||
f.dpkg_audit = "The following packages are broken\n perl\n"
|
||||
orig = f.guest
|
||||
|
||||
def guest(vmid, argv, timeout=1800):
|
||||
rc, out, err = orig(vmid, argv, timeout)
|
||||
if "-f" in argv:
|
||||
f.dpkg_audit = "The following packages are broken\n perl\n"
|
||||
return rc, out, err
|
||||
f.guest = guest
|
||||
self.refused(f, "R13")
|
||||
|
||||
|
||||
class Conffiles(unittest.TestCase):
|
||||
# dpkg's two shapes, measured live 2026-10-04: an unchanged file is UPDATED; a locally changed one is KEPT.
|
||||
def test_updated_vs_kept(self):
|
||||
f = Fake()
|
||||
f.install_out = ("Installing new version of config file /etc/debian_version ...\n"
|
||||
"Configuration file '/etc/ssh/sshd_config'\n ==> Modified (by you or by a script) since installation.\n"
|
||||
" ==> Keeping old config file as default.\n")
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 0, rep)
|
||||
self.assertIn("os-apply: CONFFILE updated /etc/debian_version (it was not changed locally)", f.logs)
|
||||
self.assertTrue(any(l.startswith("os-apply: CONFFILE kept /etc/ssh/sshd_config") for l in f.logs), f.logs)
|
||||
self.assertEqual(rep["conffiles_kept"], ["/etc/ssh/sshd_config"])
|
||||
self.assertFalse(any("kept /etc/debian_version" in l for l in f.logs), "an updated file must not be reported as kept")
|
||||
|
||||
|
||||
class Failure(unittest.TestCase):
|
||||
def test_install_failure_is_rc3_with_dpkg_state(self):
|
||||
f = Fake()
|
||||
f.install_rc = 100
|
||||
rc, rep = run(f)
|
||||
self.assertEqual(rc, 3)
|
||||
self.assertEqual(rep["failed"]["rc"], 100)
|
||||
self.assertTrue(any(l.startswith("os-apply: FAILED rc=100 step=install") for l in f.logs))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -2,6 +2,7 @@ package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"io"
|
||||
@@ -21,14 +22,18 @@ type fakeBackupAPI struct {
|
||||
vzdumpErr error
|
||||
waitErr error
|
||||
cfg proxmox.GuestConfig
|
||||
goneGuests map[int]bool // R-689: vmids whose config lookup answers "does not exist"
|
||||
aclGuests map[int]bool // R-689: vmids outside the token's ACL — PVE answers 403 "permission denied"
|
||||
cfgErr error
|
||||
content []proxmox.StorageContent
|
||||
contentErr error
|
||||
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate)
|
||||
storageErr error
|
||||
vzdumps []proxmox.VzdumpOptions
|
||||
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
|
||||
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
|
||||
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate) — DEFINITIONS: like
|
||||
// production's GET /storage, it never carries usage; ListStorage strips Avail/Used (R-685's live lesson)
|
||||
nodeStorages []proxmox.Storage // returned by NodeStorage (GET /nodes/{node}/storage — WITH usage)
|
||||
storageErr error
|
||||
vzdumps []proxmox.VzdumpOptions
|
||||
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
|
||||
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
|
||||
}
|
||||
|
||||
func (f *fakeBackupAPI) Vzdump(_ context.Context, o proxmox.VzdumpOptions) (string, error) {
|
||||
@@ -41,13 +46,30 @@ func (f *fakeBackupAPI) WaitTask(_ context.Context, _ string, _ proxmox.WaitOpti
|
||||
}
|
||||
return proxmox.TaskStatus{Status: "stopped", ExitStatus: "OK"}, f.waitErr
|
||||
}
|
||||
func (f *fakeBackupAPI) GuestConfig(_ context.Context, _ int) (proxmox.GuestConfig, error) {
|
||||
func (f *fakeBackupAPI) GuestConfig(_ context.Context, vmid int) (proxmox.GuestConfig, error) {
|
||||
if f.aclGuests[vmid] {
|
||||
return proxmox.GuestConfig{}, fmt.Errorf("proxmox: GET /nodes/n/lxc/%d/config -> HTTP 403: permission denied at /vms/%d (missing privilege VM.Audit)", vmid, vmid)
|
||||
}
|
||||
if f.goneGuests[vmid] { // R-689: PVE's answer for a deleted guest
|
||||
return proxmox.GuestConfig{}, fmt.Errorf("proxmox: GET /nodes/n/lxc/%d/config -> HTTP 500: Configuration file 'nodes/n/lxc/%d.conf' does not exist", vmid, vmid)
|
||||
}
|
||||
return f.cfg, f.cfgErr
|
||||
}
|
||||
func (f *fakeBackupAPI) StorageContent(_ context.Context, _ string) ([]proxmox.StorageContent, error) {
|
||||
return f.content, f.contentErr
|
||||
}
|
||||
func (f *fakeBackupAPI) ListStorage(_ context.Context) ([]proxmox.Storage, error) {
|
||||
out := make([]proxmox.Storage, len(f.storages))
|
||||
for i, s := range f.storages {
|
||||
s.Avail, s.Used, s.Total = 0, 0, 0 // GET /storage has no usage — a fake that had it hid R-685's defect
|
||||
out[i] = s
|
||||
}
|
||||
return out, f.storageErr
|
||||
}
|
||||
func (f *fakeBackupAPI) NodeStorage(_ context.Context) ([]proxmox.Storage, error) {
|
||||
if f.nodeStorages != nil {
|
||||
return f.nodeStorages, f.storageErr
|
||||
}
|
||||
return f.storages, f.storageErr
|
||||
}
|
||||
func (f *fakeBackupAPI) TaskLogTail(_ context.Context, _ string, _ int) ([]string, error) {
|
||||
@@ -137,13 +159,14 @@ func TestBackup_VzdumpFailureReturnsFailedRecord(t *testing.T) {
|
||||
func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
|
||||
const big = 4 << 30 // a plausible whole-guest archive
|
||||
api := &fakeBackupAPI{content: []proxmox.StorageContent{
|
||||
{VolID: "a", Content: "backup", CTime: 10, Size: big},
|
||||
{VolID: "b", Content: "backup", CTime: 99, Size: big},
|
||||
// R-689: real vzdump names with their vmid — only a backup OF A GUEST is a candidate.
|
||||
{VolID: "local:backup/vzdump-lxc-9001-a.tar.zst", VMID: 9001, Content: "backup", CTime: 10, Size: big},
|
||||
{VolID: "local:backup/vzdump-lxc-9001-b.tar.zst", VMID: 9001, Content: "backup", CTime: 99, Size: big},
|
||||
{VolID: "iso", Content: "iso", CTime: 999, Size: big}, // not a backup → ignored
|
||||
}}
|
||||
r := NewBackupRunner(api, "local", "", "", "", quiet())
|
||||
vol, err := r.PickRestoreCandidate(context.Background())
|
||||
if err != nil || vol != "b" {
|
||||
if err != nil || vol != "local:backup/vzdump-lxc-9001-b.tar.zst" {
|
||||
t.Fatalf("pick = %q,%v want newest 'b'", vol, err)
|
||||
}
|
||||
// no backups → "".
|
||||
@@ -163,12 +186,12 @@ func TestPickRestoreCandidate_NewestOrEmpty(t *testing.T) {
|
||||
// `pick = "phantom" want the newest COMPLETE archive 'real'`.
|
||||
func TestPickRestoreCandidate_SkipsImplausibleArchives(t *testing.T) {
|
||||
api := &fakeBackupAPI{content: []proxmox.StorageContent{
|
||||
{VolID: "real", Content: "backup", CTime: 10, Size: 4 << 30},
|
||||
{VolID: "phantom", Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
|
||||
{VolID: "felhom-pbs:backup/ct/9001/real", VMID: 9001, Content: "backup", CTime: 10, Size: 4 << 30},
|
||||
{VolID: "felhom-pbs:backup/ct/9001/phantom", VMID: 9001, Content: "backup", CTime: 99, Size: 1}, // newest, and impossible
|
||||
}}
|
||||
r := NewBackupRunner(api, "local", "", "", "", quiet())
|
||||
vol, err := r.PickRestoreCandidate(context.Background())
|
||||
if err != nil || vol != "real" {
|
||||
if err != nil || vol != "felhom-pbs:backup/ct/9001/real" {
|
||||
t.Fatalf("pick = %q,%v want the newest COMPLETE archive 'real'", vol, err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
)
|
||||
|
||||
// R-672 (v0.133.0): a restore-test the SPACE preflight refused is the test's RESULT — recorded for the
|
||||
// hub as pass=false with the reason, never dropped (a band skip still is) and never a pass.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT): the scheduler's pre-v0.133.0 `if res.Skipped { return }` → "a space
|
||||
// refusal never reached the host report".
|
||||
func TestR672_SpaceSkipIsReportedNotDropped(t *testing.T) {
|
||||
store := NewStore()
|
||||
rt := &fakeRTRunner{res: reconcile.RestoreTestResult{Archive: "vol", Skipped: true,
|
||||
SkipReason: "skipped: not enough space on local-lvm: restoring 21.1 GiB (vzdump log) needs 30.3 GiB free, has 21.6 GiB"}}
|
||||
s := NewScheduler(SchedulerOptions{
|
||||
Runner: rt, Pick: func(context.Context) (string, error) { return "vol", nil }, Store: store,
|
||||
Spec: func(context.Context, string) reconcile.RestoreTestSpec {
|
||||
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009}
|
||||
},
|
||||
Cadence: time.Hour, Logger: quiet(),
|
||||
})
|
||||
s.tick(context.Background())
|
||||
got := store.RestoreTests(context.Background())
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("a space refusal never reached the host report: %+v", got)
|
||||
}
|
||||
if got[0].Pass || !got[0].Skipped || !strings.HasPrefix(got[0].Error, "skipped: not enough space") {
|
||||
t.Fatalf("record = %+v — want pass=false, skipped, the reason as the error", got[0])
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,83 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-685 (v0.134.0) — a whole-box backup that cannot fit its LOCAL target is a named skip BEFORE anything
|
||||
// starts, with the numbers in the record's Error, never a vzdump that fills the disk and fails.
|
||||
|
||||
const gib = int64(1) << 30
|
||||
|
||||
func spaceAPI(avail int64, lastArchive int64, typ string) *fakeBackupAPI {
|
||||
api := &fakeBackupAPI{vzdumpUPID: "UPID:vzdump:1",
|
||||
storages: []proxmox.Storage{{Storage: "local", Type: typ, Content: "backup", Avail: avail}}}
|
||||
if lastArchive > 0 {
|
||||
api.content = []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst",
|
||||
Content: "backup", VMID: 9201, Size: lastArchive, CTime: 1790280000}}
|
||||
}
|
||||
return api
|
||||
}
|
||||
|
||||
// TestR685_BackupThatCannotFitIsSkipped — demo-hp's shape: an 8.2 GB archive, 4 GiB free. No vzdump is
|
||||
// started; the record says why, with the numbers, under a stable prefix.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT.md): drop the spaceFits call from backup() — this test fails at "a vzdump
|
||||
// was started on a target that cannot hold it".
|
||||
func TestR685_BackupThatCannotFitIsSkipped(t *testing.T) {
|
||||
api := spaceAPI(4*gib, 8182759056, "dir")
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
rec, err := r.Backup(context.Background(), 9201)
|
||||
if len(api.vzdumps) != 0 {
|
||||
t.Fatalf("a vzdump was started on a target that cannot hold it: %+v", api.vzdumps)
|
||||
}
|
||||
if err == nil || rec.Success || !strings.HasPrefix(rec.Error, BackupSkipNoSpacePrefix) {
|
||||
t.Fatalf("want a named skip, got err=%v rec=%+v", err, rec)
|
||||
}
|
||||
for _, want := range []string{"4.0 GiB free", "7.6 GiB", "10.5 GiB"} {
|
||||
if !strings.Contains(rec.Error, want) {
|
||||
t.Errorf("the reason must carry the numbers (%q missing): %s", want, rec.Error)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestR685_BackupThatFitsRuns — the same archive with 16 GiB free (demo-hp after tonight's prune) runs.
|
||||
func TestR685_BackupThatFitsRuns(t *testing.T) {
|
||||
api := spaceAPI(16*gib, 8182759056, "dir")
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
_, _ = r.Backup(context.Background(), 9201)
|
||||
if len(api.vzdumps) != 1 {
|
||||
t.Fatalf("a backup that fits must run, vzdumps=%d", len(api.vzdumps))
|
||||
}
|
||||
}
|
||||
|
||||
// TestR685_FailsOpen — never refuse on what is not KNOWN: a PBS target, a first backup (no archive to
|
||||
// size from), an unknown free figure, or a storage list that cannot be read.
|
||||
func TestR685_FailsOpen(t *testing.T) {
|
||||
cases := map[string]*fakeBackupAPI{
|
||||
"pbs target": spaceAPI(1*gib, 8*gib, "pbs"),
|
||||
"first backup": spaceAPI(1*gib, 0, "dir"),
|
||||
"avail unknown": spaceAPI(0, 8*gib, "dir"),
|
||||
"storage list error": func() *fakeBackupAPI {
|
||||
a := spaceAPI(1*gib, 8*gib, "dir")
|
||||
a.storageErr = context.DeadlineExceeded
|
||||
return a
|
||||
}(),
|
||||
}
|
||||
for name, api := range cases {
|
||||
target := "local"
|
||||
if name == "pbs target" {
|
||||
api.storages[0].Storage = "felhom-pbs"
|
||||
target = "felhom-pbs"
|
||||
}
|
||||
r := NewBackupRunner(api, target, proxmox.ModeSnapshot, "", "", quiet())
|
||||
_, _ = r.Backup(context.Background(), 9201)
|
||||
if len(api.vzdumps) != 1 {
|
||||
t.Errorf("%s: the preflight must fail OPEN, but no vzdump ran", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,112 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-689 (v0.135.0) — demo-hp keeps its golden template in `local:backup/`. It is content "backup",
|
||||
// 654 MB and so "plausibly complete", and it was the newest SETTLED entry: the restore test picked it
|
||||
// every 6 h and failed extractconfig with a 403 (measured 2026-09-24 10:36, 09-25 04:57 and 10:57),
|
||||
// while the guest's own archive — younger than the 24 h settle — went untested and nothing was proven.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT.md): drop the guestBackupArchive call from PickSettledRestoreCandidateOn —
|
||||
// this test then picks `local:backup/felhom-golden-0.236.0.tar.zst`.
|
||||
func TestR689_TheRestoreTestNeverPicksTheGolden(t *testing.T) {
|
||||
const day = int64(86400)
|
||||
now := int64(1790370000) // 2026-09-25 ~19:00Z
|
||||
api := &fakeBackupAPI{content: []proxmox.StorageContent{
|
||||
// the guest's real archive, settled (older than the cutoff below)
|
||||
{VolID: "local:backup/vzdump-lxc-9201-2026_09_22-21_59_25.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 3*day},
|
||||
// the golden: newer, settled, big, and NOT a backup of a guest
|
||||
{VolID: "local:backup/felhom-golden-0.236.0.tar.zst", Content: "backup", Size: 654115664, CTime: now - 2*day},
|
||||
// a hand-copied tarball that PVE happens to attribute to a vmid — the name is not a vzdump's
|
||||
{VolID: "local:backup/copy-of-9201.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 2*day},
|
||||
}}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got != "local:backup/vzdump-lxc-9201-2026_09_22-21_59_25.tar.zst" {
|
||||
t.Fatalf("picked %q — the restore test must prove a backup OF A GUEST", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestR689_GuestBackupArchiveShapes(t *testing.T) {
|
||||
for _, c := range []struct {
|
||||
e proxmox.StorageContent
|
||||
ok bool
|
||||
}{
|
||||
{proxmox.StorageContent{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst", VMID: 9201}, true},
|
||||
{proxmox.StorageContent{VolID: "local:backup/vzdump-qemu-300-2026_09_24-21_59_25.vma.zst", VMID: 300}, true},
|
||||
{proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z", VMID: 9201}, true},
|
||||
{proxmox.StorageContent{VolID: "felhom-pbs:backup/vm/300/2026-07-28T05:31:14Z", VMID: 300}, true},
|
||||
{proxmox.StorageContent{VolID: "local:backup/felhom-golden-0.236.0.tar.zst"}, false},
|
||||
{proxmox.StorageContent{VolID: "local:backup/vzdump-lxc-9201-x.tar.zst", VMID: 9202}, false}, // vmid disagrees with the name
|
||||
{proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-07-28T05:31:14Z"}, false}, // no vmid reported
|
||||
} {
|
||||
if ok, why := guestBackupArchive(c.e); ok != c.ok {
|
||||
t.Errorf("%s vmid=%d: ok=%v (%s), want %v", c.e.VolID, c.e.VMID, ok, why, c.ok)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// R-689 (v0.136.0) — the measured demo-hp shape right after v0.135.0: the golden (skipped), a leftover archive
|
||||
// of guest 9100 deleted in August (settled), and today's archive of 9201 (not settled yet). The pick must be
|
||||
// NOTHING — never the deleted guest's archive. With 9201's archive settled, that one.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT.md): drop the known-guest check — the pick is the 9100 leftover.
|
||||
func TestR689_AnArchiveOfADeletedGuestIsNeverPicked(t *testing.T) {
|
||||
const day = int64(86400)
|
||||
now := int64(1790476000)
|
||||
api := &fakeBackupAPI{goneGuests: map[int]bool{9100: true}, content: []proxmox.StorageContent{
|
||||
{VolID: "local:backup/felhom-golden-0.236.0.tar.zst", Content: "backup", Size: 654115664, CTime: now - 14*day},
|
||||
{VolID: "local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst", Content: "backup", VMID: 9100, Size: 656970239, CTime: now - 37*day},
|
||||
{VolID: "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 7*3600},
|
||||
}}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
|
||||
if err != nil || got != "" {
|
||||
t.Fatalf("picked %q err=%v — a deleted guest's archive proves nothing about this box", got, err)
|
||||
}
|
||||
got, _, _ = r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now, 0).UTC())
|
||||
if got != "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst" {
|
||||
t.Fatalf("with 9201's archive settled the pick is %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Any OTHER lookup failure is not "the guest is gone": the tier must read UNKNOWN (an error), never
|
||||
// "nothing to prove".
|
||||
func TestR689_AGuestLookupFailureIsUnknownNotEmpty(t *testing.T) {
|
||||
api := &fakeBackupAPI{cfgErr: fmt.Errorf("proxmox: connection refused"), content: []proxmox.StorageContent{
|
||||
{VolID: "local:backup/vzdump-lxc-9201-x.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: 10},
|
||||
}}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Time{}); err == nil {
|
||||
t.Fatal("a failed guest lookup read as a clean answer")
|
||||
}
|
||||
}
|
||||
|
||||
// v0.137.0 — THE MEASURED ANSWER: the agent's token sees only its pool, so for the deleted guest PVE says 403
|
||||
// "permission denied at /vms/9100", not "does not exist" (demo-hp, right after v0.136.0 — the local tier read
|
||||
// UNKNOWN). Such a guest is not one this agent manages: its archive is skipped, the tier is not an error.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT.md): drop the "permission denied" case — the pick errors.
|
||||
func TestR689_AGuestOutsideTheAgentsACLIsNotAKnownGuest(t *testing.T) {
|
||||
const day = int64(86400)
|
||||
now := int64(1790476000)
|
||||
api := &fakeBackupAPI{aclGuests: map[int]bool{9100: true}, content: []proxmox.StorageContent{
|
||||
{VolID: "local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst", Content: "backup", VMID: 9100, Size: 656970239, CTime: now - 37*day},
|
||||
{VolID: "local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst", Content: "backup", VMID: 9201, Size: 8 << 30, CTime: now - 7*3600},
|
||||
}}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Unix(now-day, 0).UTC())
|
||||
if err != nil || got != "" {
|
||||
t.Fatalf("picked %q err=%v — want nothing and no error (the only settled archive is not ours)", got, err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,74 @@
|
||||
package backup
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
const (
|
||||
thisBoxKey = "de:51:7a:18:cb:39:22:30:2c:84:f5:8b:d1:91:4b:7e:81:bb:69:b8:89:0f:57:ac:d3:59:e1:1a:62:25:11:2c"
|
||||
earlierBox1 = "6b:ca:5f:3f:ca:0f:e2:3f:fb:24:62:89:bf:e7:64:59:9a:41:c5:e6:e3:9f:3f:5f:e1:71:7b:a1:9d:24:67:82"
|
||||
earlierBox2 = "fe:3d:db:95:d4:df:ab:e1:7d:4a:89:fa:2b:07:53:6a:e4:d2:85:95:d1:90:27:4b:d9:c6:92:20:95:04:e5:d4"
|
||||
)
|
||||
|
||||
// R-727 (v0.138.0) — the 2026-09-30 shape, measured on a fresh box for a returning customer: the PBS
|
||||
// namespace held two archives of earlier boxes (same guest 9201, same token) and this box's own, which was not
|
||||
// settled yet. The old picker chose the earlier box's newest settled archive and failed `wrong key`.
|
||||
// The CONSEQUENCE asserted: no archive of another box is ever picked; with this box's archive settled it is picked.
|
||||
// COMPANION RED-PROOF: remove the `ownKey != "" && !EqualFold(...)` skip → the first case picks 2026-09-16T21:59:54Z.
|
||||
func TestR727_TheRestoreTestTakesOnlyThisBoxsArchives(t *testing.T) {
|
||||
day := int64(86400)
|
||||
now := int64(1790740000) // 2026-09-30 ~04:00Z
|
||||
own := proxmox.StorageContent{VolID: "felhom-pbs:backup/ct/9201/2026-09-29T19:37:07Z", Content: "backup", VMID: 9201, Size: 3490689830, CTime: 1790710627, Encrypted: thisBoxKey}
|
||||
api := &fakeBackupAPI{
|
||||
storages: []proxmox.Storage{{Storage: "felhom-pbs", Type: "pbs", EncryptionKey: thisBoxKey}},
|
||||
content: []proxmox.StorageContent{
|
||||
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T17:27:32Z", Content: "backup", VMID: 9201, Size: 4774114206, CTime: 1789579652, Encrypted: earlierBox2},
|
||||
{VolID: "felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z", Content: "backup", VMID: 9201, Size: 20811501236, CTime: 1789595994, Encrypted: earlierBox1},
|
||||
own,
|
||||
},
|
||||
}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
|
||||
// 1. The night of 2026-09-30: this box's own archive is ~6 h old, not settled (cutoff 24 h) — nothing to prove.
|
||||
got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Unix(now-day, 0).UTC())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got != "" {
|
||||
t.Fatalf("picked %q — an archive of ANOTHER box is never this box's proof (R-727)", got)
|
||||
}
|
||||
// 2. A day later this box's own archive is settled — it is the one picked.
|
||||
got, _, err = r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Unix(now+day, 0).UTC())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got != own.VolID {
|
||||
t.Fatalf("picked %q, want this box's own %q", got, own.VolID)
|
||||
}
|
||||
}
|
||||
|
||||
// An unencrypted storage (a local dir) holds only this box's vzdumps — no key filter applies.
|
||||
func TestR727_UnencryptedStorageIsNotFiltered(t *testing.T) {
|
||||
api := &fakeBackupAPI{
|
||||
storages: []proxmox.Storage{{Storage: "local", Type: "dir"}},
|
||||
content: []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_29-21_27_05.tar.zst", Content: "backup", VMID: 9201, Size: 955425507, CTime: 1790710025}},
|
||||
}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
if got, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "local", time.Time{}); err != nil || got == "" {
|
||||
t.Fatalf("got %q err %v", got, err)
|
||||
}
|
||||
}
|
||||
|
||||
// A storage-list failure makes the tier UNKNOWN (an error), never "nothing to prove".
|
||||
func TestR727_KeyLookupFailureIsUnknown(t *testing.T) {
|
||||
api := &fakeBackupAPI{storageErr: errors.New("proxmox: GET /storage -> HTTP 500"), content: []proxmox.StorageContent{{VolID: "felhom-pbs:backup/ct/9201/x", Content: "backup", VMID: 9201}}}
|
||||
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
|
||||
if _, _, err := r.PickSettledRestoreCandidateOn(context.Background(), "felhom-pbs", time.Time{}); err == nil {
|
||||
t.Fatal("a failed key lookup must surface as an error (tier UNKNOWN)")
|
||||
}
|
||||
}
|
||||
@@ -5,6 +5,7 @@ import (
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
@@ -22,6 +23,10 @@ type BackupAPI interface {
|
||||
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
|
||||
// ListStorage enumerates storages (name+type) — used to scope local-only retention (never prune PBS).
|
||||
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
|
||||
// NodeStorage is GET /nodes/{node}/storage — the storages WITH live usage (avail/used). R-685's space
|
||||
// preflight reads free space HERE: ListStorage (GET /storage) is the cluster DEFINITIONS and carries no
|
||||
// usage at all — measured live 2026-09-24 on demo-hp, where reading it let a backup through.
|
||||
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
|
||||
// TaskLogTail reads trailing task-log lines — used to read the ACTUAL vzdump mode
|
||||
// (PVE may downgrade a requested snapshot to stop for a stopped guest — spike B1).
|
||||
TaskLogTail(ctx context.Context, upid string, limit int) ([]string, error)
|
||||
@@ -170,6 +175,16 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
|
||||
rec.UncoveredVolumes = []string{}
|
||||
}
|
||||
|
||||
// R-685 (v0.134.0): will the new archive FIT on a local target? Asked before anything runs, so a
|
||||
// target that cannot hold it is a named SKIP with the numbers, not a nightly "No space left on
|
||||
// device" that only the vzdump log explains (demo-hp, every night from 2026-09-23 — R-684).
|
||||
if ok, why := r.spaceFits(ctx, vmid); !ok {
|
||||
rec.Error = BackupSkipNoSpacePrefix + why
|
||||
rec.DurationSeconds = time.Since(start).Seconds()
|
||||
r.logger.Warn("backup SKIPPED by the space preflight (R-685) — nothing was started", "vmid", vmid, "target", r.target, "reason", why)
|
||||
return rec, fmt.Errorf("backup: %s", rec.Error)
|
||||
}
|
||||
|
||||
upid, err := r.api.Vzdump(ctx, proxmox.VzdumpOptions{
|
||||
VMID: vmid, Storage: r.target, Mode: r.mode, Notes: r.notes,
|
||||
PruneBackups: r.localPruneSpec(ctx), // local target → keep-last=N; PBS/unknown → "" (no prune)
|
||||
@@ -217,6 +232,56 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
|
||||
return rec, nil
|
||||
}
|
||||
|
||||
// BackupSkipNoSpacePrefix starts a backup record's Error when the space preflight refused (R-685) — a
|
||||
// stable prefix the controller's page and the hub can key on.
|
||||
const BackupSkipNoSpacePrefix = "skipped: not enough space: "
|
||||
|
||||
// Space preflight margins (R-685): the new archive is predicted as the newest archive of this guest on the
|
||||
// target × backupSpaceGrowth, plus backupSpaceFloorBytes of headroom for the host. MEASURED 2026-09-24:
|
||||
// demo-hp 9201's archives grew 5.8 → 6.2 → 6.9 → 7.6 GB in four nights (+10 % a night at worst), so 1.25
|
||||
// covers two nights' growth. PVE prunes old archives only AFTER a successful backup, so the free space
|
||||
// must hold the new archive while every kept one still exists.
|
||||
const (
|
||||
backupSpaceGrowth = 1.25
|
||||
backupSpaceFloorBytes = int64(1) << 30
|
||||
)
|
||||
|
||||
// spaceFits answers whether a new archive of vmid fits on a LOCAL (non-PBS) target. It FAILS OPEN — a
|
||||
// backup is the thing being protected, so an unreadable storage, an unknown type or a first backup (no
|
||||
// previous archive to size from) proceeds and says so; only a POSITIVE "it does not fit" refuses.
|
||||
func (r *BackupRunner) spaceFits(ctx context.Context, vmid int) (bool, string) {
|
||||
// NodeStorage, never ListStorage: only the node view carries avail (see BackupAPI.NodeStorage).
|
||||
stores, err := r.api.NodeStorage(ctx)
|
||||
if err != nil {
|
||||
r.logger.Warn("backup: space preflight could not read storage usage — proceeding (fail-open)", "target", r.target, "err", err)
|
||||
return true, ""
|
||||
}
|
||||
var st *proxmox.Storage
|
||||
for i := range stores {
|
||||
if stores[i].Storage == r.target {
|
||||
st = &stores[i]
|
||||
break
|
||||
}
|
||||
}
|
||||
if st == nil || st.Type == "pbs" || st.Avail <= 0 {
|
||||
return true, "" // PBS dedups and has its own lifecycle; an unknown avail never refuses
|
||||
}
|
||||
_, last, err := r.latestArchive(ctx, vmid)
|
||||
if err != nil || last <= 0 {
|
||||
r.logger.Info("backup: space preflight has no previous archive to size from — proceeding", "vmid", vmid, "target", r.target)
|
||||
return true, ""
|
||||
}
|
||||
need := int64(float64(last)*backupSpaceGrowth) + backupSpaceFloorBytes
|
||||
if st.Avail >= need {
|
||||
r.logger.Info("backup: space preflight passed", "vmid", vmid, "target", r.target, "last_archive_bytes", last, "need_bytes", need, "avail_bytes", st.Avail)
|
||||
return true, ""
|
||||
}
|
||||
return false, fmt.Sprintf("%s has %s free; the last archive of guest %d was %s, so a new one needs about %s (old archives are removed only after a successful backup)",
|
||||
r.target, humanGiB(st.Avail), vmid, humanGiB(last), humanGiB(need))
|
||||
}
|
||||
|
||||
func humanGiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
|
||||
|
||||
// watchForSnapshot polls the running backup's task log until it sees the storage-snapshot marker
|
||||
// (→ onSnapshot once) or the requested mode is reported as `stop` (→ downgraded; the marker will
|
||||
// never come, so stop watching) or ctx is cancelled (backup finished). Best-effort: a log-read
|
||||
@@ -292,12 +357,58 @@ func (r *BackupRunner) PickSettledRestoreCandidateOn(ctx context.Context, target
|
||||
if err != nil {
|
||||
return "", time.Time{}, err
|
||||
}
|
||||
// R-727 (v0.138.0): on an ENCRYPTED storage, only archives written with THIS storage's key are this box's.
|
||||
// Measured 2026-09-30 on a fresh box for a returning customer: the PBS namespace still held two archives
|
||||
// of earlier boxes (same guest id 9201, same token), the newest settled one was an earlier box's, and the
|
||||
// test failed `wrong key` every evaluation. The archive carries no host id; its key fingerprint is the
|
||||
// discriminator (PVE's content `encrypted`, the storage's `encryption-key`). A lookup failure returns
|
||||
// the error — the tier reads UNKNOWN, never "nothing to prove".
|
||||
ownKey, err := r.storageKeyFingerprint(ctx, target)
|
||||
if err != nil {
|
||||
return "", time.Time{}, fmt.Errorf("reading the key fingerprint of storage %s: %w", target, err)
|
||||
}
|
||||
var best string
|
||||
var bestCTime int64 = -1
|
||||
known := map[int]bool{} // vmid → the guest exists on this node (asked once per vmid per pick)
|
||||
for _, e := range contents {
|
||||
if e.Content != "backup" {
|
||||
continue
|
||||
}
|
||||
// R-689 (v0.135.0): only a backup OF A GUEST is a restore-test candidate. demo-hp keeps its golden
|
||||
// template in `local:backup/` — content "backup", 654 MB, plausibly complete — and it was picked as
|
||||
// the newest settled archive every 6 h and failed extractconfig (403) each time, while the guest's
|
||||
// real archive went untested.
|
||||
if ok, why := guestBackupArchive(e); !ok {
|
||||
r.noteNotAGuestBackupOnce(e, why)
|
||||
continue
|
||||
}
|
||||
if ownKey != "" && !strings.EqualFold(e.Encrypted, ownKey) {
|
||||
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("written by another box (key %s, this box's key %s) — not this box's proof", shortFP(e.Encrypted), shortFP(ownKey)))
|
||||
continue
|
||||
}
|
||||
// R-689 (v0.136.0): … OF A GUEST THAT STILL EXISTS here. Measured on demo-hp 2026-09-27 right after
|
||||
// v0.135.0: with the golden skipped, the pick fell to `vzdump-lxc-9100-2026_08_21…`, a leftover of a
|
||||
// guest deleted in August — proving nothing about any guest this box runs. "Does not exist" skips the
|
||||
// archive; any OTHER lookup failure is returned, so the tier reads UNKNOWN, never "nothing to prove".
|
||||
if _, seen := known[e.VMID]; !seen {
|
||||
_, err := r.api.GuestConfig(ctx, e.VMID)
|
||||
switch {
|
||||
case err == nil:
|
||||
known[e.VMID] = true
|
||||
case strings.Contains(err.Error(), "does not exist"), strings.Contains(err.Error(), "permission denied"):
|
||||
// v0.137.0: PVE answers 403 "permission denied at /vms/<id>" — not "does not exist" — for a guest
|
||||
// outside the agent's ACL (the `felhom` pool). Measured on demo-hp after v0.136.0: the deleted
|
||||
// guest 9100's archive made the local tier UNKNOWN every evaluation. A guest the agent cannot
|
||||
// read is not one it manages; its archive is not a candidate.
|
||||
known[e.VMID] = false
|
||||
default:
|
||||
return "", time.Time{}, fmt.Errorf("checking whether guest %d still exists: %w", e.VMID, err)
|
||||
}
|
||||
}
|
||||
if !known[e.VMID] {
|
||||
r.noteNotAGuestBackupOnce(e, fmt.Sprintf("guest %d does not exist on this node or is not one this agent manages", e.VMID))
|
||||
continue
|
||||
}
|
||||
if !notAfter.IsZero() && e.CTime > notAfter.Unix() {
|
||||
continue // not settled yet — a newer archive is not a reason to re-prove an older one
|
||||
}
|
||||
@@ -520,5 +631,79 @@ func ToHubRestoreTest(res reconcile.RestoreTestResult, testedAt time.Time) hub.R
|
||||
if res.Err != nil {
|
||||
rt.Error = res.Err.Error()
|
||||
}
|
||||
if res.SkipReason != "" { // R-672: the space preflight refused — reported, never a pass
|
||||
rt.Pass = false
|
||||
rt.Skipped = true
|
||||
rt.Error = res.SkipReason
|
||||
}
|
||||
return rt
|
||||
}
|
||||
|
||||
// guestBackupArchive reports whether a storage entry is a whole-guest backup of a known guest — a
|
||||
// `vzdump-<type>-<vmid>-…` file on a dir storage, or a `backup/{ct,vm}/<vmid>/<time>` snapshot on a PBS
|
||||
// datastore — whose vmid the storage itself reports. Anything else in a backup content type (a golden
|
||||
// template, a hand-copied tarball) is not a backup of a guest and is never restore-tested (R-689).
|
||||
// Pure, so the rule is unit-tested without a storage.
|
||||
func guestBackupArchive(e proxmox.StorageContent) (bool, string) {
|
||||
if e.VMID <= 0 {
|
||||
return false, "not a backup of a guest (the storage reports no vmid)"
|
||||
}
|
||||
vol := e.VolID
|
||||
if i := strings.Index(vol, ":"); i >= 0 {
|
||||
vol = vol[i+1:]
|
||||
}
|
||||
vol = strings.TrimPrefix(vol, "backup/")
|
||||
vmid := strconv.Itoa(e.VMID)
|
||||
switch {
|
||||
case strings.HasPrefix(vol, "vzdump-lxc-"+vmid+"-"), strings.HasPrefix(vol, "vzdump-qemu-"+vmid+"-"):
|
||||
return true, ""
|
||||
case strings.HasPrefix(vol, "ct/"+vmid+"/"), strings.HasPrefix(vol, "vm/"+vmid+"/"):
|
||||
return true, ""
|
||||
}
|
||||
return false, "not a vzdump archive or a PBS snapshot of guest " + vmid
|
||||
}
|
||||
|
||||
// noteNotAGuestBackupOnce logs, once per volid, that a backup-content entry is not a restore-test
|
||||
// candidate because it is not a backup of a guest (R-689). INFO, not WARN: a golden template kept in
|
||||
// the backup directory is the operator's, and not a fault.
|
||||
func (r *BackupRunner) noteNotAGuestBackupOnce(e proxmox.StorageContent, why string) {
|
||||
r.rejectedMu.Lock()
|
||||
if r.rejected == nil {
|
||||
r.rejected = map[string]struct{}{}
|
||||
}
|
||||
_, seen := r.rejected[e.VolID]
|
||||
if !seen {
|
||||
r.rejected[e.VolID] = struct{}{}
|
||||
}
|
||||
r.rejectedMu.Unlock()
|
||||
if !seen {
|
||||
r.logger.Info("backup: restore-test skips an entry that is not a backup of a guest",
|
||||
"target", r.target, "volid", e.VolID, "size_bytes", e.Size, "reason", why)
|
||||
}
|
||||
}
|
||||
|
||||
// storageKeyFingerprint returns the named storage's client-side encryption key fingerprint ("" when the
|
||||
// storage is not encrypted — a local dir holds only this box's own vzdumps).
|
||||
func (r *BackupRunner) storageKeyFingerprint(ctx context.Context, target string) (string, error) {
|
||||
sts, err := r.api.ListStorage(ctx)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
for _, st := range sts {
|
||||
if st.Storage == target {
|
||||
return strings.TrimSpace(st.EncryptionKey), nil
|
||||
}
|
||||
}
|
||||
return "", nil
|
||||
}
|
||||
|
||||
// shortFP is the first 8 bytes of a key fingerprint, for a log line.
|
||||
func shortFP(fp string) string {
|
||||
if fp == "" {
|
||||
return "none"
|
||||
}
|
||||
if len(fp) > 23 {
|
||||
return fp[:23] + "…"
|
||||
}
|
||||
return fp
|
||||
}
|
||||
|
||||
@@ -208,9 +208,11 @@ func (s *Scheduler) tick(ctx context.Context) {
|
||||
spec := s.spec(ctx, archive)
|
||||
spec.Archive = archive
|
||||
res := s.runner.RunRestoreTest(ctx, spec)
|
||||
if res.Skipped {
|
||||
if res.Skipped && res.SkipReason == "" {
|
||||
return // already logged by the engine (no free scratch VMID)
|
||||
}
|
||||
// R-672: a SPACE refusal is the test's result — reported (pass=false, the reason as the error),
|
||||
// never dropped and never a pass. It earns no rotation credit, so the tier stays due.
|
||||
rt := ToHubRestoreTest(res, s.now())
|
||||
s.store.RecordRestoreTest(rt)
|
||||
// Rotation credit is given ONLY on success. A failing tier must keep sorting first, or a tier
|
||||
|
||||
@@ -123,6 +123,9 @@ var manifest = []Capability{
|
||||
{"dnsmasq-install", "dnsmasq package install", "/usr/bin/apt-get", []string{"install", "-y", "-q", "dnsmasq"}, false, ""},
|
||||
{"dnsmasq-write", "dnsmasq drop-in write", "/usr/bin/install", []string{"-m", "0644", "/tmp/felhom-resolver-x.conf", "/etc/dnsmasq.d/felhom-x.conf"}, false, ""},
|
||||
{"dnsmasq-enable", "dnsmasq enable", "/usr/bin/systemctl", []string{"enable", "--now", "dnsmasq"}, false, ""},
|
||||
|
||||
// ---- OS updates, guest fast lane (`11` §5.4.1; the wrapper holds every rule) ----
|
||||
{"osapply-run", "OS update wrapper (guest fast lane)", "/usr/local/sbin/felhom-os-apply", []string{"--plan", "/var/lib/felhom-agent/os/plan-x.json"}, false, ""},
|
||||
{"dnsmasq-reload", "dnsmasq reload", "/usr/bin/systemctl", []string{"reload", "dnsmasq"}, false, ""},
|
||||
{"dnsmasq-restart", "dnsmasq restart (LAN-DNS self-heal)", "/usr/bin/systemctl", []string{"restart", "dnsmasq"}, false, ""},
|
||||
{"dnsmasq-rm", "dnsmasq drop-in remove (decommission)", "/usr/bin/rm", []string{"-f", "/etc/dnsmasq.d/felhom-x.conf"}, false, ""},
|
||||
|
||||
@@ -345,6 +345,12 @@ type BackupConfig struct {
|
||||
// between an archive settling and its proof, and the retry rate of a tier whose restore-test
|
||||
// keeps failing. See defaultRestoreTestEvalInterval for the measurement it was chosen from.
|
||||
RestoreTestEvalIntervalSeconds int `json:"restore_test_eval_interval_seconds"`
|
||||
// RestoreTestSpaceFactor / RestoreTestSpaceReserveGiB are the restore-test's space margin (R-672,
|
||||
// v0.133.0): a test starts only when the target storage has free ≥ restored × factor + reserve,
|
||||
// `restored` being the UNCOMPRESSED size. 0/unset → 1.2 and 5 GiB. A test config may raise them to
|
||||
// watch the refusal (the brief's live case a).
|
||||
RestoreTestSpaceFactor float64 `json:"restore_test_space_factor,omitempty"`
|
||||
RestoreTestSpaceReserveGiB float64 `json:"restore_test_space_reserve_gib,omitempty"`
|
||||
// RestoreTestSettleSeconds is how long an archive must have sat on its tier before it is a
|
||||
// restore-test candidate (R-86); 0 → default (24h), negative → 0 (no settle requirement).
|
||||
// Restore-testing an archive a backup is still writing proves nothing about the backup that
|
||||
@@ -615,6 +621,20 @@ func (b BackupConfig) RestoreTestEvalInterval() time.Duration {
|
||||
}
|
||||
}
|
||||
|
||||
// RestoreTestSpace returns the restore-test's space margin (R-672): factor (≥ 1) and reserve bytes.
|
||||
// Unset or out-of-range → 1.2 and 5 GiB.
|
||||
func (b BackupConfig) RestoreTestSpace() (factor float64, reserveBytes int64) {
|
||||
factor = b.RestoreTestSpaceFactor
|
||||
if factor < 1 {
|
||||
factor = 1.2
|
||||
}
|
||||
reserveBytes = int64(b.RestoreTestSpaceReserveGiB * float64(1<<30))
|
||||
if reserveBytes <= 0 {
|
||||
reserveBytes = 5 << 30
|
||||
}
|
||||
return factor, reserveBytes
|
||||
}
|
||||
|
||||
// RestoreTestSettle returns how long an archive must have sat before it is a restore-test
|
||||
// candidate (R-86): a positive value as-is, negative → 0 (no settle requirement), 0 → the default.
|
||||
//
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
// Package httpx holds the one HTTP-transport default this repo may not lose.
|
||||
//
|
||||
// Every client here pins TLS — PBS and PVE by leaf-cert SHA-256, the hub by an optional CA file —
|
||||
// so none of them can use http.DefaultTransport and each hand-rolls its own. Hand-rolling silently
|
||||
// discards DefaultTransport's settings, and one of them is load-bearing:
|
||||
//
|
||||
// Transport: &http.Transport{TLSClientConfig: tlsCfg} // IdleConnTimeout == 0 == NO timeout
|
||||
//
|
||||
// A zero IdleConnTimeout means idle keep-alive connections are retained FOREVER, not "use a sane
|
||||
// default". Combined with a client that is rebuilt on a schedule and dropped (pbsTargetsFromPVE
|
||||
// builds a fresh pbs.Client per cycle), every cycle strands one connection that nothing will ever
|
||||
// close: the abandoned Transport becomes unreachable but its persistConn read-loop goroutine keeps
|
||||
// the socket alive, and an unreachable Transport does not close its connections.
|
||||
//
|
||||
// Measured cost, live: 388 established connections accumulated on ep0's PBS proxy between
|
||||
// 2026-08-18 09:51:22Z and 2026-08-20 08:02:13Z — 194 from each of the two boxes, held open on BOTH
|
||||
// sides, one per agent poll cycle, on a proxy whose descriptor ceiling is 65536. See R-344 and
|
||||
// felhom.eu/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md.
|
||||
package httpx
|
||||
|
||||
import (
|
||||
"crypto/tls"
|
||||
"net/http"
|
||||
"time"
|
||||
)
|
||||
|
||||
// DefaultIdleConnTimeout is how long an idle keep-alive connection is retained before it is closed.
|
||||
//
|
||||
// It is 90s because that is http.DefaultTransport's own value: the fix for R-344 restores a
|
||||
// standard-library default rather than inventing a number, so there is nothing here to tune and
|
||||
// nothing to justify. It is comfortably shorter than every cadence that drives these clients (the
|
||||
// 15-minute live-snapshot collect and the 6-hour verify loop), so a connection abandoned by one
|
||||
// cycle is closed long before the next.
|
||||
const DefaultIdleConnTimeout = 90 * time.Second
|
||||
|
||||
// NewTransport builds a FRESH *http.Transport pinned to tlsCfg, with the idle-connection timeout
|
||||
// applied.
|
||||
//
|
||||
// Fresh, never shared: each caller pins a different endpoint, and a shared transport would pool
|
||||
// connections across differently pinned servers. Reusing http.DefaultTransport for the same reason
|
||||
// is not an option — it would drop the pin entirely.
|
||||
//
|
||||
// idleConnTimeout <= 0 means USE THE DEFAULT. It deliberately does not mean "no timeout": no-timeout
|
||||
// is the bug this package exists to prevent, and an unset field must never be able to reintroduce
|
||||
// it. Callers pass their configured value straight through; only tests pass a short one.
|
||||
//
|
||||
// Only IdleConnTimeout is set. The other DefaultTransport settings this transport also lacks
|
||||
// (MaxIdleConns, TLSHandshakeTimeout, ExpectContinueTimeout) are deliberately left alone: none of
|
||||
// them accumulates anything, every client bounds its whole request with http.Client.Timeout, and
|
||||
// widening the change would have made the R-344 measurement unattributable.
|
||||
func NewTransport(tlsCfg *tls.Config, idleConnTimeout time.Duration) *http.Transport {
|
||||
if idleConnTimeout <= 0 {
|
||||
idleConnTimeout = DefaultIdleConnTimeout
|
||||
}
|
||||
return &http.Transport{
|
||||
TLSClientConfig: tlsCfg,
|
||||
IdleConnTimeout: idleConnTimeout,
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,74 @@
|
||||
package httpx
|
||||
|
||||
import (
|
||||
"crypto/tls"
|
||||
"net/http"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// TestNewTransport_ZeroMeansDefaultNeverForever is the whole point of this package.
|
||||
//
|
||||
// http.Transport's zero IdleConnTimeout means "retain idle connections FOREVER". Any code path that
|
||||
// can reach that zero reintroduces R-344, so an unset, zero or negative value must all land on the
|
||||
// default. If someone later "simplifies" NewTransport by passing the argument straight through,
|
||||
// this fails.
|
||||
func TestNewTransport_ZeroMeansDefaultNeverForever(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
in time.Duration
|
||||
want time.Duration
|
||||
}{
|
||||
{"zero", 0, DefaultIdleConnTimeout},
|
||||
{"negative", -time.Hour, DefaultIdleConnTimeout},
|
||||
{"explicit short value (tests)", 50 * time.Millisecond, 50 * time.Millisecond},
|
||||
{"explicit long value", time.Hour, time.Hour},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
got := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, tc.in).IdleConnTimeout
|
||||
if got != tc.want {
|
||||
t.Fatalf("IdleConnTimeout = %v, want %v", got, tc.want)
|
||||
}
|
||||
if got == 0 {
|
||||
t.Fatal("IdleConnTimeout is 0 — that is 'never expire', which is the R-344 defect itself")
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestDefaultIdleConnTimeout_MatchesTheStandardLibrary pins the number to its justification.
|
||||
//
|
||||
// 90s is not a tuned value; it is what http.DefaultTransport uses. Reading it off the standard
|
||||
// library rather than hardcoding 90 means the constant cannot drift away from the reason given for
|
||||
// it in the package doc.
|
||||
func TestDefaultIdleConnTimeout_MatchesTheStandardLibrary(t *testing.T) {
|
||||
std, ok := http.DefaultTransport.(*http.Transport)
|
||||
if !ok {
|
||||
t.Skip("http.DefaultTransport is not an *http.Transport in this Go build")
|
||||
}
|
||||
if DefaultIdleConnTimeout != std.IdleConnTimeout {
|
||||
t.Fatalf("DefaultIdleConnTimeout = %v but http.DefaultTransport uses %v — the doc comment's justification no longer holds",
|
||||
DefaultIdleConnTimeout, std.IdleConnTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewTransport_IsFreshEveryCall guards the pooling property the pinning relies on.
|
||||
//
|
||||
// Each caller pins a DIFFERENT endpoint. A shared transport would pool connections across
|
||||
// differently pinned servers, so returning a package-level singleton would be a security change
|
||||
// dressed as a tidy-up.
|
||||
func TestNewTransport_IsFreshEveryCall(t *testing.T) {
|
||||
a := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
|
||||
b := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
|
||||
if a == b {
|
||||
t.Fatal("NewTransport returned the SAME transport twice — connections would be pooled across differently pinned endpoints")
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewTransport_KeepsTheTLSConfig — the transport gains a field; it must lose nothing.
|
||||
func TestNewTransport_KeepsTheTLSConfig(t *testing.T) {
|
||||
cfg := &tls.Config{MinVersion: tls.VersionTLS12, InsecureSkipVerify: true} //nolint:gosec // test only
|
||||
if got := NewTransport(cfg, 0).TLSClientConfig; got != cfg {
|
||||
t.Fatalf("TLSClientConfig = %p, want the config passed in (%p) — the pin would be dropped", got, cfg)
|
||||
}
|
||||
}
|
||||
+31
-2
@@ -15,6 +15,8 @@ import (
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||
)
|
||||
|
||||
const reportPath = "/api/v1/host-report"
|
||||
@@ -49,8 +51,11 @@ func NewClient(cfg config.HubConfig, logger *slog.Logger) (*Client, error) {
|
||||
tlsCfg.RootCAs = pool
|
||||
}
|
||||
hc := &http.Client{
|
||||
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
|
||||
Transport: &http.Transport{TLSClientConfig: tlsCfg},
|
||||
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
|
||||
// R-344, consistency only: this client is built ONCE per process, so it never accumulated
|
||||
// and contributed nothing to the ep0 leak. It carried the same missing default, which over
|
||||
// a tunnel is how one idle connection survives long enough to fail on next use.
|
||||
Transport: httpx.NewTransport(tlsCfg, 0),
|
||||
}
|
||||
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
|
||||
}
|
||||
@@ -420,3 +425,27 @@ func (c *Client) FetchRetainedIdentityEscrow(ctx context.Context) (*RetainedEscr
|
||||
}
|
||||
return &out, nil
|
||||
}
|
||||
|
||||
// PostOSReport sends the OS-update leg's report after every run (hub v0.130.0): POST /api/v1/hosts/{id}/os-report.
|
||||
// Per-host key, self-scoped on the hub. Errors are typed like RegisterWG's and never include the bearer.
|
||||
func (c *Client) PostOSReport(ctx context.Context, body []byte) error {
|
||||
if c.hostID == "" {
|
||||
return fmt.Errorf("hub: PostOSReport requires a configured host_id")
|
||||
}
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL+"/api/v1/hosts/"+c.hostID+"/os-report", bytes.NewReader(body))
|
||||
if err != nil {
|
||||
return fmt.Errorf("hub: building os-report request: %w", err)
|
||||
}
|
||||
req.Header.Set("Authorization", "Bearer "+c.apiKey)
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
resp, err := c.hc.Do(req)
|
||||
if err != nil {
|
||||
return &TransportError{Err: err}
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
|
||||
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
|
||||
return &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -103,6 +103,7 @@ type Collector struct {
|
||||
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
|
||||
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
|
||||
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
|
||||
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
|
||||
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
|
||||
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
||||
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
||||
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
|
||||
return c
|
||||
}
|
||||
|
||||
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
|
||||
type ControllerSupervisorReporter interface {
|
||||
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
|
||||
}
|
||||
|
||||
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
|
||||
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
|
||||
c.ctrlSup = r
|
||||
return c
|
||||
}
|
||||
|
||||
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
|
||||
// omitted). Returns the collector for chaining.
|
||||
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
|
||||
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
||||
if c.pbsdr != nil {
|
||||
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
|
||||
}
|
||||
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
|
||||
if c.ctrlSup != nil {
|
||||
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
|
||||
}
|
||||
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
|
||||
if c.guestNet != nil {
|
||||
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
package hub
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The os_update block is a cross-repo contract: testdata/desired-state-osupdate.golden.json is byte-identical with
|
||||
// felhom.eu/hub/internal/api/testdata (the hub's TestOSUpdate_DesiredBlockMatchesTheGolden proves the hub SERVES
|
||||
// it). Here: the agent DECODES every field. A renamed json tag on either side fails one of the two tests.
|
||||
func TestOSUpdateGolden_Decodes(t *testing.T) {
|
||||
raw, err := os.ReadFile("testdata/desired-state-osupdate.golden.json")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var resp DesiredStateResponse
|
||||
if err := json.Unmarshal(raw, &resp); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
o := resp.DesiredState.OSUpdate
|
||||
if o == nil || o.Ring != 1 || !o.Enabled || o.Release == nil {
|
||||
t.Fatalf("os_update = %+v", o)
|
||||
}
|
||||
r := o.Release
|
||||
if r.ID != "os-20261004-120000" || r.Snapshot != "20261004T120000Z" || len(r.Packages) != 2 ||
|
||||
r.Packages[1].Name != "openssl" || r.Packages[1].Version != "3.5.7-1~deb13u3" || r.Packages[1].Origin != "Debian-Security" {
|
||||
t.Fatalf("release = %+v", r)
|
||||
}
|
||||
}
|
||||
@@ -124,6 +124,39 @@ type HostReport struct {
|
||||
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
|
||||
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
|
||||
OOB *OOBStatus `json:"oob,omitempty"`
|
||||
|
||||
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
|
||||
// how many times the agent restarted a dead controller, when last and why, whether it gave up
|
||||
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
|
||||
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
|
||||
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
|
||||
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
|
||||
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
|
||||
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
|
||||
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
|
||||
}
|
||||
|
||||
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
|
||||
type ControllerSupervisorStatus struct {
|
||||
Guests []ControllerSupervisorGuest `json:"guests"`
|
||||
}
|
||||
|
||||
// ControllerSupervisorGuest is one supervised guest.
|
||||
type ControllerSupervisorGuest struct {
|
||||
VMID int `json:"vmid"`
|
||||
RestartsTotal int `json:"restarts_total"`
|
||||
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
|
||||
LastReason string `json:"last_reason,omitempty"`
|
||||
Crashloop bool `json:"crashloop"`
|
||||
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
|
||||
Parked bool `json:"parked"`
|
||||
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
|
||||
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
|
||||
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
|
||||
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
|
||||
Restarts24h int `json:"restarts_24h"`
|
||||
SlowCrashloop bool `json:"slow_crashloop"`
|
||||
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
|
||||
}
|
||||
|
||||
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
|
||||
@@ -437,6 +470,10 @@ type RestoreTest struct {
|
||||
// mount layout, not just booted. Additive — a hub that predates them ignores the unknown keys.
|
||||
MountParity string `json:"mount_parity,omitempty"`
|
||||
MountInventory []string `json:"mount_inventory,omitempty"`
|
||||
// Skipped (R-672, v0.133.0): the test did NOT run — the space preflight refused, and Error says
|
||||
// why ("skipped: not enough space on …"). Pass is false. A hub that predates the key reads a
|
||||
// failed test with that error, which is the honest reading.
|
||||
Skipped bool `json:"skipped,omitempty"`
|
||||
}
|
||||
|
||||
// PBSSnapshot is one PBS (offsite) snapshot's inventory + integrity state (doc 03 §8, slice
|
||||
@@ -511,6 +548,32 @@ type WireDesiredState struct {
|
||||
RestoreDirective *WireRestoreDirective `json:"restore_directive,omitempty"` // slice 10D (forward-compat)
|
||||
Wireguard *WireWireguard `json:"wireguard,omitempty"` // S3 (doc 06 §3.2; golden-pinned)
|
||||
PBSDR *WirePBSDR `json:"pbs_dr,omitempty"` // PBS DR tier (slice 2 consumer)
|
||||
OSUpdate *WireOSUpdate `json:"os_update,omitempty"` // OS updates, guest fast lane (agent v0.140.0)
|
||||
}
|
||||
|
||||
// WireOSUpdate is the hub-OWNED OS-update block (hub v0.130.0, `11-os-updates.md` §5.3), merged into the served
|
||||
// document at read time. Ring 0 installs every pending Debian / Debian-Security fix; ring 1 installs exactly the
|
||||
// newest approved release. Absent (older hub) → the agent treats the box as ring 1, ON, no release: it reports
|
||||
// and installs nothing. Golden: testdata/desired-state-osupdate.golden.json (byte-identical with the hub's).
|
||||
type WireOSUpdate struct {
|
||||
Ring int `json:"ring"`
|
||||
Enabled bool `json:"enabled"`
|
||||
Release *WireOSRelease `json:"release,omitempty"`
|
||||
}
|
||||
|
||||
// WireOSRelease is an approved version set; Snapshot is the approval time (YYYYMMDDTHHMMSSZ) the wrapper uses
|
||||
// for snapshot.debian.org when Debian has already replaced a version (decision 79).
|
||||
type WireOSRelease struct {
|
||||
ID string `json:"id"`
|
||||
Snapshot string `json:"snapshot"`
|
||||
Packages []WireOSPackage `json:"packages"`
|
||||
}
|
||||
|
||||
// WireOSPackage is one approved name=version and its origin ("Debian" | "Debian-Security").
|
||||
type WireOSPackage struct {
|
||||
Name string `json:"name"`
|
||||
Version string `json:"version"`
|
||||
Origin string `json:"origin"`
|
||||
}
|
||||
|
||||
// WirePBSDR is the hub's PBS-DR-tier descriptor (PBS DR slice 1, hub/internal/web/pbsdr.go
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"generation": 1,
|
||||
"desired_state": {
|
||||
"os_update": {
|
||||
"ring": 1,
|
||||
"enabled": true,
|
||||
"release": {
|
||||
"id": "os-20261004-120000",
|
||||
"snapshot": "20261004T120000Z",
|
||||
"packages": [
|
||||
{"name": "libc6", "version": "2.41-12+deb13u4", "origin": "Debian"},
|
||||
{"name": "openssl", "version": "3.5.7-1~deb13u3", "origin": "Debian-Security"}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||
)
|
||||
|
||||
// The OS leg (agent v0.140.0) runs after a SUCCESSFUL primary backup, and only then; and it runs BEFORE the
|
||||
// host-wide heavy-op gate is released, so a restore-test cannot start in the middle of it (`11` C10).
|
||||
// Red-proof: drop the `b.Success &&` guard and the failed-backup sub-case fails; move the call after release()
|
||||
// and the gate sub-case fails.
|
||||
func TestAfterPrimaryBackup(t *testing.T) {
|
||||
run := func(t *testing.T, failErr string) (calls []int, gateHeld bool) {
|
||||
gate := &backup.InFlight{}
|
||||
b := &fakeBackups{failErr: failErr}
|
||||
srv := newTestServerS(t, &fakeGuests{}, b, &fakeStore{}, nil)
|
||||
srv.inFlight = gate
|
||||
var mu sync.Mutex
|
||||
done := make(chan struct{}, 1)
|
||||
srv.SetAfterPrimaryBackup(func(_ context.Context, vmid int) {
|
||||
rel, _, ok := gate.TryAcquire("probe")
|
||||
mu.Lock()
|
||||
calls = append(calls, vmid)
|
||||
gateHeld = !ok
|
||||
mu.Unlock()
|
||||
if ok {
|
||||
rel()
|
||||
}
|
||||
done <- struct{}{}
|
||||
})
|
||||
h := srv.Handler()
|
||||
if do(t, h, "POST", "/backup", "A", "").Code != http.StatusAccepted {
|
||||
t.Fatal("POST /backup not accepted")
|
||||
}
|
||||
select {
|
||||
case <-done:
|
||||
case <-time.After(500 * time.Millisecond):
|
||||
}
|
||||
time.Sleep(20 * time.Millisecond)
|
||||
mu.Lock()
|
||||
defer mu.Unlock()
|
||||
return calls, gateHeld
|
||||
}
|
||||
t.Run("success runs the leg under the gate", func(t *testing.T) {
|
||||
calls, held := run(t, "")
|
||||
if len(calls) != 1 {
|
||||
t.Fatalf("the leg ran %d time(s), want 1", len(calls))
|
||||
}
|
||||
if !held {
|
||||
t.Fatal("the heavy-op gate was free while the leg ran — a restore-test could overlap it")
|
||||
}
|
||||
})
|
||||
t.Run("a failed backup runs nothing", func(t *testing.T) {
|
||||
if calls, _ := run(t, "vzdump exploded"); len(calls) != 0 {
|
||||
t.Fatalf("the leg ran after a FAILED backup: %v", calls)
|
||||
}
|
||||
})
|
||||
}
|
||||
@@ -0,0 +1,133 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"log/slog"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
|
||||
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
|
||||
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
|
||||
|
||||
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
|
||||
t.Helper()
|
||||
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
|
||||
var resp struct {
|
||||
Data BackupStatusResponse `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||
t.Fatalf("decode: %v (%s)", err, w.Body.String())
|
||||
}
|
||||
return resp.Data.Tiers
|
||||
}
|
||||
|
||||
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
|
||||
t.Helper()
|
||||
srv, err := NewServer(Options{
|
||||
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
|
||||
Storage: fakeStorage{targets: targets},
|
||||
Tokens: staticTokens{"A": 8200},
|
||||
BackupTiers: []BackupTier{
|
||||
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
|
||||
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
|
||||
},
|
||||
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
srv.baseCtx = context.Background()
|
||||
srv.now = func() time.Time { return testNow }
|
||||
return srv
|
||||
}
|
||||
|
||||
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
|
||||
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
|
||||
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
|
||||
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
|
||||
st := &fakeStore{backups: []hub.Backup{
|
||||
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
|
||||
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
|
||||
}}
|
||||
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
|
||||
tiers := tierStatesOf(t, srv)
|
||||
if len(tiers) != 2 {
|
||||
t.Fatalf("want 2 tiers, got %+v", tiers)
|
||||
}
|
||||
local, pbs := tiers[0], tiers[1]
|
||||
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
|
||||
t.Fatalf("local tier lost its successful backup: %+v", local)
|
||||
}
|
||||
if pbs.LastSuccess != nil {
|
||||
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
|
||||
}
|
||||
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
|
||||
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
|
||||
}
|
||||
if pbs.Storage != StoragePresenceAbsent {
|
||||
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
|
||||
}
|
||||
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
|
||||
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
|
||||
var resp struct {
|
||||
Data BackupStatusResponse `json:"data"`
|
||||
}
|
||||
_ = json.Unmarshal(w.Body.Bytes(), &resp)
|
||||
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
|
||||
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
|
||||
}
|
||||
}
|
||||
|
||||
// /backup/tiers advertises storage presence, tri-state.
|
||||
func TestBackupTiers_StoragePresence(t *testing.T) {
|
||||
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
|
||||
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
|
||||
var resp struct {
|
||||
Data BackupTiersResponse `json:"data"`
|
||||
}
|
||||
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
got := map[string]string{}
|
||||
for _, ti := range resp.Data.Tiers {
|
||||
got[ti.Target] = ti.Storage
|
||||
}
|
||||
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
|
||||
t.Fatalf("storage presence wrong: %v", got)
|
||||
}
|
||||
// An unreadable storage view is "unknown", never "absent".
|
||||
srv.storage = tierErrStorage{}
|
||||
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
|
||||
t.Fatalf("unreadable storage view must be unknown, got %q", p)
|
||||
}
|
||||
}
|
||||
|
||||
type tierErrStorage struct{}
|
||||
|
||||
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
|
||||
return nil, context.DeadlineExceeded
|
||||
}
|
||||
|
||||
// A targeted request keeps the pre-R-517 bytes (no tiers array).
|
||||
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
|
||||
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
|
||||
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
|
||||
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
|
||||
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
|
||||
}
|
||||
}
|
||||
|
||||
func containsKey(b []byte, key string) bool {
|
||||
var m struct {
|
||||
Data map[string]json.RawMessage `json:"data"`
|
||||
}
|
||||
_ = json.Unmarshal(b, &m)
|
||||
_, ok := m.Data[key]
|
||||
return ok
|
||||
}
|
||||
@@ -0,0 +1,453 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// R-523 — the in-guest controller supervisor.
|
||||
//
|
||||
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
|
||||
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
|
||||
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
|
||||
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
|
||||
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
|
||||
// dashboard answered 502 for 33 minutes until the box was power-cycled.
|
||||
//
|
||||
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
|
||||
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
|
||||
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
|
||||
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
|
||||
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
|
||||
// No new privilege.
|
||||
//
|
||||
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
|
||||
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
|
||||
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
|
||||
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
|
||||
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
|
||||
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
|
||||
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
|
||||
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
|
||||
// again after 30 minutes.
|
||||
//
|
||||
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
|
||||
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
|
||||
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
|
||||
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
|
||||
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
|
||||
// record) can never read as a new restart.
|
||||
|
||||
const (
|
||||
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
|
||||
// so a killed controller is restarted 30–60 s after it died.
|
||||
controllerSupervisorInterval = 30 * time.Second
|
||||
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
|
||||
controllerSupervisorConfirm = 2
|
||||
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
|
||||
// controllerCrashloopPause.
|
||||
controllerCrashloopMax = 3
|
||||
controllerCrashloopWindow = 15 * time.Minute
|
||||
controllerCrashloopPause = 30 * time.Minute
|
||||
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
|
||||
// watchdog is indistinguishable from a dead one (standing rule 3).
|
||||
controllerSupervisorHeartbeatEvery = 20
|
||||
|
||||
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
|
||||
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
|
||||
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
|
||||
// the controller container, and a customer inside the guest cannot park the supervisor.
|
||||
ControllerParkedMarker = "controller-parked"
|
||||
|
||||
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
|
||||
|
||||
// R-539 (operator ruling 3 of 2026-09-16) — the SLOW crash loop. The 3-in-15-minutes brake above
|
||||
// cannot see a controller that dies every 20 minutes: no two restarts share its window, so it is
|
||||
// restarted for ever and the only trace is an info event that mails nobody (measured 2026-09-16,
|
||||
// R-531). A second counter over 24 hours raises a WARNING at the fifth restart. It does NOT stop
|
||||
// restarting — the fast brake stays the only brake, unchanged. Every restart the supervisor
|
||||
// performs counts, including one that follows a deliberate operator `docker kill` (measured
|
||||
// 2026-09-15: the supervisor cannot tell a kill from a crash, and a controller that is killed five
|
||||
// times a day is worth a line to the operator either way).
|
||||
controllerSlowCrashloopWindow = 24 * time.Hour
|
||||
controllerSlowCrashloopMax = 5
|
||||
// controllerSlowCounterFile holds the 24-hour restart times and the last raise, per guest, beside
|
||||
// the parked marker.
|
||||
controllerSlowCounterFile = "controller-restarts-24h.json"
|
||||
)
|
||||
|
||||
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
|
||||
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
|
||||
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
|
||||
type controllerSupState struct {
|
||||
notRunningSeen int
|
||||
restarts []time.Time // restart times inside the crash-loop window (pruned)
|
||||
restartsTotal int
|
||||
lastRestartAt time.Time
|
||||
lastReason string
|
||||
crashloopSince time.Time // zero = not in a crash-loop pause
|
||||
parked bool
|
||||
|
||||
// R-539 — the slow counter. PERSISTED, unlike everything above, and the precedent's reason does not
|
||||
// apply to it: persisting the fast record could carry a stale "give up" across the restart that
|
||||
// fixed it, but this record never gives anything up — it only warns. Losing it on an agent restart,
|
||||
// on the other hand, would hide exactly the box it exists for (one whose agent restarts too).
|
||||
restarts24h []time.Time
|
||||
slowCrashloopSince time.Time // the last raise; kept after it ages out, the hub keys on it MOVING
|
||||
}
|
||||
|
||||
type controllerSupervisor struct {
|
||||
mu sync.Mutex
|
||||
guests map[int]*controllerSupState
|
||||
sweeps int
|
||||
}
|
||||
|
||||
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
|
||||
// (staleLock) or the guest executor is not wired.
|
||||
func (s *Server) WatchControllers(ctx context.Context) {
|
||||
if s.staleLock == nil || s.guestExec == nil {
|
||||
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
|
||||
return
|
||||
}
|
||||
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
|
||||
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
|
||||
"crashloop_window", controllerCrashloopWindow.String(),
|
||||
"slow_crashloop_max", controllerSlowCrashloopMax, "slow_crashloop_window", controllerSlowCrashloopWindow.String(),
|
||||
"guests_dir", s.guestsStateDir())
|
||||
t := time.NewTicker(controllerSupervisorInterval)
|
||||
defer t.Stop()
|
||||
for {
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-t.C:
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (s *Server) guestsStateDir() string {
|
||||
if s.guestsDir != "" {
|
||||
return s.guestsDir
|
||||
}
|
||||
return defaultGuestsStateDir
|
||||
}
|
||||
|
||||
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
|
||||
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
|
||||
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
|
||||
// not stat the file within.
|
||||
func (s *Server) provisionedGuest(vmid int) bool {
|
||||
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
|
||||
return err == nil && fi.IsDir()
|
||||
}
|
||||
|
||||
func (s *Server) controllerParked(vmid int) bool {
|
||||
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
|
||||
return err == nil
|
||||
}
|
||||
|
||||
func (s *Server) supState(vmid int) *controllerSupState {
|
||||
if s.ctrlSup.guests == nil {
|
||||
s.ctrlSup.guests = map[int]*controllerSupState{}
|
||||
}
|
||||
st := s.ctrlSup.guests[vmid]
|
||||
if st == nil {
|
||||
st = &controllerSupState{}
|
||||
s.loadSlowCounter(vmid, st)
|
||||
s.ctrlSup.guests[vmid] = st
|
||||
}
|
||||
return st
|
||||
}
|
||||
|
||||
// slowCounterRecord is the on-disk shape of the R-539 counter.
|
||||
type slowCounterRecord struct {
|
||||
Restarts []time.Time `json:"restarts"`
|
||||
SlowCrashloopSince time.Time `json:"slow_crashloop_since,omitempty"`
|
||||
}
|
||||
|
||||
func (s *Server) slowCounterPath(vmid int) string {
|
||||
return filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), controllerSlowCounterFile)
|
||||
}
|
||||
|
||||
// loadSlowCounter restores the persisted counter into a fresh state. Absent = a clean start; unreadable
|
||||
// or corrupt = a clean start with a WARN (a warning counter must never block supervision).
|
||||
func (s *Server) loadSlowCounter(vmid int, st *controllerSupState) {
|
||||
b, err := os.ReadFile(s.slowCounterPath(vmid))
|
||||
if err != nil {
|
||||
if !os.IsNotExist(err) {
|
||||
s.logger.Warn("controller-supervisor: slow counter unreadable — starting it from zero", "vmid", vmid, "err", err)
|
||||
}
|
||||
return
|
||||
}
|
||||
var rec slowCounterRecord
|
||||
if err := json.Unmarshal(b, &rec); err != nil {
|
||||
s.logger.Warn("controller-supervisor: slow counter corrupt — starting it from zero", "vmid", vmid, "err", err)
|
||||
return
|
||||
}
|
||||
st.restarts24h = pruneBefore(rec.Restarts, s.clock().Add(-controllerSlowCrashloopWindow))
|
||||
st.slowCrashloopSince = rec.SlowCrashloopSince
|
||||
if len(st.restarts24h) > 0 || !st.slowCrashloopSince.IsZero() {
|
||||
s.logger.Info("controller-supervisor: slow counter restored from disk", "vmid", vmid,
|
||||
"restarts_24h", len(st.restarts24h), "slow_crashloop_since", st.slowCrashloopSince.Format(time.RFC3339))
|
||||
}
|
||||
}
|
||||
|
||||
// saveSlowCounter writes the counter atomically (tmp + rename, 0600). A failure is logged and the
|
||||
// in-memory counter carries on — the next restart retries the write.
|
||||
func (s *Server) saveSlowCounter(vmid int, rec slowCounterRecord) {
|
||||
path := s.slowCounterPath(vmid)
|
||||
b, err := json.Marshal(rec)
|
||||
if err == nil {
|
||||
tmp := path + ".tmp"
|
||||
if err = os.WriteFile(tmp, b, 0o600); err == nil {
|
||||
err = os.Rename(tmp, path)
|
||||
}
|
||||
}
|
||||
if err != nil {
|
||||
s.logger.Warn("controller-supervisor: could not persist the slow counter (kept in memory)", "vmid", vmid, "path", path, "err", err)
|
||||
}
|
||||
}
|
||||
|
||||
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
|
||||
// cycle without waiting on the ticker.
|
||||
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
|
||||
if s.staleLock == nil || s.guestExec == nil {
|
||||
return
|
||||
}
|
||||
guests, err := s.staleLock.Guests(ctx)
|
||||
if err != nil {
|
||||
// Ownership unproven ⇒ touch nothing (the guest-power rule).
|
||||
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
|
||||
return
|
||||
}
|
||||
var evaluated, down int
|
||||
for _, g := range guests {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
if !s.provisionedGuest(g.VMID) {
|
||||
continue
|
||||
}
|
||||
evaluated++
|
||||
if !s.superviseOneController(ctx, g.VMID, g.Status) {
|
||||
down++
|
||||
}
|
||||
}
|
||||
s.ctrlSup.mu.Lock()
|
||||
s.ctrlSup.sweeps++
|
||||
sweeps := s.ctrlSup.sweeps
|
||||
s.ctrlSup.mu.Unlock()
|
||||
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
|
||||
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
|
||||
"guests_evaluated", evaluated, "controllers_not_running", down)
|
||||
}
|
||||
}
|
||||
|
||||
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
|
||||
// known=false means the question could not be answered (pct exec failed for a reason other than a
|
||||
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
|
||||
// running": `docker rm` of the controller is the same outage as a kill.
|
||||
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
|
||||
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
|
||||
if err != nil {
|
||||
msg := strings.ToLower(err.Error() + " " + out)
|
||||
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
|
||||
return false, true, "absent"
|
||||
}
|
||||
return false, false, ""
|
||||
}
|
||||
status = strings.TrimSpace(out)
|
||||
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
|
||||
return status == "running" || status == "restarting", true, status
|
||||
}
|
||||
|
||||
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
|
||||
// allows. Returns false when the controller was observed not running.
|
||||
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
|
||||
now := s.clock()
|
||||
if guestStatus != "running" {
|
||||
s.resetNotRunning(vmid)
|
||||
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
|
||||
}
|
||||
|
||||
running, known, status := s.controllerRunning(ctx, vmid)
|
||||
if !known {
|
||||
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
|
||||
s.resetNotRunning(vmid)
|
||||
return true
|
||||
}
|
||||
parked := s.controllerParked(vmid)
|
||||
s.ctrlSup.mu.Lock()
|
||||
st := s.supState(vmid)
|
||||
st.parked = parked
|
||||
if running {
|
||||
st.notRunningSeen = 0
|
||||
s.ctrlSup.mu.Unlock()
|
||||
return true
|
||||
}
|
||||
st.notRunningSeen++
|
||||
seen := st.notRunningSeen
|
||||
s.ctrlSup.mu.Unlock()
|
||||
|
||||
if parked {
|
||||
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
|
||||
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
|
||||
return false
|
||||
}
|
||||
s.swapMu.Lock()
|
||||
swapping := s.swapInFlight[vmid]
|
||||
s.swapMu.Unlock()
|
||||
if swapping {
|
||||
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
|
||||
"vmid", vmid, "status", status)
|
||||
s.resetNotRunning(vmid)
|
||||
return false
|
||||
}
|
||||
if seen < controllerSupervisorConfirm {
|
||||
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
|
||||
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
|
||||
return false
|
||||
}
|
||||
lock, _, err := s.staleLock.Lock(ctx, vmid)
|
||||
if err != nil {
|
||||
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
|
||||
return false
|
||||
}
|
||||
if lock != "" {
|
||||
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
|
||||
return false
|
||||
}
|
||||
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
|
||||
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
|
||||
"vmid", vmid, "backup_running", busy, "err", berr)
|
||||
return false
|
||||
}
|
||||
|
||||
// Backoff.
|
||||
s.ctrlSup.mu.Lock()
|
||||
st = s.supState(vmid)
|
||||
if !st.crashloopSince.IsZero() {
|
||||
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
|
||||
s.ctrlSup.mu.Unlock()
|
||||
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
|
||||
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
|
||||
return false
|
||||
}
|
||||
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
|
||||
// crash-loop (the hub keys on it moving, not on it clearing).
|
||||
st.restarts = nil
|
||||
st.crashloopSince = time.Time{}
|
||||
}
|
||||
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
|
||||
if len(st.restarts) >= controllerCrashloopMax {
|
||||
st.crashloopSince = now
|
||||
n := len(st.restarts)
|
||||
s.ctrlSup.mu.Unlock()
|
||||
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
|
||||
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
|
||||
return false
|
||||
}
|
||||
s.ctrlSup.mu.Unlock()
|
||||
|
||||
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
|
||||
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
|
||||
"vmid", vmid, "status", status, "unit", bootstrapUnit)
|
||||
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
|
||||
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
|
||||
reason += "; restart FAILED: " + err.Error()
|
||||
}
|
||||
s.ctrlSup.mu.Lock()
|
||||
st = s.supState(vmid)
|
||||
st.restarts = append(st.restarts, now)
|
||||
st.restartsTotal++
|
||||
st.lastRestartAt = now
|
||||
st.lastReason = reason
|
||||
st.notRunningSeen = 0
|
||||
// R-539: the slow counter. Raise at most once per 24 hours — the hub mails on the raise MOVING.
|
||||
st.restarts24h = append(pruneBefore(st.restarts24h, now.Add(-controllerSlowCrashloopWindow)), now)
|
||||
n24 := len(st.restarts24h)
|
||||
raised := false
|
||||
if n24 >= controllerSlowCrashloopMax && (st.slowCrashloopSince.IsZero() || now.Sub(st.slowCrashloopSince) >= controllerSlowCrashloopWindow) {
|
||||
st.slowCrashloopSince = now
|
||||
raised = true
|
||||
}
|
||||
rec := slowCounterRecord{Restarts: append([]time.Time(nil), st.restarts24h...), SlowCrashloopSince: st.slowCrashloopSince}
|
||||
s.ctrlSup.mu.Unlock()
|
||||
s.saveSlowCounter(vmid, rec)
|
||||
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason, "restarts_24h", n24)
|
||||
if raised {
|
||||
s.logger.Warn("controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop",
|
||||
"vmid", vmid, "restarts_24h", n24, "window", controllerSlowCrashloopWindow.String(), "threshold", controllerSlowCrashloopMax)
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func (s *Server) resetNotRunning(vmid int) {
|
||||
s.ctrlSup.mu.Lock()
|
||||
defer s.ctrlSup.mu.Unlock()
|
||||
if st := s.ctrlSup.guests[vmid]; st != nil {
|
||||
st.notRunningSeen = 0
|
||||
}
|
||||
}
|
||||
|
||||
func (s *Server) clock() time.Time {
|
||||
if s.now != nil {
|
||||
return s.now()
|
||||
}
|
||||
return time.Now().UTC()
|
||||
}
|
||||
|
||||
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
|
||||
out := ts[:0]
|
||||
for _, t := range ts {
|
||||
if !t.Before(cutoff) {
|
||||
out = append(out, t)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
|
||||
// Nil when the supervisor is not wired, so the stanza is omitted.
|
||||
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
|
||||
if s.staleLock == nil || s.guestExec == nil {
|
||||
return nil
|
||||
}
|
||||
s.ctrlSup.mu.Lock()
|
||||
defer s.ctrlSup.mu.Unlock()
|
||||
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
|
||||
for vmid, st := range s.ctrlSup.guests {
|
||||
g := hub.ControllerSupervisorGuest{
|
||||
VMID: vmid,
|
||||
RestartsTotal: st.restartsTotal,
|
||||
LastReason: st.lastReason,
|
||||
Parked: st.parked,
|
||||
Crashloop: !st.crashloopSince.IsZero(),
|
||||
}
|
||||
if !st.lastRestartAt.IsZero() {
|
||||
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
|
||||
}
|
||||
if !st.crashloopSince.IsZero() {
|
||||
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
|
||||
}
|
||||
now := s.clock()
|
||||
g.Restarts24h = len(pruneBefore(append([]time.Time(nil), st.restarts24h...), now.Add(-controllerSlowCrashloopWindow)))
|
||||
if !st.slowCrashloopSince.IsZero() {
|
||||
g.SlowCrashloopSince = st.slowCrashloopSince.UTC().Format(time.RFC3339)
|
||||
g.SlowCrashloop = now.Sub(st.slowCrashloopSince) < controllerSlowCrashloopWindow
|
||||
}
|
||||
out.Guests = append(out.Guests, g)
|
||||
}
|
||||
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,350 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"io"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
|
||||
// again", and each guard is pinned by the case where acting would be wrong.
|
||||
|
||||
type supExec struct {
|
||||
mu sync.Mutex
|
||||
status map[int]string // docker .State.Status per vmid; "" = container absent
|
||||
inspectErr error // non-nil = pct exec itself failed (unknown)
|
||||
restarts map[int]int
|
||||
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
|
||||
// stays "exited").
|
||||
onRestart string
|
||||
}
|
||||
|
||||
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
switch {
|
||||
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
|
||||
if f.inspectErr != nil {
|
||||
return "", f.inspectErr
|
||||
}
|
||||
st, ok := f.status[vmid]
|
||||
if !ok || st == "" {
|
||||
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
|
||||
}
|
||||
return st + "\n", nil
|
||||
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
|
||||
if f.restarts == nil {
|
||||
f.restarts = map[int]int{}
|
||||
}
|
||||
f.restarts[vmid]++
|
||||
if f.onRestart != "" {
|
||||
f.status[vmid] = f.onRestart
|
||||
}
|
||||
return "", nil
|
||||
}
|
||||
return "", errors.New("supExec: unexpected args")
|
||||
}
|
||||
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
|
||||
return "", errors.New("supExec: no stdin exec expected")
|
||||
}
|
||||
func (f *supExec) count(vmid int) int {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
return f.restarts[vmid]
|
||||
}
|
||||
|
||||
type supClock struct{ t time.Time }
|
||||
|
||||
func (c *supClock) now() time.Time { return c.t }
|
||||
|
||||
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
for _, v := range provisioned {
|
||||
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
|
||||
s := &Server{
|
||||
staleLock: ctl,
|
||||
guestExec: ex,
|
||||
guestsDir: dir,
|
||||
swapInFlight: map[int]bool{},
|
||||
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
|
||||
now: clk.now,
|
||||
}
|
||||
return s, clk, dir
|
||||
}
|
||||
|
||||
func runningGuest(vmid int) *fakeGuestPowerCtl {
|
||||
return &fakeGuestPowerCtl{
|
||||
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
|
||||
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
|
||||
}
|
||||
}
|
||||
|
||||
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
|
||||
// the first (the bootstrap's own rm-f/run window must never be raced).
|
||||
//
|
||||
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
|
||||
// stays 0 → "the killed controller was NOT restarted — this is R-523".
|
||||
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 0 {
|
||||
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
|
||||
}
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 1 {
|
||||
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
|
||||
}
|
||||
st := s.ControllerSupervisorStatus(context.Background())
|
||||
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
|
||||
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
|
||||
}
|
||||
// Healthy again → no further restart.
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 1 {
|
||||
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
|
||||
}
|
||||
}
|
||||
|
||||
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{}, onRestart: "running"}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if n := ex.count(9201); n != 1 {
|
||||
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
|
||||
}
|
||||
}
|
||||
|
||||
func TestControllerSupervisor_Guards(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
|
||||
}{
|
||||
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
|
||||
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
}},
|
||||
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
|
||||
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
|
||||
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.backupRun = map[int]bool{9201: true}
|
||||
}},
|
||||
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.backupErr = errors.New("tasks unreadable")
|
||||
}},
|
||||
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.guests[0].Status = "stopped"
|
||||
}},
|
||||
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
|
||||
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
|
||||
}},
|
||||
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||
ctl.guestsErr = errors.New("api down")
|
||||
}},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||
ctl := runningGuest(9201)
|
||||
s, _, dir := supServer(t, ex, ctl, 9201)
|
||||
tc.setup(s, ex, ctl, dir)
|
||||
for i := 0; i < 4; i++ {
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
}
|
||||
if n := ex.count(9201); n != 0 {
|
||||
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
|
||||
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9202: "exited"}}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
|
||||
for i := 0; i < 3; i++ {
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
}
|
||||
if n := ex.count(9202); n != 0 {
|
||||
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
|
||||
}
|
||||
}
|
||||
|
||||
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
|
||||
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
|
||||
// again.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
|
||||
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
|
||||
// failed at "crash-looping controller restarted 10 times".
|
||||
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
|
||||
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
ctx := context.Background()
|
||||
|
||||
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
clk.t = clk.t.Add(controllerSupervisorInterval)
|
||||
}
|
||||
if n := ex.count(9201); n != controllerCrashloopMax {
|
||||
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
|
||||
}
|
||||
st := s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||
if !st.Crashloop || st.CrashloopSince == "" {
|
||||
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
|
||||
}
|
||||
// Still paused 25 minutes later.
|
||||
clk.t = clk.t.Add(15 * time.Minute)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
if n := ex.count(9201); n != controllerCrashloopMax {
|
||||
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
|
||||
}
|
||||
// After the pause: tries again.
|
||||
clk.t = clk.t.Add(controllerCrashloopPause)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
s.ControllerSupervisorTick(ctx)
|
||||
if n := ex.count(9201); n != controllerCrashloopMax+1 {
|
||||
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
|
||||
}
|
||||
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
|
||||
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
|
||||
}
|
||||
}
|
||||
|
||||
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
|
||||
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
|
||||
var m map[string][]map[string]any
|
||||
if err := json.Unmarshal(b, &m); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
g := m["guests"][0]
|
||||
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked", "restarts_24h", "slow_crashloop"} {
|
||||
if _, ok := g[k]; !ok {
|
||||
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
// ---- R-539 (operator ruling 3 of 2026-09-16): the SLOW crash loop ----------------------------------
|
||||
|
||||
// supKillOnce kills the controller and lets the supervisor restart it (two confirming sweeps), then
|
||||
// moves the clock on by gap. The container comes back "running", so each restart is a separate act.
|
||||
func supKillOnce(t *testing.T, s *Server, ex *supExec, clk *supClock, gap time.Duration) {
|
||||
t.Helper()
|
||||
before := ex.count(9201)
|
||||
ex.mu.Lock()
|
||||
ex.status[9201] = "exited"
|
||||
ex.mu.Unlock()
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
clk.t = clk.t.Add(controllerSupervisorInterval)
|
||||
s.ControllerSupervisorTick(context.Background())
|
||||
if ex.count(9201) != before+1 {
|
||||
t.Fatalf("kill was not followed by exactly one restart (restarts %d → %d)", before, ex.count(9201))
|
||||
}
|
||||
clk.t = clk.t.Add(gap)
|
||||
}
|
||||
|
||||
// The consequence: a controller that dies every 20 minutes — never three times inside the 15-minute
|
||||
// brake — raises slow_crashloop on the FIFTH restart in 24 hours, and the raise does not move again on
|
||||
// the sixth (the hub mails on movement; once per 24 hours is the ruling).
|
||||
//
|
||||
// RED-PROOF: without the slow counter the stanza never sets slow_crashloop → "five restarts 20 minutes
|
||||
// apart did not raise slow_crashloop — this is R-539".
|
||||
func TestControllerSupervisor_SlowCrashloop(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
|
||||
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
ctx := context.Background()
|
||||
|
||||
for i := 1; i <= 4; i++ {
|
||||
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||
}
|
||||
g := s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||
if g.Crashloop {
|
||||
t.Fatalf("the 15-minute brake fired on restarts 20 minutes apart — the fixture is wrong: %+v", g)
|
||||
}
|
||||
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
|
||||
t.Fatalf("slow_crashloop raised after only 4 restarts: %+v", g)
|
||||
}
|
||||
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||
g = s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||
if !g.SlowCrashloop || g.SlowCrashloopSince == "" || g.Restarts24h != 5 {
|
||||
t.Fatalf("five restarts 20 minutes apart did not raise slow_crashloop — this is R-539: %+v", g)
|
||||
}
|
||||
first := g.SlowCrashloopSince
|
||||
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||
g = s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||
if g.SlowCrashloopSince != first {
|
||||
t.Fatalf("the raise moved again on the 6th restart (%q → %q) — the operator would be mailed per restart", first, g.SlowCrashloopSince)
|
||||
}
|
||||
if !g.SlowCrashloop {
|
||||
t.Fatalf("slow_crashloop cleared while the loop continues: %+v", g)
|
||||
}
|
||||
}
|
||||
|
||||
// The negative control: restarts that never reach five inside any 24 hours never raise it.
|
||||
func TestControllerSupervisor_SpreadRestartsNeverSlowCrashloop(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
|
||||
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||
for i := 0; i < 8; i++ { // eight restarts, 7 hours apart: at most 4 inside any 24 hours
|
||||
supKillOnce(t, s, ex, clk, 7*time.Hour)
|
||||
}
|
||||
g := s.ControllerSupervisorStatus(context.Background()).Guests[0]
|
||||
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
|
||||
t.Fatalf("restarts 7 hours apart raised slow_crashloop: %+v", g)
|
||||
}
|
||||
if g.Restarts24h > 4 {
|
||||
t.Fatalf("restarts_24h=%d — the 24-hour window is not pruning", g.Restarts24h)
|
||||
}
|
||||
}
|
||||
|
||||
// An agent restart must not reset the slow counter (the ruling; a box whose AGENT also restarts would
|
||||
// otherwise never reach five). The same state directory, a fresh Server.
|
||||
//
|
||||
// RED-PROOF: keep the counter in memory only → the second Server starts at 0 → "the agent restart
|
||||
// reset the slow counter".
|
||||
func TestControllerSupervisor_SlowCounterSurvivesAgentRestart(t *testing.T) {
|
||||
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
|
||||
s, clk, dir := supServer(t, ex, runningGuest(9201), 9201)
|
||||
for i := 0; i < 4; i++ {
|
||||
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||
}
|
||||
s2 := &Server{
|
||||
staleLock: s.staleLock,
|
||||
guestExec: ex,
|
||||
guestsDir: dir,
|
||||
swapInFlight: map[int]bool{},
|
||||
logger: s.logger,
|
||||
now: clk.now,
|
||||
}
|
||||
supKillOnce(t, s2, ex, clk, 20*time.Minute)
|
||||
g := s2.ControllerSupervisorStatus(context.Background()).Guests[0]
|
||||
if g.Restarts24h != 5 || !g.SlowCrashloop {
|
||||
t.Fatalf("the agent restart reset the slow counter: %+v", g)
|
||||
}
|
||||
}
|
||||
@@ -160,6 +160,10 @@ type Options struct {
|
||||
// NetStorage is the privileged network-mount (NAS) surface (Part A1). OPTIONAL — when nil, the
|
||||
// /netstorage endpoints report "not configured". Satisfied by *storage.SudoHostOps.
|
||||
NetStorage NetworkStorageOps
|
||||
// AfterPrimaryBackup (agent v0.140.0, `11-os-updates.md` §8 step 2) runs right after a SUCCESSFUL backup on the
|
||||
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
|
||||
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
|
||||
AfterPrimaryBackup func(ctx context.Context, vmid int)
|
||||
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
|
||||
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
|
||||
Privileged PrivilegedRunner
|
||||
@@ -185,6 +189,9 @@ type Options struct {
|
||||
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
|
||||
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
|
||||
StaleLock StaleLockController
|
||||
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
|
||||
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
|
||||
GuestsStateDir string
|
||||
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
|
||||
ControllerSwapStateDir string
|
||||
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
|
||||
@@ -273,6 +280,7 @@ type Server struct {
|
||||
tiers []BackupTier
|
||||
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
|
||||
inFlight *backup.InFlight
|
||||
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
|
||||
logger *slog.Logger
|
||||
now func() time.Time
|
||||
|
||||
@@ -362,6 +370,12 @@ type Server struct {
|
||||
swapMu sync.Mutex
|
||||
swapInFlight map[int]bool
|
||||
|
||||
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
|
||||
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
|
||||
guestExec GuestExecutor
|
||||
guestsDir string
|
||||
ctrlSup controllerSupervisor
|
||||
|
||||
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
|
||||
// detached pipeline runs through (tests inject; production defaults set in NewServer).
|
||||
netVerifyMu sync.Mutex
|
||||
@@ -460,6 +474,7 @@ func NewServer(o Options) (*Server, error) {
|
||||
// the primary is always first, because that is what the untargeted endpoints act on.
|
||||
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
|
||||
s.inFlight = o.InFlight
|
||||
s.afterPrimaryBackup = o.AfterPrimaryBackup
|
||||
if s.backups == nil && len(s.tiers) > 0 {
|
||||
s.backups = s.tiers[0].Service
|
||||
}
|
||||
@@ -484,7 +499,9 @@ func NewServer(o Options) (*Server, error) {
|
||||
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
|
||||
if o.ControllerSwap != nil {
|
||||
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
|
||||
s.guestExec = o.ControllerSwap
|
||||
}
|
||||
s.guestsDir = o.GuestsStateDir
|
||||
return s, nil
|
||||
}
|
||||
|
||||
@@ -883,6 +900,10 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
|
||||
}
|
||||
s.store.RecordBackup(b)
|
||||
s.finishJob(key, jobID, b)
|
||||
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
|
||||
if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
|
||||
s.afterPrimaryBackup(base, vmid)
|
||||
}
|
||||
}()
|
||||
writeStatus(w, http.StatusAccepted, true, BackupResponse{VMID: vmid, JobID: jobID, Phase: PhaseRunning}, "")
|
||||
}
|
||||
@@ -1084,6 +1105,10 @@ type BackupTierInfo struct {
|
||||
Target string `json:"target"`
|
||||
CadenceSeconds int64 `json:"cadence_seconds"`
|
||||
Primary bool `json:"primary"`
|
||||
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
|
||||
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
|
||||
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
|
||||
Storage string `json:"storage,omitempty"`
|
||||
}
|
||||
|
||||
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||
@@ -1093,11 +1118,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
|
||||
Target: t.TargetID,
|
||||
CadenceSeconds: int64(t.Cadence.Seconds()),
|
||||
Primary: t.Primary,
|
||||
Storage: s.storagePresence(r.Context(), t.TargetID),
|
||||
})
|
||||
}
|
||||
writeOK(w, resp)
|
||||
}
|
||||
|
||||
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
|
||||
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
|
||||
// successful read that does not list the storage is "absent".
|
||||
func (s *Server) storagePresence(ctx context.Context, target string) string {
|
||||
if s.storage == nil || target == "" {
|
||||
return StoragePresenceUnknown
|
||||
}
|
||||
targets, err := s.storage.Observe(ctx)
|
||||
if err != nil {
|
||||
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
|
||||
return StoragePresenceUnknown
|
||||
}
|
||||
for _, t := range targets {
|
||||
if t.Name == target {
|
||||
return StoragePresencePresent
|
||||
}
|
||||
}
|
||||
return StoragePresenceAbsent
|
||||
}
|
||||
|
||||
const (
|
||||
StoragePresencePresent = "present"
|
||||
StoragePresenceAbsent = "absent"
|
||||
StoragePresenceUnknown = "unknown"
|
||||
)
|
||||
|
||||
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
|
||||
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
|
||||
// result ("presence is not success").
|
||||
type TierBackupState struct {
|
||||
Target string `json:"target"`
|
||||
Primary bool `json:"primary"`
|
||||
Storage string `json:"storage"` // present | absent | unknown
|
||||
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
|
||||
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
|
||||
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
|
||||
// LastSuccessSource is "storage".
|
||||
LastSuccess *hub.Backup `json:"last_success,omitempty"`
|
||||
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
|
||||
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
|
||||
}
|
||||
|
||||
// TierAttempt is the newest recorded attempt on a tier, successful or not.
|
||||
type TierAttempt struct {
|
||||
StartedAt string `json:"started_at"`
|
||||
Success bool `json:"success"`
|
||||
Error string `json:"error,omitempty"`
|
||||
}
|
||||
|
||||
// tierBackupStates builds the per-tier view for one guest.
|
||||
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
|
||||
out := make([]TierBackupState, 0, len(s.tiers))
|
||||
for _, t := range s.tiers {
|
||||
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
|
||||
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
|
||||
st.LastSuccess, st.LastSuccessSource = b, "record"
|
||||
} else if st.Storage != StoragePresenceAbsent {
|
||||
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
|
||||
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
|
||||
StartedAt: when.UTC().Format(time.RFC3339)}
|
||||
st.LastSuccessSource = "storage"
|
||||
}
|
||||
}
|
||||
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
|
||||
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
|
||||
}
|
||||
out = append(out, st)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// tierFromRequest resolves the `?target=` query parameter to a tier.
|
||||
//
|
||||
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
|
||||
@@ -1144,6 +1241,10 @@ type BackupStatusResponse struct {
|
||||
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
|
||||
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
|
||||
Target string `json:"target,omitempty"`
|
||||
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
|
||||
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
|
||||
// response exactly as before.
|
||||
Tiers []TierBackupState `json:"tiers,omitempty"`
|
||||
}
|
||||
|
||||
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||
@@ -1155,6 +1256,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
|
||||
// across ANY target (echo == "" → pickLatestBackup's match-any path).
|
||||
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
|
||||
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
|
||||
if echo == "" {
|
||||
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
|
||||
}
|
||||
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
|
||||
resp.Phase = job.Phase
|
||||
resp.JobID = job.JobID
|
||||
@@ -1371,3 +1475,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
|
||||
w.WriteHeader(code)
|
||||
_ = json.NewEncoder(w).Encode(apiResponse{OK: ok, Data: data, Error: errMsg})
|
||||
}
|
||||
|
||||
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
|
||||
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn }
|
||||
|
||||
@@ -0,0 +1,459 @@
|
||||
// Package osupdate is the agent's OS-update leg for the customer GUEST (`11-os-updates.md` §8 step 2, agent v0.140.0).
|
||||
//
|
||||
// It runs right after the night's successful whole-guest backup, while the backup goroutine still holds the host-wide
|
||||
// heavy-op gate (so it never overlaps a backup or a restore-test, `11` C10), at most once per night. All root work is
|
||||
// the wrapper `felhom-os-apply` (configs/, its own tests); this package only builds plans, calls the wrapper through
|
||||
// sudo, judges health and reports to the hub.
|
||||
//
|
||||
// NO AUTOMATIC UNDO in this release (R-837, measured 2026-10-04): a customer guest cannot be snapshotted — PVE
|
||||
// refuses any snapshot not named `vzdump` when the guest has host-path binds (mp8, mp9). A failed health check
|
||||
// therefore stops, reports `health_failed` and the hub mails the operator; the whole-guest backup taken minutes
|
||||
// earlier is the undo, by hand.
|
||||
package osupdate
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// WrapperPath is the pinned sudoers vector (configs/felhom-agent.sudoers FELHOM_OSAPPLY).
|
||||
const WrapperPath = "/usr/local/sbin/felhom-os-apply"
|
||||
|
||||
// DefaultPlanDir is where plans are written (the sudoers glob names it).
|
||||
const DefaultPlanDir = "/var/lib/felhom-agent/os"
|
||||
|
||||
// Package is one name=version with its origin.
|
||||
type Package struct {
|
||||
Name string `json:"name"`
|
||||
Version string `json:"version"`
|
||||
Origin string `json:"origin"`
|
||||
}
|
||||
|
||||
// Pending is one update the guest's sources offer (origin as apt names it, possibly several).
|
||||
type Pending struct {
|
||||
Name string `json:"name"`
|
||||
From string `json:"from"`
|
||||
To string `json:"to"`
|
||||
Origin []string `json:"origin"`
|
||||
}
|
||||
|
||||
// Container is one container's state as the wrapper saw it.
|
||||
type Container struct {
|
||||
State string `json:"state"`
|
||||
Health string `json:"health"` // healthy | unhealthy | starting | none
|
||||
}
|
||||
|
||||
// Health is the guest's health snapshot (the wrapper's `health` object).
|
||||
type Health struct {
|
||||
DockerOK bool `json:"docker_ok"`
|
||||
NetworkOK bool `json:"network_ok"`
|
||||
Controller string `json:"controller"`
|
||||
Containers map[string]Container `json:"containers"`
|
||||
}
|
||||
|
||||
// WrapperReport is the wrapper's OSAPPLY-REPORT object.
|
||||
type WrapperReport struct {
|
||||
Mode string `json:"mode"`
|
||||
Refused json.RawMessage `json:"refused"`
|
||||
Failed json.RawMessage `json:"failed"`
|
||||
Upgraded []Package `json:"upgraded"`
|
||||
Installed []Package `json:"installed"`
|
||||
Pending []Pending `json:"pending"`
|
||||
RestartNeeded []string `json:"restart_needed"`
|
||||
DockerRestartNeeded bool `json:"docker_restart_needed"`
|
||||
RebootNeeded bool `json:"reboot_needed"`
|
||||
HealthBefore *Health `json:"health_before"`
|
||||
HealthAfter *Health `json:"health_after"`
|
||||
Health *Health `json:"health"`
|
||||
}
|
||||
|
||||
func (w WrapperReport) refused() bool { return len(w.Refused) > 0 && string(w.Refused) != "null" }
|
||||
func (w WrapperReport) failed() bool { return len(w.Failed) > 0 && string(w.Failed) != "null" }
|
||||
|
||||
// Report is what the hub receives (hub osupdates.Report — field-exact).
|
||||
type Report struct {
|
||||
RunID string `json:"run_id"`
|
||||
Trigger string `json:"trigger"`
|
||||
Mode string `json:"mode"`
|
||||
Ring int `json:"ring"`
|
||||
ReleaseID string `json:"release_id"`
|
||||
Outcome string `json:"outcome"`
|
||||
Healthy bool `json:"healthy"`
|
||||
HealthReason string `json:"health_reason,omitempty"`
|
||||
VMID int `json:"vmid"`
|
||||
Upgraded []Package `json:"upgraded,omitempty"`
|
||||
Installed []Package `json:"installed,omitempty"`
|
||||
Pending []Pending `json:"pending,omitempty"`
|
||||
NotCovered []string `json:"not_covered,omitempty"`
|
||||
RestartNeeded []string `json:"restart_needed,omitempty"`
|
||||
DockerRestartNeeded bool `json:"docker_restart_needed,omitempty"`
|
||||
RebootNeeded bool `json:"reboot_needed,omitempty"`
|
||||
Refused json.RawMessage `json:"refused,omitempty"`
|
||||
}
|
||||
|
||||
// Reporter posts a report to the hub (*hub.Client).
|
||||
type Reporter interface {
|
||||
PostOSReport(ctx context.Context, body []byte) error
|
||||
}
|
||||
|
||||
// Leg runs one OS-update pass for the customer guest.
|
||||
type Leg struct {
|
||||
Runner proxmox.Runner
|
||||
Hub Reporter
|
||||
Logger *slog.Logger
|
||||
PlanDir string
|
||||
StatePath string // last night run (once per night)
|
||||
HealthWait time.Duration // how long health may take to come back (default 5 min)
|
||||
HealthPoll time.Duration // default 15 s
|
||||
MinGap time.Duration // between night runs (default 20 h)
|
||||
Now func() time.Time
|
||||
Sleep func(context.Context, time.Duration)
|
||||
|
||||
mu sync.Mutex
|
||||
block *hub.WireOSUpdate
|
||||
have bool
|
||||
}
|
||||
|
||||
// OnDesiredState stores the hub's os_update block (desired.RawConsumer — store only, never block).
|
||||
func (l *Leg) OnDesiredState(_ context.Context, resp *hub.DesiredStateResponse) {
|
||||
if resp == nil {
|
||||
return
|
||||
}
|
||||
l.mu.Lock()
|
||||
defer l.mu.Unlock()
|
||||
l.block, l.have = resp.DesiredState.OSUpdate, true
|
||||
}
|
||||
|
||||
// Block returns the newest os_update block. No block (an older hub, or nothing fetched yet) = ring 1, ON, no
|
||||
// release: the box reports and installs nothing.
|
||||
func (l *Leg) Block() hub.WireOSUpdate {
|
||||
l.mu.Lock()
|
||||
defer l.mu.Unlock()
|
||||
if l.block == nil {
|
||||
return hub.WireOSUpdate{Ring: 1, Enabled: true}
|
||||
}
|
||||
return *l.block
|
||||
}
|
||||
|
||||
// SetBlock sets the block directly (the selftest fetches the desired state itself).
|
||||
func (l *Leg) SetBlock(b *hub.WireOSUpdate) {
|
||||
l.mu.Lock()
|
||||
defer l.mu.Unlock()
|
||||
l.block, l.have = b, true
|
||||
}
|
||||
|
||||
func (l *Leg) now() time.Time {
|
||||
if l.Now != nil {
|
||||
return l.Now()
|
||||
}
|
||||
return time.Now()
|
||||
}
|
||||
|
||||
func (l *Leg) log() *slog.Logger {
|
||||
if l.Logger != nil {
|
||||
return l.Logger
|
||||
}
|
||||
return slog.Default()
|
||||
}
|
||||
|
||||
func (l *Leg) sleep(ctx context.Context, d time.Duration) {
|
||||
if l.Sleep != nil {
|
||||
l.Sleep(ctx, d)
|
||||
return
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
case <-time.After(d):
|
||||
}
|
||||
}
|
||||
|
||||
// IsFast reports whether every origin apt names is Debian / Debian-Security (the fast lane, `11` C3).
|
||||
func IsFast(origins []string) bool {
|
||||
if len(origins) == 0 {
|
||||
return false
|
||||
}
|
||||
for _, o := range origins {
|
||||
if o != "Debian" && o != "Debian-Security" {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func fastOrigin(origins []string) string {
|
||||
for _, o := range origins {
|
||||
if o == "Debian-Security" {
|
||||
return o
|
||||
}
|
||||
}
|
||||
return "Debian"
|
||||
}
|
||||
|
||||
// HealthVerdict is THE health rule (`11` §5.4.1; pinned by TestHealthVerdict_*): after the run, docker answers, the
|
||||
// guest's network resolves, the controller's own health check is `healthy`, and every container that was running
|
||||
// before is running again — and healthy again if it was healthy before. "starting" is not yet healthy.
|
||||
func HealthVerdict(before, after *Health) (bool, string) {
|
||||
if after == nil {
|
||||
return false, "no health reading"
|
||||
}
|
||||
if !after.DockerOK {
|
||||
return false, "docker does not answer"
|
||||
}
|
||||
if !after.NetworkOK {
|
||||
return false, "the guest cannot resolve deb.debian.org"
|
||||
}
|
||||
if after.Controller != "healthy" {
|
||||
return false, "the controller is " + after.Controller
|
||||
}
|
||||
if before == nil {
|
||||
return true, ""
|
||||
}
|
||||
names := make([]string, 0, len(before.Containers))
|
||||
for n := range before.Containers {
|
||||
names = append(names, n)
|
||||
}
|
||||
sort.Strings(names)
|
||||
for _, n := range names {
|
||||
b := before.Containers[n]
|
||||
if b.State != "running" {
|
||||
continue
|
||||
}
|
||||
a, ok := after.Containers[n]
|
||||
if !ok || a.State != "running" {
|
||||
return false, n + " was running and is not"
|
||||
}
|
||||
if b.Health == "healthy" && a.Health != "healthy" {
|
||||
return false, n + " was healthy and is " + a.Health
|
||||
}
|
||||
}
|
||||
return true, ""
|
||||
}
|
||||
|
||||
// MergeBaseline is the health baseline from two readings: a container counts as running (and healthy) if EITHER
|
||||
// reading saw it so. nil-safe. Pinned by TestHealth_BaselineIsTheStartOfTheLeg.
|
||||
func MergeBaseline(a, b *Health) *Health {
|
||||
if a == nil {
|
||||
return b
|
||||
}
|
||||
if b == nil {
|
||||
return a
|
||||
}
|
||||
out := &Health{DockerOK: a.DockerOK && b.DockerOK, NetworkOK: a.NetworkOK && b.NetworkOK, Controller: b.Controller,
|
||||
Containers: map[string]Container{}}
|
||||
for _, h := range []*Health{a, b} {
|
||||
for n, c := range h.Containers {
|
||||
cur, seen := out.Containers[n]
|
||||
if !seen || (cur.State != "running" && c.State == "running") {
|
||||
out.Containers[n] = c
|
||||
continue
|
||||
}
|
||||
if c.State == "running" && c.Health == "healthy" {
|
||||
out.Containers[n] = c
|
||||
}
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// call writes the plan and runs the wrapper once.
|
||||
func (l *Leg) call(ctx context.Context, runID, mode string, vmid int, rel hub.WireOSRelease, pkgs []Package) (WrapperReport, error) {
|
||||
dir := l.PlanDir
|
||||
if dir == "" {
|
||||
dir = DefaultPlanDir
|
||||
}
|
||||
if err := os.MkdirAll(dir, 0o700); err != nil {
|
||||
return WrapperReport{}, fmt.Errorf("osupdate: plan dir: %w", err)
|
||||
}
|
||||
plan := map[string]any{"release_id": rel.ID, "layer": "guest", "lane": "fast", "vmid": vmid, "mode": mode,
|
||||
"snapshot": rel.Snapshot, "packages": pkgs}
|
||||
if pkgs == nil {
|
||||
plan["packages"] = []Package{}
|
||||
}
|
||||
b, _ := json.Marshal(plan)
|
||||
path := filepath.Join(dir, "plan-"+runID+"-"+mode+".json")
|
||||
if err := os.WriteFile(path, b, 0o600); err != nil {
|
||||
return WrapperReport{}, fmt.Errorf("osupdate: write plan: %w", err)
|
||||
}
|
||||
defer os.Remove(path)
|
||||
stdout, stderr, err := l.Runner.Run(ctx, WrapperPath, "--plan", path)
|
||||
for _, line := range strings.Split(strings.TrimSpace(string(stderr)), "\n") {
|
||||
if strings.HasPrefix(line, "os-apply: ") {
|
||||
l.log().Info("osupdate: wrapper", "line", line)
|
||||
}
|
||||
}
|
||||
var rep WrapperReport
|
||||
found := false
|
||||
for _, line := range strings.Split(string(stdout), "\n") {
|
||||
if strings.HasPrefix(line, "OSAPPLY-REPORT ") {
|
||||
if jerr := json.Unmarshal([]byte(strings.TrimPrefix(line, "OSAPPLY-REPORT ")), &rep); jerr == nil {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
return rep, fmt.Errorf("osupdate: wrapper gave no report (err %v): %s", err, strings.TrimSpace(string(stderr)))
|
||||
}
|
||||
return rep, nil // a refusal / failure is IN the report (exit 2 / 3), not an error here
|
||||
}
|
||||
|
||||
// Run is one pass: inventory → (switch, ring, plan) → apply → health → report. trigger is "night" or "debug".
|
||||
func (l *Leg) Run(ctx context.Context, vmid int, trigger string) Report {
|
||||
runID := l.now().UTC().Format("20060102T150405Z")
|
||||
blk := l.Block()
|
||||
rel := hub.WireOSRelease{ID: "ring0-" + runID}
|
||||
if blk.Ring == 1 {
|
||||
rel = hub.WireOSRelease{}
|
||||
if blk.Release != nil {
|
||||
rel = *blk.Release
|
||||
}
|
||||
}
|
||||
rep := Report{RunID: runID, Trigger: trigger, Ring: blk.Ring, VMID: vmid, ReleaseID: rel.ID, Mode: "inventory"}
|
||||
lg := l.log().With("run", runID, "vmid", vmid, "ring", blk.Ring, "trigger", trigger)
|
||||
|
||||
if trigger == "night" && l.StatePath != "" {
|
||||
gap := l.MinGap
|
||||
if gap == 0 {
|
||||
gap = 20 * time.Hour
|
||||
}
|
||||
if b, err := os.ReadFile(l.StatePath); err == nil {
|
||||
if last, perr := time.Parse(time.RFC3339, strings.TrimSpace(string(b))); perr == nil && l.now().Sub(last) < gap {
|
||||
lg.Info("osupdate: skipped — already ran tonight", "last", last.UTC().Format(time.RFC3339))
|
||||
rep.Outcome = "skipped"
|
||||
return rep
|
||||
}
|
||||
}
|
||||
}
|
||||
lg.Info("osupdate: START", "enabled", blk.Enabled, "release", rel.ID)
|
||||
|
||||
inv, err := l.call(ctx, runID, "inventory", vmid, rel, nil)
|
||||
if err != nil || inv.refused() || inv.failed() {
|
||||
rep.Outcome, rep.Refused = "refused", inv.Refused
|
||||
if err != nil {
|
||||
rep.Outcome, rep.HealthReason = "failed", err.Error()
|
||||
}
|
||||
return l.finish(ctx, lg, rep)
|
||||
}
|
||||
var plan []Package
|
||||
switch {
|
||||
case !blk.Enabled:
|
||||
rep.Outcome = "inventory"
|
||||
lg.Info("osupdate: switched OFF for this box — reporting only")
|
||||
case blk.Ring == 0:
|
||||
for _, p := range inv.Pending {
|
||||
if IsFast(p.Origin) {
|
||||
plan = append(plan, Package{Name: p.Name, Version: p.To, Origin: fastOrigin(p.Origin)})
|
||||
}
|
||||
}
|
||||
default:
|
||||
for _, p := range rel.Packages {
|
||||
plan = append(plan, Package{Name: p.Name, Version: p.Version, Origin: p.Origin})
|
||||
}
|
||||
}
|
||||
pkgNames := map[string]bool{}
|
||||
for _, p := range plan {
|
||||
pkgNames[p.Name] = true
|
||||
}
|
||||
if blk.Enabled && len(plan) == 0 {
|
||||
rep.Outcome = "nothing"
|
||||
}
|
||||
final := inv
|
||||
if blk.Enabled && len(plan) > 0 {
|
||||
rep.Mode = "apply"
|
||||
ap, err := l.call(ctx, runID, "apply", vmid, rel, plan)
|
||||
switch {
|
||||
case err != nil:
|
||||
rep.Outcome, rep.HealthReason = "failed", err.Error()
|
||||
return l.finish(ctx, lg, rep)
|
||||
case ap.refused():
|
||||
rep.Outcome, rep.Refused = "refused", ap.Refused
|
||||
return l.finish(ctx, lg, rep)
|
||||
case ap.failed():
|
||||
rep.Outcome, rep.Refused = "failed", ap.Failed
|
||||
}
|
||||
final = ap
|
||||
rep.Upgraded = ap.Upgraded
|
||||
if rep.Outcome == "" {
|
||||
if len(ap.Upgraded) == 0 {
|
||||
rep.Outcome = "nothing"
|
||||
} else {
|
||||
rep.Outcome = "applied"
|
||||
}
|
||||
}
|
||||
// Health: compare with what the guest looked like BEFORE the run; give restarted services time.
|
||||
wait, poll := l.HealthWait, l.HealthPoll
|
||||
if wait == 0 {
|
||||
wait = 5 * time.Minute
|
||||
}
|
||||
if poll == 0 {
|
||||
poll = 15 * time.Second
|
||||
}
|
||||
deadline := l.now().Add(wait)
|
||||
cur := ap.HealthAfter
|
||||
// The baseline is the guest as it was at the START of the leg (the inventory's reading) merged with the
|
||||
// apply's own "before": an app that stops at any point during the run counts. Found live 2026-10-04: an app
|
||||
// stopped between the inventory and the apply's own reading was taken as "stopped before" and ignored.
|
||||
base := MergeBaseline(inv.HealthAfter, ap.HealthBefore)
|
||||
for {
|
||||
ok, why := HealthVerdict(base, cur)
|
||||
rep.Healthy, rep.HealthReason = ok, why
|
||||
if ok || !l.now().Before(deadline) || ctx.Err() != nil {
|
||||
break
|
||||
}
|
||||
l.sleep(ctx, poll)
|
||||
hr, herr := l.call(ctx, runID, "health", vmid, rel, nil)
|
||||
if herr == nil && hr.Health != nil {
|
||||
cur = hr.Health
|
||||
}
|
||||
}
|
||||
if !rep.Healthy && rep.Outcome == "applied" {
|
||||
rep.Outcome = "health_failed"
|
||||
}
|
||||
} else {
|
||||
ok, why := HealthVerdict(nil, inv.HealthAfter)
|
||||
rep.Healthy, rep.HealthReason = ok, why
|
||||
}
|
||||
rep.Installed, rep.Pending = final.Installed, final.Pending
|
||||
rep.RestartNeeded, rep.DockerRestartNeeded, rep.RebootNeeded = final.RestartNeeded, final.DockerRestartNeeded, final.RebootNeeded
|
||||
rep.NotCovered = notCovered(final.Pending, blk.Ring, pkgNames)
|
||||
if trigger == "night" && l.StatePath != "" {
|
||||
_ = os.WriteFile(l.StatePath, []byte(l.now().UTC().Format(time.RFC3339)), 0o600)
|
||||
}
|
||||
return l.finish(ctx, lg, rep)
|
||||
}
|
||||
|
||||
// notCovered lists pending updates no approved release covers: in ring 0 everything outside the fast lane; in
|
||||
// ring 1 also every fast-lane update the release did not name.
|
||||
func notCovered(pending []Pending, ring int, planned map[string]bool) []string {
|
||||
var out []string
|
||||
for _, p := range pending {
|
||||
if !IsFast(p.Origin) || (ring == 1 && !planned[p.Name]) {
|
||||
out = append(out, p.Name)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func (l *Leg) finish(ctx context.Context, lg *slog.Logger, rep Report) Report {
|
||||
lg.Info("osupdate: DONE", "outcome", rep.Outcome, "healthy", rep.Healthy, "reason", rep.HealthReason,
|
||||
"upgraded", len(rep.Upgraded), "pending", len(rep.Pending), "not_covered", len(rep.NotCovered), "restart_needed", len(rep.RestartNeeded))
|
||||
if l.Hub != nil {
|
||||
body, _ := json.Marshal(rep)
|
||||
rctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), time.Minute)
|
||||
defer cancel()
|
||||
if err := l.Hub.PostOSReport(rctx, body); err != nil {
|
||||
lg.Warn("osupdate: reporting to the hub failed (the run itself is done)", "err", err)
|
||||
}
|
||||
}
|
||||
return rep
|
||||
}
|
||||
@@ -0,0 +1,281 @@
|
||||
package osupdate
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
)
|
||||
|
||||
// fakeWrapper plays /usr/local/sbin/felhom-os-apply: it reads the plan the leg wrote and answers per mode.
|
||||
type fakeWrapper struct {
|
||||
t *testing.T
|
||||
pending []Pending
|
||||
applyRep WrapperReport
|
||||
healthSeq []*Health // answers to successive "health" calls
|
||||
plans []map[string]any
|
||||
}
|
||||
|
||||
func (f *fakeWrapper) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
|
||||
if name != WrapperPath || len(args) != 2 || args[0] != "--plan" {
|
||||
f.t.Fatalf("unexpected command %s %v", name, args)
|
||||
}
|
||||
b, err := os.ReadFile(args[1])
|
||||
if err != nil {
|
||||
f.t.Fatal(err)
|
||||
}
|
||||
var plan map[string]any
|
||||
json.Unmarshal(b, &plan)
|
||||
f.plans = append(f.plans, plan)
|
||||
var rep WrapperReport
|
||||
healthy := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
|
||||
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "running", Health: "healthy"}}}
|
||||
switch plan["mode"] {
|
||||
case "inventory":
|
||||
rep = WrapperReport{Mode: "inventory", Pending: f.pending, HealthAfter: healthy,
|
||||
Installed: []Package{{Name: "libc6", Version: "u3", Origin: "Debian"}}}
|
||||
case "apply":
|
||||
rep = f.applyRep
|
||||
rep.Mode = "apply"
|
||||
if rep.HealthBefore == nil {
|
||||
rep.HealthBefore = healthy
|
||||
}
|
||||
if rep.HealthAfter == nil {
|
||||
rep.HealthAfter = healthy
|
||||
}
|
||||
case "health":
|
||||
if len(f.healthSeq) > 0 {
|
||||
rep.Health, f.healthSeq = f.healthSeq[0], f.healthSeq[1:]
|
||||
} else {
|
||||
rep.Health = healthy
|
||||
}
|
||||
}
|
||||
out, _ := json.Marshal(rep)
|
||||
return []byte("OSAPPLY-REPORT " + string(out) + "\n"), []byte("os-apply: DONE rc=0\n"), nil
|
||||
}
|
||||
|
||||
func (f *fakeWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, args ...string) ([]byte, []byte, error) {
|
||||
return f.Run(ctx, name, args...)
|
||||
}
|
||||
|
||||
type fakeHub struct{ reports []Report }
|
||||
|
||||
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
|
||||
var r Report
|
||||
json.Unmarshal(body, &r)
|
||||
h.reports = append(h.reports, r)
|
||||
return nil
|
||||
}
|
||||
|
||||
func newLeg(t *testing.T, w *fakeWrapper, blk *hub.WireOSUpdate) (*Leg, *fakeHub) {
|
||||
h := &fakeHub{}
|
||||
now := time.Date(2026, 10, 4, 4, 0, 0, 0, time.UTC)
|
||||
l := &Leg{Runner: w, Hub: h, PlanDir: t.TempDir(), StatePath: filepath.Join(t.TempDir(), "last"),
|
||||
HealthWait: time.Minute, HealthPoll: 10 * time.Second,
|
||||
Now: func() time.Time { return now },
|
||||
Sleep: func(_ context.Context, d time.Duration) { now = now.Add(d) }}
|
||||
if blk != nil {
|
||||
l.SetBlock(blk)
|
||||
}
|
||||
return l, h
|
||||
}
|
||||
|
||||
var pend = []Pending{
|
||||
{Name: "libc6", From: "u3", To: "u4", Origin: []string{"Debian"}},
|
||||
{Name: "openssl", From: "u1", To: "u3", Origin: []string{"Debian-Security", "Debian"}},
|
||||
{Name: "docker-ce", From: "29.7", To: "29.8", Origin: []string{"Docker CE"}},
|
||||
}
|
||||
|
||||
func modes(w *fakeWrapper) string {
|
||||
var m []string
|
||||
for _, p := range w.plans {
|
||||
m = append(m, p["mode"].(string))
|
||||
}
|
||||
return strings.Join(m, ",")
|
||||
}
|
||||
|
||||
// Ring 0 plans every pending FAST-LANE update (Docker excluded, `11` C3) and reports what it now runs.
|
||||
func TestRing0_PlansTheFastLaneOnly(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6", Version: "u4"}, {Name: "openssl", Version: "u3"}}, Pending: pend[2:]}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
rep := l.Run(context.Background(), 9201, "night")
|
||||
if rep.Outcome != "applied" || !rep.Healthy {
|
||||
t.Fatalf("rep = %+v", rep)
|
||||
}
|
||||
pk := w.plans[1]["packages"].([]any)
|
||||
if len(pk) != 2 {
|
||||
t.Fatalf("ring-0 plan = %v, want libc6 + openssl only", pk)
|
||||
}
|
||||
if o := pk[1].(map[string]any)["origin"]; o != "Debian-Security" {
|
||||
t.Fatalf("openssl origin = %v", o)
|
||||
}
|
||||
if w.plans[1]["snapshot"] != "" {
|
||||
t.Fatalf("ring 0 installs from live sources, snapshot = %v", w.plans[1]["snapshot"])
|
||||
}
|
||||
if len(rep.NotCovered) != 1 || rep.NotCovered[0] != "docker-ce" {
|
||||
t.Fatalf("not covered = %v", rep.NotCovered)
|
||||
}
|
||||
if len(h.reports) != 1 || h.reports[0].Outcome != "applied" {
|
||||
t.Fatalf("hub got %+v", h.reports)
|
||||
}
|
||||
}
|
||||
|
||||
// Ring 1 installs EXACTLY the approved release (its versions, its snapshot), nothing it computed itself.
|
||||
func TestRing1_InstallsExactlyTheRelease(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6", Version: "u4-approved"}}, Pending: pend[1:]}}
|
||||
rel := &hub.WireOSRelease{ID: "os-1", Snapshot: "20261004T080000Z", Packages: []hub.WireOSPackage{{Name: "libc6", Version: "u4-approved", Origin: "Debian"}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true, Release: rel})
|
||||
rep := l.Run(context.Background(), 9201, "night")
|
||||
if rep.Outcome != "applied" || rep.ReleaseID != "os-1" {
|
||||
t.Fatalf("rep = %+v", rep)
|
||||
}
|
||||
ap := w.plans[1]
|
||||
pk := ap["packages"].([]any)
|
||||
if len(pk) != 1 || pk[0].(map[string]any)["version"] != "u4-approved" || ap["snapshot"] != "20261004T080000Z" || ap["release_id"] != "os-1" {
|
||||
t.Fatalf("ring-1 plan = %v", ap)
|
||||
}
|
||||
// openssl is pending and fast-lane but NOT in the release → not covered; docker-ce is never covered.
|
||||
if strings.Join(rep.NotCovered, ",") != "openssl,docker-ce" {
|
||||
t.Fatalf("not covered = %v", rep.NotCovered)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRing1_NoReleaseInstallsNothing(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 1, Enabled: true})
|
||||
rep := l.Run(context.Background(), 9201, "night")
|
||||
if rep.Outcome != "nothing" || modes(w) != "inventory" {
|
||||
t.Fatalf("rep=%+v modes=%s", rep, modes(w))
|
||||
}
|
||||
}
|
||||
|
||||
// No block from the hub (an older hub): ring 1, ON, no release → reports, installs nothing.
|
||||
func TestNoBlock_IsRing1Nothing(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend}
|
||||
l, _ := newLeg(t, w, nil)
|
||||
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "nothing" || rep.Ring != 1 || modes(w) != "inventory" {
|
||||
t.Fatalf("rep=%+v modes=%s", rep, modes(w))
|
||||
}
|
||||
}
|
||||
|
||||
// Switched OFF: the box reports but installs nothing (the brief, Part D 2).
|
||||
func TestSwitchOff_ReportsOnly(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: false})
|
||||
rep := l.Run(context.Background(), 9201, "night")
|
||||
if rep.Outcome != "inventory" || modes(w) != "inventory" || len(h.reports) != 1 || len(rep.Pending) != 3 {
|
||||
t.Fatalf("rep=%+v modes=%s", rep, modes(w))
|
||||
}
|
||||
}
|
||||
|
||||
// Unhealthy after the run, and still unhealthy at the end of the wait → health_failed (the hub mails the operator).
|
||||
func TestHealth_FailsAfterTheWait(t *testing.T) {
|
||||
bad := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
|
||||
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "exited"}}}
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: bad},
|
||||
healthSeq: []*Health{bad, bad, bad, bad, bad, bad, bad, bad}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
rep := l.Run(context.Background(), 9201, "night")
|
||||
if rep.Outcome != "health_failed" || rep.Healthy || !strings.Contains(rep.HealthReason, "app was running") {
|
||||
t.Fatalf("rep = %+v", rep)
|
||||
}
|
||||
if h.reports[0].Outcome != "health_failed" {
|
||||
t.Fatal("the hub was not told")
|
||||
}
|
||||
}
|
||||
|
||||
// A service that takes a moment to come back is not a failure: the poll sees it recover inside the wait.
|
||||
func TestHealth_RecoversInsideTheWait(t *testing.T) {
|
||||
starting := &Health{DockerOK: true, NetworkOK: true, Controller: "starting"}
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}}, HealthAfter: starting},
|
||||
healthSeq: []*Health{starting}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "applied" || !rep.Healthy {
|
||||
t.Fatalf("rep = %+v", rep)
|
||||
}
|
||||
}
|
||||
|
||||
func TestHealthVerdict(t *testing.T) {
|
||||
ok := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{"a": {State: "running", Health: "healthy"}}}
|
||||
cases := []struct {
|
||||
name string
|
||||
before *Health
|
||||
after *Health
|
||||
want bool
|
||||
}{
|
||||
{"all good", ok, ok, true},
|
||||
{"no reading", ok, nil, false},
|
||||
{"docker down", ok, &Health{NetworkOK: true, Controller: "healthy"}, false},
|
||||
{"no network", ok, &Health{DockerOK: true, Controller: "healthy"}, false},
|
||||
{"controller starting", ok, &Health{DockerOK: true, NetworkOK: true, Controller: "starting"}, false},
|
||||
{"only the controller differs", ok, &Health{DockerOK: true, NetworkOK: true, Controller: "unhealthy", Containers: map[string]Container{"a": {State: "running", Health: "healthy"}}}, false},
|
||||
{"app gone", ok, &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{}}, false},
|
||||
{"app unhealthy", ok, &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{"a": {State: "running", Health: "unhealthy"}}}, false},
|
||||
{"stopped before stays stopped", &Health{Containers: map[string]Container{"x": {State: "exited"}}}, &Health{DockerOK: true, NetworkOK: true, Controller: "healthy"}, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got, why := HealthVerdict(c.before, c.after); got != c.want {
|
||||
t.Errorf("%s: got %v (%s), want %v", c.name, got, why, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestOncePerNight(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}}}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
l.Run(context.Background(), 9201, "night")
|
||||
n := len(w.plans)
|
||||
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "skipped" || len(w.plans) != n {
|
||||
t.Fatalf("a second night run in the same night ran: %+v", rep)
|
||||
}
|
||||
if rep := l.Run(context.Background(), 9201, "debug"); rep.Outcome == "skipped" {
|
||||
t.Fatal("the debug action must not be throttled")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRefusedIsReported(t *testing.T) {
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Refused: json.RawMessage(`{"code":"R6","reason":"x"}`)}}
|
||||
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
if rep := l.Run(context.Background(), 9201, "night"); rep.Outcome != "refused" || h.reports[0].Outcome != "refused" {
|
||||
t.Fatalf("rep = %+v", rep)
|
||||
}
|
||||
}
|
||||
|
||||
// The root wrapper's own suite (configs/test_felhom_os_apply.py) runs with `go test ./...` so CI covers it.
|
||||
func TestWrapperSuite(t *testing.T) {
|
||||
py, err := exec.LookPath("python3")
|
||||
if err != nil {
|
||||
t.Skip("python3 not available")
|
||||
}
|
||||
cmd := exec.Command(py, "../../configs/test_felhom_os_apply.py")
|
||||
out, err := cmd.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("wrapper suite failed: %v\n%s", err, out)
|
||||
}
|
||||
if !strings.Contains(string(out), "OK") {
|
||||
t.Fatalf("wrapper suite did not report OK:\n%s", out)
|
||||
}
|
||||
}
|
||||
|
||||
// An app that stops BETWEEN the start of the leg and the apply's own "before" reading still fails the run.
|
||||
// Measured live 2026-10-04 on demo-hp (privatebin stopped 1 s after the apply plan was written: the old rule
|
||||
// passed). Red-proof: use ap.HealthBefore alone as the baseline and this fails.
|
||||
func TestHealth_BaselineIsTheStartOfTheLeg(t *testing.T) {
|
||||
stoppedEarly := &Health{DockerOK: true, NetworkOK: true, Controller: "healthy", Containers: map[string]Container{
|
||||
"felhom-controller": {State: "running", Health: "healthy"}, "app": {State: "exited"}}}
|
||||
w := &fakeWrapper{t: t, pending: pend, applyRep: WrapperReport{Upgraded: []Package{{Name: "libc6"}},
|
||||
HealthBefore: stoppedEarly, HealthAfter: stoppedEarly},
|
||||
healthSeq: []*Health{stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly, stoppedEarly}}
|
||||
l, _ := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||
rep := l.Run(context.Background(), 9201, "night")
|
||||
if rep.Outcome != "health_failed" || !strings.Contains(rep.HealthReason, "app was running") {
|
||||
t.Fatalf("an app that stopped during the run passed: %+v", rep)
|
||||
}
|
||||
}
|
||||
+14
-2
@@ -10,6 +10,8 @@ import (
|
||||
"net/url"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||
)
|
||||
|
||||
// Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no
|
||||
@@ -31,6 +33,11 @@ type Config struct {
|
||||
Secret string // token secret (from <id>.pw)
|
||||
Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy.
|
||||
Timeout time.Duration
|
||||
|
||||
// IdleConnTimeout bounds how long this client's idle keep-alive connections are retained.
|
||||
// Zero means httpx.DefaultIdleConnTimeout (90s) — it does NOT mean "no timeout", which is the
|
||||
// R-344 defect. Production leaves it unset; only tests set it, to avoid a 90-second wait.
|
||||
IdleConnTimeout time.Duration
|
||||
}
|
||||
|
||||
// NewClient builds a fingerprint-pinned, token-authed PBS client.
|
||||
@@ -55,8 +62,13 @@ func NewClient(cfg Config) (*Client, error) {
|
||||
authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret,
|
||||
namespace: cfg.Namespace,
|
||||
http: &http.Client{
|
||||
Timeout: timeout,
|
||||
Transport: &http.Transport{TLSClientConfig: tlsCfg},
|
||||
Timeout: timeout,
|
||||
// R-344: this transport MUST come from httpx. pbsTargetsFromPVE (cmd/felhom-agent/
|
||||
// main.go) builds a fresh Client every cycle and drops the previous one, so a
|
||||
// transport with no idle timeout strands one connection per cycle, forever, on both
|
||||
// sides. That leaked 388 sockets onto ep0 in 46 hours. Pinned by
|
||||
// TestAbandonedClientsReleaseTheirConnections — do not inline an http.Transport here.
|
||||
Transport: httpx.NewTransport(tlsCfg, cfg.IdleConnTimeout),
|
||||
},
|
||||
}, nil
|
||||
}
|
||||
|
||||
@@ -0,0 +1,173 @@
|
||||
package pbs
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||
)
|
||||
|
||||
// connCounter is the SERVER-side observer. It counts what the server actually holds, which is the
|
||||
// only thing that answers the question this file exists for: a client that believes it closed a
|
||||
// connection, and a server still holding the socket, is precisely the R-344 shape. Asserting on
|
||||
// anything client-side would be asserting the mechanism instead of the consequence.
|
||||
type connCounter struct {
|
||||
mu sync.Mutex
|
||||
open int
|
||||
total int // every connection ever accepted — how many times the client DIALLED
|
||||
}
|
||||
|
||||
func (c *connCounter) hook(_ net.Conn, s http.ConnState) {
|
||||
c.mu.Lock()
|
||||
defer c.mu.Unlock()
|
||||
switch s {
|
||||
case http.StateNew:
|
||||
c.open++
|
||||
c.total++
|
||||
case http.StateClosed, http.StateHijacked:
|
||||
c.open--
|
||||
}
|
||||
}
|
||||
|
||||
func (c *connCounter) counts() (open, total int) {
|
||||
c.mu.Lock()
|
||||
defer c.mu.Unlock()
|
||||
return c.open, c.total
|
||||
}
|
||||
|
||||
// waitForOpen polls until the server holds want connections, or fails naming what it still holds.
|
||||
func (c *connCounter) waitForOpen(t *testing.T, want int, within time.Duration, what string) {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(within)
|
||||
for {
|
||||
open, total := c.counts()
|
||||
if open == want {
|
||||
return
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
t.Fatalf("%s: after %s the server still holds %d open connection(s), want %d (%d dialled in total)",
|
||||
what, within, open, want, total)
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
}
|
||||
|
||||
// newCountingPBSServer is newPBSTestServer plus a ConnState hook. Kept separate rather than
|
||||
// changing the shared helper, so the existing tests are untouched by this file.
|
||||
func newCountingPBSServer(t *testing.T, fn http.HandlerFunc) (*httptest.Server, string, *connCounter) {
|
||||
t.Helper()
|
||||
cc := &connCounter{}
|
||||
ts := httptest.NewUnstartedServer(fn)
|
||||
ts.Config.ConnState = cc.hook
|
||||
ts.StartTLS()
|
||||
t.Cleanup(ts.Close)
|
||||
return ts, fingerprintOf(ts), cc
|
||||
}
|
||||
|
||||
// TestAbandonedClientsReleaseTheirConnections is Scenario A, and it is the load-bearing test for
|
||||
// R-344.
|
||||
//
|
||||
// It models what pbsTargetsFromPVE actually does — build a client, use it once, drop it on the
|
||||
// floor without closing anything — and asserts the CONSEQUENCE on the server: the connections go
|
||||
// away. Before the fix every one of these stayed established forever on both sides; 388 of them
|
||||
// accumulated on ep0 in 46 hours.
|
||||
//
|
||||
// Deliberately NOT asserted: that err == nil, or that IdleConnTimeout holds some value. Both were
|
||||
// true of the leaking code.
|
||||
func TestAbandonedClientsReleaseTheirConnections(t *testing.T) {
|
||||
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
|
||||
w.Write([]byte(`{"data":[]}`))
|
||||
})
|
||||
host, port := hostPort(t, ts.URL)
|
||||
|
||||
const cycles = 5
|
||||
for i := 0; i < cycles; i++ {
|
||||
// One fresh client per "cycle", exactly as pbsTargetsFromPVE builds one per collect.
|
||||
c, err := NewClient(Config{
|
||||
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
|
||||
IdleConnTimeout: 50 * time.Millisecond, // production uses the 90s default
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
|
||||
t.Fatalf("cycle %d: %v", i, err)
|
||||
}
|
||||
_ = c // dropped here — nothing closes it, nothing can
|
||||
}
|
||||
|
||||
if _, total := cc.counts(); total != cycles {
|
||||
t.Fatalf("setup is not modelling the leak: want %d separate dials (one per abandoned client), got %d", cycles, total)
|
||||
}
|
||||
cc.waitForOpen(t, 0, 5*time.Second, "abandoned pbs.Clients")
|
||||
}
|
||||
|
||||
// TestAbandonedClientsReleaseTheirConnections_ProductionDefaultIsUsable pins the value that ships.
|
||||
//
|
||||
// The field being settable is exactly how it could silently become zero again — and zero used to
|
||||
// mean "never expire". This asserts the production path (Config leaving it unset) lands on the
|
||||
// standard-library default, so the leak cannot return through an unset field.
|
||||
func TestPBSClient_UnsetIdleTimeoutUsesTheDefault(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
cfg time.Duration
|
||||
want time.Duration
|
||||
}{
|
||||
{"unset — the production path", 0, httpx.DefaultIdleConnTimeout},
|
||||
{"explicit zero is NOT no-timeout", 0, httpx.DefaultIdleConnTimeout},
|
||||
{"negative is NOT no-timeout", -time.Second, httpx.DefaultIdleConnTimeout},
|
||||
{"an explicit value is honoured", 3 * time.Second, 3 * time.Second},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
c, err := NewClient(Config{
|
||||
Server: "pbs.example", Fingerprint: strings.Repeat("ab", 32),
|
||||
TokenID: "u@pbs!t", Secret: "s", IdleConnTimeout: tc.cfg,
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
tr, ok := c.http.Transport.(*http.Transport)
|
||||
if !ok {
|
||||
t.Fatalf("transport is %T, not *http.Transport — the httpx wiring was replaced", c.http.Transport)
|
||||
}
|
||||
if tr.IdleConnTimeout != tc.want {
|
||||
t.Fatalf("IdleConnTimeout = %v, want %v (zero would mean connections are retained FOREVER — that is R-344)", tr.IdleConnTimeout, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestPBSClient_KeepAliveStillReuses is Scenario C, and it is the guard against a "fix" that is
|
||||
// worse than the bug.
|
||||
//
|
||||
// Disabling keep-alive entirely would also make the leak go away — by dialling a fresh connection
|
||||
// for every single request, which on a box polling ~40,000 times a day is strictly worse than what
|
||||
// we started with. The fix must retire IDLE connections without stopping reuse.
|
||||
func TestPBSClient_KeepAliveStillReuses(t *testing.T) {
|
||||
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
|
||||
w.Write([]byte(`{"data":[]}`))
|
||||
})
|
||||
host, port := hostPort(t, ts.URL)
|
||||
|
||||
c, err := NewClient(Config{
|
||||
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
|
||||
IdleConnTimeout: 30 * time.Second, // long enough that reuse is what is being measured
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for i := 0; i < 3; i++ {
|
||||
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
|
||||
t.Fatalf("request %d: %v", i, err)
|
||||
}
|
||||
}
|
||||
if _, total := cc.counts(); total != 1 {
|
||||
t.Fatalf("one client made 3 sequential requests over %d connection(s), want 1 — keep-alive reuse is broken, which would make the poll load WORSE than the leak", total)
|
||||
}
|
||||
}
|
||||
@@ -10,6 +10,8 @@ import (
|
||||
"net/url"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||
)
|
||||
|
||||
// doer is the minimal HTTP surface the client needs; *http.Client satisfies it.
|
||||
@@ -65,8 +67,10 @@ func NewClient(cfg Config) (*Client, error) {
|
||||
timeout = 30 * time.Second
|
||||
}
|
||||
hc := &http.Client{
|
||||
Timeout: timeout,
|
||||
Transport: &http.Transport{TLSClientConfig: tlsCfg},
|
||||
Timeout: timeout,
|
||||
// R-344, consistency only: built ONCE per process, so it never accumulated and contributed
|
||||
// nothing to the ep0 leak. Same missing default, corrected for the same reason.
|
||||
Transport: httpx.NewTransport(tlsCfg, 0),
|
||||
}
|
||||
return &Client{
|
||||
base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json",
|
||||
|
||||
@@ -219,15 +219,18 @@ type Storage struct {
|
||||
UsedFraction float64 `json:"used_fraction,omitempty"`
|
||||
|
||||
// Type-specific config (durable_id sources).
|
||||
Server string `json:"server,omitempty"` // nfs/cifs/pbs server host
|
||||
Export string `json:"export,omitempty"` // nfs export path
|
||||
Share string `json:"share,omitempty"` // cifs share name
|
||||
Datastore string `json:"datastore,omitempty"` // pbs datastore name
|
||||
Fingerprint string `json:"fingerprint,omitempty"` // pbs server cert fingerprint
|
||||
Username string `json:"username,omitempty"` // pbs auth id, e.g. "felhom@pbs!n100"
|
||||
Namespace string `json:"namespace,omitempty"` // pbs namespace ("" = root; per-customer tenancy = S4)
|
||||
VGName string `json:"vgname,omitempty"` // lvm/lvmthin volume group
|
||||
ThinPool string `json:"thinpool,omitempty"` // lvmthin pool LV name
|
||||
Server string `json:"server,omitempty"` // nfs/cifs/pbs server host
|
||||
// EncryptionKey is the storage's own client-side key FINGERPRINT (pbs; the key itself stays in
|
||||
// /etc/pve/priv). R-727: an archive encrypted with any other key was written by another box.
|
||||
EncryptionKey string `json:"encryption-key,omitempty"`
|
||||
Export string `json:"export,omitempty"` // nfs export path
|
||||
Share string `json:"share,omitempty"` // cifs share name
|
||||
Datastore string `json:"datastore,omitempty"` // pbs datastore name
|
||||
Fingerprint string `json:"fingerprint,omitempty"` // pbs server cert fingerprint
|
||||
Username string `json:"username,omitempty"` // pbs auth id, e.g. "felhom@pbs!n100"
|
||||
Namespace string `json:"namespace,omitempty"` // pbs namespace ("" = root; per-customer tenancy = S4)
|
||||
VGName string `json:"vgname,omitempty"` // lvm/lvmthin volume group
|
||||
ThinPool string `json:"thinpool,omitempty"` // lvmthin pool LV name
|
||||
}
|
||||
|
||||
// StorageContent is one entry of GET /nodes/{node}/storage/{store}/content
|
||||
@@ -239,4 +242,7 @@ type StorageContent struct {
|
||||
Size int64 `json:"size"`
|
||||
CTime int64 `json:"ctime"`
|
||||
VMID int `json:"vmid,omitempty"`
|
||||
// Encrypted is the fingerprint of the key a PBS archive was encrypted with ("" = not encrypted). R-727:
|
||||
// the restore test reads it to tell THIS box's archives from an earlier box's in the same namespace.
|
||||
Encrypted string `json:"encrypted,omitempty"`
|
||||
}
|
||||
|
||||
@@ -302,6 +302,19 @@ func (e *Engine) runBringUp(ctx context.Context, spec BringUpSpec, res *BringUpR
|
||||
}
|
||||
}
|
||||
|
||||
// R-834: DR keeps the archive's `onboot: 1`, binds the host's REAL drives (4d) and STARTS the guest
|
||||
// — right on a replaced host, where the original is gone. Beside a LIVE original it would be a
|
||||
// second controller for the same household on the same drives. So DR refuses when this host
|
||||
// still carries the original (the archive's source VMID) or any guest that binds the drives.
|
||||
// A copy beside the original is the restore-test's job (onboot=0, throwaway stand-ins, torn
|
||||
// down) or the runbook's beside-restore. Pinned by TestRunBringUp_DRRefusesBesideALiveOriginal.
|
||||
if spec.Mode == ModeDRGuestLoss {
|
||||
if why := e.liveOriginalBeside(ctx, lxc, spec.Archive); why != "" {
|
||||
res.Err = fmt.Errorf("reconcile: dr bring-up refused: %s — a DR restore beside a live original would run two boxes on the same drives (R-834)", why)
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
base := JournalEntry{OpID: e.bringUpOpID(spec.VMID), VMID: spec.VMID, Kind: bringUpKind, Rollback: true}
|
||||
|
||||
// OWN the rollback BEFORE any mutation. From here a crash leaves an in-flight Rollback
|
||||
@@ -734,3 +747,47 @@ func net0MAC(cfg proxmox.GuestConfig) string {
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// archiveSourceVMID reads the source guest's VMID from a backup volid: a vzdump file
|
||||
// (`…/vzdump-lxc-<vmid>-<date>.tar.zst`) or a PBS snapshot (`…:backup/ct/<vmid>/<time>`). 0 = unknown.
|
||||
func archiveSourceVMID(archive string) int {
|
||||
if i := strings.Index(archive, "vzdump-lxc-"); i >= 0 {
|
||||
rest := archive[i+len("vzdump-lxc-"):]
|
||||
if j := strings.Index(rest, "-"); j > 0 {
|
||||
if n, err := strconv.Atoi(rest[:j]); err == nil {
|
||||
return n
|
||||
}
|
||||
}
|
||||
}
|
||||
if i := strings.Index(archive, "ct/"); i >= 0 {
|
||||
rest := archive[i+len("ct/"):]
|
||||
if j := strings.Index(rest, "/"); j > 0 {
|
||||
if n, err := strconv.Atoi(rest[:j]); err == nil {
|
||||
return n
|
||||
}
|
||||
}
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// liveOriginalBeside says why a DR bring-up would land beside a live original ("" = it would not):
|
||||
// the archive's source guest still exists here, or a guest binds the drives parent. Fails CLOSED: a
|
||||
// guest whose config cannot be read cannot be ruled out.
|
||||
func (e *Engine) liveOriginalBeside(ctx context.Context, lxc []proxmox.Guest, archive string) string {
|
||||
src := archiveSourceVMID(archive)
|
||||
for _, g := range lxc {
|
||||
if src > 0 && g.VMID == src {
|
||||
return fmt.Sprintf("the archive's source guest %d still exists on this host (status %s)", g.VMID, g.Status)
|
||||
}
|
||||
cfg, err := e.api.GuestConfig(ctx, g.VMID)
|
||||
if err != nil {
|
||||
return fmt.Sprintf("guest %d's config could not be read to rule out a live original: %v", g.VMID, err)
|
||||
}
|
||||
for slot, v := range cfg.MountPoints() {
|
||||
if source, _, _ := strings.Cut(v, ","); source == structuralParentDir {
|
||||
return fmt.Sprintf("guest %d binds the household drives (%s %s)", g.VMID, slot, structuralParentDir)
|
||||
}
|
||||
}
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
package reconcile
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-834: a whole-guest restore BESIDE a live original must never come up as a second box on the same
|
||||
// drives. The DR route keeps `onboot: 1`, binds the real drives and starts the guest, so it refuses
|
||||
// when the original (or any guest binding the drives) is still on this host; on a replaced host it
|
||||
// proceeds and keeps its binds. Red-proof: make liveOriginalBeside return "" and the refusals pass
|
||||
// the restore through (the "refused" sub-tests fail).
|
||||
func TestRunBringUp_DRRefusesBesideALiveOriginal(t *testing.T) {
|
||||
const target = 9299
|
||||
drivesBind := proxmox.GuestConfig{Extra: map[string]json.RawMessage{
|
||||
"mp8": json.RawMessage(`"/mnt/felhom-drives,mp=/mnt/felhom-drives"`),
|
||||
}}
|
||||
cases := []struct {
|
||||
name string
|
||||
archive string
|
||||
lxc []proxmox.Guest
|
||||
cfg map[int]proxmox.GuestConfig
|
||||
refuse string // substring of the refusal; "" = must proceed
|
||||
}{
|
||||
{"the source guest still exists", "local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst",
|
||||
[]proxmox.Guest{{VMID: 9201, Status: "running"}}, map[int]proxmox.GuestConfig{9201: scratchCfg()}, "source guest 9201"},
|
||||
{"another guest binds the drives (PBS archive)", "felhom-pbs:backup/ct/9201/2026-10-04T02:34:55Z",
|
||||
[]proxmox.Guest{{VMID: 9300, Status: "stopped"}}, map[int]proxmox.GuestConfig{9300: drivesBind}, "binds the household drives"},
|
||||
{"a guest whose config cannot be read", "local:backup/vzdump-lxc-9201-x.tar.zst",
|
||||
[]proxmox.Guest{{VMID: 9400, Status: "running"}}, map[int]proxmox.GuestConfig{}, "could not be read"},
|
||||
{"replaced host: only an unrelated scratch guest", "local:backup/vzdump-lxc-9201-x.tar.zst",
|
||||
[]proxmox.Guest{{VMID: 9202, Status: "running"}}, map[int]proxmox.GuestConfig{9202: scratchCfg()}, ""},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
cfg := c.cfg
|
||||
cfg[target] = scratchCfg()
|
||||
api := &fakeAPI{lxc: c.lxc, cfg: cfg}
|
||||
e, fr, _, q := newDREngine(t, api)
|
||||
defer q.Close()
|
||||
res := e.RunBringUp(context.Background(), BringUpSpec{
|
||||
Mode: ModeDRGuestLoss, Archive: c.archive, VMID: target, RestoreStorage: "local-lvm", KeepMAC: true,
|
||||
})
|
||||
if c.refuse != "" {
|
||||
if res.Err == nil || !strings.Contains(res.Err.Error(), c.refuse) || !strings.Contains(res.Err.Error(), "R-834") {
|
||||
t.Fatalf("want a refusal naming %q, got %+v", c.refuse, res)
|
||||
}
|
||||
if len(api.restores) != 0 || len(api.starts) != 0 || len(fr.cmds) != 0 {
|
||||
t.Fatalf("a refused DR touched the host: restores=%d starts=%v cmds=%v", len(api.restores), api.starts, fr.cmds)
|
||||
}
|
||||
return
|
||||
}
|
||||
if res.Err != nil || !res.Pass {
|
||||
t.Fatalf("a DR on a replaced host must proceed, got %+v", res)
|
||||
}
|
||||
// … and there it keeps the REAL drives bind (the right binds on a replaced host).
|
||||
joined := strings.Join(fr.cmds, "\n")
|
||||
if !strings.Contains(joined, "-mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives") {
|
||||
t.Fatalf("the DR guest lost its drives bind: %v", fr.cmds)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// Provisioning restores the GOLDEN (no drives, onboot set by the back-half on purpose): a drives-
|
||||
// binding guest on the host does not block it — the rule is DR's alone.
|
||||
func TestRunBringUp_ProvisionNotBlockedByADrivesBind(t *testing.T) {
|
||||
api := &fakeAPI{
|
||||
lxc: []proxmox.Guest{{VMID: 9201, Status: "running"}},
|
||||
cfg: map[int]proxmox.GuestConfig{
|
||||
9201: {Extra: map[string]json.RawMessage{"mp8": json.RawMessage(`"/mnt/felhom-drives,mp=/mnt/felhom-drives"`)}},
|
||||
9203: scratchCfg(),
|
||||
},
|
||||
}
|
||||
e, _, q := newEngine(t, api, EmptyProvider{})
|
||||
defer q.Close()
|
||||
res := e.RunBringUp(context.Background(), BringUpSpec{Mode: ModeProvision, Archive: "local:vztmpl/felhom-golden.tar.zst", VMID: 9203, RestoreStorage: "local-lvm"})
|
||||
if res.Err != nil || !res.Pass {
|
||||
t.Fatalf("provision must proceed, got %+v", res)
|
||||
}
|
||||
}
|
||||
|
||||
func TestArchiveSourceVMID(t *testing.T) {
|
||||
for in, want := range map[string]int{
|
||||
"local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst": 9201,
|
||||
"felhom-pbs:backup/ct/9201/2026-10-04T02:34:55Z": 9201,
|
||||
"tmp-dooplex-copy:backup/ct/9201/2026-10-03T19:00:00Z": 9201,
|
||||
"local:vztmpl/felhom-golden.tar.zst": 0,
|
||||
"vol": 0,
|
||||
} {
|
||||
if got := archiveSourceVMID(in); got != want {
|
||||
t.Errorf("archiveSourceVMID(%q) = %d, want %d", in, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// R-834, the restore-test route: its scratch guest sits BESIDE the live original by design, so it must
|
||||
// carry no host-path bind — the archive's mp8 (the household's drives) and mp9 (the original's
|
||||
// bootstrap) are replaced by throwaway volumes AT restore time. Measured live 2026-10-04 on demo-hp
|
||||
// (`audits/backup-close-2026-10-04/partA/`): onboot 0 and no host bind on every poll. Red-proof:
|
||||
// make drRestoreOverrides return the archive's own mp8 value and this fails.
|
||||
func TestRestoreTest_NoHostPathBindBesideTheOriginal(t *testing.T) {
|
||||
api := &fakeAPI{
|
||||
cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()},
|
||||
extractCfg: "hostname: demo-hp\nonboot: 1\nrootfs: local-lvm:vm-9201-disk-0,size=16G\n" +
|
||||
"mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G\n" +
|
||||
"mp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\n" +
|
||||
"mp9: /var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1\n",
|
||||
}
|
||||
e, _, q := newEngine(t, api, EmptyProvider{})
|
||||
defer q.Close()
|
||||
_ = e.RunRestoreTest(context.Background(), RestoreTestSpec{
|
||||
Archive: "local:backup/vzdump-lxc-9201-x.tar.zst", RestoreStorage: "local-lvm",
|
||||
ScratchMin: 990000, ScratchMax: 990009, SourceTier: "local",
|
||||
})
|
||||
if len(api.restores) != 1 {
|
||||
t.Fatalf("want one restore, got %+v", api.restores)
|
||||
}
|
||||
r := api.restores[0]
|
||||
for _, slot := range []string{"mp8", "mp9"} {
|
||||
v, ok := r.MountOverrides[slot]
|
||||
if !ok || strings.HasPrefix(v, "/") {
|
||||
t.Fatalf("%s = %q (present=%v): the scratch beside the original must get a throwaway volume, never the host path", slot, v, ok)
|
||||
}
|
||||
}
|
||||
for slot, v := range r.MountOverrides {
|
||||
if strings.HasPrefix(v, "/") {
|
||||
t.Fatalf("%s carries a host path %q into the scratch guest", slot, v)
|
||||
}
|
||||
}
|
||||
if r.ConfigOverrides["onboot"] != "0" {
|
||||
t.Fatalf("onboot = %q, want 0", r.ConfigOverrides["onboot"])
|
||||
}
|
||||
}
|
||||
@@ -52,7 +52,7 @@ func newDREngine(t *testing.T, api GuestAPI) (*Engine, *fakeRunner, string, *Que
|
||||
t.Cleanup(q.Close)
|
||||
fr := &fakeRunner{}
|
||||
sd := t.TempDir()
|
||||
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: EmptyProvider{}, HostRunner: fr, StateDir: sd})
|
||||
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: EmptyProvider{}, HostRunner: fr, StateDir: sd, RestoreSpace: roomySpace{}})
|
||||
return e, fr, sd, q
|
||||
}
|
||||
|
||||
|
||||
@@ -44,6 +44,17 @@ type Engine struct {
|
||||
|
||||
opSeq uint64 // atomic; makes each op id unique per attempt
|
||||
|
||||
// restoreSpace + spacePolicy are the restore-test's space preflight (R-672). nil space REFUSES
|
||||
// every restore-test (fail-closed) — see restoretest_space.go.
|
||||
restoreSpace RestoreSpace
|
||||
spacePolicy SpacePolicy
|
||||
|
||||
// scratchMu guards activeScratch (the vmids a running restore-test owns — the retry timer never
|
||||
// touches those) and teardownTries (failed timer retries per journal op, R-672 rule 3).
|
||||
scratchMu sync.Mutex
|
||||
activeScratch map[int]bool
|
||||
teardownTries map[string]int
|
||||
|
||||
// lastRes records the most recent successful Reconcile Result (v0.90.0, R-28 fast-tick source).
|
||||
// The fast-tick reads it to decide convergence: actionable drift is Planned − Pending > 0 (a
|
||||
// destructive pending_signature refusal is EXPECTED state, not drift to hammer on). lastOK is
|
||||
@@ -86,6 +97,10 @@ type EngineOptions struct {
|
||||
HostRunner proxmox.Runner
|
||||
// StateDir is the agent state dir ("" → /var/lib/felhom-agent); only the 4d swap reads it.
|
||||
StateDir string
|
||||
// RestoreSpace is the restore-test's space preflight (R-672). nil → every restore-test is REFUSED
|
||||
// with its reason (fail-closed). SpacePolicy zero → DefaultSpacePolicy.
|
||||
RestoreSpace RestoreSpace
|
||||
SpacePolicy SpacePolicy
|
||||
}
|
||||
|
||||
// NewEngine builds an Engine. The Queue is shared (the single §10 choke point); the
|
||||
@@ -124,9 +139,24 @@ func NewEngine(opts EngineOptions) *Engine {
|
||||
logger: logger,
|
||||
hostRun: opts.HostRunner,
|
||||
stateDir: stateDir,
|
||||
|
||||
restoreSpace: opts.RestoreSpace,
|
||||
spacePolicy: policyOrDefault(opts.SpacePolicy),
|
||||
activeScratch: map[int]bool{},
|
||||
teardownTries: map[string]int{},
|
||||
}
|
||||
}
|
||||
|
||||
func policyOrDefault(p SpacePolicy) SpacePolicy {
|
||||
if p.Factor < 1 {
|
||||
p.Factor = DefaultSpacePolicy.Factor
|
||||
}
|
||||
if p.ReserveBytes <= 0 {
|
||||
p.ReserveBytes = DefaultSpacePolicy.ReserveBytes
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// Result summarizes one Reconcile pass.
|
||||
type Result struct {
|
||||
Planned int
|
||||
|
||||
@@ -213,7 +213,7 @@ func newEngine(t *testing.T, api GuestAPI, provider DesiredProvider) (*Engine, *
|
||||
t.Cleanup(func() { j.Close() })
|
||||
q := NewQueue()
|
||||
t.Cleanup(q.Close)
|
||||
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: provider})
|
||||
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: provider, RestoreSpace: roomySpace{}})
|
||||
return e, j, q
|
||||
}
|
||||
|
||||
|
||||
@@ -50,10 +50,18 @@ type RestoreTestResult struct {
|
||||
ScratchVMID int
|
||||
Pass bool
|
||||
Verified string // "boot+running" this slice
|
||||
Skipped bool // no free scratch VMID in band → test not run
|
||||
Err error
|
||||
StartedAt time.Time
|
||||
Duration time.Duration
|
||||
Skipped bool // test not run: no free scratch VMID in band, or the space preflight refused (R-672)
|
||||
// SkipReason is set when the SPACE PREFLIGHT refused (R-672): the test did not run, and this is
|
||||
// reported to the hub as the test's result (pass=false), never as a pass. Empty for a band skip.
|
||||
SkipReason string
|
||||
// TargetStorage is where the restore went (rule 2 may move it off the tested guest's pool);
|
||||
// RequiredBytes/AvailBytes are rule 1's figures.
|
||||
TargetStorage string
|
||||
RequiredBytes int64
|
||||
AvailBytes int64
|
||||
Err error
|
||||
StartedAt time.Time
|
||||
Duration time.Duration
|
||||
// StartWarnings holds the warning line(s) the guest-start task emitted (e.g. the
|
||||
// systemd-nesting advisory). Populated only when the start exited "WARNINGS: N";
|
||||
// always surfaced, NEVER used to decide pass/fail (the verdict is liveness — waitRunning).
|
||||
@@ -136,6 +144,29 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
|
||||
return res
|
||||
}
|
||||
|
||||
// R-672: the space preflight, BEFORE anything is journaled or created.
|
||||
rawCfg, err := e.api.ExtractArchiveConfig(ctx, spec.Archive)
|
||||
if err != nil {
|
||||
res.Err = fmt.Errorf("reconcile: restore-test extract archive config: %w", err)
|
||||
return res
|
||||
}
|
||||
v := PreflightRestoreSpace(ctx, e.restoreSpace, e.spacePolicy, spec.Archive, rawCfg, spec.RestoreStorage)
|
||||
res.TargetStorage, res.RequiredBytes, res.AvailBytes = v.Storage, v.Required, v.Avail
|
||||
if !v.OK {
|
||||
res.Skipped = true
|
||||
res.SkipReason = "skipped: " + v.Reason
|
||||
e.logger.Warn("restore-test SKIPPED by the space preflight (R-672) — nothing was created",
|
||||
"archive", spec.Archive, "storage", v.Storage, "required_bytes", v.Required, "avail_bytes", v.Avail, "reason", v.Reason)
|
||||
res.Duration = time.Since(now)
|
||||
return res
|
||||
}
|
||||
if v.Storage != spec.RestoreStorage {
|
||||
e.logger.Info("restore-test: restoring OFF the tested guest's own pool (R-672 rule 2)",
|
||||
"configured", spec.RestoreStorage, "avoided", v.Avoided, "target", v.Storage)
|
||||
}
|
||||
e.logger.Info("restore-test: space preflight passed", "storage", v.Storage, "required_bytes", v.Required, "avail_bytes", v.Avail)
|
||||
spec.RestoreStorage = v.Storage
|
||||
|
||||
lxc, err := e.api.ListLXC(ctx)
|
||||
if err != nil {
|
||||
res.Err = fmt.Errorf("reconcile: restore-test list guests: %w", err)
|
||||
@@ -164,7 +195,9 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
|
||||
// Serialize on the scratch VMID's lane (inherits §10), and capture the result.
|
||||
var vmidOccupied bool
|
||||
ch := e.queue.Submit(vmid, func() error {
|
||||
vmidOccupied = e.runScratchTest(ctx, vmid, spec, &res)
|
||||
e.markScratch(vmid, true)
|
||||
defer e.markScratch(vmid, false)
|
||||
vmidOccupied = e.runScratchTest(ctx, vmid, spec, rawCfg, &res)
|
||||
return res.Err
|
||||
})
|
||||
<-ch
|
||||
@@ -182,7 +215,7 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
|
||||
// runScratchTest is the journaled body (runs on vmid's queue lane). The occupied return is true
|
||||
// ONLY when PVE synchronously refused the restore because the vmid already holds a guest (one
|
||||
// the pool-blind band scan couldn't see) — the caller then advances to the next band vmid (F2).
|
||||
func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestSpec, res *RestoreTestResult) (occupied bool) {
|
||||
func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestSpec, rawCfg string, res *RestoreTestResult) (occupied bool) {
|
||||
base := JournalEntry{OpID: e.scratchOpID(vmid), VMID: vmid, Kind: scratchKind, Scratch: true}
|
||||
|
||||
// OWN the scratch guest's cleanup BEFORE any mutation. From here, a crash is recoverable.
|
||||
@@ -215,11 +248,7 @@ func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestS
|
||||
// genuinely EXTRACTED — full fidelity; the added runtime IS the verification), the two
|
||||
// structural binds → throwaway stand-ins. An unreadable archive config or an unknown
|
||||
// topology REFUSES up front — never restore a partial guest to "verify" it.
|
||||
rawCfg, err := e.api.ExtractArchiveConfig(ctx, spec.Archive)
|
||||
if err != nil {
|
||||
res.Err = fmt.Errorf("reconcile: restore-test extract archive config: %w", err)
|
||||
return false
|
||||
}
|
||||
// The archive's config was read ONCE, by the space preflight (R-672), and is passed in.
|
||||
mountOverrides, err := drRestoreOverrides(rawCfg, spec.RestoreStorage)
|
||||
if err != nil {
|
||||
res.Err = fmt.Errorf("reconcile: restore-test: %w", err)
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
package reconcile
|
||||
|
||||
import "context"
|
||||
|
||||
// ── A failed scratch teardown is retried on a TIMER, not only at agent start (R-672 rule 3) ────────
|
||||
//
|
||||
// MEASURED 2026-09-24 on demo-hp: the scheduled restore-test's teardown failed (`lvremove … contains a
|
||||
// filesystem in use`, a transient hold) and logged "left for Recover" — and Recover runs ONLY at agent
|
||||
// start, so the 22 GiB scratch guest sat in the full pool for 2.5 hours until an agent restart. The
|
||||
// timer calls RetryScratchTeardown every 10 minutes: the SAME resolution as Recover (recoverScratch —
|
||||
// the gate's benign scratch destroy, idempotent when the guest is already gone), restricted to Scratch
|
||||
// entries that carry a launch-proof UPID and that no running restore-test owns. After
|
||||
// MaxTeardownTries failed attempts for one entry the operator is told (the caller reports it); the
|
||||
// timer keeps trying.
|
||||
|
||||
// MaxTeardownTries is how many failed timer retries of one scratch entry happen before the operator
|
||||
// is told.
|
||||
const MaxTeardownTries = 3
|
||||
|
||||
// ScratchRetryResult summarizes one timer pass.
|
||||
type ScratchRetryResult struct {
|
||||
Examined int
|
||||
Destroyed int
|
||||
Clean int // already gone
|
||||
Failed int
|
||||
// GaveUp lists the scratch vmids whose failed tries reached MaxTeardownTries IN THIS PASS — each
|
||||
// is reported exactly once (the caller tells the operator).
|
||||
GaveUp []int
|
||||
}
|
||||
|
||||
func (e *Engine) markScratch(vmid int, active bool) {
|
||||
e.scratchMu.Lock()
|
||||
defer e.scratchMu.Unlock()
|
||||
if active {
|
||||
e.activeScratch[vmid] = true
|
||||
} else {
|
||||
delete(e.activeScratch, vmid)
|
||||
}
|
||||
}
|
||||
|
||||
func (e *Engine) scratchActive(vmid int) bool {
|
||||
e.scratchMu.Lock()
|
||||
defer e.scratchMu.Unlock()
|
||||
return e.activeScratch[vmid]
|
||||
}
|
||||
|
||||
// RetryScratchTeardown is the timer's pass. It never touches a non-Scratch entry (unlike Recover,
|
||||
// which also resolves generic in-flight operations and must therefore run only at start), never an
|
||||
// entry without a launch-proof UPID (nothing was created), and never a vmid a running test owns.
|
||||
func (e *Engine) RetryScratchTeardown(ctx context.Context) ScratchRetryResult {
|
||||
var out ScratchRetryResult
|
||||
if e.journal == nil {
|
||||
return out
|
||||
}
|
||||
for _, entry := range e.journal.InFlight() {
|
||||
if !entry.Scratch || entry.UPID == "" || e.scratchActive(entry.VMID) {
|
||||
continue
|
||||
}
|
||||
out.Examined++
|
||||
var r RecoverResult
|
||||
e.recoverScratch(ctx, entry, &r)
|
||||
switch {
|
||||
case r.ScratchDestroyed > 0:
|
||||
out.Destroyed++
|
||||
e.forgetTries(entry.OpID)
|
||||
case r.ScratchClean > 0:
|
||||
out.Clean++
|
||||
e.forgetTries(entry.OpID)
|
||||
default:
|
||||
out.Failed++
|
||||
n := e.addTry(entry.OpID)
|
||||
e.logger.Warn("restore-test: scratch teardown retry failed (timer)", "vmid", entry.VMID, "op_id", entry.OpID, "try", n)
|
||||
if n == MaxTeardownTries {
|
||||
out.GaveUp = append(out.GaveUp, entry.VMID)
|
||||
e.logger.Error("restore-test: scratch guest still NOT torn down after repeated retries — telling the operator",
|
||||
"vmid", entry.VMID, "tries", n)
|
||||
}
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func (e *Engine) addTry(op string) int {
|
||||
e.scratchMu.Lock()
|
||||
defer e.scratchMu.Unlock()
|
||||
e.teardownTries[op]++
|
||||
return e.teardownTries[op]
|
||||
}
|
||||
|
||||
func (e *Engine) forgetTries(op string) {
|
||||
e.scratchMu.Lock()
|
||||
defer e.scratchMu.Unlock()
|
||||
delete(e.teardownTries, op)
|
||||
}
|
||||
@@ -0,0 +1,73 @@
|
||||
package reconcile
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-672 rule 3 (v0.133.0): a failed scratch teardown is retried on a TIMER. The consequence asserted:
|
||||
// the leaked scratch guest is destroyed by a timer pass (not only by a restart's Recover), the operator
|
||||
// is told exactly once after MaxTeardownTries failures, and a vmid a running test owns is never touched.
|
||||
|
||||
// leakScratch runs a restore-test whose teardown fails, leaving scratch 990000 in-flight — the
|
||||
// 2026-09-24 shape ("lvremove … contains a filesystem in use").
|
||||
func leakScratch(t *testing.T) (*Engine, *fakeAPI, *Journal) {
|
||||
t.Helper()
|
||||
api := &fakeAPI{cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}, restoreUPID: "UPID:r", destroyErr: errors.New("lvremove: contains a filesystem in use")}
|
||||
e, j := spaceEngine(t, api, roomySpace{})
|
||||
e.RunRestoreTest(context.Background(), RestoreTestSpec{Archive: "local:backup/x.tar.zst", RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009})
|
||||
if len(j.InFlight()) != 1 {
|
||||
t.Fatalf("setup: want the scratch left in-flight after a failed teardown, got %+v", j.InFlight())
|
||||
}
|
||||
api.lxc = []proxmox.Guest{{VMID: 990000}}
|
||||
return e, api, j
|
||||
}
|
||||
|
||||
// COMPANION RED-PROOF (REPORT): make RetryScratchTeardown return without touching the journal (the
|
||||
// v0.132.0 shape — only Recover at start resolved a leak) → "the leaked scratch was not destroyed by
|
||||
// the timer".
|
||||
func TestRetry_TheTimerDestroysALeakedScratch(t *testing.T) {
|
||||
e, api, j := leakScratch(t)
|
||||
api.destroyErr = nil // the transient hold is gone
|
||||
before := len(api.destroys)
|
||||
r := e.RetryScratchTeardown(context.Background())
|
||||
if r.Destroyed != 1 || len(api.destroys) != before+1 || api.destroys[len(api.destroys)-1] != 990000 {
|
||||
t.Fatalf("the leaked scratch was not destroyed by the timer: result=%+v destroys=%v", r, api.destroys)
|
||||
}
|
||||
if len(j.InFlight()) != 0 {
|
||||
t.Fatalf("the entry is still in flight after a successful retry: %+v", j.InFlight())
|
||||
}
|
||||
if r2 := e.RetryScratchTeardown(context.Background()); r2.Examined != 0 {
|
||||
t.Fatalf("a resolved entry was examined again: %+v", r2)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRetry_OperatorToldOnceAfterThreeFailures(t *testing.T) {
|
||||
e, _, _ := leakScratch(t)
|
||||
var gave [][]int
|
||||
for i := 0; i < MaxTeardownTries+2; i++ {
|
||||
gave = append(gave, e.RetryScratchTeardown(context.Background()).GaveUp)
|
||||
}
|
||||
for i, g := range gave {
|
||||
want := 0
|
||||
if i == MaxTeardownTries-1 {
|
||||
want = 1
|
||||
}
|
||||
if len(g) != want {
|
||||
t.Fatalf("pass %d gave up on %v — want the operator told exactly once, on pass %d", i+1, g, MaxTeardownTries)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRetry_NeverTouchesARunningTest(t *testing.T) {
|
||||
e, api, _ := leakScratch(t)
|
||||
api.destroyErr = nil
|
||||
e.markScratch(990000, true) // a restore-test is (again) working on this vmid
|
||||
before := len(api.destroys)
|
||||
if r := e.RetryScratchTeardown(context.Background()); r.Examined != 0 || len(api.destroys) != before {
|
||||
t.Fatalf("the timer touched a scratch a running test owns: %+v destroys=%v", r, api.destroys)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,176 @@
|
||||
package reconcile
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"sort"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// ── The restore-test's space preflight (R-672, agent v0.133.0) ─────────────────────────────────────
|
||||
//
|
||||
// MEASURED 2026-09-24 on demo-hp: the scheduled restore-test restored 9201's archive into `local-lvm`
|
||||
// — the SAME thin pool that holds 9201 — with no free-space check. The pool reached 100 %
|
||||
// (`out_of_data_space`, `error_if_no_space`), and 9201's rootfs and data volume remounted READ-ONLY.
|
||||
// Evidence: felhom.eu `documentation/audits/night-2026-09-24/C-02…C-07`, `audits/r672-2026-09-24/`.
|
||||
//
|
||||
// THE THREE RULES, all decided here before ANY mutation (before the scratch entry is even journaled):
|
||||
// 1. SPACE FIRST. The target storage must have free data ≥ restored × factor + reserve (defaults 1.2 and
|
||||
// 5 GiB, `backup.restore_test_space_factor` / `backup.restore_test_space_reserve_gib`), and a thin
|
||||
// pool's metadata must have room for the same share. `restored` is the UNCOMPRESSED size — the
|
||||
// archive FILE is the wrong number: 9201's archive was 6.9 GB and its restore wrote 22.6 GB, so
|
||||
// "file × 1.2 + 5 GiB" (14.3 GB) would have let the 2026-09-24 test run into a pool with 23 GB free.
|
||||
// 2. KEEP OFF THE TESTED GUEST'S POOL when another eligible storage (active, takes `rootdir`, and the
|
||||
// agent holds Datastore.AllocateSpace on it) passes rule 1. With only one, rule 1 decides.
|
||||
// 3. UNKNOWN REFUSES. An unreadable size, an unreadable storage or an unknown thin-pool metadata fill is
|
||||
// a skip with its reason, never a guess — the fail-safe direction of every guard in this project.
|
||||
// A refusal is reported to the hub as the test's RESULT ("skipped: …", pass=false), never as a pass.
|
||||
// Pinned by restoretest_space_test.go.
|
||||
|
||||
// RestoreSpace is the preflight's seam onto the host. Production: internal/restorespace.
|
||||
type RestoreSpace interface {
|
||||
// RestoredBytes is how many bytes restoring `archive` will write (uncompressed), and where that
|
||||
// figure came from (for the log and the refusal).
|
||||
RestoredBytes(ctx context.Context, archive string) (bytes int64, source string, err error)
|
||||
// Free reports the storage's free data bytes and, for a thin pool, its metadata-used fraction.
|
||||
Free(ctx context.Context, storage string) (StorageFree, error)
|
||||
// Eligible lists the storages a restore-test may target: active, content `rootdir`, and the agent
|
||||
// holds Datastore.AllocateSpace there.
|
||||
Eligible(ctx context.Context) ([]string, error)
|
||||
}
|
||||
|
||||
// StorageFree is one storage's free space as the preflight judges it.
|
||||
type StorageFree struct {
|
||||
AvailBytes int64
|
||||
UsedBytes int64
|
||||
Thin bool
|
||||
// MetaUsedFraction is the thin pool's metadata use (0..1); MetaKnown false = could not be read.
|
||||
MetaUsedFraction float64
|
||||
MetaKnown bool
|
||||
}
|
||||
|
||||
// SpacePolicy is rule 1's margin.
|
||||
type SpacePolicy struct {
|
||||
Factor float64 // ≥ 1
|
||||
ReserveBytes int64
|
||||
}
|
||||
|
||||
// DefaultSpacePolicy is 1.2 × restored + 5 GiB.
|
||||
var DefaultSpacePolicy = SpacePolicy{Factor: 1.2, ReserveBytes: 5 << 30}
|
||||
|
||||
// SpaceVerdict is the preflight's answer.
|
||||
type SpaceVerdict struct {
|
||||
OK bool
|
||||
Storage string // the storage the restore goes to (when OK) or was judged (when not)
|
||||
Required int64
|
||||
Avail int64
|
||||
Reason string // empty when OK
|
||||
// Avoided is the tested guest's own storage, when rule 2 moved the restore off it.
|
||||
Avoided string
|
||||
}
|
||||
|
||||
// requiredBytes is rule 1's figure.
|
||||
func (p SpacePolicy) requiredBytes(restored int64) int64 {
|
||||
f := p.Factor
|
||||
if f < 1 {
|
||||
f = DefaultSpacePolicy.Factor
|
||||
}
|
||||
return int64(float64(restored)*f) + p.ReserveBytes
|
||||
}
|
||||
|
||||
// fits judges one storage against rule 1 (data AND thin metadata). An unknown metadata fill on a thin
|
||||
// pool refuses (rule 3).
|
||||
func fits(fr StorageFree, required int64) (bool, string) {
|
||||
if fr.AvailBytes < required {
|
||||
return false, fmt.Sprintf("needs %s free, has %s", gib(required), gib(fr.AvailBytes))
|
||||
}
|
||||
if fr.Thin {
|
||||
if !fr.MetaKnown {
|
||||
return false, "thin-pool metadata fill unknown"
|
||||
}
|
||||
// The metadata a restore of `required` bytes needs, in the pool's own proportion of metadata to
|
||||
// data. A pool with no data yet has no proportion to read → only the absolute ceiling applies.
|
||||
need := 0.0
|
||||
if fr.UsedBytes > 0 {
|
||||
need = fr.MetaUsedFraction * float64(required) / float64(fr.UsedBytes)
|
||||
}
|
||||
if fr.MetaUsedFraction+need > 0.9 {
|
||||
return false, fmt.Sprintf("thin-pool metadata would reach %.0f%% (now %.0f%%)", 100*(fr.MetaUsedFraction+need), 100*fr.MetaUsedFraction)
|
||||
}
|
||||
}
|
||||
return true, ""
|
||||
}
|
||||
|
||||
func gib(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
|
||||
|
||||
// sourceStorages returns the storage ids that hold the ARCHIVED guest's volumes (rootfs and every mpN
|
||||
// that names a `storage:volume`), read from the archive's own embedded config — the guest under test.
|
||||
// Bind mounts (a leading "/") carry no storage.
|
||||
func sourceStorages(rawCfg string) map[string]bool {
|
||||
out := map[string]bool{}
|
||||
for k, v := range archiveCurrentConfig(rawCfg) {
|
||||
if k != "rootfs" && !(strings.HasPrefix(k, "mp") && len(k) > 2 && k[2] >= '0' && k[2] <= '9') {
|
||||
continue
|
||||
}
|
||||
vol := strings.TrimSpace(strings.SplitN(strings.TrimSpace(v), ",", 2)[0])
|
||||
if vol == "" || strings.HasPrefix(vol, "/") {
|
||||
continue
|
||||
}
|
||||
if st, _, ok := strings.Cut(vol, ":"); ok && st != "" {
|
||||
out[st] = true
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// PreflightRestoreSpace applies the three rules. `configured` is `backup.restore_storage`.
|
||||
func PreflightRestoreSpace(ctx context.Context, space RestoreSpace, policy SpacePolicy, archive, rawCfg, configured string) SpaceVerdict {
|
||||
if space == nil {
|
||||
return SpaceVerdict{Storage: configured, Reason: "no space check is wired — refusing (fail-closed)"}
|
||||
}
|
||||
restored, src, err := space.RestoredBytes(ctx, archive)
|
||||
if err != nil || restored <= 0 {
|
||||
return SpaceVerdict{Storage: configured, Reason: fmt.Sprintf("cannot tell how much the restore writes (%v)", err)}
|
||||
}
|
||||
required := policy.requiredBytes(restored)
|
||||
own := sourceStorages(rawCfg)
|
||||
|
||||
// Rule 2: the configured storage holds the guest under test → try the others first.
|
||||
var order []string
|
||||
avoided := ""
|
||||
if own[configured] {
|
||||
eligible, eerr := space.Eligible(ctx)
|
||||
if eerr == nil {
|
||||
sort.Strings(eligible)
|
||||
for _, s := range eligible {
|
||||
if s != configured && !own[s] {
|
||||
order = append(order, s)
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(order) > 0 {
|
||||
avoided = configured
|
||||
}
|
||||
}
|
||||
order = append(order, configured)
|
||||
|
||||
var last SpaceVerdict
|
||||
for _, s := range order {
|
||||
fr, ferr := space.Free(ctx, s)
|
||||
if ferr != nil {
|
||||
last = SpaceVerdict{Storage: s, Required: required, Reason: fmt.Sprintf("cannot read free space on %s (%v)", s, ferr)}
|
||||
continue
|
||||
}
|
||||
ok, why := fits(fr, required)
|
||||
v := SpaceVerdict{OK: ok, Storage: s, Required: required, Avail: fr.AvailBytes}
|
||||
if ok {
|
||||
if s != configured {
|
||||
v.Avoided = avoided
|
||||
}
|
||||
return v
|
||||
}
|
||||
v.Reason = fmt.Sprintf("not enough space on %s: restoring %s (%s) %s", s, gib(restored), src, why)
|
||||
last = v
|
||||
}
|
||||
return last
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
package reconcile
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-672 (v0.133.0) — the restore-test's space preflight. Every test asserts the CONSEQUENCE: whether the
|
||||
// Proxmox API was asked to restore anything, where to, and what the result says — never only the verdict.
|
||||
|
||||
const gb = int64(1000 * 1000 * 1000)
|
||||
|
||||
// fakeSpace is a configurable RestoreSpace.
|
||||
type fakeSpace struct {
|
||||
restored int64
|
||||
restoredErr error
|
||||
free map[string]StorageFree
|
||||
freeErr map[string]error
|
||||
eligible []string
|
||||
}
|
||||
|
||||
func (f fakeSpace) RestoredBytes(context.Context, string) (int64, string, error) {
|
||||
return f.restored, "fake", f.restoredErr
|
||||
}
|
||||
func (f fakeSpace) Free(_ context.Context, s string) (StorageFree, error) {
|
||||
if err := f.freeErr[s]; err != nil {
|
||||
return StorageFree{}, err
|
||||
}
|
||||
fr, ok := f.free[s]
|
||||
if !ok {
|
||||
return StorageFree{}, errors.New("unknown storage")
|
||||
}
|
||||
return fr, nil
|
||||
}
|
||||
func (f fakeSpace) Eligible(context.Context) ([]string, error) { return f.eligible, nil }
|
||||
|
||||
// thin9201 is demo-hp's local-lvm at 10:29 on 2026-09-24, just before the restore-test that filled it:
|
||||
// 23.2 GB free, 33.3 GB used, metadata 2.65 %.
|
||||
var thin9201 = StorageFree{AvailBytes: 23210892 * 1024, UsedBytes: 33277043 * 1024, Thin: true, MetaUsedFraction: 0.0265, MetaKnown: true}
|
||||
|
||||
// archive9201 is 9201's archive config: both volumes on local-lvm.
|
||||
const archive9201 = "hostname: demo-hp\nrootfs: local-lvm:vm-9201-disk-0,size=32G\nmp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G\nmp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\n"
|
||||
|
||||
func spaceEngine(t *testing.T, api *fakeAPI, sp RestoreSpace) (*Engine, *Journal) {
|
||||
t.Helper()
|
||||
j, err := OpenJournal(filepath.Join(t.TempDir(), "journal.log"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(func() { j.Close() })
|
||||
q := NewQueue()
|
||||
t.Cleanup(q.Close)
|
||||
return NewEngine(EngineOptions{API: api, Queue: q, Journal: j, RestoreSpace: sp}), j
|
||||
}
|
||||
|
||||
func run9201(e *Engine) RestoreTestResult {
|
||||
return e.RunRestoreTest(context.Background(), RestoreTestSpec{
|
||||
Archive: "local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst", RestoreStorage: "local-lvm",
|
||||
ScratchMin: 990000, ScratchMax: 990009, SourceTier: "local",
|
||||
})
|
||||
}
|
||||
|
||||
// TestSpace_The2026_09_24TestIsRefused replays R-672: 9201's archive restores 22.6 GB (its vzdump log),
|
||||
// the pool has 23.2 GB free. The test must NOT start — no restore call, no journaled scratch — and the
|
||||
// result must say why, as a non-pass.
|
||||
//
|
||||
// COMPANION RED-PROOFS (REPORT): (1) the preflight removed (v0.132.0's shape) → a restore into local-lvm
|
||||
// is issued; (2) `restored` taken from the archive FILE (6.9 GB, the brief's "archive × 1.2 + 5 GiB") →
|
||||
// 6.9×1.2+5.4 = 13.7 GB < 23.2 GB free, so the test starts — the defect the uncompressed size exists for.
|
||||
func TestSpace_The2026_09_24TestIsRefused(t *testing.T) {
|
||||
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
|
||||
e, j := spaceEngine(t, api, fakeSpace{restored: 22607360000, free: map[string]StorageFree{"local-lvm": thin9201}})
|
||||
res := run9201(e)
|
||||
if len(api.restores) != 0 {
|
||||
t.Fatalf("a restore was issued into a pool that cannot take it: %+v", api.restores)
|
||||
}
|
||||
if len(j.InFlight()) != 0 {
|
||||
t.Fatalf("a scratch entry was journaled for a test that must not start: %+v", j.InFlight())
|
||||
}
|
||||
if res.Pass || !res.Skipped || !strings.Contains(res.SkipReason, "not enough space on local-lvm") {
|
||||
t.Fatalf("result = pass=%v skipped=%v reason=%q — want a non-pass skip naming the storage", res.Pass, res.Skipped, res.SkipReason)
|
||||
}
|
||||
if res.RequiredBytes < 32*gb || res.AvailBytes != thin9201.AvailBytes {
|
||||
t.Fatalf("required=%d avail=%d — want ≥ 32 GB required (22.6 × 1.2 + 5 GiB) against 23.2 GB", res.RequiredBytes, res.AvailBytes)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSpace_KeepsOffTheTestedGuestsPool — rule 2: another eligible storage that fits takes the restore.
|
||||
func TestSpace_KeepsOffTheTestedGuestsPool(t *testing.T) {
|
||||
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
|
||||
e, _ := spaceEngine(t, api, fakeSpace{restored: 22607360000, eligible: []string{"local-lvm", "big-dir"},
|
||||
free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: 10 * gb, Thin: true, MetaKnown: true}, "big-dir": {AvailBytes: 500 * gb}}})
|
||||
res := run9201(e)
|
||||
if len(api.restores) != 1 || api.restores[0].Storage != "big-dir" {
|
||||
t.Fatalf("restores = %+v — want ONE restore onto big-dir, off 9201's own pool", api.restores)
|
||||
}
|
||||
if res.TargetStorage != "big-dir" {
|
||||
t.Fatalf("target=%q", res.TargetStorage)
|
||||
}
|
||||
for k, v := range api.restores[0].MountOverrides {
|
||||
if strings.HasPrefix(v, "local-lvm:") {
|
||||
t.Fatalf("%s still lands on the tested guest's pool: %s", k, v)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestSpace_OnlyOnePool_RuleOneDecides — no other eligible storage: the tested guest's pool is used when it
|
||||
// fits (demo-hp's real shape: nvme-scratch takes rootdir but the agent holds no AllocateSpace there).
|
||||
func TestSpace_OnlyOnePool_RuleOneDecides(t *testing.T) {
|
||||
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
|
||||
e, _ := spaceEngine(t, api, fakeSpace{restored: 2 * gb, eligible: []string{"local-lvm"}, free: map[string]StorageFree{"local-lvm": thin9201}})
|
||||
res := run9201(e)
|
||||
if len(api.restores) != 1 || api.restores[0].Storage != "local-lvm" || res.Skipped || res.TargetStorage != "local-lvm" {
|
||||
t.Fatalf("restores=%+v skipped=%v — a 2 GB restore fits 23 GB free on the only pool", api.restores, res.Skipped)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSpace_UnknownRefuses — rule 3, one case per unknown. Nothing is restored in any of them.
|
||||
func TestSpace_UnknownRefuses(t *testing.T) {
|
||||
cases := map[string]RestoreSpace{
|
||||
"no space check wired": nil,
|
||||
"restore size unknown": fakeSpace{restoredErr: errors.New("no vzdump log"), free: map[string]StorageFree{"local-lvm": thin9201}},
|
||||
"free space unreadable": fakeSpace{restored: gb, freeErr: map[string]error{"local-lvm": errors.New("api down")}},
|
||||
"thin metadata unknown": fakeSpace{restored: gb, free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: gb, Thin: true}}},
|
||||
"metadata would overrun": fakeSpace{restored: 10 * gb, free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: 10 * gb, Thin: true, MetaUsedFraction: 0.5, MetaKnown: true}}},
|
||||
}
|
||||
for name, sp := range cases {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
|
||||
var e *Engine
|
||||
if sp == nil {
|
||||
e, _ = spaceEngine(t, api, nil)
|
||||
} else {
|
||||
e, _ = spaceEngine(t, api, sp)
|
||||
}
|
||||
res := run9201(e)
|
||||
if len(api.restores) != 0 || res.Pass || !res.Skipped || res.SkipReason == "" {
|
||||
t.Fatalf("restores=%d pass=%v skipped=%v reason=%q — an unknown must refuse before anything moves",
|
||||
len(api.restores), res.Pass, res.Skipped, res.SkipReason)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSpace_SourceStorages reads the tested guest's pools from the ARCHIVE's config, binds excluded.
|
||||
func TestSpace_SourceStorages(t *testing.T) {
|
||||
got := sourceStorages(archive9201 + "mp1: other:vm-9201-disk-2,mp=/x,size=1G\n[snap]\nrootfs: snapstore:x\n")
|
||||
if !got["local-lvm"] || !got["other"] || got["snapstore"] || len(got) != 2 {
|
||||
t.Fatalf("sourceStorages = %v — want local-lvm + other, binds and snapshot sections excluded", got)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,16 @@
|
||||
package reconcile
|
||||
|
||||
import "context"
|
||||
|
||||
// roomySpace is the permissive RestoreSpace the pre-R-672 restore-test tests run with: 1 GiB restored,
|
||||
// 1 TiB free, thin metadata known and low. The space rules themselves are pinned in
|
||||
// restoretest_space_test.go.
|
||||
type roomySpace struct{}
|
||||
|
||||
func (roomySpace) RestoredBytes(context.Context, string) (int64, string, error) {
|
||||
return 1 << 30, "test", nil
|
||||
}
|
||||
func (roomySpace) Free(context.Context, string) (StorageFree, error) {
|
||||
return StorageFree{AvailBytes: 1 << 40, UsedBytes: 1 << 30, Thin: true, MetaUsedFraction: 0.01, MetaKnown: true}, nil
|
||||
}
|
||||
func (roomySpace) Eligible(context.Context) ([]string, error) { return nil, nil }
|
||||
@@ -0,0 +1,176 @@
|
||||
// Package restorespace is the production seam behind reconcile.RestoreSpace (R-672, agent v0.133.0):
|
||||
// how much a restore of an archive writes, how much a storage has free, and which storages a
|
||||
// restore-test may target. Every read that cannot answer returns an error — the preflight then
|
||||
// REFUSES (reconcile/restoretest_space.go rule 3); nothing here guesses.
|
||||
package restorespace
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"path"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
|
||||
)
|
||||
|
||||
// API is the Proxmox subset the provider reads.
|
||||
type API interface {
|
||||
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
|
||||
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
|
||||
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
|
||||
Permissions(ctx context.Context, aclPath string) (map[string]int, error)
|
||||
}
|
||||
|
||||
// Provider implements reconcile.RestoreSpace.
|
||||
type Provider struct {
|
||||
API API
|
||||
// ThinMeta reads a thin pool's metadata-used fraction (storage.HostOps.ThinPoolMetadata).
|
||||
ThinMeta func(ctx context.Context, vg, pool string) (float64, bool)
|
||||
// ReadFile reads a vzdump log; nil → os.ReadFile.
|
||||
ReadFile func(name string) ([]byte, error)
|
||||
}
|
||||
|
||||
var _ reconcile.RestoreSpace = (*Provider)(nil)
|
||||
|
||||
// totalWrittenRe is vzdump's own count of the bytes tar wrote into the archive — the UNCOMPRESSED size,
|
||||
// i.e. what a restore writes back ("INFO: Total bytes written: 22607360000 (22GiB, 49MiB/s)").
|
||||
var totalWrittenRe = regexp.MustCompile(`Total bytes written:\s*(\d+)`)
|
||||
|
||||
// archiveExts are the vzdump archive suffixes; the log is the archive name without it + ".log".
|
||||
var archiveExts = []string{".tar.zst", ".tar.gz", ".tar.lzo", ".tgz", ".tar"}
|
||||
|
||||
func (p *Provider) readFile(name string) ([]byte, error) {
|
||||
if p.ReadFile != nil {
|
||||
return p.ReadFile(name)
|
||||
}
|
||||
return os.ReadFile(name)
|
||||
}
|
||||
|
||||
// RestoredBytes: a file-backed archive → its vzdump log's "Total bytes written"; a PBS archive → the
|
||||
// size Proxmox reports for the snapshot (its logical, uncompressed size). The archive FILE size is never
|
||||
// used: it is compressed (6.9 GB for a 22.6 GB restore, measured 2026-09-24).
|
||||
func (p *Provider) RestoredBytes(ctx context.Context, archive string) (int64, string, error) {
|
||||
id, vol, ok := strings.Cut(archive, ":")
|
||||
if !ok || id == "" || vol == "" {
|
||||
return 0, "", fmt.Errorf("not a storage volid: %q", archive)
|
||||
}
|
||||
st, err := p.storageConfig(ctx, id)
|
||||
if err != nil {
|
||||
return 0, "", err
|
||||
}
|
||||
switch st.Type {
|
||||
case "pbs":
|
||||
items, err := p.API.StorageContent(ctx, id)
|
||||
if err != nil {
|
||||
return 0, "", fmt.Errorf("list %s: %w", id, err)
|
||||
}
|
||||
for _, it := range items {
|
||||
if it.VolID == archive && it.Size > 0 {
|
||||
return it.Size, "pbs snapshot size", nil
|
||||
}
|
||||
}
|
||||
return 0, "", fmt.Errorf("archive %s not listed on %s with a size", archive, id)
|
||||
default:
|
||||
if st.Path == "" {
|
||||
return 0, "", fmt.Errorf("storage %s (%s) has no path to read a vzdump log from", id, st.Type)
|
||||
}
|
||||
base := path.Base(vol) // "backup/vzdump-lxc-…tar.zst" → "vzdump-lxc-…tar.zst"
|
||||
stem := ""
|
||||
for _, ext := range archiveExts {
|
||||
if strings.HasSuffix(base, ext) {
|
||||
stem = strings.TrimSuffix(base, ext)
|
||||
break
|
||||
}
|
||||
}
|
||||
if stem == "" {
|
||||
return 0, "", fmt.Errorf("unknown archive suffix: %s", base)
|
||||
}
|
||||
logPath := path.Join(st.Path, "dump", stem+".log")
|
||||
b, err := p.readFile(logPath)
|
||||
if err != nil {
|
||||
return 0, "", fmt.Errorf("read vzdump log %s: %w", logPath, err)
|
||||
}
|
||||
m := totalWrittenRe.FindSubmatch(b)
|
||||
if m == nil {
|
||||
return 0, "", fmt.Errorf("vzdump log %s carries no \"Total bytes written\"", logPath)
|
||||
}
|
||||
n, err := strconv.ParseInt(string(m[1]), 10, 64)
|
||||
if err != nil || n <= 0 {
|
||||
return 0, "", fmt.Errorf("vzdump log %s: bad byte count %q", logPath, m[1])
|
||||
}
|
||||
return n, "vzdump log: total bytes written", nil
|
||||
}
|
||||
}
|
||||
|
||||
func (p *Provider) storageConfig(ctx context.Context, id string) (proxmox.Storage, error) {
|
||||
all, err := p.API.ListStorage(ctx)
|
||||
if err != nil {
|
||||
return proxmox.Storage{}, fmt.Errorf("list storage config: %w", err)
|
||||
}
|
||||
for _, s := range all {
|
||||
if s.Storage == id {
|
||||
return s, nil
|
||||
}
|
||||
}
|
||||
return proxmox.Storage{}, fmt.Errorf("storage %s not configured", id)
|
||||
}
|
||||
|
||||
// Free reads the node's live usage for `storage`; a thin pool adds its metadata fill.
|
||||
func (p *Provider) Free(ctx context.Context, storage string) (reconcile.StorageFree, error) {
|
||||
live, err := p.API.NodeStorage(ctx)
|
||||
if err != nil {
|
||||
return reconcile.StorageFree{}, fmt.Errorf("node storage: %w", err)
|
||||
}
|
||||
for _, s := range live {
|
||||
if s.Storage != storage {
|
||||
continue
|
||||
}
|
||||
if s.Active != 1 {
|
||||
return reconcile.StorageFree{}, fmt.Errorf("storage %s is not active", storage)
|
||||
}
|
||||
fr := reconcile.StorageFree{AvailBytes: s.Avail, UsedBytes: s.Used, Thin: s.Type == "lvmthin"}
|
||||
if fr.Thin {
|
||||
cfg, cerr := p.storageConfig(ctx, storage)
|
||||
if cerr == nil && p.ThinMeta != nil && cfg.VGName != "" && cfg.ThinPool != "" {
|
||||
fr.MetaUsedFraction, fr.MetaKnown = p.ThinMeta(ctx, cfg.VGName, cfg.ThinPool)
|
||||
}
|
||||
}
|
||||
return fr, nil
|
||||
}
|
||||
return reconcile.StorageFree{}, fmt.Errorf("storage %s not reported by the node", storage)
|
||||
}
|
||||
|
||||
// Eligible: active, content includes `rootdir`, and the agent holds Datastore.AllocateSpace on
|
||||
// /storage/<id>. The permission is read for the SPECIFIC privilege — the box-wide grant answers every
|
||||
// path with inherited privileges (proxmox.Client.Permissions).
|
||||
func (p *Provider) Eligible(ctx context.Context) ([]string, error) {
|
||||
live, err := p.API.NodeStorage(ctx)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("node storage: %w", err)
|
||||
}
|
||||
var out []string
|
||||
for _, s := range live {
|
||||
if s.Active != 1 || !hasContent(s.Content, "rootdir") {
|
||||
continue
|
||||
}
|
||||
privs, perr := p.API.Permissions(ctx, "/storage/"+s.Storage)
|
||||
if perr != nil || privs["Datastore.AllocateSpace"] != 1 {
|
||||
continue
|
||||
}
|
||||
out = append(out, s.Storage)
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
func hasContent(list, want string) bool {
|
||||
for _, c := range strings.Split(list, ",") {
|
||||
if strings.TrimSpace(c) == want {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,133 @@
|
||||
package restorespace
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-672 (v0.133.0). The provider behind the restore-test's space preflight. No test reaches a real
|
||||
// Proxmox or a real file: the API and ReadFile are fakes.
|
||||
|
||||
type fakeAPI struct {
|
||||
cfg []proxmox.Storage
|
||||
live []proxmox.Storage
|
||||
content map[string][]proxmox.StorageContent
|
||||
perms map[string]map[string]int
|
||||
}
|
||||
|
||||
func (f fakeAPI) ListStorage(context.Context) ([]proxmox.Storage, error) { return f.cfg, nil }
|
||||
func (f fakeAPI) NodeStorage(context.Context) ([]proxmox.Storage, error) { return f.live, nil }
|
||||
func (f fakeAPI) StorageContent(_ context.Context, s string) ([]proxmox.StorageContent, error) {
|
||||
return f.content[s], nil
|
||||
}
|
||||
func (f fakeAPI) Permissions(_ context.Context, p string) (map[string]int, error) {
|
||||
if m, ok := f.perms[p]; ok {
|
||||
return m, nil
|
||||
}
|
||||
return map[string]int{}, nil
|
||||
}
|
||||
|
||||
const archive = "local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst"
|
||||
|
||||
// The real log's tail (demo-hp, 2026-09-23): 22.6 GB written, a 6.91 GB archive file.
|
||||
const vzdumpLog = "2026-09-23 07:02:47 INFO: Total bytes written: 22607360000 (22GiB, 49MiB/s)\n2026-09-23 07:02:47 INFO: archive file size: 6.91GB\n"
|
||||
|
||||
// demoHP is demo-hp's storage layout: `local` (dir, backups), `local-lvm` (thin), `nvme-scratch` (dir,
|
||||
// rootdir — but the agent holds NO grant there), a pbs.
|
||||
func demoHP() fakeAPI {
|
||||
return fakeAPI{
|
||||
cfg: []proxmox.Storage{
|
||||
{Storage: "local", Type: "dir", Path: "/var/lib/vz"},
|
||||
{Storage: "local-lvm", Type: "lvmthin", VGName: "pve", ThinPool: "data"},
|
||||
{Storage: "nvme-scratch", Type: "dir", Path: "/mnt/hdd_1"},
|
||||
{Storage: "felhom-pbs", Type: "pbs"},
|
||||
},
|
||||
live: []proxmox.Storage{
|
||||
{Storage: "local", Type: "dir", Content: "vztmpl,backup,iso,import", Active: 1, Avail: 4 << 30, Used: 34 << 30},
|
||||
{Storage: "local-lvm", Type: "lvmthin", Content: "images,rootdir", Active: 1, Avail: 23210892 * 1024, Used: 33277043 * 1024},
|
||||
{Storage: "nvme-scratch", Type: "dir", Content: "images,rootdir", Active: 1, Avail: 800 << 30},
|
||||
{Storage: "felhom-pbs", Type: "pbs", Content: "backup", Active: 0},
|
||||
},
|
||||
content: map[string][]proxmox.StorageContent{
|
||||
"local": {{VolID: archive, Size: 7417540996}},
|
||||
"felhom-pbs": {{VolID: "felhom-pbs:backup/ct/9201/2026-09-23T02:00:00Z", Size: 21 << 30}},
|
||||
},
|
||||
perms: map[string]map[string]int{
|
||||
"/storage/local": {"Datastore.Audit": 1, "Datastore.AllocateSpace": 1},
|
||||
"/storage/local-lvm": {"Datastore.Audit": 1, "Datastore.AllocateSpace": 1},
|
||||
// nvme-scratch: only the inherited box-wide Datastore.Audit — the trap proxmox.Permissions names.
|
||||
"/storage/nvme-scratch": {"Datastore.Audit": 1},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestRestoredBytes_ReadsTheUncompressedSize — the vzdump log's "Total bytes written", never the archive
|
||||
// FILE size (6.9 GB for a 22.6 GB restore).
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT): return the storage content's Size for a dir storage → 7417540996, and
|
||||
// this test fails at "the compressed file size was used".
|
||||
func TestRestoredBytes_ReadsTheUncompressedSize(t *testing.T) {
|
||||
var asked string
|
||||
p := &Provider{API: demoHP(), ReadFile: func(n string) ([]byte, error) { asked = n; return []byte(vzdumpLog), nil }}
|
||||
n, src, err := p.RestoredBytes(context.Background(), archive)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if n == 7417540996 {
|
||||
t.Fatal("the compressed file size was used — the restore writes 3× that")
|
||||
}
|
||||
if n != 22607360000 || asked != "/var/lib/vz/dump/vzdump-lxc-9201-2026_09_23-06_55_25.log" {
|
||||
t.Fatalf("n=%d from %q (log %q)", n, src, asked)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRestoredBytes_UnknownIsAnError(t *testing.T) {
|
||||
cases := map[string]*Provider{
|
||||
"no log": {API: demoHP(), ReadFile: func(string) ([]byte, error) { return nil, errors.New("ENOENT") }},
|
||||
"log without size": {API: demoHP(), ReadFile: func(string) ([]byte, error) { return []byte("ERROR: failed\n"), nil }},
|
||||
}
|
||||
for name, p := range cases {
|
||||
if n, _, err := p.RestoredBytes(context.Background(), archive); err == nil {
|
||||
t.Fatalf("%s: got %d, want an error (the preflight then refuses)", name, n)
|
||||
}
|
||||
}
|
||||
if _, _, err := (&Provider{API: demoHP()}).RestoredBytes(context.Background(), "local:backup/weird.vma"); err == nil {
|
||||
t.Fatal("an unknown archive suffix must be an error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRestoredBytes_PBS(t *testing.T) {
|
||||
p := &Provider{API: demoHP()}
|
||||
n, _, err := p.RestoredBytes(context.Background(), "felhom-pbs:backup/ct/9201/2026-09-23T02:00:00Z")
|
||||
if err != nil || n != 21<<30 {
|
||||
t.Fatalf("n=%d err=%v", n, err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestEligible_NeedsTheSpecificGrant — nvme-scratch takes rootdir but the agent holds only the inherited
|
||||
// Datastore.Audit there, so it is NOT eligible; `local` holds no rootdir.
|
||||
func TestEligible_NeedsTheSpecificGrant(t *testing.T) {
|
||||
got, err := (&Provider{API: demoHP()}).Eligible(context.Background())
|
||||
if err != nil || len(got) != 1 || got[0] != "local-lvm" {
|
||||
t.Fatalf("eligible = %v (%v) — want only local-lvm on demo-hp", got, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFree_ThinCarriesMetadata(t *testing.T) {
|
||||
p := &Provider{API: demoHP(), ThinMeta: func(_ context.Context, vg, pool string) (float64, bool) {
|
||||
if vg != "pve" || pool != "data" {
|
||||
t.Fatalf("metadata read for %s/%s", vg, pool)
|
||||
}
|
||||
return 0.0265, true
|
||||
}}
|
||||
fr, err := p.Free(context.Background(), "local-lvm")
|
||||
if err != nil || !fr.Thin || !fr.MetaKnown || fr.MetaUsedFraction != 0.0265 || fr.AvailBytes != 23210892*1024 {
|
||||
t.Fatalf("free = %+v err=%v", fr, err)
|
||||
}
|
||||
if _, err := p.Free(context.Background(), "felhom-pbs"); err == nil {
|
||||
t.Fatal("an inactive storage must be an error")
|
||||
}
|
||||
}
|
||||
@@ -6,6 +6,7 @@ import (
|
||||
"log/slog"
|
||||
"regexp"
|
||||
"strings"
|
||||
"sync"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
@@ -35,6 +36,45 @@ type Observer struct {
|
||||
host HostReader
|
||||
ops HostOps
|
||||
logger *slog.Logger
|
||||
|
||||
// R-672 (v0.133.0): a thin pool crossing thinPoolAlarmFraction (data OR metadata) requests an
|
||||
// out-of-band host report at once, so the hub's storage-fill alarm sees it in seconds instead of at
|
||||
// the next 15-minute report. Rising edge per pool; re-armed below thinPoolRearmFraction.
|
||||
highMu sync.Mutex
|
||||
high map[string]bool
|
||||
onThinHigh func()
|
||||
}
|
||||
|
||||
const (
|
||||
thinPoolAlarmFraction = 0.90
|
||||
thinPoolRearmFraction = 0.85
|
||||
)
|
||||
|
||||
// SetThinHighTrigger wires the out-of-band report request (main: the storage trigger channel).
|
||||
func (o *Observer) SetThinHighTrigger(f func()) { o.onThinHigh = f }
|
||||
|
||||
// noteThinFill is the edge detector. key separates data from metadata so each has its own edge.
|
||||
func (o *Observer) noteThinFill(key string, frac float64) {
|
||||
o.highMu.Lock()
|
||||
if o.high == nil {
|
||||
o.high = map[string]bool{}
|
||||
}
|
||||
fire := false
|
||||
switch {
|
||||
case frac >= thinPoolAlarmFraction && !o.high[key]:
|
||||
o.high[key] = true
|
||||
fire = true
|
||||
case frac < thinPoolRearmFraction && o.high[key]:
|
||||
o.high[key] = false
|
||||
}
|
||||
o.highMu.Unlock()
|
||||
if fire {
|
||||
o.logger.Error("storage: thin pool crossed 90% — requesting an immediate host report (the hub alarms)",
|
||||
"pool", key, "fraction", frac)
|
||||
if o.onThinHigh != nil {
|
||||
o.onThinHigh()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// NewObserver builds an Observer. host defaults to a ProcHostReader; logger to the
|
||||
@@ -260,6 +300,7 @@ func (o *Observer) build(s proxmox.Storage, mounts []Mount) observed {
|
||||
o.logger.Warn("storage: lvmthin pool data fill is high (a full pool corrupts every guest on it)",
|
||||
"storage", s.Storage, "data_used_fraction", frac)
|
||||
}
|
||||
o.noteThinFill(s.Storage+"/data", frac)
|
||||
}
|
||||
|
||||
// SMART-only device hint (v0.95.0): a dir-storage that lives INSIDE a shared filesystem (the
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
package storage
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||
)
|
||||
|
||||
// R-672 (v0.133.0): a thin pool crossing 90 % requests an out-of-band host report ONCE, through the
|
||||
// watchdog's own read path (Known — every few seconds), so the hub's storage-fill alarm sees the pool in
|
||||
// seconds, not at the next 15-minute report. Re-armed below 85 %.
|
||||
//
|
||||
// COMPANION RED-PROOF (REPORT): remove the noteThinFill call from the data path → "no report was
|
||||
// requested when the pool crossed 90 %".
|
||||
func TestThinHigh_RequestsOneReportPerCrossing(t *testing.T) {
|
||||
pool := proxmox.Storage{Storage: "local-lvm", Type: "lvmthin", Content: "rootdir,images", Total: 1000, Active: 1}
|
||||
api := &fakeStorageAPI{node: "n", cluster: []proxmox.Storage{pool}}
|
||||
o := NewObserver(api, &fakeHostReader{}, nil, quietLogger())
|
||||
asked := 0
|
||||
o.SetThinHighTrigger(func() { asked++ })
|
||||
for i, used := range []int64{800, 910, 950, 1000, 840, 920} {
|
||||
p := pool
|
||||
p.Used, p.Avail, p.UsedFraction = used, 1000-used, float64(used)/1000
|
||||
api.nodeSt = []proxmox.Storage{p}
|
||||
if _, err := o.Known(context.Background()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
want := map[int]int{0: 0, 1: 1, 2: 1, 3: 1, 4: 1, 5: 2}[i]
|
||||
if asked != want {
|
||||
t.Fatalf("after %d/1000 used: %d report requests, want %d", used, asked, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -43,6 +43,9 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
|
||||
SHARED_INSTRUCTIONS = os.path.join(
|
||||
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
|
||||
# R-389 — shared, like the two above: it lives in felhom.eu/scripts/ and is never copied.
|
||||
SHARED_OBSERVATIONS = os.path.join(
|
||||
os.path.dirname(ROOT), "felhom.eu", "scripts", "observations_gate.py")
|
||||
|
||||
# (label, absolute script path, args, fast)
|
||||
GATES = [
|
||||
@@ -53,6 +56,8 @@ GATES = [
|
||||
# missing TAG is what actually broke every install, and the pre-push hook is the earliest place
|
||||
# that can catch it.
|
||||
("release-complete", os.path.join(ROOT, "scripts", "check-release-complete.py"), [], True),
|
||||
# R-389 — a REPORT.md observation with no register row behind it. Fast: stdlib file reads.
|
||||
("observations", SHARED_OBSERVATIONS, [ROOT], True),
|
||||
]
|
||||
|
||||
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
|
||||
|
||||
Reference in New Issue
Block a user