Compare commits
8 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| adaf86ad57 | |||
| 74b5eae5b0 | |||
| 7e82f325b8 | |||
| acccb66bd3 | |||
| de812bc027 | |||
| 3e8ebeb96c | |||
| cefdc731a4 | |||
| e56dcb8a4c |
@@ -21,6 +21,7 @@ fast, and wrong.
|
|||||||
|
|
||||||
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
|
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
|
||||||
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
|
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
|
||||||
path-scoped rules existed. The single source is now felhom.eu/CLAUDE.md "Code quality rules"; this
|
path-scoped rules existed. Deliberate scoped copies now live in felhom.eu/.claude/rules/hub.md and
|
||||||
file is the scoped copy that loads exactly where health checks are written. (2026-08-06)
|
felhom-controller/.claude/rules/gates.md (hub.md's comment names them); none is the single source. This
|
||||||
|
file is the copy that loads exactly where agent health checks are written. (2026-08-06; corrected 2026-10-06)
|
||||||
-->
|
-->
|
||||||
|
|||||||
@@ -0,0 +1,69 @@
|
|||||||
|
---
|
||||||
|
unconditional: true
|
||||||
|
---
|
||||||
|
# Unprompted work — rules for any session without a task file
|
||||||
|
|
||||||
|
> Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from
|
||||||
|
> `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the
|
||||||
|
> part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives
|
||||||
|
> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace
|
||||||
|
> root's unversioned `.claude/rules/`; change all five or none.
|
||||||
|
|
||||||
|
## 1. What you may pick up on your own
|
||||||
|
|
||||||
|
- A register row **you or another CC session filed**, with owner CC, at P3 or a bounded P2, that
|
||||||
|
needs **no operator decision**, touches **no customer data by design**, and introduces **no
|
||||||
|
mechanism nobody has measured**. Smallest first.
|
||||||
|
- A defect you find while exercising the product, filed as a row **before** you fix it — **unless it is small**:
|
||||||
|
a small finding is fixed in the session and never filed (the size rule, `OPEN-ITEMS.md` „How a row is filed").
|
||||||
|
- Hygiene: register compression, stale citations, rows with no owner, documents that contradict
|
||||||
|
live source.
|
||||||
|
|
||||||
|
**Not yours, ever, without a task file or an operator word:** money; anything that changes risk to
|
||||||
|
customer data; anything that changes a promise the product makes to a customer; anything that
|
||||||
|
reverses a documented design decision (`documentation/architecture/` — a design decision is not a
|
||||||
|
defect, R-370); anything on DooPlex or ep0; baking or vouching a golden; promoting a
|
||||||
|
catalog version; a new external dependency.
|
||||||
|
|
||||||
|
## 2. When you may decide instead of ask (operator grant, 2026-09-14)
|
||||||
|
|
||||||
|
You may take a decision yourself when **all** of these hold: the architecture folder and the register
|
||||||
|
give a clear direction; your choice follows that direction; it is reversible without customer-data
|
||||||
|
risk; and you can write it in the `09-update-architecture.md` §3 shape — one answerable sentence, the
|
||||||
|
options, what each costs, why this one. **Then record it** as a dated decision in `CONTEXT.md` and
|
||||||
|
the owning architecture document, tagged *decided by CC unattended — operator may reverse*, and put
|
||||||
|
it **first** in the morning note. A decision you cannot write in that shape is one you do not take.
|
||||||
|
|
||||||
|
## 3. The discipline a task file used to carry
|
||||||
|
|
||||||
|
1. **Baselines first.** Read each repo's `main` hash and version from live source before touching it.
|
||||||
|
2. **Read the architecture document for the area, and name it** in the report, before any claim.
|
||||||
|
3. **Red-proof every correctness fix.** A test never seen failing has not been shown to test anything.
|
||||||
|
4. **Live-validate on a Tier-0 box** through the endpoints the UI invokes. `demo-hp` is `ssh hp`.
|
||||||
|
Throwaway apps only; the standing apps and `bentopdf` stay.
|
||||||
|
5. **Evidence off the machine at the end of each phase**, before any revert (R-320).
|
||||||
|
6. **One release per repo per session**, with a CHANGELOG entry (controller: with its `MinAgent`
|
||||||
|
line), REPORT overwritten, floor raised to deliver it. **No golden unless a drill or fresh install
|
||||||
|
needs one** (the waiver, R-468). **No `--no-verify`.**
|
||||||
|
7. **An enumerated gap becomes a row in the same session — or, if it is small, is fixed in it** (the size rule).
|
||||||
|
Prose is not a record.
|
||||||
|
8. **Hungarian text is searched with ASCII fragments**, with a positive and a negative control.
|
||||||
|
9. **Never leave a half-state.** If time runs out, revert to clean and say what was reverted.
|
||||||
|
10. **Teardown, three layers, stated** — machine, host, hub — or "provisioned nothing".
|
||||||
|
|
||||||
|
## 4. The morning note
|
||||||
|
|
||||||
|
One screen, plain language, in this order: **decisions you took** (§2) first; what you exercised;
|
||||||
|
what broke and whether you fixed it; rows opened and closed with the register size before and after;
|
||||||
|
what needs the operator, each with what happens if they do nothing. No file paths, no function
|
||||||
|
names, no row numbers as the subject of a sentence.
|
||||||
|
|
||||||
|
## 5. Instruction files
|
||||||
|
|
||||||
|
**Instruction files (`CLAUDE.md`, `.claude/rules/*`) are kept true by the session that finds them wrong**
|
||||||
|
(operator ruling 2026-10-06, `09` §3 decision 150). A session MAY, without asking: correct a stale fact (a command, a
|
||||||
|
count, a version, a path, a description of what a gate does), add a fact it proved, and remove a reference to something
|
||||||
|
that no longer exists. Each edit is named in the report (file, line, before, after, why). A session MAY NOT, without the
|
||||||
|
operator's word: loosen a safety rule, a fence, a „never", a protected machine, a secret rule, or a review step; or
|
||||||
|
remove a rule. When in doubt, it is a rule change, and it goes to the operator. If Claude Code's own permission check
|
||||||
|
asks before such an edit, wait for the operator's click; if it refuses, record that and file the exact line.
|
||||||
+36
-1
@@ -1,4 +1,39 @@
|
|||||||
## unreleased
|
## Unreleased (2026-10-06 night, later) — after a restart the agent remembers the last backup per tier (R-894); three more SMART counters on the wire (R-330)
|
||||||
|
|
||||||
|
Ships with the memory-kill check below as v0.150.0, AFTER the 2026-10-07 night read-back. Nothing delivered tonight.
|
||||||
|
|
||||||
|
- **The defect (measured 2026-10-05 on demo-hp):** the agent restarted at 04:57; at 06:25 the off-site storage answered *Can't connect*; the per-tier backup record is in memory only, so the due-check fell back to an EMPTY record and the 7-day tier (last copy 4 days old) read DUE; the controller asked and vzdump failed.
|
||||||
|
- New `internal/backup/backup_state.go` `BackupSuccessState`: the newest SUCCESSFUL backup per tier and guest, on disk (`<oob state dir>/backup-success-state.json`, atomic tmp+rename, 0600). Only successes are written; a corrupt file reads as nothing known.
|
||||||
|
- `internal/localapi` `handleBackupDue`: when the tier's storage CANNOT be read, the saved copy stands in for the in-memory record. A fresh copy → not due („… (storage unreadable — age from the last success saved on disk)"); a copy older than the cadence → DUE; no copy → the old answer (DUE, age unknown). A storage that answers stays the ground truth: an archive absent there is due even when the file remembers one.
|
||||||
|
- Wired in `buildLocalAPIServer` (`LastKnownBackups`); the local API's backup job saves each success.
|
||||||
|
- Tests: `TestBackupDue_R894_*` (restart = a new server and a new state from the same file; fresh / old / none / storage answers / failed backup not saved), `TestBackupSuccessState_*`, `TestR894_LastKnownBackupsIsWiredIntoTheDaemon` (AST). Four red-proofs observed (`felhom.eu/documentation/audits/night-burndown-2026-10-06/s4/`).
|
||||||
|
- **R-330 (disk health Phase 2, the wire only):** the SMART summary carries three more SATA raw counters — `reported_uncorrect` (187), `command_timeout` (188, carried as the vendor reports it; some pack several counters), `udma_crc_errors` (199). Pointer + omitempty: an attribute the drive does not report is OMITTED (unknown), never 0. No verdict reads them yet. Tests `TestParseSMART_R330_*` (two red-proofs, `felhom.eu/documentation/audits/night-burndown-2026-10-06/r330/`).
|
||||||
|
|
||||||
|
## Unreleased (2026-10-06 night) — the Docker step proves the engine reports a memory kill (`09` §3 decision 157, R-528)
|
||||||
|
|
||||||
|
To be released as v0.150.0 with its config bundle AFTER the 2026-10-07 night read-back (the night of 2026-10-06 runs v0.149.0 on purpose).
|
||||||
|
|
||||||
|
- `configs/felhom-os-apply`: after a docker-layer APPLY (after `health_after`) the wrapper runs `oom_check()`: a throwaway container from the image the running controller uses (`--pull never`, `--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; „pass" only when `OOMKilled=true` AND the `oom` event; it waits 2 s and reads the events window to the guest's epoch + 1 (measured: a window closed in the same second missed the event); the container is always removed. Reported as `oom_check`; it never changes the step's outcome or health. A wrapper-only mode `oom-check` runs the check alone (no apt, no engine change), by hand as root.
|
||||||
|
- `internal/osupdate`: `WrapperReport` and `Report` carry `oom_check` verbatim, on the normal pass and on the kept-copy path (R-868).
|
||||||
|
- `configs/test_felhom_os_apply.py`: its `unittest.main()` sat in the middle of the file, so 11 tests (UnsentReport, SaveReportOnDisk, AgentDiesMidPass, CrashLeftTheJournal) never ran — moved to the end; all pass.
|
||||||
|
- Tests: the OOMCheck class (pass, OOMKilled=false, no event, unreadable image, removal on an inspect error, not on other layers or in health mode, the events window after the settle wait, mode oom-check alone and its refusals); TestDocker_OOMCheckReachesTheHubUnchanged, TestR868_KeptCopyCarriesTheOOMCheck. 12 red-proofs in `felhom.eu/documentation/audits/readback-2026-10-07/F/`.
|
||||||
|
|
||||||
|
## Unreleased (2026-10-06 evening) — the shared rule file (`09` §3 decision 152); no code change
|
||||||
|
|
||||||
|
- `.claude/rules/unprompted-work.md` added, byte-identical to the copies in felhom.eu, felhom-controller, app-catalog-felhom.eu and the workspace root (checked with `diff` against the controller's copy and one md5 across all five). Its copies line names five copies.
|
||||||
|
|
||||||
|
## Unreleased (2026-10-06 afternoon) — instruction files kept true (`09` §3 decision 150); no code change
|
||||||
|
|
||||||
|
- `CLAUDE.md` „Gates — ONE entry point": the runner runs every gate in its `GATES` table (five: three shared, `published`, `release-complete`); `--fast` skips `published` (network). It said two gates and „all of them".
|
||||||
|
- `CLAUDE.md`: the decoy gate and its audit are named with their `felhom.eu/` prefix (they do not exist in this repo).
|
||||||
|
- `.claude/rules/health-checks.md` (comment): the health-check rule's copies live in felhom.eu `hub.md` and the controller's `gates.md`; it named felhom.eu `CLAUDE.md` „Code quality rules", which holds no such rule.
|
||||||
|
|
||||||
|
## v0.149.0 — a weekly disk trim of each customer guest, the crash-boot fact for the controller, the phantom WARN names its runbook (R-444, R-856, R-99; operator rulings `09` §3 139, 143, 140) (2026-10-06)
|
||||||
|
|
||||||
|
Released by `scripts/release-agent.sh`: binary sha256 `6bcae9c2eb5d97e8285316583870059835793893299e291891a53a4ce505585f`
|
||||||
|
config bundle sha256 `e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66` (tag `v0.149.0` = `f277e61`).
|
||||||
|
**The bundle carries the new sudoers rule for the trim (`FELHOM_FSTRIM`) — deliver it with the binary:** signed
|
||||||
|
`agent_update`, then signed `agent_config_update`.
|
||||||
|
|
||||||
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
|
- R-856 (`09` §3 decision 143): new local-API route `GET /host/crash-guard` — passes the host crash guard's last-boot record (present, last_boot_at, last_boot_unclean, tripped) from /var/lib/felhom-crash-guard/state.json to the controller, which waits ~15 min with app mails after a crash boot. Read-only, no Proxmox call, guest-token authed; a missing/unreadable/garbled file answers 200 present:false (never an error page). An older agent answers 404, which the controller reads as unknown (normal 90 s grace) — no controller MinAgent raise needed.
|
||||||
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
|
- R-444 (`09` §3 decision 139): weekly guest disk trim. New sudoers alias FELHOM_FSTRIM with ONE exact rule `/usr/sbin/pct ^fstrim [0-9]+$` (rides the signed config bundle; decoys pinned by TestSudoersFstrimRuleIsExact) and capability guest-fstrim (non-critical). New internal/fstrim job: each owned RUNNING guest gets `pct fstrim <vmid>` once a week - due Wednesday from 10:00 host-local, starts only 10:00-20:59 (never the 01:00-06:59 night), holds the one-heavy-op gate so it never runs beside a backup or restore-test (busy -> deferred to the next hourly tick; a box that was off catches up at its next daytime hour); a failed trim WARNs and is retried at most 3 times that week; bytes parsed from `pct fstrim`'s "(N bytes) trimmed" lines; positive log `fstrim: guest N trimmed X GiB in Ys`; last result per guest persisted in <state_dir>/guest-disk-trim.json and reported as the new omitempty host-report stanza `guest_disk_trim`. Opt-out: agent.json "disk_trim": {"disable": true}.
|
||||||
|
|||||||
@@ -52,11 +52,11 @@ This is in the core because breaching it is how this component stops being audit
|
|||||||
|
|
||||||
## Gates — ONE entry point
|
## Gates — ONE entry point
|
||||||
|
|
||||||
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs this
|
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs every
|
||||||
repo's gates — `reuse_refs_check` and `instructions_gate`, both the **shared** copies in
|
gate in its `GATES` table (that table is the list); the shared ones — `reuse_refs_check`,
|
||||||
`felhom.eu/scripts/`, never copied into this repo (a copy recreates the drift they detect; an absent
|
`instructions_gate`, `observations_gate` — are the copies in `felhom.eu/scripts/`, never copied into
|
||||||
sibling clone FAILS). `--fast` selects the gates touching no network and no container runtime; today
|
this repo (a copy recreates the drift they detect; an absent sibling clone FAILS). `--fast` selects the
|
||||||
that is all of them. **A missing gate is a FAILURE, never a skip.**
|
gates touching no network and no container runtime, and skips `published` (network), naming it. **A missing gate is a FAILURE, never a skip.**
|
||||||
|
|
||||||
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
|
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
|
||||||
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
|
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
|
||||||
@@ -104,8 +104,8 @@ the mechanism are exempt.
|
|||||||
|
|
||||||
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
|
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
|
||||||
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
|
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
|
||||||
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
|
lacks. `felhom.eu/scripts/decoy_coverage_gate.py` (run by felhom.eu's `repo_gates.py`, for all four repos) refuses a new gate that has neither a decoy nor a named
|
||||||
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
|
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
|
||||||
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
|
decoys withdrawn as illegitimate: `felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
|
||||||
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
|
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
|
||||||
`os.listdir`, and a glob over a hand-maintained list.
|
`os.listdir`, and a glob over a hand-maintained list.
|
||||||
|
|||||||
@@ -1,16 +1,7 @@
|
|||||||
# REPORT — agent v0.148.0 (2026-10-06, the burn-down night)
|
# REPORT — the shared rule file (2026-10-06 evening)
|
||||||
|
|
||||||
Full session report: `felhom.eu/REPORT-burndown3-2026-10-06.md`. Baseline `208fac8` (v0.147.0).
|
Operator ruling 2026-10-06 14:24 (`09` §3 decision 152): the agent repo gets its copy of the shared rule file. Added
|
||||||
|
`.claude/rules/unprompted-work.md`, byte-identical to the other copies (`diff` against felhom-controller's copy: no
|
||||||
**Released:** v0.148.0 (tag = `861d32a`; binary sha256 `3e68a087…`, bundle `a6fa4f58…`, verified by download). R-349
|
output; md5 `c1e6c881…` across all five before the copies line changed, one md5 after). No code changed; no release.
|
||||||
(the host report carries `agent_sha256` — the hub's host pages read „matches vouched” for all three boxes) and R-25's
|
`agent_gates.py --fast`: reuse-refs, instructions, release-complete, observations OK. The session report is
|
||||||
agent half (the format answer carries the verified `fs_uuid`). **Delivered** by signed `agent_update` (all three on
|
`felhom.eu/REPORT.md`.
|
||||||
0.148.0 by 22:50Z) and `agent_config_update` (root files 0.148.0 by 22:58Z) to demo-hp, demo-felhom, Tester 1. Tester 2:
|
|
||||||
nothing sent (off).
|
|
||||||
|
|
||||||
**On main, unreleased:** R-426 decoys (new `scripts/test_gate_decoys.py`: published against a fake Gitea,
|
|
||||||
release-complete, the shared reuse-refs/instructions/observations).
|
|
||||||
|
|
||||||
**Said plainly:** `go test ./internal/osupdate` was red on DooPlex from 21:25 to 01:55 — `configs/test_felhom_config_bundle.py`
|
|
||||||
read the installer 1.32.0 KEPT names as installer-written files. The v0.148.0 binary was released inside that window;
|
|
||||||
its code is unaffected (a test-only sibling coupling; CI has no Go). Fixed in `b2b82ae`.
|
|
||||||
|
|||||||
@@ -163,6 +163,7 @@
|
|||||||
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
|
| `fstrim.Trimmer` (R-444) | internal/fstrim/fstrim.go | `New(runner, guests, gate, statePath, logger)` / `Pass(ctx)` / `GuestDiskTrimStatus(ctx)` / `ParseTrimmed(out)` | the weekly `pct fstrim <vmid>` of owned running guests (Wednesday from 10:00 local, starts 10:00-20:59 only), under `backup.InFlight`; last result per guest persisted and reported as `guest_disk_trim` | A busy gate DEFERS to the next hourly tick, never waits; a failed trim retries at most `MaxAttemptsPerWeek`; the report reads the persisted record, it never runs pct |
|
||||||
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
|
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
|
||||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
||||||
|
| `backup.BackupSuccessState` | internal/backup/backup_state.go | `RecordBackupSuccess(target, b)` / `LastKnownSuccess(target, vmid)` | Newest SUCCESSFUL backup per tier+guest, persisted (atomic tmp+rename) — the due-check's fallback when the tier's storage cannot be read after a restart (R-894) | **Read ONLY when the storage cannot be read** — a storage that answers is the ground truth (R-84), and an archive absent there must make the tier due even when this file remembers one. A saved copy older than the cadence still reads due. Only successes are written. |
|
||||||
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
||||||
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
|
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
|
||||||
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
|
| `localapi.BackupTier` + `normalizeBackupTiers` / `config.BackupConfig.BackupTiers` | internal/localapi/backup_tiers.go, internal/config/config.go | `normalizeBackupTiers(tiers, legacy, cadence) []BackupTier`; `BackupTiers() ([]BackupTier, []string)` | THE R-82 multi-tier resolution — one runner per tier, primary first | **The untargeted local-API contract is FROZEN**: no `?target=` ⇒ primary tier ⇒ pre-R-82 response BYTES (Target is `omitempty` and stays empty). Never default a missing cadence — reject it and log the warning at ERROR. Never share one retention knob between tiers. Jobs are keyed by (vmid,target). |
|
||||||
|
|||||||
@@ -1901,12 +1901,15 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
|||||||
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
|
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
|
||||||
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
|
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
|
||||||
Store: store,
|
Store: store,
|
||||||
Storage: observer,
|
// R-894: the newest success per tier on disk — the due-check's fallback when the storage cannot be
|
||||||
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
|
// read right after a restart. Same state dir as restore-test-state.json.
|
||||||
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
|
LastKnownBackups: backup.NewBackupSuccessState(filepath.Join(cfg.OOB.WithDefaults().StateDir, "backup-success-state.json")),
|
||||||
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
|
Storage: observer,
|
||||||
Tokens: tokens,
|
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
|
||||||
BackupCadence: cfg.Backup.BackupCadence(),
|
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
|
||||||
|
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
|
||||||
|
Tokens: tokens,
|
||||||
|
BackupCadence: cfg.Backup.BackupCadence(),
|
||||||
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
|
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
|
||||||
Disks: hostOps,
|
Disks: hostOps,
|
||||||
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
|
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
|
||||||
|
|||||||
@@ -0,0 +1,74 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"go/ast"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-894 — the on-disk backup record is WIRED on the daemon path (the built-but-never-wired class).
|
||||||
|
// main → runDaemon → buildLocalAPIServer, and inside it the localapi.Options literal carries
|
||||||
|
// LastKnownBackups built by backup.NewBackupSuccessState. An AST walk, not a string match, for the
|
||||||
|
// reasons in escrow_recover_wiring_test.go.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (observed): delete the `LastKnownBackups:` line from buildLocalAPIServer → this
|
||||||
|
// fails with "localapi.Options in buildLocalAPIServer has no LastKnownBackups field". Restored.
|
||||||
|
func TestR894_LastKnownBackupsIsWiredIntoTheDaemon(t *testing.T) {
|
||||||
|
_, f := parseMain(t)
|
||||||
|
if !callsWithin(f, "main")["runDaemon"] || !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
|
||||||
|
t.Fatal("main → runDaemon → buildLocalAPIServer is broken — the path this test asserts is not the live one")
|
||||||
|
}
|
||||||
|
var field, built bool
|
||||||
|
for _, d := range f.Decls {
|
||||||
|
fd, ok := d.(*ast.FuncDecl)
|
||||||
|
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||||
|
cl, ok := n.(*ast.CompositeLit)
|
||||||
|
if !ok {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
sel, ok := cl.Type.(*ast.SelectorExpr)
|
||||||
|
if !ok {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
if pkg, _ := sel.X.(*ast.Ident); pkg == nil || pkg.Name+"."+sel.Sel.Name != "localapi.Options" {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
for _, el := range cl.Elts {
|
||||||
|
kv, ok := el.(*ast.KeyValueExpr)
|
||||||
|
if !ok {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "LastKnownBackups" {
|
||||||
|
field = true
|
||||||
|
if callsIn(kv.Value)["backup.NewBackupSuccessState"] {
|
||||||
|
built = true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
})
|
||||||
|
}
|
||||||
|
if !field {
|
||||||
|
t.Fatal("localapi.Options in buildLocalAPIServer has no LastKnownBackups field")
|
||||||
|
}
|
||||||
|
if !built {
|
||||||
|
t.Fatal("LastKnownBackups is not built by backup.NewBackupSuccessState")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func callsIn(n ast.Node) map[string]bool {
|
||||||
|
out := map[string]bool{}
|
||||||
|
ast.Inspect(n, func(n ast.Node) bool {
|
||||||
|
if ce, ok := n.(*ast.CallExpr); ok {
|
||||||
|
if fn, ok := ce.Fun.(*ast.SelectorExpr); ok {
|
||||||
|
if x, ok := fn.X.(*ast.Ident); ok {
|
||||||
|
out[x.Name+"."+fn.Sel.Name] = true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
})
|
||||||
|
return out
|
||||||
|
}
|
||||||
+104
-1
@@ -32,6 +32,14 @@
|
|||||||
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
|
# held packages, kernel taint, the crash guard; guest Debian, Docker engine, containerd, live-restore.
|
||||||
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
|
# live-restore-on (v0.142.0, layer guest) the ONE-TIME step of `09` decision 87: merge `"live-restore": true`
|
||||||
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
|
# into the guest's /etc/docker/daemon.json and `systemctl reload docker`. NEVER a restart (R-835).
|
||||||
|
# oom-check (R-528, `09` decision 157; layer docker, lane slow) ONLY the memory-kill check below: no apt, no engine
|
||||||
|
# change, no authority needed. Wrapper-only — the agent never writes this plan; run it by hand as root.
|
||||||
|
# R-528 (`09` decision 157): a docker-layer APPLY also runs `oom_check()` after health_after and reports it as
|
||||||
|
# "oom_check": {"result": "pass"|"fail"|"error", "oom_killed": bool, "oom_event": bool, "exit_code": int|null,
|
||||||
|
# "image": str|null, "detail": str}
|
||||||
|
# — one throwaway container (the controller's own image, no network / volume / port, 64 MB cap) is made to exceed its
|
||||||
|
# memory; "pass" only when the engine says OOMKilled=true AND emits the `oom` event. It never changes the step's
|
||||||
|
# outcome or health: the hub decides whether the engine set can be approved. Pinned by the OOMCheck tests.
|
||||||
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
|
# Output: log lines on stderr and the journal (tag felhom-os-apply); the LAST stdout line is
|
||||||
# OSAPPLY-REPORT <one JSON object>
|
# OSAPPLY-REPORT <one JSON object>
|
||||||
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
|
# which is what the agent parses. Exit 0 = done; 2 = refused (nothing changed); 3 = failed during install.
|
||||||
@@ -68,6 +76,16 @@ JOURNAL_MARK = "@@FELHOM-DPKG-JOURNAL@@"
|
|||||||
DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true"
|
DPKG_STATE_SCRIPT = "dpkg --audit; echo " + JOURNAL_MARK + "; ls -A /var/lib/dpkg/updates 2>/dev/null; true"
|
||||||
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
|
# The installer's ROOT-OWNED record (felhom-host-install.sh `state_set mode`); the agent cannot write it.
|
||||||
INSTALL_STATE = "/var/lib/felhom-install/state.json"
|
INSTALL_STATE = "/var/lib/felhom-install/state.json"
|
||||||
|
# R-528: the memory-kill check (oom_check). One 200 MB block under a 64 MB cap: measured on Docker 29.8.2 to be
|
||||||
|
# OOM-killed with OOMKilled=true and an `oom` event. Every call is bounded: timeouts (the clock read twice) + the
|
||||||
|
# settle wait stay within 90 s.
|
||||||
|
OOMCHECK_PREFIX = "felhom-oomcheck-"
|
||||||
|
OOMCHECK_SCRIPT = "dd if=/dev/zero of=/dev/null bs=200M count=1"
|
||||||
|
OOMCHECK_TIMEOUTS = {"image": 10, "clock": 5, "run": 30, "inspect": 10, "events": 10, "rm": 15}
|
||||||
|
# Measured on demo-hp 9201 (Docker 29.8.2, 2026-10-06, audits/readback-2026-10-07/F/F1, F2): an `--until` taken right
|
||||||
|
# after the run MISSED the oom event although OOMKilled=true; after a 2 s wait and `--until` = guest epoch + 1 it is
|
||||||
|
# seen. Pinned by test_events_window_ends_after_the_settle_wait.
|
||||||
|
OOMCHECK_SETTLE = 2
|
||||||
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
|
# Kernel, boot and firmware packages are the SLOW lane on the host whatever their origin (`11` C3, §5.2): a host
|
||||||
# reboot is needed for them to take effect, and a bad one can stop the box from booting.
|
# reboot is needed for them to take effect, and a bad one can stop the box from booting.
|
||||||
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
|
HOST_SLOW_RE = re.compile(r"^(linux-(image|headers|kbuild|modules|base)|proxmox-kernel|proxmox-default-kernel|pve-kernel|"
|
||||||
@@ -405,7 +423,7 @@ class Apply:
|
|||||||
|
|
||||||
def check_plan(self, plan):
|
def check_plan(self, plan):
|
||||||
mode = plan.get("mode", "apply")
|
mode = plan.get("mode", "apply")
|
||||||
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update"):
|
if mode not in ("apply", "inventory", "health", "facts", "live-restore-on", "bundle", "agent_update", "oom-check"):
|
||||||
raise Refused("R11", f"unknown mode {mode!r}")
|
raise Refused("R11", f"unknown mode {mode!r}")
|
||||||
if mode == "agent_update":
|
if mode == "agent_update":
|
||||||
if plan.get("layer") != "host":
|
if plan.get("layer") != "host":
|
||||||
@@ -427,6 +445,8 @@ class Apply:
|
|||||||
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
|
raise Refused("R11", "facts is a host-layer mode (it reads the host and the guest)")
|
||||||
if mode == "live-restore-on" and layer != "guest":
|
if mode == "live-restore-on" and layer != "guest":
|
||||||
raise Refused("R11", "live-restore-on is a guest-layer mode")
|
raise Refused("R11", "live-restore-on is a guest-layer mode")
|
||||||
|
if mode == "oom-check" and layer != "docker":
|
||||||
|
raise Refused("R11", "oom-check is a docker-layer mode (it checks the guest's Docker engine)")
|
||||||
if plan.get("undo") and layer != "docker":
|
if plan.get("undo") and layer != "docker":
|
||||||
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
|
raise Refused("R5", "an undo (downgrade) exists only for the Docker layer, inside a signed job")
|
||||||
vmid = plan.get("vmid")
|
vmid = plan.get("vmid")
|
||||||
@@ -951,6 +971,10 @@ class Apply:
|
|||||||
log = self.r.log
|
log = self.r.log
|
||||||
if self.mode == "live-restore-on":
|
if self.mode == "live-restore-on":
|
||||||
return self.live_restore_on()
|
return self.live_restore_on()
|
||||||
|
if self.mode == "oom-check":
|
||||||
|
# R-528: the check alone — no apt, no engine change; check_guest above still applies.
|
||||||
|
self.report["oom_check"] = self.oom_check()
|
||||||
|
return 0
|
||||||
self.who, self.allow_downgrade = ("fast", False)
|
self.who, self.allow_downgrade = ("fast", False)
|
||||||
if self.layer == "docker" and self.mode == "apply":
|
if self.layer == "docker" and self.mode == "apply":
|
||||||
self.who, self.allow_downgrade = self.docker_authority(plan)
|
self.who, self.allow_downgrade = self.docker_authority(plan)
|
||||||
@@ -990,8 +1014,87 @@ class Apply:
|
|||||||
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
|
self.report["docker_engine"] = out_v.strip() if rc_v == 0 and out_v.strip() else "unknown"
|
||||||
self.report["reboot_scanned"] = "reboot_needed" in self.report
|
self.report["reboot_scanned"] = "reboot_needed" in self.report
|
||||||
self.report["health_after"] = self.health()
|
self.report["health_after"] = self.health()
|
||||||
|
if self.layer == "docker" and self.mode == "apply":
|
||||||
|
# R-528 (`09` decision 157): does the engine report a memory kill? Reported only — never the outcome.
|
||||||
|
self.report["oom_check"] = self.oom_check()
|
||||||
return 0
|
return 0
|
||||||
|
|
||||||
|
# ---------- R-528: the memory-kill check ----------
|
||||||
|
def oom_check(self):
|
||||||
|
"""Run one throwaway container over its memory cap in the guest and read what the engine says about it.
|
||||||
|
Never raises: any failure becomes result "error". The container is ALWAYS removed (finally), and a failed
|
||||||
|
removal is named in the detail."""
|
||||||
|
t = OOMCHECK_TIMEOUTS
|
||||||
|
res = {"result": "error", "oom_killed": False, "oom_event": False, "exit_code": None, "image": None, "detail": ""}
|
||||||
|
log = self.r.log
|
||||||
|
try:
|
||||||
|
rc, out, err = self.g(["docker", "inspect", "-f", "{{.Config.Image}}", "felhom-controller"], timeout=t["image"])
|
||||||
|
except Exception as e:
|
||||||
|
rc, out, err = -1, "", str(e)
|
||||||
|
img = out.strip() if rc == 0 else ""
|
||||||
|
if not img or any(c.isspace() for c in img):
|
||||||
|
res["detail"] = f"the controller's image could not be read (rc={rc}): {(err or out).strip()[:200]}"
|
||||||
|
log(f"os-apply: OOM-CHECK error — {res['detail']}")
|
||||||
|
return res
|
||||||
|
res["image"] = img
|
||||||
|
name = f"{OOMCHECK_PREFIX}{os.getpid()}-{os.urandom(4).hex()}"
|
||||||
|
notes = []
|
||||||
|
try:
|
||||||
|
t0 = self.guest_epoch()
|
||||||
|
if t0 is None:
|
||||||
|
raise RuntimeError("the guest clock could not be read")
|
||||||
|
rrc, rout, rerr = self.g(["docker", "run", "--name", name, "--pull", "never", "--network", "none",
|
||||||
|
"--memory", "64m", "--memory-swap", "64m", "--label", "felhom.oomcheck=1",
|
||||||
|
"--entrypoint", "sh", img, "-c", OOMCHECK_SCRIPT], timeout=t["run"])
|
||||||
|
self.r.sleep(OOMCHECK_SETTLE) # the engine publishes the oom event a moment after the run returns (F1/F2)
|
||||||
|
t1 = self.guest_epoch()
|
||||||
|
if t1 is None:
|
||||||
|
t1 = t0 + t["run"] + OOMCHECK_SETTLE + 1
|
||||||
|
irc, iout, ierr = self.g(["docker", "inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}", name], timeout=t["inspect"])
|
||||||
|
if irc != 0:
|
||||||
|
raise RuntimeError(f"the check container could not be inspected (run rc={rrc}: {(rerr or rout).strip()[:120]}; "
|
||||||
|
f"inspect rc={irc}: {(ierr or iout).strip()[:120]})")
|
||||||
|
f = iout.split()
|
||||||
|
res["oom_killed"] = bool(f) and f[0] == "true"
|
||||||
|
try:
|
||||||
|
res["exit_code"] = int(f[1]) if len(f) > 1 else None
|
||||||
|
except ValueError:
|
||||||
|
res["exit_code"] = None
|
||||||
|
erc, eout, eerr = self.g(["docker", "events", "--since", str(t0 - 1), "--until", str(t1 + 1),
|
||||||
|
"--filter", f"container={name}", "--filter", "event=oom",
|
||||||
|
"--format", "{{.Action}}"], timeout=t["events"])
|
||||||
|
if erc != 0:
|
||||||
|
notes.append(f"the event read failed (rc={erc}): {(eerr or eout).strip()[:120]}")
|
||||||
|
res["oom_event"] = erc == 0 and any(l.strip() == "oom" for l in eout.splitlines())
|
||||||
|
if res["oom_killed"] and res["oom_event"]:
|
||||||
|
res["result"] = "pass"
|
||||||
|
notes.insert(0, "the engine reported the memory kill: OOMKilled=true and the oom event")
|
||||||
|
else:
|
||||||
|
res["result"] = "fail"
|
||||||
|
miss = [w for w, ok in (("OOMKilled=true", res["oom_killed"]), ("the oom event", res["oom_event"])) if not ok]
|
||||||
|
notes.insert(0, f"the engine did not report the memory kill: missing {' and '.join(miss)} (exit code {res['exit_code']})")
|
||||||
|
except Exception as e:
|
||||||
|
res["result"] = "error"
|
||||||
|
notes.insert(0, f"the check could not finish: {type(e).__name__}: {str(e)[:200]}")
|
||||||
|
finally:
|
||||||
|
try:
|
||||||
|
mrc, mout, merr = self.g(["docker", "rm", "-f", name], timeout=t["rm"])
|
||||||
|
if mrc != 0 and "no such container" not in (merr + mout).lower():
|
||||||
|
notes.append(f"the check container {name} could not be removed (rc={mrc}): {(merr or mout).strip()[:120]}")
|
||||||
|
except Exception as e:
|
||||||
|
notes.append(f"the check container {name} could not be removed: {type(e).__name__}: {str(e)[:120]}")
|
||||||
|
res["detail"] = "; ".join(notes)
|
||||||
|
log(f"os-apply: OOM-CHECK result={res['result']} oom_killed={res['oom_killed']} oom_event={res['oom_event']} "
|
||||||
|
f"exit={res['exit_code']} image={img} — {res['detail']}")
|
||||||
|
return res
|
||||||
|
|
||||||
|
def guest_epoch(self):
|
||||||
|
rc, out, _ = self.g(["date", "+%s"], timeout=OOMCHECK_TIMEOUTS["clock"])
|
||||||
|
try:
|
||||||
|
return int(out.strip()) if rc == 0 else None
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
|
||||||
def dpkg_state(self):
|
def dpkg_state(self):
|
||||||
"""`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an
|
"""`dpkg --audit` AND dpkg's update journal, in ONE call (R-876, agent v0.145.0). A crash in the middle of an
|
||||||
install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp
|
install can leave `/var/lib/dpkg/updates/` non-empty while `--audit` reads clean — measured on demo-hp
|
||||||
|
|||||||
@@ -83,7 +83,7 @@ class Fake:
|
|||||||
return self.clock
|
return self.clock
|
||||||
|
|
||||||
def sleep(self, s):
|
def sleep(self, s):
|
||||||
pass
|
self.sleeps = getattr(self, "sleeps", []) + [(len(self.calls), s)] # (calls made before it, seconds)
|
||||||
|
|
||||||
def verify_sig(self, signers, key_id, ns, blob, sig):
|
def verify_sig(self, signers, key_id, ns, blob, sig):
|
||||||
self.verified = (signers, key_id, ns, blob, sig)
|
self.verified = (signers, key_id, ns, blob, sig)
|
||||||
@@ -196,6 +196,27 @@ class Fake:
|
|||||||
return 0, self.engine + "\n", ""
|
return 0, self.engine + "\n", ""
|
||||||
if cmd == "docker" and a[1:3] == ["ps", "-q"]:
|
if cmd == "docker" and a[1:3] == ["ps", "-q"]:
|
||||||
return 0, "".join(i + "\n" for i in self.ids), ""
|
return 0, "".join(i + "\n" for i in self.ids), ""
|
||||||
|
# R-528: the memory-kill check's engine (oom_image None = unreadable; oom_state / oom_event the engine's answer)
|
||||||
|
if cmd == "date" and a[1:] == ["+%s"]:
|
||||||
|
# each read is 3 s later than the last, so the order of the reads is visible in the values
|
||||||
|
self.oom_epochs = getattr(self, "oom_epochs", []) + [int(self.clock) + 3 * len(getattr(self, "oom_epochs", []))]
|
||||||
|
return 0, f"{self.oom_epochs[-1]}\n", ""
|
||||||
|
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.Config.Image}}"]:
|
||||||
|
img = getattr(self, "oom_image", "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
|
||||||
|
return (0, img + "\n", "") if img is not None else (1, "", "Error: No such object: felhom-controller")
|
||||||
|
if cmd == "docker" and a[1] == "run":
|
||||||
|
self.oom_runs = getattr(self, "oom_runs", []) + [a]
|
||||||
|
return 137, "", ""
|
||||||
|
if cmd == "docker" and a[1:4] == ["inspect", "-f", "{{.State.OOMKilled}} {{.State.ExitCode}}"]:
|
||||||
|
if getattr(self, "oom_inspect_raises", False):
|
||||||
|
raise subprocess.TimeoutExpired(a, 10)
|
||||||
|
return 0, getattr(self, "oom_state", "true 137") + "\n", ""
|
||||||
|
if cmd == "docker" and a[1] == "events":
|
||||||
|
self.oom_events_argv = a
|
||||||
|
return 0, ("oom\n" if getattr(self, "oom_event", True) else ""), ""
|
||||||
|
if cmd == "docker" and a[1:3] == ["rm", "-f"]:
|
||||||
|
self.oom_removed = getattr(self, "oom_removed", []) + a[3:]
|
||||||
|
return 0, a[3] + "\n", ""
|
||||||
if cmd == "docker" and a[1] == "inspect":
|
if cmd == "docker" and a[1] == "inspect":
|
||||||
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
|
mounts = {"aaa111": "/felhom-controller|/var/run/docker.sock;/app/data;", "bbb222": "/app|/data;"}
|
||||||
return 0, mounts.get(a[-1], "/other|;") + "\n", ""
|
return 0, mounts.get(a[-1], "/other|;") + "\n", ""
|
||||||
@@ -997,8 +1018,6 @@ class RealSignatureCheck(unittest.TestCase):
|
|||||||
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
|
self.assertNotEqual(r.verify_sig(self.signers, "someone-else", "felhom-op-v1", blob, sig), 0)
|
||||||
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
|
self.assertNotEqual(r.verify_sig(self.signers, "felhom-op-1", "felhom-op-v1", blob, self.sign(blob, ns="other-ns")), 0)
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
unittest.main()
|
|
||||||
|
|
||||||
|
|
||||||
class UnsentReport(unittest.TestCase):
|
class UnsentReport(unittest.TestCase):
|
||||||
@@ -1179,3 +1198,131 @@ class CrashLeftTheJournal(unittest.TestCase):
|
|||||||
self.assertEqual(rc, 0, rep)
|
self.assertEqual(rc, 0, rep)
|
||||||
self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs)
|
self.assertTrue(any("INTERRUPTED" in l for l in f.logs), f.logs)
|
||||||
self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs)
|
self.assertTrue(any(l.startswith("os-apply: REPAIR ") and l.endswith("forced") for l in f.logs), f.logs)
|
||||||
|
|
||||||
|
|
||||||
|
class OOMCheck(unittest.TestCase):
|
||||||
|
"""R-528 (`09` decision 157): after a Docker engine step the wrapper proves the engine reports a memory kill
|
||||||
|
(OOMKilled=true AND the `oom` event). Reported only; the hub decides. Red-proofs: audits/readback-2026-10-07/F/."""
|
||||||
|
|
||||||
|
def apply(self, **kw):
|
||||||
|
f = docker_fake(signed=signed_job())
|
||||||
|
for k, v in kw.items():
|
||||||
|
setattr(f, k, v)
|
||||||
|
rc, rep = run(f)
|
||||||
|
self.assertEqual(rc, 0, rep)
|
||||||
|
return f, rep
|
||||||
|
|
||||||
|
def test_pass_when_oomkilled_and_the_event(self):
|
||||||
|
f, rep = self.apply()
|
||||||
|
oc = rep["oom_check"]
|
||||||
|
self.assertEqual(oc["result"], "pass", oc)
|
||||||
|
self.assertEqual((oc["oom_killed"], oc["oom_event"], oc["exit_code"]), (True, True, 137))
|
||||||
|
self.assertEqual(oc["image"], "gitea.dooplex.hu/admin/felhom-controller:0.300.0")
|
||||||
|
self.assertEqual(sorted(oc), ["detail", "exit_code", "image", "oom_event", "oom_killed", "result"])
|
||||||
|
run_argv = f.oom_runs[0]
|
||||||
|
name = run_argv[run_argv.index("--name") + 1]
|
||||||
|
self.assertTrue(re.match(r"^felhom-oomcheck-[0-9]+-[0-9a-f]{8}$", name), name)
|
||||||
|
for flag, val in (("--pull", "never"), ("--network", "none"), ("--memory", "64m"), ("--memory-swap", "64m"),
|
||||||
|
("--label", "felhom.oomcheck=1"), ("--entrypoint", "sh")):
|
||||||
|
self.assertEqual(run_argv[run_argv.index(flag) + 1], val, flag)
|
||||||
|
self.assertNotIn("-v", run_argv)
|
||||||
|
self.assertNotIn("-p", run_argv)
|
||||||
|
self.assertEqual(run_argv[-3:], ["gitea.dooplex.hu/admin/felhom-controller:0.300.0", "-c", osapply.OOMCHECK_SCRIPT])
|
||||||
|
self.assertIn(f"container={name}", f.oom_events_argv)
|
||||||
|
self.assertIn("event=oom", f.oom_events_argv)
|
||||||
|
self.assertEqual(f.oom_removed, [name], "the check container must be removed")
|
||||||
|
self.assertTrue(rep["health_after"], "health is read before the check")
|
||||||
|
# bounded: every call has a timeout and the clock is read twice — the worst case stays within 90 s
|
||||||
|
self.assertLessEqual(sum(osapply.OOMCHECK_TIMEOUTS.values()) + osapply.OOMCHECK_TIMEOUTS["clock"]
|
||||||
|
+ osapply.OOMCHECK_SETTLE, 90)
|
||||||
|
|
||||||
|
def test_events_window_ends_after_the_settle_wait(self):
|
||||||
|
# Measured (F1/F2): an --until taken right after the run missed the oom event. The window must end after a
|
||||||
|
# wait of >= 2 s that comes AFTER the run, at the guest epoch read after that wait, + 1.
|
||||||
|
f, rep = self.apply()
|
||||||
|
run_i = next(i for i, c in enumerate(f.calls) if c[-1][:2] == ["docker", "run"])
|
||||||
|
date_i = [i for i, c in enumerate(f.calls) if c[-1] == ["date", "+%s"]]
|
||||||
|
waits = [(i, s) for i, s in getattr(f, "sleeps", []) if i > run_i]
|
||||||
|
self.assertTrue(waits and waits[0][1] >= 2, f"no settle wait after the run: {getattr(f, 'sleeps', None)}")
|
||||||
|
self.assertTrue(date_i[-1] >= waits[0][0], "the end epoch must be read after the wait")
|
||||||
|
ev = f.oom_events_argv
|
||||||
|
since, until = int(ev[ev.index("--since") + 1]), int(ev[ev.index("--until") + 1])
|
||||||
|
self.assertEqual(until, f.oom_epochs[-1] + 1, "until = the guest epoch read after the wait, + 1")
|
||||||
|
self.assertGreater(until, f.oom_epochs[0] + 1)
|
||||||
|
self.assertLess(since, f.oom_epochs[0] + 1)
|
||||||
|
|
||||||
|
def test_oomkilled_false_is_fail(self):
|
||||||
|
f, rep = self.apply(oom_state="false 0")
|
||||||
|
oc = rep["oom_check"]
|
||||||
|
self.assertEqual(oc["result"], "fail", oc)
|
||||||
|
self.assertIn("OOMKilled=true", oc["detail"])
|
||||||
|
self.assertTrue(oc["oom_event"])
|
||||||
|
self.assertEqual(len(f.oom_removed), 1)
|
||||||
|
|
||||||
|
def test_no_event_is_fail(self):
|
||||||
|
f, rep = self.apply(oom_event=False)
|
||||||
|
oc = rep["oom_check"]
|
||||||
|
self.assertEqual(oc["result"], "fail", oc)
|
||||||
|
self.assertIn("the oom event", oc["detail"])
|
||||||
|
self.assertTrue(oc["oom_killed"])
|
||||||
|
|
||||||
|
def test_image_unreadable_is_error(self):
|
||||||
|
f, rep = self.apply(oom_image=None)
|
||||||
|
oc = rep["oom_check"]
|
||||||
|
self.assertEqual(oc["result"], "error", oc)
|
||||||
|
self.assertIsNone(oc["image"])
|
||||||
|
self.assertIn("image could not be read", oc["detail"])
|
||||||
|
self.assertFalse(hasattr(f, "oom_runs"), "no container is started without an image")
|
||||||
|
|
||||||
|
def test_container_removed_even_when_inspect_raises(self):
|
||||||
|
f, rep = self.apply(oom_inspect_raises=True)
|
||||||
|
oc = rep["oom_check"]
|
||||||
|
self.assertEqual(oc["result"], "error", oc)
|
||||||
|
self.assertEqual(len(f.oom_removed), 1, "the container must be removed even when inspect raised")
|
||||||
|
self.assertTrue(f.oom_removed[0].startswith(osapply.OOMCHECK_PREFIX))
|
||||||
|
self.assertNotIn("failed", rep, "the check never turns the step into a failure")
|
||||||
|
|
||||||
|
def test_never_on_guest_or_host_or_in_health_mode(self):
|
||||||
|
g = Fake()
|
||||||
|
rc, rep = run(g)
|
||||||
|
self.assertEqual(rc, 0, rep)
|
||||||
|
h = Fake()
|
||||||
|
h.plan["layer"] = "host"
|
||||||
|
h.plan["packages"] = [{"name": "bash", "version": "5.2.37-2+b10", "origin": "Debian"}]
|
||||||
|
h.installed["bash"] = "5.2.37-2+b9"
|
||||||
|
h.live["bash"] = {"5.2.37-2+b10"}
|
||||||
|
rc_h, rep_h = run(h)
|
||||||
|
self.assertEqual(rc_h, 0, rep_h)
|
||||||
|
d = docker_fake(signed=signed_job())
|
||||||
|
d.plan["mode"] = "health"
|
||||||
|
rc_d, rep_d = run(d)
|
||||||
|
self.assertEqual(rc_d, 0, rep_d)
|
||||||
|
for f, r in ((g, rep), (h, rep_h), (d, rep_d)):
|
||||||
|
self.assertNotIn("oom_check", r)
|
||||||
|
self.assertFalse(hasattr(f, "oom_runs"), r.get("layer"))
|
||||||
|
self.assertFalse(any(c[-1][:2] == ["docker", "run"] for c in f.calls))
|
||||||
|
|
||||||
|
def test_mode_oom_check_runs_only_the_check(self):
|
||||||
|
f = docker_fake() # no authority: the check changes nothing, so it needs none
|
||||||
|
f.plan["mode"], f.plan["packages"] = "oom-check", []
|
||||||
|
rc, rep = run(f)
|
||||||
|
self.assertEqual(rc, 0, rep)
|
||||||
|
self.assertEqual(rep["oom_check"]["result"], "pass", rep)
|
||||||
|
self.assertFalse(any("apt-get" in c[-1] or "dpkg-query" in c[-1] for c in f.calls), f.calls)
|
||||||
|
self.assertEqual(len(f.oom_removed), 1)
|
||||||
|
|
||||||
|
def test_mode_oom_check_keeps_the_refusals(self):
|
||||||
|
f = docker_fake()
|
||||||
|
f.plan["mode"], f.plan["packages"] = "oom-check", []
|
||||||
|
f.files["/etc/pve/lxc/9201.conf"] = "arch: amd64\n"
|
||||||
|
rc, rep = run(f)
|
||||||
|
self.assertEqual((rc, rep["refused"]["code"]), (2, "R10"), rep)
|
||||||
|
self.assertFalse(hasattr(f, "oom_runs"))
|
||||||
|
g = Fake()
|
||||||
|
g.plan["mode"] = "oom-check"
|
||||||
|
rc, rep = run(g)
|
||||||
|
self.assertEqual((rc, rep["refused"]["code"]), (2, "R11"), rep)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
|
|||||||
@@ -0,0 +1,131 @@
|
|||||||
|
package backup
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"sort"
|
||||||
|
"strconv"
|
||||||
|
"sync"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||||
|
)
|
||||||
|
|
||||||
|
// BackupSuccessState persists the newest SUCCESSFUL whole-guest backup per tier and guest (R-894).
|
||||||
|
//
|
||||||
|
// Why it exists. The due-check (`localapi` handleBackupDue) asks the tier's storage when a backup last
|
||||||
|
// landed (R-84) and falls back to the in-memory record when the storage cannot be read. The in-memory
|
||||||
|
// record is empty after an agent restart (Store, R-348), so "storage unreadable" right after a restart
|
||||||
|
// read as "no record — DUE". Measured 2026-10-05 on demo-hp: the agent restarted at 04:57, the off-site
|
||||||
|
// storage answered "Can't connect" at 06:25, the 7-day tier — last copy 2026-10-01 — read DUE, the
|
||||||
|
// controller asked, and vzdump failed. This file is the last known copy the fallback reads instead.
|
||||||
|
//
|
||||||
|
// It is read ONLY when the storage cannot be read. A storage that answers is the ground truth and wins,
|
||||||
|
// in both directions: an archive found there counts, and an archive absent there is absent even when
|
||||||
|
// this file remembers a success (a pruned or deleted archive must make the tier due — the same reason
|
||||||
|
// R-84 chose the storage over a persisted record). Pinned by
|
||||||
|
// TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers.
|
||||||
|
//
|
||||||
|
// Only SUCCESSES are written (the RestoreTestState rule): a failure must stay due and be retried, so a
|
||||||
|
// record of a failure has no reader.
|
||||||
|
type BackupSuccessState struct {
|
||||||
|
path string
|
||||||
|
mu sync.Mutex
|
||||||
|
last map[string]savedSuccess // key(target, vmid) → the newest success
|
||||||
|
}
|
||||||
|
|
||||||
|
type savedSuccess struct {
|
||||||
|
target string
|
||||||
|
vmid int
|
||||||
|
at time.Time
|
||||||
|
}
|
||||||
|
|
||||||
|
// backupSuccessJSON is one entry on disk.
|
||||||
|
type backupSuccessJSON struct {
|
||||||
|
Target string `json:"target"`
|
||||||
|
VMID int `json:"vmid"`
|
||||||
|
StartedAt string `json:"started_at"`
|
||||||
|
}
|
||||||
|
|
||||||
|
func backupStateKey(target string, vmid int) string { return target + "/" + strconv.Itoa(vmid) }
|
||||||
|
|
||||||
|
// NewBackupSuccessState opens (or creates) the state at path. A missing or unreadable file degrades to
|
||||||
|
// "nothing known" — the pre-R-894 behaviour, which is DUE — and never wedges the daemon.
|
||||||
|
func NewBackupSuccessState(path string) *BackupSuccessState {
|
||||||
|
s := &BackupSuccessState{path: path, last: map[string]savedSuccess{}}
|
||||||
|
data, err := os.ReadFile(path)
|
||||||
|
if err != nil {
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
var entries []backupSuccessJSON
|
||||||
|
if json.Unmarshal(data, &entries) != nil {
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
for _, e := range entries {
|
||||||
|
t, perr := time.Parse(time.RFC3339, e.StartedAt)
|
||||||
|
if perr != nil {
|
||||||
|
continue // one unreadable entry must not lose the others
|
||||||
|
}
|
||||||
|
s.last[backupStateKey(e.Target, e.VMID)] = savedSuccess{target: e.Target, vmid: e.VMID, at: t.UTC()}
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
// RecordBackupSuccess saves b when it is a success newer than the one on file. target is the tier the
|
||||||
|
// job ran on (the due-check's key); a failure or an unparseable time is ignored.
|
||||||
|
func (s *BackupSuccessState) RecordBackupSuccess(target string, b hub.Backup) error {
|
||||||
|
if s == nil || !b.Success {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
t, err := time.Parse(time.RFC3339, b.StartedAt)
|
||||||
|
if err != nil {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
s.mu.Lock()
|
||||||
|
defer s.mu.Unlock()
|
||||||
|
k := backupStateKey(target, b.VMID)
|
||||||
|
if old, ok := s.last[k]; ok && !t.After(old.at) {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
s.last[k] = savedSuccess{target: target, vmid: b.VMID, at: t.UTC()}
|
||||||
|
return s.saveLocked()
|
||||||
|
}
|
||||||
|
|
||||||
|
// LastKnownSuccess returns the newest saved success for this tier and guest (ok=false = none on file).
|
||||||
|
func (s *BackupSuccessState) LastKnownSuccess(target string, vmid int) (time.Time, bool) {
|
||||||
|
if s == nil {
|
||||||
|
return time.Time{}, false
|
||||||
|
}
|
||||||
|
s.mu.Lock()
|
||||||
|
defer s.mu.Unlock()
|
||||||
|
e, ok := s.last[backupStateKey(target, vmid)]
|
||||||
|
return e.at, ok
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *BackupSuccessState) saveLocked() error {
|
||||||
|
entries := make([]backupSuccessJSON, 0, len(s.last))
|
||||||
|
for _, e := range s.last {
|
||||||
|
entries = append(entries, backupSuccessJSON{Target: e.target, VMID: e.vmid, StartedAt: e.at.Format(time.RFC3339)})
|
||||||
|
}
|
||||||
|
// Deterministic file content (Go's map order is random).
|
||||||
|
sort.Slice(entries, func(i, j int) bool {
|
||||||
|
if entries[i].Target != entries[j].Target {
|
||||||
|
return entries[i].Target < entries[j].Target
|
||||||
|
}
|
||||||
|
return entries[i].VMID < entries[j].VMID
|
||||||
|
})
|
||||||
|
data, err := json.MarshalIndent(entries, "", " ")
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
if err := os.MkdirAll(filepath.Dir(s.path), 0o755); err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
tmp := s.path + ".tmp"
|
||||||
|
if err := os.WriteFile(tmp, data, 0o600); err != nil {
|
||||||
|
os.Remove(tmp)
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
return os.Rename(tmp, s.path)
|
||||||
|
}
|
||||||
@@ -0,0 +1,59 @@
|
|||||||
|
package backup
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-894: the on-disk newest success per tier survives a restart (a new state from the same file).
|
||||||
|
func TestBackupSuccessState_SurvivesRestart(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||||
|
s := NewBackupSuccessState(path)
|
||||||
|
at := time.Date(2026, 10, 1, 20, 15, 0, 0, time.UTC)
|
||||||
|
if err := s.RecordBackupSuccess("felhom-pbs", hub.Backup{VMID: 9201, Success: true, StartedAt: at.Format(time.RFC3339)}); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
got, ok := NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 9201)
|
||||||
|
if !ok || !got.Equal(at) {
|
||||||
|
t.Fatalf("after a restart the saved copy must read back; got %v ok=%v", got, ok)
|
||||||
|
}
|
||||||
|
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("local", 9201); ok {
|
||||||
|
t.Fatal("another tier must not borrow this tier's copy")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Only a NEWER success replaces the saved one; failures and unparseable times are ignored.
|
||||||
|
func TestBackupSuccessState_KeepsNewestSuccessOnly(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "s.json")
|
||||||
|
s := NewBackupSuccessState(path)
|
||||||
|
newer := time.Date(2026, 10, 5, 0, 0, 0, 0, time.UTC)
|
||||||
|
older := newer.Add(-48 * time.Hour)
|
||||||
|
for _, b := range []hub.Backup{
|
||||||
|
{VMID: 1, Success: true, StartedAt: newer.Format(time.RFC3339)},
|
||||||
|
{VMID: 1, Success: true, StartedAt: older.Format(time.RFC3339)}, // older: ignored
|
||||||
|
{VMID: 1, Success: false, StartedAt: newer.Add(time.Hour).Format(time.RFC3339)}, // failure: ignored
|
||||||
|
{VMID: 1, Success: true, StartedAt: "not-a-time"}, // unparseable: ignored
|
||||||
|
} {
|
||||||
|
if err := s.RecordBackupSuccess("t", b); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if got, _ := NewBackupSuccessState(path).LastKnownSuccess("t", 1); !got.Equal(newer) {
|
||||||
|
t.Fatalf("want the newest success %v, got %v", newer, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A corrupt file degrades to "nothing known" (the pre-R-894 DUE answer), never a crash.
|
||||||
|
func TestBackupSuccessState_CorruptFileIsEmpty(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "s.json")
|
||||||
|
if err := os.WriteFile(path, []byte("{not json"), 0o600); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if _, ok := NewBackupSuccessState(path).LastKnownSuccess("t", 1); ok {
|
||||||
|
t.Fatal("a corrupt file must read as nothing known")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -30,6 +30,9 @@ import (
|
|||||||
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
|
// 2026-08-20, two consecutive host-reports with `0 backups` while `pvesm list` showed archives on both tiers. What
|
||||||
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
|
// is unaffected is the hub's VERDICT: it looks back 7 days over stored reports (felhom.eu hub/internal/monitor/
|
||||||
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
|
// deadline.go backupEvidenceLookback) and the storage stays the ground truth (R-84).
|
||||||
|
// The due-check's fallback for an UNREADABLE storage no longer reads this store alone (R-894): the newest
|
||||||
|
// success per tier is also on disk (BackupSuccessState), so a restart followed by an unreachable storage
|
||||||
|
// reads the last known copy, not "never".
|
||||||
type Store struct {
|
type Store struct {
|
||||||
mu sync.Mutex
|
mu sync.Mutex
|
||||||
byTarget map[string]hub.Backup // latest backup per target id
|
byTarget map[string]hub.Backup // latest backup per target id
|
||||||
|
|||||||
@@ -454,6 +454,17 @@ type SmartSummary struct {
|
|||||||
ReallocatedSectors *int `json:"reallocated_sectors"`
|
ReallocatedSectors *int `json:"reallocated_sectors"`
|
||||||
PendingSectors *int `json:"pending_sectors"`
|
PendingSectors *int `json:"pending_sectors"`
|
||||||
OfflineUncorrectable *int `json:"offline_uncorrectable"`
|
OfflineUncorrectable *int `json:"offline_uncorrectable"`
|
||||||
|
// R-330 (disk health Phase 2): three more SATA raw counters. omitempty + pointer: absent (an
|
||||||
|
// older agent, an NVMe/USB device, or a drive that does not report the attribute) is OMITTED —
|
||||||
|
// unknown, never a zero (S-39). Wire only: no verdict reads them yet.
|
||||||
|
// 187 Reported_Uncorrect — the failing drive's most telling counter (normalized 1 vs thresh 0,
|
||||||
|
// raw 1001) while SMART still said PASSED.
|
||||||
|
// 188 Command_Timeout — some vendors PACK several counters into the 48-bit raw value, so the
|
||||||
|
// number is carried as reported and must not be compared across vendors.
|
||||||
|
// 199 UDMA_CRC_Error_Count — cabling / link errors, not the medium.
|
||||||
|
ReportedUncorrect *int64 `json:"reported_uncorrect,omitempty"`
|
||||||
|
CommandTimeout *int64 `json:"command_timeout,omitempty"`
|
||||||
|
UDMACRCErrors *int64 `json:"udma_crc_errors,omitempty"`
|
||||||
|
|
||||||
// NVMe attributes.
|
// NVMe attributes.
|
||||||
CriticalWarning *int `json:"critical_warning"`
|
CriticalWarning *int `json:"critical_warning"`
|
||||||
|
|||||||
@@ -0,0 +1,144 @@
|
|||||||
|
package localapi
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"io"
|
||||||
|
"log/slog"
|
||||||
|
"net/http"
|
||||||
|
"path/filepath"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-894 — after an agent restart, an UNREADABLE storage must fall back to the last success saved on
|
||||||
|
// disk, not to "never". Measured 2026-10-05 on demo-hp: a restart at 04:57, the off-site storage
|
||||||
|
// unreachable at 06:25, the 7-day tier (last copy 4 days old) read DUE, vzdump failed.
|
||||||
|
//
|
||||||
|
// Every test here builds a NEW server and a NEW BackupSuccessState from the same file — that is the
|
||||||
|
// restart. The in-memory store (fakeStore) is always fresh, as after a real restart.
|
||||||
|
|
||||||
|
// r894Server builds a two-tier server whose off-site tier answers the storage listing with lister.
|
||||||
|
func r894Server(t *testing.T, path string, pbsSvc BackupService) *Server {
|
||||||
|
t.Helper()
|
||||||
|
srv, err := NewServer(Options{
|
||||||
|
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: &fakeStore{},
|
||||||
|
Storage: fakeStorage{targets: []hub.StorageTarget{{Name: "local"}, {Name: "felhom-pbs"}}},
|
||||||
|
Tokens: staticTokens{"A": 8200},
|
||||||
|
BackupTiers: []BackupTier{
|
||||||
|
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
|
||||||
|
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: pbsSvc},
|
||||||
|
},
|
||||||
|
LastKnownBackups: backup.NewBackupSuccessState(path),
|
||||||
|
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
srv.baseCtx = context.Background()
|
||||||
|
srv.now = func() time.Time { return testNow }
|
||||||
|
return srv
|
||||||
|
}
|
||||||
|
|
||||||
|
// unreadable is the off-site storage as demo-hp saw it: "Can't connect to 10.77.0.1:8007".
|
||||||
|
func unreadable() archiveLister {
|
||||||
|
return archiveLister{fakeBackups: &fakeBackups{}, err: errStorageRead}
|
||||||
|
}
|
||||||
|
|
||||||
|
// THE R-894 CASE, end to end. Agent 1 takes an off-site backup through POST /backup (the fake
|
||||||
|
// runner's success is 12 h before testNow). The agent restarts. The storage cannot be read. The tier
|
||||||
|
// must read NOT due, from the copy saved on disk.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (observed): delete the `lookup == archiveUnknown && s.lastKnown != nil` block in
|
||||||
|
// handleBackupDue → this fails with "after a restart an unreadable storage must fall back to the saved
|
||||||
|
// copy (12 h old, 7-day tier) — NOT due; got {… Due:true … AgeState:unknown …}". Restored.
|
||||||
|
func TestBackupDue_R894_RestartThenUnreadableStorage_FreshSavedCopyIsNotDue(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||||
|
|
||||||
|
// Agent 1: a real backup job through the endpoint the controller calls.
|
||||||
|
first := r894Server(t, path, &fakeBackups{})
|
||||||
|
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
|
||||||
|
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
|
||||||
|
}
|
||||||
|
waitFor(t, func() bool {
|
||||||
|
_, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200)
|
||||||
|
return ok
|
||||||
|
})
|
||||||
|
|
||||||
|
// Agent 2: a restart (new server, new state from the same file), and the storage is unreachable.
|
||||||
|
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||||
|
if got.Due {
|
||||||
|
t.Fatalf("after a restart an unreadable storage must fall back to the saved copy (12 h old, 7-day tier) — NOT due; got %+v", got)
|
||||||
|
}
|
||||||
|
if got.AgeState != AgeStateKnown || got.AgeSecs == nil || *got.AgeSecs != int64((12*time.Hour).Seconds()) {
|
||||||
|
t.Fatalf("the age must come from the saved copy (12 h, known); got %+v", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The deliberate rule stays: an unreadable storage must not suppress a backup that IS due. A saved
|
||||||
|
// copy older than the cadence reads DUE.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (observed): make the fallback answer not-due whenever a saved copy exists
|
||||||
|
// (`if fromDisk { …Due:false… }` before the cadence check) → this fails with "a saved copy 9 days old
|
||||||
|
// under a 7-day cadence MUST read due". Restored.
|
||||||
|
func TestBackupDue_R894_RestartThenUnreadableStorage_OldSavedCopyIsDue(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||||
|
st := backup.NewBackupSuccessState(path)
|
||||||
|
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, 9*24*time.Hour, true)); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||||
|
if !got.Due {
|
||||||
|
t.Fatalf("a saved copy 9 days old under a 7-day cadence MUST read due; got %+v", got)
|
||||||
|
}
|
||||||
|
if got.AgeState != AgeStateKnown {
|
||||||
|
t.Fatalf("the age is known (from disk); got %+v", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// No saved copy → the pre-R-894 answer, byte for byte: DUE, age UNKNOWN (never ABSENT — the controller
|
||||||
|
// fires its window-gate valve only on absent, R-88).
|
||||||
|
func TestBackupDue_R894_RestartThenUnreadableStorage_NoSavedCopyIsDueUnknown(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||||
|
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||||
|
if !got.Due || got.AgeState != AgeStateUnknown || got.AgeSecs != nil {
|
||||||
|
t.Fatalf("no saved copy + unreadable storage must stay DUE with age unknown; got %+v", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A storage that ANSWERS is the ground truth: an archive absent there makes the tier due even when the
|
||||||
|
// file remembers a fresh success (a pruned or deleted copy must be made again).
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (observed): drop `lookup == archiveUnknown &&` from the fallback condition → this
|
||||||
|
// fails with "the storage answered 'no archive' — the saved copy must NOT stand in for it". Restored.
|
||||||
|
func TestBackupDue_R894_SavedCopyIgnoredWhenStorageAnswers(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||||
|
st := backup.NewBackupSuccessState(path)
|
||||||
|
if err := st.RecordBackupSuccess("felhom-pbs", backupAt("felhom-pbs", 8200, time.Hour, true)); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
absent := archiveLister{fakeBackups: &fakeBackups{}, found: false}
|
||||||
|
got := dueFor(t, r894Server(t, path, absent).Handler(), "felhom-pbs")
|
||||||
|
if !got.Due {
|
||||||
|
t.Fatalf("the storage answered 'no archive' — the saved copy must NOT stand in for it; got %+v", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A FAILED backup is never saved: it must not make a tier look fresh after a restart.
|
||||||
|
func TestBackupDue_R894_FailedBackupIsNotSaved(t *testing.T) {
|
||||||
|
path := filepath.Join(t.TempDir(), "backup-success-state.json")
|
||||||
|
first := r894Server(t, path, &fakeBackups{failErr: "could not activate storage 'felhom-pbs'"})
|
||||||
|
if rr := do(t, first.Handler(), "POST", "/backup?target=felhom-pbs", "A", ""); rr.Code != http.StatusAccepted {
|
||||||
|
t.Fatalf("POST /backup: %d %s", rr.Code, rr.Body.String())
|
||||||
|
}
|
||||||
|
waitFor(t, func() bool { return len(first.store.Backups(context.Background())) == 1 })
|
||||||
|
if _, ok := backup.NewBackupSuccessState(path).LastKnownSuccess("felhom-pbs", 8200); ok {
|
||||||
|
t.Fatal("a failed backup must never be saved as a success")
|
||||||
|
}
|
||||||
|
got := dueFor(t, r894Server(t, path, unreadable()).Handler(), "felhom-pbs")
|
||||||
|
if !got.Due {
|
||||||
|
t.Fatalf("after a failed backup and a restart the tier must still be due; got %+v", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
+61
-22
@@ -85,6 +85,13 @@ type BackupStore interface {
|
|||||||
RestoreTests(ctx context.Context) []hub.RestoreTest
|
RestoreTests(ctx context.Context) []hub.RestoreTest
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// LastKnownBackupStore (R-894) is the on-disk newest-success-per-tier record. Satisfied by
|
||||||
|
// *backup.BackupSuccessState.
|
||||||
|
type LastKnownBackupStore interface {
|
||||||
|
RecordBackupSuccess(target string, b hub.Backup) error
|
||||||
|
LastKnownSuccess(target string, vmid int) (time.Time, bool)
|
||||||
|
}
|
||||||
|
|
||||||
// StorageView yields the host's observed storage targets (for mapping a mount's storage id →
|
// StorageView yields the host's observed storage targets (for mapping a mount's storage id →
|
||||||
// fast/slow class). Satisfied by *storage.Observer.
|
// fast/slow class). Satisfied by *storage.Observer.
|
||||||
type StorageView interface {
|
type StorageView interface {
|
||||||
@@ -164,6 +171,10 @@ type Options struct {
|
|||||||
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
|
// PRIMARY tier, inside the backup goroutine and BEFORE the host-wide heavy-op gate is released — so the OS leg
|
||||||
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
|
// that it starts can never overlap another backup or a restore-test (`11` C10). OPTIONAL — nil → nothing runs.
|
||||||
AfterPrimaryBackup func(ctx context.Context, vmid int)
|
AfterPrimaryBackup func(ctx context.Context, vmid int)
|
||||||
|
// LastKnownBackups (R-894) keeps the newest successful backup per tier ON DISK, so the due-check's
|
||||||
|
// fallback for an UNREADABLE storage after an agent restart is the last known copy, not "never".
|
||||||
|
// nil = the pre-R-894 behaviour (in-memory record only).
|
||||||
|
LastKnownBackups LastKnownBackupStore
|
||||||
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
|
// Privileged runs the fenced root wrappers (E-2a: felhom-backup-target-apply). OPTIONAL — when
|
||||||
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
|
// nil, POST /backup/target reports "not configured". Satisfied by *proxmox.ExecRunner.
|
||||||
Privileged PrivilegedRunner
|
Privileged PrivilegedRunner
|
||||||
@@ -229,7 +240,6 @@ type Options struct {
|
|||||||
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
|
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
|
||||||
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
|
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
|
||||||
EscrowRecovery EscrowRecoverer
|
EscrowRecovery EscrowRecoverer
|
||||||
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// defaultBackupCadence is the fallback /backup/due window when none is configured.
|
// defaultBackupCadence is the fallback /backup/due window when none is configured.
|
||||||
@@ -279,21 +289,22 @@ type Server struct {
|
|||||||
// is the pre-R-82 shape.
|
// is the pre-R-82 shape.
|
||||||
tiers []BackupTier
|
tiers []BackupTier
|
||||||
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
|
// inFlight (R-85) is shared with the restore-test scheduler so the two never run together.
|
||||||
inFlight *backup.InFlight
|
inFlight *backup.InFlight
|
||||||
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
|
afterPrimaryBackup func(ctx context.Context, vmid int) // the OS leg (agent v0.140.0); nil = none
|
||||||
logger *slog.Logger
|
lastKnown LastKnownBackupStore // R-894: on-disk newest success per tier; nil = none
|
||||||
now func() time.Time
|
logger *slog.Logger
|
||||||
|
now func() time.Time
|
||||||
|
|
||||||
disks DiskOps // slice 8C (optional)
|
disks DiskOps // slice 8C (optional)
|
||||||
diskGate StorageGate // slice 8C (optional)
|
diskGate StorageGate // slice 8C (optional)
|
||||||
guestList GuestLister // slice 8C (optional)
|
guestList GuestLister // slice 8C (optional)
|
||||||
guestAttach GuestAttacher // slice 10 P2 (optional)
|
guestAttach GuestAttacher // slice 10 P2 (optional)
|
||||||
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
|
mem MemoryOps // v0.90.0 R-24 guest RAM resize (optional)
|
||||||
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
|
memMu sync.Mutex // single-flight around a resize apply (one customer per host)
|
||||||
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
|
netStorage NetworkStorageOps // Part A1: NAS network mounts (optional)
|
||||||
netMountRoot string // the user-data namespace root for the network-mount role gate
|
netMountRoot string // the user-data namespace root for the network-mount role gate
|
||||||
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
|
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
|
||||||
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
|
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
|
||||||
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
|
// crashGuardStatePath (R-856) is the host crash guard's state file read by GET /host/crash-guard;
|
||||||
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
|
// empty = defaultCrashGuardStatePath. A seam: tests point it at a fixture.
|
||||||
crashGuardStatePath string
|
crashGuardStatePath string
|
||||||
@@ -301,11 +312,11 @@ type Server struct {
|
|||||||
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
|
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
|
||||||
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
|
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
|
||||||
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
|
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
|
||||||
escrowRecovery EscrowRecoverer
|
escrowRecovery EscrowRecoverer
|
||||||
intent IntentRecorder // slice 10 P3 (optional)
|
intent IntentRecorder // slice 10 P3 (optional)
|
||||||
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
|
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
|
||||||
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
|
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
|
||||||
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
|
staleLock StaleLockController // F2-b startup stale-lock recovery (optional)
|
||||||
// guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog.
|
// guestPower (F-REBOOT) is per-guest start-attempt state for the guest-power watchdog.
|
||||||
// Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState).
|
// Guarded by guestPowerMu in guestpower.go; in-memory on purpose (see guestPowerState).
|
||||||
guestPower map[int]guestPowerState
|
guestPower map[int]guestPowerState
|
||||||
@@ -478,6 +489,7 @@ func NewServer(o Options) (*Server, error) {
|
|||||||
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
|
s.tiers = normalizeBackupTiers(o.BackupTiers, o.Backups, cadence)
|
||||||
s.inFlight = o.InFlight
|
s.inFlight = o.InFlight
|
||||||
s.afterPrimaryBackup = o.AfterPrimaryBackup
|
s.afterPrimaryBackup = o.AfterPrimaryBackup
|
||||||
|
s.lastKnown = o.LastKnownBackups
|
||||||
if s.backups == nil && len(s.tiers) > 0 {
|
if s.backups == nil && len(s.tiers) > 0 {
|
||||||
s.backups = s.tiers[0].Service
|
s.backups = s.tiers[0].Service
|
||||||
}
|
}
|
||||||
@@ -905,6 +917,13 @@ func (s *Server) handleBackup(w http.ResponseWriter, r *http.Request, vmid int)
|
|||||||
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
|
s.logger.Info("local-api: backup job complete", "vmid", vmid, "target", tier.TargetID, "job", jobID, "archive", b.Archive)
|
||||||
}
|
}
|
||||||
s.store.RecordBackup(b)
|
s.store.RecordBackup(b)
|
||||||
|
if b.Success && s.lastKnown != nil {
|
||||||
|
if err := s.lastKnown.RecordBackupSuccess(tier.TargetID, b); err != nil {
|
||||||
|
// Not fatal: the backup exists. Only the fallback after a restart loses this copy.
|
||||||
|
s.logger.Warn("local-api: could not save the backup on disk for the due-check fallback (R-894)",
|
||||||
|
"vmid", vmid, "target", tier.TargetID, "err", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
s.finishJob(key, jobID, b)
|
s.finishJob(key, jobID, b)
|
||||||
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
|
// OS leg (agent v0.140.0): after the night's whole-guest copy exists, still holding the heavy-op gate.
|
||||||
if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
|
if b.Success && tier.Primary && s.afterPrimaryBackup != nil {
|
||||||
@@ -1067,6 +1086,20 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
|
|||||||
newest, haveNewest = t, true
|
newest, haveNewest = t, true
|
||||||
unparseable = false // ground truth supersedes an unreadable in-memory timestamp
|
unparseable = false // ground truth supersedes an unreadable in-memory timestamp
|
||||||
}
|
}
|
||||||
|
// R-894: the storage could not be read → the last success saved ON DISK stands in for the in-memory
|
||||||
|
// record a restart emptied. ONLY on archiveUnknown: a storage that answers is the ground truth, and an
|
||||||
|
// archive absent there must make the tier due even when the file remembers one (a pruned copy).
|
||||||
|
// A saved copy older than the cadence still reads DUE below — an unreadable storage never suppresses
|
||||||
|
// a backup that is due.
|
||||||
|
fromDisk := false
|
||||||
|
if lookup == archiveUnknown && s.lastKnown != nil {
|
||||||
|
if saved, ok := s.lastKnown.LastKnownSuccess(tier.TargetID, vmid); ok && (!haveNewest || saved.After(newest)) {
|
||||||
|
newest, haveNewest, fromDisk = saved, true, true
|
||||||
|
unparseable = false
|
||||||
|
s.logger.Info("local-api: backup storage unreadable — due-check uses the last success saved on disk (R-894)",
|
||||||
|
"vmid", vmid, "target", tier.TargetID, "saved", saved.UTC().Format(time.RFC3339))
|
||||||
|
}
|
||||||
|
}
|
||||||
if !haveNewest {
|
if !haveNewest {
|
||||||
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
|
// R-88 Part 2: THREE distinct reasons for a nil age, each with its own state. Only ABSENT is a
|
||||||
// positive claim of "never backed up"; only that one may license the controller to bypass its
|
// positive claim of "never backed up"; only that one may license the controller to bypass its
|
||||||
@@ -1089,13 +1122,17 @@ func (s *Server) handleBackupDue(w http.ResponseWriter, r *http.Request, vmid in
|
|||||||
}
|
}
|
||||||
age := s.now().Sub(newest)
|
age := s.now().Sub(newest)
|
||||||
ageSecs := int64(age.Seconds())
|
ageSecs := int64(age.Seconds())
|
||||||
|
suffix := ""
|
||||||
|
if fromDisk {
|
||||||
|
suffix = " (storage unreadable — age from the last success saved on disk)"
|
||||||
|
}
|
||||||
if age >= tier.Cadence {
|
if age >= tier.Cadence {
|
||||||
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
|
writeOK(w, BackupDueResponse{VMID: vmid, Due: true, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
|
||||||
Reason: "older than cadence", Target: echo})
|
Reason: "older than cadence" + suffix, Target: echo})
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
|
writeOK(w, BackupDueResponse{VMID: vmid, Due: false, AgeSecs: &ageSecs, AgeState: AgeStateKnown,
|
||||||
Reason: "within cadence window", Target: echo})
|
Reason: "within cadence window" + suffix, Target: echo})
|
||||||
}
|
}
|
||||||
|
|
||||||
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
|
// BackupTiersResponse is GET /backup/tiers (R-82): the tiers this agent serves, primary first.
|
||||||
@@ -1483,4 +1520,6 @@ func writeStatus(w http.ResponseWriter, code int, ok bool, data any, errMsg stri
|
|||||||
}
|
}
|
||||||
|
|
||||||
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
|
// SetAfterPrimaryBackup wires the hook that runs after a successful primary-tier backup (the OS leg, agent v0.140.0).
|
||||||
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) { s.afterPrimaryBackup = fn }
|
func (s *Server) SetAfterPrimaryBackup(fn func(ctx context.Context, vmid int)) {
|
||||||
|
s.afterPrimaryBackup = fn
|
||||||
|
}
|
||||||
|
|||||||
@@ -106,6 +106,9 @@ type WrapperReport struct {
|
|||||||
LiveRestore json.RawMessage `json:"live_restore"`
|
LiveRestore json.RawMessage `json:"live_restore"`
|
||||||
Facts json.RawMessage `json:"facts"`
|
Facts json.RawMessage `json:"facts"`
|
||||||
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
|
Bundle json.RawMessage `json:"bundle"` // the config bundle's result (R-840, mode "bundle")
|
||||||
|
// OOMCheck (R-528, `09` decision 157): the docker layer's memory-kill check, {result, oom_killed, oom_event,
|
||||||
|
// exit_code, image, detail}. Carried to the hub UNCHANGED; the agent never reads it.
|
||||||
|
OOMCheck json.RawMessage `json:"oom_check"`
|
||||||
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
|
// R-868 (v0.144.0): the agent's ids, echoed from the plan, so a report kept on disk can be sent without the
|
||||||
// agent process that started the pass. ReleaseID / VMID were always in the report.
|
// agent process that started the pass. ReleaseID / VMID were always in the report.
|
||||||
RunID string `json:"run_id"`
|
RunID string `json:"run_id"`
|
||||||
@@ -143,6 +146,9 @@ type Report struct {
|
|||||||
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
|
DockerEngine string `json:"docker_engine,omitempty"` // docker layer: the engine after the step
|
||||||
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
|
Authority string `json:"authority,omitempty"` // docker layer: ring0 | signed
|
||||||
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
|
Undo bool `json:"undo,omitempty"` // docker layer: a signed undo (downgrade)
|
||||||
|
// OOMCheck: docker layer — the wrapper's oom_check object, byte-for-byte (R-528; the hub decides approval on it).
|
||||||
|
// Pinned by TestDocker_OOMCheckReachesTheHubUnchanged and TestR868_KeptCopyCarriesTheOOMCheck.
|
||||||
|
OOMCheck json.RawMessage `json:"oom_check,omitempty"`
|
||||||
|
|
||||||
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
|
unsent string // R-868: the wrapper's kept copy of this pass's report — deleted once the hub has it
|
||||||
}
|
}
|
||||||
@@ -630,6 +636,7 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
|||||||
}
|
}
|
||||||
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
|
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
|
||||||
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
|
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
|
||||||
|
rep.OOMCheck = rawOrNil(wr.OOMCheck)
|
||||||
if rep.Outcome == "" {
|
if rep.Outcome == "" {
|
||||||
switch {
|
switch {
|
||||||
case rep.Mode == "inventory" && !blk.Enabled:
|
case rep.Mode == "inventory" && !blk.Enabled:
|
||||||
@@ -707,6 +714,14 @@ func (l *Leg) runLayer(ctx context.Context, runID, layer string, vmid int, trigg
|
|||||||
return l.finish(ctx, lg, rep)
|
return l.finish(ctx, lg, rep)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// rawOrNil: a wrapper field that is absent or JSON null stays out of the hub report (omitempty).
|
||||||
|
func rawOrNil(m json.RawMessage) json.RawMessage {
|
||||||
|
if len(m) == 0 || string(m) == "null" {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
return m
|
||||||
|
}
|
||||||
|
|
||||||
func onlyDocker(in []Package) []Package {
|
func onlyDocker(in []Package) []Package {
|
||||||
var out []Package
|
var out []Package
|
||||||
for _, p := range in {
|
for _, p := range in {
|
||||||
|
|||||||
@@ -98,9 +98,13 @@ func (f *fakeWrapper) RunStdin(ctx context.Context, _ io.Reader, name string, ar
|
|||||||
return f.Run(ctx, name, args...)
|
return f.Run(ctx, name, args...)
|
||||||
}
|
}
|
||||||
|
|
||||||
type fakeHub struct{ reports []Report }
|
type fakeHub struct {
|
||||||
|
reports []Report
|
||||||
|
bodies [][]byte // the exact bytes posted (R-528: the oom_check object must arrive unchanged)
|
||||||
|
}
|
||||||
|
|
||||||
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
|
func (h *fakeHub) PostOSReport(_ context.Context, body []byte) error {
|
||||||
|
h.bodies = append(h.bodies, append([]byte(nil), body...))
|
||||||
var r Report
|
var r Report
|
||||||
json.Unmarshal(body, &r)
|
json.Unmarshal(body, &r)
|
||||||
h.reports = append(h.reports, r)
|
h.reports = append(h.reports, r)
|
||||||
@@ -504,3 +508,42 @@ func TestHealthVerdict_ControllerBlindToDockerFails(t *testing.T) {
|
|||||||
t.Fatal("an older wrapper (no field) must not fail")
|
t.Fatal("an older wrapper (no field) must not fail")
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// R-528 (`09` decision 157): the wrapper's oom_check object reaches the hub's docker report byte-for-byte; the guest
|
||||||
|
// and host reports carry none. COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in runLayer → "no
|
||||||
|
// oom_check in the docker report".
|
||||||
|
const oomCheckWire = `{"detail":"the engine reported the memory kill: OOMKilled=true and the oom event","exit_code":137,"image":"gitea.dooplex.hu/admin/felhom-controller:0.300.0","oom_event":true,"oom_killed":true,"result":"pass"}`
|
||||||
|
|
||||||
|
func TestDocker_OOMCheckReachesTheHubUnchanged(t *testing.T) {
|
||||||
|
w := &fakeWrapper{t: t, applyRep: map[string]WrapperReport{LayerDocker: {
|
||||||
|
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}},
|
||||||
|
DockerEngine: "29.8.2", Authority: "ring0", OOMCheck: json.RawMessage(oomCheckWire)}}}
|
||||||
|
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||||
|
l.Run(context.Background(), 9201, "night")
|
||||||
|
found := false
|
||||||
|
for _, b := range h.bodies {
|
||||||
|
var m map[string]json.RawMessage
|
||||||
|
if err := json.Unmarshal(b, &m); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
var layer string
|
||||||
|
json.Unmarshal(m["layer"], &layer)
|
||||||
|
oc, has := m["oom_check"]
|
||||||
|
if layer != LayerDocker {
|
||||||
|
if has {
|
||||||
|
t.Fatalf("the %s report carries an oom_check: %s", layer, oc)
|
||||||
|
}
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
found = true
|
||||||
|
if !has {
|
||||||
|
t.Fatalf("no oom_check in the docker report: %s", b)
|
||||||
|
}
|
||||||
|
if string(oc) != oomCheckWire {
|
||||||
|
t.Fatalf("oom_check changed on the way:\n got %s\nwant %s", oc, oomCheckWire)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if !found {
|
||||||
|
t.Fatalf("no docker report posted: %s", calls(w))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -144,6 +144,7 @@ func (l *Leg) reportFromKept(ctx context.Context, wr WrapperReport, path string)
|
|||||||
}
|
}
|
||||||
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
|
rep.Upgraded, rep.PassSeconds = wr.Upgraded, wr.PassSeconds
|
||||||
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
|
rep.DockerEngine, rep.Authority, rep.Undo = wr.DockerEngine, wr.Authority, wr.Undo
|
||||||
|
rep.OOMCheck = rawOrNil(wr.OOMCheck)
|
||||||
wantEngine := ""
|
wantEngine := ""
|
||||||
for _, u := range wr.Upgraded {
|
for _, u := range wr.Upgraded {
|
||||||
if u.Name == "docker-ce" {
|
if u.Name == "docker-ce" {
|
||||||
|
|||||||
@@ -150,3 +150,24 @@ func must(t *testing.T, err error) {
|
|||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// R-528: a kept docker report (the agent was killed) still carries the oom_check object to the hub unchanged.
|
||||||
|
// COMPANION RED-PROOF: drop `rep.OOMCheck = rawOrNil(wr.OOMCheck)` in reportFromKept → "oom_check lost".
|
||||||
|
func TestR868_KeptCopyCarriesTheOOMCheck(t *testing.T) {
|
||||||
|
w := &fakeWrapper{t: t}
|
||||||
|
l, h := newLeg(t, w, &hub.WireOSUpdate{Ring: 0, Enabled: true})
|
||||||
|
ring := 0
|
||||||
|
kept := WrapperReport{Mode: "apply", Layer: LayerDocker, RunID: "20261007T020000Z", Trigger: "night", Ring: &ring, VMID: 9201,
|
||||||
|
ReleaseID: "ring0-20261007T020000Z", HealthBefore: guestOK(), HealthAfter: guestOK(), DockerEngine: "29.8.2",
|
||||||
|
Upgraded: []Package{{Name: "docker-ce", Version: "5:29.8.2-1~debian.13~trixie"}}, OOMCheck: json.RawMessage(oomCheckWire)}
|
||||||
|
b, _ := json.Marshal(kept)
|
||||||
|
must(t, os.WriteFile(reportFile(l.PlanDir, kept.RunID, LayerDocker, "apply"), b, 0o600))
|
||||||
|
if n := l.SendUnsent(context.Background()); n != 1 || len(h.bodies) != 1 {
|
||||||
|
t.Fatalf("sent %d, bodies %d", n, len(h.bodies))
|
||||||
|
}
|
||||||
|
var m map[string]json.RawMessage
|
||||||
|
must(t, json.Unmarshal(h.bodies[0], &m))
|
||||||
|
if string(m["oom_check"]) != oomCheckWire {
|
||||||
|
t.Fatalf("oom_check lost or changed: %q", m["oom_check"])
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,64 @@
|
|||||||
|
package storage
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-330 (disk health Phase 2) — attributes 187, 188 and 199 ride the wire.
|
||||||
|
//
|
||||||
|
// The failing drive of 2026-08-14 carried 187 Reported_Uncorrect at raw 1001 while SMART said PASSED;
|
||||||
|
// none of the three reached the controller. The values are carried as RAW counters; an absent
|
||||||
|
// attribute is OMITTED from the JSON (unknown), never sent as 0 (S-39).
|
||||||
|
|
||||||
|
// The 2026-08-14 shape: PASSED, 187 raw 1001, plus 188/199 and the existing three.
|
||||||
|
const r330SATA = `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[
|
||||||
|
{"id":5,"raw":{"value":0}},
|
||||||
|
{"id":187,"raw":{"value":1001}},
|
||||||
|
{"id":188,"raw":{"value":4295032833}},
|
||||||
|
{"id":197,"raw":{"value":8}},
|
||||||
|
{"id":198,"raw":{"value":8}},
|
||||||
|
{"id":199,"raw":{"value":3}}]}}`
|
||||||
|
|
||||||
|
// COMPANION RED-PROOF (observed): delete the three R-330 cases in parseSMART → this fails with
|
||||||
|
// "187 Reported_Uncorrect must be carried (raw 1001); got <nil>". Restored.
|
||||||
|
func TestParseSMART_R330_CarriesTheThreeCounters(t *testing.T) {
|
||||||
|
s := parseSMART([]byte(r330SATA))
|
||||||
|
if s.ReportedUncorrect == nil || *s.ReportedUncorrect != 1001 {
|
||||||
|
t.Fatalf("187 Reported_Uncorrect must be carried (raw 1001); got %v", s.ReportedUncorrect)
|
||||||
|
}
|
||||||
|
// 188's raw value is vendor-packed on some drives (this one is 0x100010001): carried as reported, not truncated.
|
||||||
|
if s.CommandTimeout == nil || *s.CommandTimeout != 4295032833 {
|
||||||
|
t.Fatalf("188 Command_Timeout must be carried as the full raw value; got %v", s.CommandTimeout)
|
||||||
|
}
|
||||||
|
if s.UDMACRCErrors == nil || *s.UDMACRCErrors != 3 {
|
||||||
|
t.Fatalf("199 UDMA_CRC_Error_Count must be carried (raw 3); got %v", s.UDMACRCErrors)
|
||||||
|
}
|
||||||
|
if s.PendingSectors == nil || *s.PendingSectors != 8 {
|
||||||
|
t.Fatalf("the existing counters must be unchanged; pending=%v", s.PendingSectors)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A drive (or an NVMe device) that does not report the attributes leaves them nil, and the JSON
|
||||||
|
// OMITS the keys — the receiver reads "unknown", never a measured zero.
|
||||||
|
//
|
||||||
|
// COMPANION RED-PROOF (observed): drop `,omitempty` from the three tags in hub.SmartSummary → this
|
||||||
|
// fails with "an unreported attribute must be omitted, not sent: … reported_uncorrect …". Restored.
|
||||||
|
func TestParseSMART_R330_AbsentIsOmittedNotZero(t *testing.T) {
|
||||||
|
for name, raw := range map[string]string{
|
||||||
|
"sata without the three": `{"smart_status":{"passed":true},"ata_smart_attributes":{"table":[{"id":5,"raw":{"value":0}}]}}`,
|
||||||
|
"nvme": `{"smart_status":{"passed":true},"nvme_smart_health_information_log":{"critical_warning":0,"media_errors":0,"percentage_used":3}}`,
|
||||||
|
} {
|
||||||
|
s := parseSMART([]byte(raw))
|
||||||
|
if s.ReportedUncorrect != nil || s.CommandTimeout != nil || s.UDMACRCErrors != nil {
|
||||||
|
t.Fatalf("%s: unreported attributes must stay nil; got %v %v %v", name, s.ReportedUncorrect, s.CommandTimeout, s.UDMACRCErrors)
|
||||||
|
}
|
||||||
|
b, _ := json.Marshal(s)
|
||||||
|
for _, k := range []string{"reported_uncorrect", "command_timeout", "udma_crc_errors"} {
|
||||||
|
if strings.Contains(string(b), `"`+k+`"`) {
|
||||||
|
t.Fatalf("%s: an unreported attribute must be omitted, not sent: %s", name, b)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -44,6 +44,9 @@ const (
|
|||||||
ataReallocatedSectorCt = 5
|
ataReallocatedSectorCt = 5
|
||||||
ataCurrentPending = 197
|
ataCurrentPending = 197
|
||||||
ataOfflineUncorrect = 198
|
ataOfflineUncorrect = 198
|
||||||
|
ataReportedUncorrect = 187 // R-330
|
||||||
|
ataCommandTimeout = 188 // R-330
|
||||||
|
ataUDMACRCErrorCount = 199 // R-330
|
||||||
)
|
)
|
||||||
|
|
||||||
// parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading
|
// parseSMART maps smartctl JSON to a hub.SmartSummary, handling SATA + NVMe and degrading
|
||||||
@@ -85,6 +88,12 @@ func parseSMART(raw []byte) hub.SmartSummary {
|
|||||||
s.PendingSectors = intPtr(int(a.Raw.Value))
|
s.PendingSectors = intPtr(int(a.Raw.Value))
|
||||||
case ataOfflineUncorrect:
|
case ataOfflineUncorrect:
|
||||||
s.OfflineUncorrectable = intPtr(int(a.Raw.Value))
|
s.OfflineUncorrectable = intPtr(int(a.Raw.Value))
|
||||||
|
case ataReportedUncorrect:
|
||||||
|
s.ReportedUncorrect = int64Ptr(a.Raw.Value)
|
||||||
|
case ataCommandTimeout:
|
||||||
|
s.CommandTimeout = int64Ptr(a.Raw.Value)
|
||||||
|
case ataUDMACRCErrorCount:
|
||||||
|
s.UDMACRCErrors = int64Ptr(a.Raw.Value)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -142,3 +151,5 @@ func parseThinPoolMetadata(raw []byte) (float64, bool) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func intPtr(v int) *int { return &v }
|
func intPtr(v int) *int { return &v }
|
||||||
|
|
||||||
|
func int64Ptr(v int64) *int64 { return &v }
|
||||||
|
|||||||
Reference in New Issue
Block a user