Compare commits
16 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 18d03bd437 | |||
| e98b857684 | |||
| dcdeb3d16d | |||
| 610804b98d | |||
| 4586f0f7f6 | |||
| 205e22babe | |||
| 058b945064 | |||
| 40d857b527 | |||
| 7ae6990bac | |||
| 7569f34aeb | |||
| ede49b610d | |||
| f17ed11599 | |||
| 1db56bf837 | |||
| 53d047a6c1 | |||
| 28ba8593b8 | |||
| 6981450110 |
+285
@@ -1,3 +1,286 @@
|
|||||||
|
## Unreleased (to become v0.132.0) — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
|
||||||
|
|
||||||
|
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
|
||||||
|
`controller_slow_crashloop`; an older hub ignores them.
|
||||||
|
|
||||||
|
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
|
||||||
|
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
|
||||||
|
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
|
||||||
|
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
|
||||||
|
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
|
||||||
|
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
|
||||||
|
does **not** stop restarting — the fast brake remains the only brake.
|
||||||
|
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
|
||||||
|
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
|
||||||
|
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
|
||||||
|
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
|
||||||
|
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
|
||||||
|
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
|
||||||
|
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
|
||||||
|
|
||||||
|
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
|
||||||
|
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
|
||||||
|
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
|
||||||
|
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
|
||||||
|
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
|
||||||
|
|
||||||
|
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
|
||||||
|
|
||||||
|
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
|
||||||
|
|
||||||
|
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
|
||||||
|
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
|
||||||
|
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
|
||||||
|
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
|
||||||
|
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
|
||||||
|
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
|
||||||
|
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
|
||||||
|
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
|
||||||
|
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
|
||||||
|
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
|
||||||
|
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
|
||||||
|
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
|
||||||
|
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
|
||||||
|
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
|
||||||
|
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
|
||||||
|
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
|
||||||
|
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
|
||||||
|
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
|
||||||
|
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
|
||||||
|
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
|
||||||
|
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
|
||||||
|
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
|
||||||
|
after an agent restart — `last_success_source: storage`), the last attempt kept apart
|
||||||
|
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
|
||||||
|
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
|
||||||
|
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
|
||||||
|
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
|
||||||
|
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
|
||||||
|
last success".
|
||||||
|
|
||||||
|
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
|
||||||
|
|
||||||
|
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
|
||||||
|
|
||||||
|
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
|
||||||
|
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
|
||||||
|
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
|
||||||
|
absent). **All four found by accident.** The gates enforce everything else here and were the one part
|
||||||
|
nothing had checked.
|
||||||
|
|
||||||
|
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
|
||||||
|
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
|
||||||
|
|
||||||
|
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
|
||||||
|
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
|
||||||
|
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
|
||||||
|
proves the cause was the listing rather than the decoy.
|
||||||
|
|
||||||
|
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
|
||||||
|
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
|
||||||
|
|
||||||
|
**In this repo:** no gate changed, and that is the result. `release-complete` was decoyed and is
|
||||||
|
SOUND — the sweep's attempt (a non-version heading on top of `CHANGELOG.md`) was WITHDRAWN as
|
||||||
|
illegitimate, because `HEAD_RE.search` scans the whole file and still names `v0.130.0`. The three
|
||||||
|
shared gates are covered from `felhom.eu`; the remaining two are named in the decoy-coverage
|
||||||
|
exemption list (R-426) as UNTESTED, not as sound.
|
||||||
|
|
||||||
|
## v0.130.0 — the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
|
||||||
|
|
||||||
|
> **RELEASED 2026-08-20**, on the operator's word, after the fix was proved on both boxes.
|
||||||
|
> `sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, 14,141,158 bytes,
|
||||||
|
> tag `v0.130.0` at `7569f34`. Reproducible: a rebuild with `-trimpath -buildvcs=false` matches the
|
||||||
|
> published artifact byte for byte (R-186's property, checked rather than assumed).
|
||||||
|
>
|
||||||
|
> The heading read `## UNRELEASED — v0.130.0 candidate` until this point, deliberately: while the fix
|
||||||
|
> was hand-installed on `demo-hp` only, publishing would have pushed it onto `demo-felhom` through
|
||||||
|
> self-update and destroyed the control the proof rested on. **`release-complete` convicted on the
|
||||||
|
> release heading and was right to** — the answer was to stop claiming a release, not to bypass the
|
||||||
|
> gate. See R-347.
|
||||||
|
>
|
||||||
|
> **The fleet runs these exact bytes.** Both demo boxes were first given a hand build made during the
|
||||||
|
> proof — same source, same version string, **different bytes** (`256e0829…`), because
|
||||||
|
> `release-agent.sh` builds with `-trimpath -buildvcs=false` and a hand build does not. Nothing would
|
||||||
|
> have corrected that: the boxes already reported `0.130.0`, so self-update saw the vouched version as
|
||||||
|
> installed and would have done nothing, forever. Both were reinstalled from the **downloaded package**
|
||||||
|
> and now report `a56a92a7…`. Filed as **R-349**, because every prove-then-publish train hits it.
|
||||||
|
|
||||||
|
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
|
||||||
|
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
|
||||||
|
proxy whose descriptor ceiling is 65536 and whose runway at that rate was ~323 days. The connections
|
||||||
|
were held open on **both** sides: ep0 showed 388 while the two boxes showed 194 + 194, at two separate
|
||||||
|
instants, with the same source ports on each side, and **not one closed in a 31-minute window**.
|
||||||
|
`ss -tnp` on the boxes named the holder: **`felhom-agent`**, 194 of 194 on each, one PID.
|
||||||
|
|
||||||
|
**`pvestatd` and `proxmox-backup-client` made 162,404 requests in that window and leaked zero.** They
|
||||||
|
are 99.5% of the traffic to that endpoint and 0% of the leak. The agent made 811 requests — of which
|
||||||
|
387 were `GET .../snapshots` — and leaked 388 sockets. One per call, within one.
|
||||||
|
|
||||||
|
**The defect, and it is two things compounding.**
|
||||||
|
|
||||||
|
- `internal/pbs/client.go` built its transport as a composite literal:
|
||||||
|
`&http.Transport{TLSClientConfig: tlsCfg}`. That takes **`IdleConnTimeout` zero, which does not mean
|
||||||
|
"use a sane default" — it means retain idle keep-alive connections FOREVER.**
|
||||||
|
`http.DefaultTransport` sets 90s; hand-rolling the transport (which every client here must do,
|
||||||
|
because they all pin TLS) silently discards it.
|
||||||
|
- `pbsTargetsFromPVE` (`cmd/felhom-agent/main.go`) builds **a fresh `pbs.Client` every cycle**, as its
|
||||||
|
own doc comment says, and drops the previous one. An abandoned `http.Transport` does **not** close
|
||||||
|
its connections — it becomes unreachable while its `persistConn` read-loop goroutine keeps the
|
||||||
|
socket alive. So each cycle stranded exactly one connection that nothing could ever close.
|
||||||
|
|
||||||
|
The cadences reconcile without fitting: a 900s live-snapshot collect (184.7 cycles in the window) plus
|
||||||
|
a 6-hour verify loop (7.7) predicts 192.4 per box against **194 observed**.
|
||||||
|
|
||||||
|
**The fix is one field, restored to the standard library's own value.** New leaf package
|
||||||
|
`internal/httpx` owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is what
|
||||||
|
`http.DefaultTransport` uses, so there is nothing invented here to justify or tune — and
|
||||||
|
`NewTransport(tlsCfg, idleConnTimeout)`, which returns a **fresh** transport (never shared: each caller
|
||||||
|
pins a different endpoint) and treats a zero or negative timeout as **use the default, never "no
|
||||||
|
timeout"**. `pbs.Config` gains an `IdleConnTimeout` field that production leaves unset; only tests set
|
||||||
|
it, to avoid a 90-second wait.
|
||||||
|
|
||||||
|
**`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and
|
||||||
|
were corrected in the same pass — but neither contributed to the ep0 leak, and this entry must not be
|
||||||
|
read as three leaks having been found.** Both are built **once per process**, so they held one idle
|
||||||
|
connection for the life of the daemon rather than accumulating, and neither talks to ep0:8007.
|
||||||
|
|
||||||
|
**Tests, and what they deliberately do not assert.** `internal/pbs/client_leak_test.go` counts
|
||||||
|
connections **server-side** and models what `pbsTargetsFromPVE` actually does — build a client, use it
|
||||||
|
once, drop it on the floor — then asserts the connections go away. It does not assert `err == nil` and
|
||||||
|
it does not assert that some field holds some value; both were true of the leaking code.
|
||||||
|
|
||||||
|
- **Red-proof 1 (the fix):** removing `IdleConnTimeout` from `NewTransport` fails the test with
|
||||||
|
*"after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is
|
||||||
|
in the message, so the failure cannot be mistaken for a timeout with another cause. Reverted.
|
||||||
|
- **Red-proof 2 (the fix that would be worse than the bug):** setting `DisableKeepAlives: true` also
|
||||||
|
makes the leak vanish — by dialling fresh for every request, which on a box polling ~40,000 times a
|
||||||
|
day is strictly worse than what we started with. **The leak test PASSES under that mutation**;
|
||||||
|
`TestPBSClient_KeepAliveStillReuses` is what catches it, failing with *"3 sequential requests over 3
|
||||||
|
connection(s), want 1"*. Reverted.
|
||||||
|
|
||||||
|
**What this release does NOT do.** It does not reduce the poll rate (**R-336 stays open, but re-scoped
|
||||||
|
— it was never the cause of this leak**), it does not refactor `pbsTargetsFromPVE` to cache or reuse
|
||||||
|
clients (a one-line default restores the standard behaviour; a lifecycle refactor adds
|
||||||
|
cache-invalidation questions for no measurable gain), and it adds no `CloseIdleConnections` call.
|
||||||
|
|
||||||
|
**One sentence in this entry was written before the deploy and was WRONG, and it is corrected here
|
||||||
|
rather than quietly edited.** It read: *"does not clear the 388 descriptors already stuck on ep0 —
|
||||||
|
those persist until that proxy restarts."* **Measured: they clear the moment the AGENT restarts.**
|
||||||
|
Replacing the binary on `demo-hp` released exactly its 199 descriptors within one second
|
||||||
|
(415 → 216 fd), and replacing it on `demo-felhom` released the remaining 203 (**220 → 17 fd in under
|
||||||
|
two seconds**). **17 is precisely ep0's `t0` baseline** of 2026-08-18 09:51:22Z. ep0 was read-only
|
||||||
|
throughout and its proxy PID never changed. The accumulated leak was never ep0's to hold on to — it
|
||||||
|
was held on both sides, and closing either side ends it.
|
||||||
|
|
||||||
|
## v0.129.0 — a correct code for an earlier package stops being called wrong (2026-08-12, R-311)
|
||||||
|
|
||||||
|
**The measurement this fixes.** On 2026-08-12 a recovery code that provably opens a RETAINED package
|
||||||
|
— unsealed by hand, and it restored planted files byte-identical from a store the box itself could no
|
||||||
|
longer open — was answered by this agent with *"the recovery code did not open the sealed bundle"*.
|
||||||
|
The code was correct. Nothing had ever tried the retained packages, so the engine could not tell a
|
||||||
|
correct-but-earlier code from a mistype, and the screen said so out loud: a true sentence about our
|
||||||
|
own incuriosity, read by the customer as a statement about their code.
|
||||||
|
|
||||||
|
**`OffsiteKeyRecoverer` gains an optional `FetchRetained`.** It is consulted ONLY after the current
|
||||||
|
package has refused, so the ordinary recovery pays nothing for it and cannot fail because of it. When
|
||||||
|
one of the retained packages opens, the recoverer returns `ErrCodeOpensRetained` wrapped in a
|
||||||
|
`RetainedOpenedError` carrying the supersession date — no material, no code, no password.
|
||||||
|
|
||||||
|
**The local API answers 422** ("your code is correct, it belongs to an EARLIER sealed package") — a
|
||||||
|
FIFTH status added to the R-224 switch, not a restructuring of it. 422 rather than 400 because the
|
||||||
|
request was well-formed AND the credential valid; a 400 would put it in the same bucket as a mistype,
|
||||||
|
which is the defect.
|
||||||
|
|
||||||
|
**Fail-safe in every direction.** A nil fetcher, a hub too old to have the route (404 is a clean
|
||||||
|
"none"), a transport failure, a malformed package: each leaves the original refusal standing,
|
||||||
|
unchanged. The worst outcome of this feature breaking is the behaviour we had before it existed.
|
||||||
|
Attempts are bounded (`MaxRetainedTried`, default 6) because each unwrap is ~1 s of scrypt by design
|
||||||
|
and an unbounded loop would turn one wrong code into a minutes-long hang.
|
||||||
|
|
||||||
|
**New hub client call:** `FetchRetainedIdentityEscrow` → `GET /api/v1/hosts/<id>/escrow/retained`
|
||||||
|
(hub >= v0.103.0), self-scoped by the same per-host key.
|
||||||
|
|
||||||
|
Seven tests with REAL age crypto, because the two situations are indistinguishable AT THE UNWRAP and a
|
||||||
|
faked unwrap would prove nothing about what was broken. Red-proof, asserted applied: removing the
|
||||||
|
retained lookup returns the fail-closed wrong-code error — **the lie comes back, in those words.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Gates only — 2026-08-09 (no release, no version bump, no binary published)
|
||||||
|
|
||||||
|
**Two guards, both owed since the 2026-08-09 install outage (R-273/R-287). Nothing that runs on a
|
||||||
|
customer's box changed; `scripts/` only, and the agent stays v0.128.0.**
|
||||||
|
|
||||||
|
- **`scripts/retention-policy.json` — THE retention number, in one file.** The registry stopped
|
||||||
|
serving `felhom-agent` 0.120.0 and older while `check-published-versions.py` demanded that every
|
||||||
|
tag still be downloadable. Both rules are sensible; together they are impossible, and CI went red
|
||||||
|
at a commit whose own run had been green the day before. The check now **reads the number from
|
||||||
|
this file** and bounds its assertion to the newest N generic versions.
|
||||||
|
**What CI no longer covers, said plainly rather than left to be discovered:** a released version
|
||||||
|
older than the retention window is **no longer asserted downloadable**. Its git tag and its configs
|
||||||
|
are still asserted — only the binary's presence is dropped. The check **prints exactly which
|
||||||
|
versions it stopped covering** on every run, so the narrowing cannot become permanent by accident.
|
||||||
|
**The number is an OBSERVED state, not a located ruling** — see the file's own header and R-287.
|
||||||
|
A missing or unreadable policy file is **INCONCLUSIVE (exit 2), never silently unbounded.**
|
||||||
|
- **`scripts/check-release-complete.py` — the tag half, as a machine.** Asserts that the version at
|
||||||
|
the head of `CHANGELOG.md` is tagged, that the tag points into this history, and that its package
|
||||||
|
is published. `release-agent.sh` already warned about this in as many words and the step was still
|
||||||
|
missed on 2026-08-08, which is why this is a gate and not a reminder. Legs 1–2 need no network and
|
||||||
|
therefore run in `--fast`, so the pre-push hook catches a missing tag at the earliest moment.
|
||||||
|
Registered in `agent_gates.py`; red-proved by pointing the CHANGELOG head at an unreleased
|
||||||
|
v0.129.0 — both legs convicted and each named its fix command.
|
||||||
|
|
||||||
|
## v0.128.0 — the escrow seed is asserted every tick, not remembered once (2026-08-08, R-221)
|
||||||
|
|
||||||
|
**A rebuilt box could not run the escrow ceremony at all, and there was no way forward from inside
|
||||||
|
the product.** The preflight refuses on `escrow.pbs_storage_id`; the pbsdr bridge writes that key;
|
||||||
|
and it wrote it from exactly one place — `finishConverged`, reached only on the paths that actually
|
||||||
|
converge.
|
||||||
|
|
||||||
|
**The two things live in different places and die at different times.** The convergence marker is
|
||||||
|
host-side (`<agent-state>/pbsdr/marker.json`). The key it seeds is in `agent.json`, which
|
||||||
|
`step_agent_config` renders from `base = {}` unless an explicit `--preserve-from` is passed
|
||||||
|
(`felhom.eu/scripts/felhom-host-install.sh:2396`, the render at `:2449`, the `O_TRUNC` write at
|
||||||
|
`:2579`; the flag at `:1246`, defaulting empty at `:256`) — **and the render never writes an
|
||||||
|
`escrow` section at all.** So a rebuild keeps the marker and takes the key: same descriptor, same
|
||||||
|
hash, early return, and the seed never runs again into a config that no longer has it.
|
||||||
|
|
||||||
|
**Fixed by asserting rather than remembering.** `Apply` now re-asserts the seed *before* the
|
||||||
|
idempotent early return. `seedEscrowStorageID` is unchanged and still never clobbers a different
|
||||||
|
existing value — an operator's own choice outranks the descriptor's, with a warning naming both.
|
||||||
|
|
||||||
|
**The early return is KEPT.** It exists so a converged box does not re-run Proxmox operations every
|
||||||
|
60 s, and `TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner
|
||||||
|
calls on that tick — so a "fix" that simply deleted the return would fail. Cost of the re-assert: one
|
||||||
|
small file read plus a JSON parse per tick, no exec, no network, and an early return once the value
|
||||||
|
matches.
|
||||||
|
|
||||||
|
**A seed failure can never un-converge the box:** Warn plus a message on the published status,
|
||||||
|
exactly as `finishConverged` does it — no marker write, no state change. Pinned by
|
||||||
|
`TestSeedReassertFailure_DoesNotUnconverge`.
|
||||||
|
|
||||||
|
Tests drive the real `Apply` with a real temp-dir `agent.json` and a call-recording runner; calling
|
||||||
|
`seedEscrowStorageID` directly cannot see the early return, which IS the defect. The production
|
||||||
|
wiring (`pbsdr.NewManager(..., cfg.SourcePath, ...)`) is asserted by walking `main.go`'s **AST**, not
|
||||||
|
by `strings.Contains`, which a commented-out call also satisfies.
|
||||||
|
|
||||||
|
Red-proofs: removing the new call makes Scenario A fail on today's tree; removing the early return
|
||||||
|
makes the zero-Proxmox-calls assertion fail.
|
||||||
|
|
||||||
|
## (no version bump) — a comment that claimed the hub reads a field it has no field for (2026-08-08, R-260)
|
||||||
|
|
||||||
|
Comment-only; no behaviour, no wire change, nothing to rebuild.
|
||||||
|
|
||||||
|
`HostReport.SelfUpdatePending` / `SelfUpdatePendingVersion` carried the sentence *"The hub reads an
|
||||||
|
absent field as pending=false, the correct default."* **The hub has no field for either**, so it reads
|
||||||
|
nothing — present or absent — and `encoding/json` discards them on arrival. The sentence described an
|
||||||
|
intent rather than the code and read as settled for long enough that a class sweep had to find it.
|
||||||
|
|
||||||
|
The emission is correct and stays: the agent reports the truth, and the fault is entirely in the
|
||||||
|
receiving. The missing consumer is tracked as **R-264** (OPEN), and
|
||||||
|
`felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape while recording the
|
||||||
|
existing ones as explicit, reasoned allowlist entries rather than silence.
|
||||||
|
|
||||||
## v0.127.0 — a mount Felhom itself made is not "something else" (2026-08-06, R-220)
|
## v0.127.0 — a mount Felhom itself made is not "something else" (2026-08-06, R-220)
|
||||||
|
|
||||||
**After a rebuild the customer's own drives could not be re-attached, and the refusal named an action
|
**After a rebuild the customer's own drives could not be re-attached, and the refusal named an action
|
||||||
@@ -5280,3 +5563,5 @@ client, signing, or storage/backup orchestration yet (later slices).
|
|||||||
read-only `--selftest` against the demo host with TLS fingerprint pinning.
|
read-only `--selftest` against the demo host with TLS fingerprint pinning.
|
||||||
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
|
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
|
||||||
token) is provisioned out-of-band; the agent only consumes the token.
|
token) is provisioned out-of-band; the agent only consumes the token.
|
||||||
|
|
||||||
|
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->
|
||||||
|
|||||||
@@ -101,3 +101,11 @@ the mechanism are exempt.
|
|||||||
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
|
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
|
||||||
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
|
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
|
||||||
assumption, not an observation.
|
assumption, not an observation.
|
||||||
|
|
||||||
|
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
|
||||||
|
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
|
||||||
|
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
|
||||||
|
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
|
||||||
|
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
|
||||||
|
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
|
||||||
|
`os.listdir`, and a glob over a hand-maintained list.
|
||||||
|
|||||||
@@ -1,41 +1,47 @@
|
|||||||
# REPORT — felhom-agent v0.127.0: a mount Felhom made is not foreign (R-220)
|
# REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15)
|
||||||
|
|
||||||
**Scope: the host half of R-220.** The customer-facing refusal message is the controller's half and
|
Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md`
|
||||||
ships as felhom-controller v0.203.0.
|
§4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6.
|
||||||
|
|
||||||
## What changed
|
## Measured first (A.1)
|
||||||
|
Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and
|
||||||
|
`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix.
|
||||||
|
|
||||||
| File | Change |
|
## What shipped
|
||||||
|
- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests
|
||||||
|
that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap
|
||||||
|
in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest,
|
||||||
|
3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the
|
||||||
|
events. Same GuestExecutor and sudoers grants as the swap — no new privilege.
|
||||||
|
- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage
|
||||||
|
after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`.
|
||||||
|
- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`.
|
||||||
|
|
||||||
|
## Red-proofs (each seen failing, then restored)
|
||||||
|
- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)`
|
||||||
|
- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3`
|
||||||
|
- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success`
|
||||||
|
|
||||||
|
## Release and delivery
|
||||||
|
`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`,
|
||||||
|
verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the
|
||||||
|
box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for
|
||||||
|
`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and
|
||||||
|
Peti's box stay on 0.130.0.
|
||||||
|
|
||||||
|
## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0)
|
||||||
|
| moment | result |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `internal/storage/claim.go` | `claimFacts.felhomOwnedMounts`; `classifyClaim` forgives a non-managed mountpoint **only when corroborated**; `felhomOwnedMounts()` + `procMounts()` |
|
| idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill |
|
||||||
| `internal/storage/hostops.go` | `mountTable` seam (nil ⇒ real `/proc/mounts`) |
|
| parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** |
|
||||||
| `internal/storage/claim_r220_test.go` | new — the own-drive case, the fence, and the corroboration's four edges |
|
| kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z |
|
||||||
|
| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` |
|
||||||
|
| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z |
|
||||||
|
Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200"
|
||||||
|
line during the pause that the container state contradicts.
|
||||||
|
|
||||||
## The shape chosen, and why (§7.3)
|
## Teardown
|
||||||
|
Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker
|
||||||
|
removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release).
|
||||||
|
|
||||||
**Candidate (b): the claimed check distinguishes a mount Felhom made from a foreign one** — the task
|
Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`.
|
||||||
called it "nearer the truth" and it is, because the host and its knowledge survive the rebuild while
|
|
||||||
the guest's registry does not. Candidate (a) — having the rebuild path clear the raw mounts — would
|
|
||||||
have made correctness depend on a cleanup step running, and a cleanup that does not run leaves exactly
|
|
||||||
today's defect.
|
|
||||||
|
|
||||||
**The discriminator is corroboration, not a path prefix**: the same device must ALSO be mounted under
|
|
||||||
`/mnt/felhom-drives`. Only enrolment produces that pairing.
|
|
||||||
|
|
||||||
**`/proc/mounts` rather than `lsblk MOUNTPOINTS`**, because the lsblk invocation is pinned verbatim in
|
|
||||||
the sudoers file; changing it would have coupled this fix to a config rollout. `/proc/mounts` is
|
|
||||||
world-readable and needs neither.
|
|
||||||
|
|
||||||
## Green gate
|
|
||||||
|
|
||||||
`go build` · `go vet` clean · `go test ./...` → **29 packages ok** · `agent_gates.py --fast` → all OK.
|
|
||||||
|
|
||||||
| Red-proof | Result |
|
|
||||||
|---|---|
|
|
||||||
| remove the `felhomOwnedMounts` exemption | **FAILS** — "device is mounted at /mnt/adatok (sdb)", the pre-fix refusal |
|
|
||||||
| over-widen the exemption to any `/mnt/*` | **FAILS** — "/mnt/someone-elses-disk was offered for formatting" |
|
|
||||||
|
|
||||||
## Not changed
|
|
||||||
|
|
||||||
No sudoers, no allowlisted command, no PVE surface, no format path. Every other claim signal
|
|
||||||
(system disk, read-only, LVM PV, ZFS member, member FSTYPEs, empty-topology backstop) is untouched.
|
|
||||||
|
|||||||
@@ -74,6 +74,7 @@
|
|||||||
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
|
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
|
||||||
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
|
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
|
||||||
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
|
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
|
||||||
|
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
|
||||||
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
|
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
|
||||||
|
|
||||||
### Proxmox client / hub / PBS / provisioning
|
### Proxmox client / hub / PBS / provisioning
|
||||||
@@ -86,6 +87,7 @@
|
|||||||
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
|
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
|
||||||
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
|
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
|
||||||
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
|
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
|
||||||
|
| `httpx.NewTransport` | internal/httpx/transport.go | `NewTransport(tlsCfg, idleConnTimeout) *http.Transport` | **EVERY** hand-rolled `http.Transport` in this repo — pbs, hub and proxmox all pin TLS, so none can use `http.DefaultTransport` | **R-344: never inline `&http.Transport{TLSClientConfig: ...}` again.** A composite literal takes `IdleConnTimeout` **zero, which means retain idle connections FOREVER** — `http.DefaultTransport` sets 90s and a literal does not inherit it. Combined with a client rebuilt per cycle and dropped (`pbsTargetsFromPVE`), that stranded **388 sockets on ep0 in 46 h**, held open on BOTH sides. `idleConnTimeout <= 0` means **use the default**, never "no timeout". Returns a **FRESH** transport every call — a shared one would pool connections across differently pinned endpoints. Pinned by `internal/pbs/client_leak_test.go` (server-side connection counting) + `internal/httpx/transport_test.go` |
|
||||||
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
|
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
|
||||||
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
|
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
|
||||||
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
|
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
|
||||||
|
|||||||
+45
-15
@@ -59,7 +59,7 @@ import (
|
|||||||
|
|
||||||
// version is the agent version. Overridable at build time with
|
// version is the agent version. Overridable at build time with
|
||||||
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
|
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
|
||||||
var version = "0.92.1"
|
var version = "0.131.0"
|
||||||
|
|
||||||
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
|
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
|
||||||
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
|
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
|
||||||
@@ -1401,6 +1401,10 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
|||||||
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
|
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
|
||||||
// running" signal, so a deliberately stopped guest is never touched.
|
// running" signal, so a deliberately stopped guest is never touched.
|
||||||
go localSrv.WatchGuestPower(ctx)
|
go localSrv.WatchGuestPower(ctx)
|
||||||
|
// R-523: a controller container that is simply not running (killed, stopped, a failed
|
||||||
|
// self-update) is restarted through its bootstrap unit — nothing else watches it.
|
||||||
|
collector.SetControllerSupervisorReporter(localSrv)
|
||||||
|
go localSrv.WatchControllers(ctx)
|
||||||
go func() { errc <- localSrv.Run(ctx) }()
|
go func() { errc <- localSrv.Run(ctx) }()
|
||||||
}
|
}
|
||||||
if lanLoop != nil {
|
if lanLoop != nil {
|
||||||
@@ -1769,23 +1773,48 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
|||||||
}
|
}
|
||||||
return blob, true, nil
|
return blob, true, nil
|
||||||
},
|
},
|
||||||
|
// R-311 — the RETAINED packages, wired here and ONLY here, on the same self-scoped hub client.
|
||||||
|
// Consulted only after the current package has refused the code (see tryRetained), so the
|
||||||
|
// ordinary recovery pays nothing for it and cannot fail because of it.
|
||||||
|
FetchRetained: func(ctx context.Context) ([]escrow.RetainedBlob, int, error) {
|
||||||
|
resp, ferr := hubClient.FetchRetainedIdentityEscrow(ctx)
|
||||||
|
if ferr != nil {
|
||||||
|
return nil, 0, ferr
|
||||||
|
}
|
||||||
|
out := make([]escrow.RetainedBlob, 0, len(resp.Packages))
|
||||||
|
for _, p := range resp.Packages {
|
||||||
|
blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64)
|
||||||
|
if derr != nil || len(blob) == 0 {
|
||||||
|
// One malformed package must not sink the rest — the customer's code may open a
|
||||||
|
// later one, and a skipped entry is strictly better than a refusal we cannot justify.
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
out = append(out, escrow.RetainedBlob{
|
||||||
|
Blob: blob,
|
||||||
|
SupersededAt: p.SupersededAt,
|
||||||
|
KeyFingerprint: p.KeyFingerprint,
|
||||||
|
Index: p.Index,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
return out, resp.UnopenableCount, nil
|
||||||
|
},
|
||||||
}
|
}
|
||||||
srv, err := localapi.NewServer(localapi.Options{
|
srv, err := localapi.NewServer(localapi.Options{
|
||||||
EscrowRecovery: escrowRecoverer,
|
EscrowRecovery: escrowRecoverer,
|
||||||
ListenAddr: cfg.LocalAPI.ListenAddr,
|
ListenAddr: cfg.LocalAPI.ListenAddr,
|
||||||
Cert: cert,
|
Cert: cert,
|
||||||
AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel
|
AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel
|
||||||
Guests: px,
|
Guests: px,
|
||||||
Backups: runner,
|
Backups: runner,
|
||||||
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
|
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
|
||||||
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
|
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
|
||||||
Store: store,
|
Store: store,
|
||||||
Storage: observer,
|
Storage: observer,
|
||||||
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
|
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
|
||||||
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
|
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
|
||||||
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
|
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
|
||||||
Tokens: tokens,
|
Tokens: tokens,
|
||||||
BackupCadence: cfg.Backup.BackupCadence(),
|
BackupCadence: cfg.Backup.BackupCadence(),
|
||||||
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
|
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
|
||||||
Disks: hostOps,
|
Disks: hostOps,
|
||||||
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
|
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
|
||||||
@@ -1801,6 +1830,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
|||||||
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
|
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
|
||||||
SmbCredsDir: cfg.Privileged.SmbCredsDir,
|
SmbCredsDir: cfg.Privileged.SmbCredsDir,
|
||||||
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
|
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
|
||||||
|
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
|
||||||
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
|
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
|
||||||
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
|
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
|
||||||
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
|
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
|
||||||
|
|||||||
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
|
|||||||
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
|
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
|
||||||
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
|
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
|
||||||
# view, and the docker socket. The controller reaches the agent's local API for disk management.
|
# view, and the docker socket. The controller reaches the agent's local API for disk management.
|
||||||
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
|
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
|
||||||
|
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
|
||||||
|
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
|
||||||
|
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
|
||||||
|
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
|
||||||
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
|
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
|
||||||
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
|
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
|
||||||
-v felhom-controller-data:/opt/docker/felhom-controller \
|
-v felhom-controller-data:/opt/docker/felhom-controller \
|
||||||
|
|||||||
+116
-1
@@ -47,17 +47,80 @@ var (
|
|||||||
// retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is
|
// retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is
|
||||||
// not sent hunting for a mistyped recovery code that was typed correctly.
|
// not sent hunting for a mistyped recovery code that was typed correctly.
|
||||||
ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)")
|
ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)")
|
||||||
|
// ErrCodeOpensRetained — the code did NOT open the package the hub currently holds, and DID open a
|
||||||
|
// RETAINED (earlier) one. R-311.
|
||||||
|
//
|
||||||
|
// ⚠ THIS IS NOT A FAILURE OF THE CUSTOMER'S. It is the single most important distinction on this
|
||||||
|
// path, because until 2026-08-12 it was indistinguishable from a mistype and was reported as one.
|
||||||
|
// The screen could only say "it may be a typo, or it may be an older code, and we cannot tell them
|
||||||
|
// apart from here" — and it could not tell them apart because NOTHING EVER LOOKED. Now something
|
||||||
|
// looks, so the sentence can stop hedging.
|
||||||
|
//
|
||||||
|
// It carries no material and no code: only WHICH earlier package opened, by its supersession date,
|
||||||
|
// which is the one fact the customer needs to recognise it.
|
||||||
|
ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one")
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it
|
||||||
|
// carries no secret — not the code, not the bundle, not the repository password.
|
||||||
|
type RetainedMatch struct {
|
||||||
|
// SupersededAt is when this package stopped being the current one (hub-supplied, RFC3339-ish).
|
||||||
|
// It is what the recovery screen shows so the customer can recognise which code they are holding.
|
||||||
|
SupersededAt string
|
||||||
|
// KeyFingerprint is the escrow key fingerprint of that package — operator-log material only.
|
||||||
|
KeyFingerprint string
|
||||||
|
// Index is the hub's position label within ONE response. Not durable; do not persist it.
|
||||||
|
Index int
|
||||||
|
// HasResticPassword is false when the retained package opened but carries no repository password
|
||||||
|
// (a pre-fork-4 seal). The code is still CORRECT; the history behind it still cannot be reopened.
|
||||||
|
// Collapsing this into "recoverable" would repeat R-202's mistake on a new surface.
|
||||||
|
HasResticPassword bool
|
||||||
|
}
|
||||||
|
|
||||||
|
// RetainedOpenedError wraps ErrCodeOpensRetained with the match. Callers classify with errors.Is on
|
||||||
|
// the sentinel and read the detail with errors.As.
|
||||||
|
type RetainedOpenedError struct {
|
||||||
|
Match RetainedMatch
|
||||||
|
}
|
||||||
|
|
||||||
|
func (e *RetainedOpenedError) Error() string {
|
||||||
|
return ErrCodeOpensRetained.Error() + " (superseded_at=" + e.Match.SupersededAt + ")"
|
||||||
|
}
|
||||||
|
func (e *RetainedOpenedError) Unwrap() error { return ErrCodeOpensRetained }
|
||||||
|
|
||||||
// BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none".
|
// BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none".
|
||||||
// An interface-free func field keeps this package free of any dependency on the hub client.
|
// An interface-free func field keeps this package free of any dependency on the hub client.
|
||||||
type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error)
|
type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error)
|
||||||
|
|
||||||
|
// RetainedBlob is one retained sealed package as the recoverer sees it: opaque bytes plus the labels
|
||||||
|
// needed to name it. No secret.
|
||||||
|
type RetainedBlob struct {
|
||||||
|
Blob []byte
|
||||||
|
SupersededAt string
|
||||||
|
KeyFingerprint string
|
||||||
|
Index int
|
||||||
|
}
|
||||||
|
|
||||||
|
// RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty
|
||||||
|
// slice is a clean "none". R-311.
|
||||||
|
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, unopenable int, err error)
|
||||||
|
|
||||||
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
|
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
|
||||||
type OffsiteKeyRecoverer struct {
|
type OffsiteKeyRecoverer struct {
|
||||||
Fetch BlobFetcher
|
Fetch BlobFetcher
|
||||||
|
// FetchRetained is OPTIONAL and consulted ONLY after the current package has refused the code.
|
||||||
|
// nil keeps the pre-R-311 behaviour exactly: a refusal stays a refusal. That is deliberate — an
|
||||||
|
// agent wired without it must not behave differently from one that has no retained packages.
|
||||||
|
FetchRetained RetainedFetcher
|
||||||
|
// MaxRetainedTried bounds the scrypt work a single wrong code can cost. Each attempt is ~1 s of
|
||||||
|
// KDF by design, so an unbounded loop over a long supersession history would turn one wrong code
|
||||||
|
// into a minutes-long hang on the customer's screen. 0 means the built-in default.
|
||||||
|
MaxRetainedTried int
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// defaultMaxRetainedTried — six attempts is ~6 s worst case, which is a slow screen and not a hang.
|
||||||
|
const defaultMaxRetainedTried = 6
|
||||||
|
|
||||||
// RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password.
|
// RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password.
|
||||||
//
|
//
|
||||||
// A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits
|
// A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits
|
||||||
@@ -86,10 +149,62 @@ func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, rec
|
|||||||
}
|
}
|
||||||
bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode)
|
bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return "", err // already the fail-closed "the recovery code did not unwrap…" message; no secret in it
|
// R-311 — BEFORE calling this a wrong code, ask whether it is the RIGHT code for an EARLIER
|
||||||
|
// package. The engine fails closed identically either way, so the two are indistinguishable
|
||||||
|
// from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and
|
||||||
|
// the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence
|
||||||
|
// about our own incuriosity, read by the customer as a statement about their code.
|
||||||
|
if m, ok := r.tryRetained(ctx, recoveryCode); ok {
|
||||||
|
return "", &RetainedOpenedError{Match: m}
|
||||||
|
}
|
||||||
|
return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it
|
||||||
}
|
}
|
||||||
if bundle.ResticRepoPassword == "" {
|
if bundle.ResticRepoPassword == "" {
|
||||||
return "", ErrNoResticPassword
|
return "", ErrNoResticPassword
|
||||||
}
|
}
|
||||||
return bundle.ResticRepoPassword, nil
|
return bundle.ResticRepoPassword, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// tryRetained reports whether the code opens one of this host's RETAINED packages, and which.
|
||||||
|
//
|
||||||
|
// FAILURE HERE IS SILENT AND MEANS "NO", NEVER "YES" and never a different verdict for the caller. A
|
||||||
|
// hub that cannot answer, a route an older hub does not have, a malformed blob — each leaves the
|
||||||
|
// original refusal standing, unchanged. That is the fail-safe direction: the worst outcome of this
|
||||||
|
// function breaking is the behaviour we had before it existed.
|
||||||
|
//
|
||||||
|
// NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password.
|
||||||
|
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool) {
|
||||||
|
if r.FetchRetained == nil {
|
||||||
|
return RetainedMatch{}, false
|
||||||
|
}
|
||||||
|
blobs, _, err := r.FetchRetained(ctx)
|
||||||
|
if err != nil || len(blobs) == 0 {
|
||||||
|
return RetainedMatch{}, false
|
||||||
|
}
|
||||||
|
limit := r.MaxRetainedTried
|
||||||
|
if limit <= 0 {
|
||||||
|
limit = defaultMaxRetainedTried
|
||||||
|
}
|
||||||
|
for i, rb := range blobs {
|
||||||
|
if i >= limit {
|
||||||
|
break
|
||||||
|
}
|
||||||
|
if len(rb.Blob) == 0 {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode)
|
||||||
|
if uerr != nil {
|
||||||
|
continue // this one is not the customer's; try the next
|
||||||
|
}
|
||||||
|
return RetainedMatch{
|
||||||
|
SupersededAt: rb.SupersededAt,
|
||||||
|
KeyFingerprint: rb.KeyFingerprint,
|
||||||
|
Index: rb.Index,
|
||||||
|
// A retained package can itself predate the repository-password field. The code is still
|
||||||
|
// correct and must be told so — but the history behind it still cannot be reopened, and
|
||||||
|
// saying otherwise would be a promise this path cannot keep.
|
||||||
|
HasResticPassword: bundle.ResticRepoPassword != "",
|
||||||
|
}, true
|
||||||
|
}
|
||||||
|
return RetainedMatch{}, false
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,230 @@
|
|||||||
|
package escrow
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-311 — a correct code for an EARLIER package must stop being reported as a wrong code.
|
||||||
|
//
|
||||||
|
// These use REAL age crypto, like the R-199 tests beside them, because the whole point is that the
|
||||||
|
// two situations are indistinguishable AT THE UNWRAP: both fail closed on the current package. A
|
||||||
|
// faked unwrap would prove nothing about the thing that was actually broken.
|
||||||
|
|
||||||
|
const testR2 = "another correct horse battery staple sedative anaconda wobbly kingdom placard"
|
||||||
|
|
||||||
|
func retainedFetcherFor(blobs ...RetainedBlob) RetainedFetcher {
|
||||||
|
return func(context.Context) ([]RetainedBlob, int, error) { return blobs, 0, nil }
|
||||||
|
}
|
||||||
|
|
||||||
|
// THE ONE THAT MATTERS. The customer holds the code for a package we superseded. Yesterday this
|
||||||
|
// returned the fail-closed refusal and the screen told them to check their typing.
|
||||||
|
//
|
||||||
|
// RED-PROOF: remove the `if m, ok := r.tryRetained(...)` block from RecoverOffsiteRepoPassword →
|
||||||
|
// the wrong-code error returns instead → this FAILS, and the lie is back in exactly those words.
|
||||||
|
func TestRecover_CodeOpensRetainedPackage_IsNotAWrongCode(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
const oldPW = "aaaa567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
|
||||||
|
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: oldPW}, testR)
|
||||||
|
|
||||||
|
_, err := OffsiteKeyRecoverer{
|
||||||
|
Fetch: fetcherFor(current),
|
||||||
|
FetchRetained: retainedFetcherFor(RetainedBlob{
|
||||||
|
Blob: retained, SupersededAt: "2026-08-12 15:18:55", KeyFingerprint: "7e:a6:af", Index: 0,
|
||||||
|
}),
|
||||||
|
}.RecoverOffsiteRepoPassword(context.Background(), testR) // the OLD code
|
||||||
|
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("recovery succeeded — it must NOT return a password for a retained package on this path")
|
||||||
|
}
|
||||||
|
if !errors.Is(err, ErrCodeOpensRetained) {
|
||||||
|
t.Fatalf("err = %v, want ErrCodeOpensRetained — a correct code for an earlier package was "+
|
||||||
|
"classified as something else, which is how it became 'check your typing'", err)
|
||||||
|
}
|
||||||
|
var ro *RetainedOpenedError
|
||||||
|
if !errors.As(err, &ro) {
|
||||||
|
t.Fatalf("err does not carry a RetainedOpenedError: %v", err)
|
||||||
|
}
|
||||||
|
if ro.Match.SupersededAt != "2026-08-12 15:18:55" {
|
||||||
|
t.Errorf("SupersededAt = %q — the screen needs this date to name the package", ro.Match.SupersededAt)
|
||||||
|
}
|
||||||
|
if !ro.Match.HasResticPassword {
|
||||||
|
t.Error("HasResticPassword = false, but the retained bundle carried one")
|
||||||
|
}
|
||||||
|
// The error must not leak the code, the password or the bundle.
|
||||||
|
for _, secret := range []string{testR, oldPW} {
|
||||||
|
if containsStr(err.Error(), secret) {
|
||||||
|
t.Fatalf("the error text leaks a secret")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SCENARIO A — the ordinary recovery is untouched, and it must not even ASK for retained packages.
|
||||||
|
// If the current package opens, the customer is not in this story at all.
|
||||||
|
//
|
||||||
|
// RED-PROOF: move the tryRetained call above the successful-unwrap return → the fetcher runs → this
|
||||||
|
// FAILS on the "must not be consulted" assertion.
|
||||||
|
func TestRecover_CurrentPackageOpens_RetainedNeverConsulted(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
const pw = "bbbb567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: pw}, testR)
|
||||||
|
|
||||||
|
consulted := false
|
||||||
|
got, err := OffsiteKeyRecoverer{
|
||||||
|
Fetch: fetcherFor(current),
|
||||||
|
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
|
||||||
|
consulted = true
|
||||||
|
return nil, 0, nil
|
||||||
|
},
|
||||||
|
}.RecoverOffsiteRepoPassword(context.Background(), testR)
|
||||||
|
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("the ordinary recovery broke: %v", err)
|
||||||
|
}
|
||||||
|
if got != pw {
|
||||||
|
t.Fatalf("recovered password is not the sealed one")
|
||||||
|
}
|
||||||
|
if consulted {
|
||||||
|
t.Error("the retained packages were fetched on the SUCCESS path — the ordinary recovery must pay nothing for R-311")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SCENARIO C — a genuinely wrong code opens nothing, and must still be a plain refusal. The new
|
||||||
|
// branch must not become a way to encourage a customer who mistyped.
|
||||||
|
//
|
||||||
|
// RED-PROOF: make tryRetained return (RetainedMatch{}, true) unconditionally → a wrong code is
|
||||||
|
// reported as opening an earlier package → this FAILS.
|
||||||
|
func TestRecover_WrongCode_StaysAPlainRefusal(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
|
||||||
|
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: "dddd567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
|
||||||
|
|
||||||
|
_, err := OffsiteKeyRecoverer{
|
||||||
|
Fetch: fetcherFor(current),
|
||||||
|
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-01 00:00:00"}),
|
||||||
|
}.RecoverOffsiteRepoPassword(context.Background(), "totally wrong words that open nothing at all here")
|
||||||
|
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("a wrong code succeeded")
|
||||||
|
}
|
||||||
|
if errors.Is(err, ErrCodeOpensRetained) {
|
||||||
|
t.Fatal("a WRONG code was reported as opening a retained package — that would encourage a mistype")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// FAIL-SAFE — if the retained lookup itself fails, the original refusal must stand UNCHANGED. The
|
||||||
|
// worst outcome of this feature breaking is the behaviour we had before it.
|
||||||
|
//
|
||||||
|
// RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer
|
||||||
|
// gets a new, unexplained failure mode → this FAILS.
|
||||||
|
func TestRecover_RetainedFetchFails_OriginalRefusalStands(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "eeee567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
|
||||||
|
|
||||||
|
_, err := OffsiteKeyRecoverer{
|
||||||
|
Fetch: fetcherFor(current),
|
||||||
|
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
|
||||||
|
return nil, 0, fmt.Errorf("hub exploded")
|
||||||
|
},
|
||||||
|
}.RecoverOffsiteRepoPassword(context.Background(), testR2)
|
||||||
|
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected a refusal")
|
||||||
|
}
|
||||||
|
if errors.Is(err, ErrCodeOpensRetained) {
|
||||||
|
t.Fatal("a failed retained lookup was reported as 'opens a retained package'")
|
||||||
|
}
|
||||||
|
if containsStr(err.Error(), "hub exploded") {
|
||||||
|
t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be
|
||||||
|
// indistinguishable from one whose host has no retained packages.
|
||||||
|
//
|
||||||
|
// RED-PROOF: remove the `if r.FetchRetained == nil` guard → nil-deref panic → this FAILS.
|
||||||
|
func TestRecover_NilRetainedFetcher_IsPreR311Behaviour(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "ffff567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
|
||||||
|
|
||||||
|
_, err := OffsiteKeyRecoverer{Fetch: fetcherFor(current)}.RecoverOffsiteRepoPassword(context.Background(), testR2)
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected a refusal")
|
||||||
|
}
|
||||||
|
if errors.Is(err, ErrCodeOpensRetained) {
|
||||||
|
t.Fatal("a recoverer with no retained fetcher claimed a retained package opened")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A retained package that predates the repository-password field: the code is CORRECT and must be
|
||||||
|
// said to be correct, but HasResticPassword must be false so the screen does not promise a recovery
|
||||||
|
// that cannot produce a password (the R-202 lesson, on a new surface).
|
||||||
|
//
|
||||||
|
// RED-PROOF: hardcode HasResticPassword: true → this FAILS.
|
||||||
|
func TestRecover_RetainedOpensButPredatesTheField(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "1111567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
|
||||||
|
// No ResticRepoPassword at all — the pre-fork-4 shape.
|
||||||
|
retained := sealBundle(t, IdentityBundle{TunnelToken: "T", PBSToken: "P"}, testR)
|
||||||
|
|
||||||
|
_, err := OffsiteKeyRecoverer{
|
||||||
|
Fetch: fetcherFor(current),
|
||||||
|
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-04 07:20:08"}),
|
||||||
|
}.RecoverOffsiteRepoPassword(context.Background(), testR)
|
||||||
|
|
||||||
|
if !errors.Is(err, ErrCodeOpensRetained) {
|
||||||
|
t.Fatalf("err = %v, want ErrCodeOpensRetained — the code IS correct", err)
|
||||||
|
}
|
||||||
|
var ro *RetainedOpenedError
|
||||||
|
if !errors.As(err, &ro) {
|
||||||
|
t.Fatalf("no RetainedOpenedError: %v", err)
|
||||||
|
}
|
||||||
|
if ro.Match.HasResticPassword {
|
||||||
|
t.Error("HasResticPassword = true for a bundle carrying no repository password — the screen would promise a recovery that cannot happen")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The attempt count is BOUNDED. Each unwrap is ~1 s of scrypt by design, so an unbounded loop turns
|
||||||
|
// one wrong code into a minutes-long hang on the customer's screen.
|
||||||
|
//
|
||||||
|
// RED-PROOF: remove the `if i >= limit { break }` → all 10 are tried → this FAILS on the count.
|
||||||
|
func TestRecover_RetainedAttemptsAreBounded(t *testing.T) {
|
||||||
|
ensureAge(t)
|
||||||
|
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "2222567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
|
||||||
|
junk := sealBundle(t, IdentityBundle{ResticRepoPassword: "3333567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
|
||||||
|
|
||||||
|
tried := 0
|
||||||
|
blobs := make([]RetainedBlob, 0, 10)
|
||||||
|
for i := 0; i < 10; i++ {
|
||||||
|
blobs = append(blobs, RetainedBlob{Blob: junk, SupersededAt: "2026-08-01 00:00:00", Index: i})
|
||||||
|
}
|
||||||
|
rec := OffsiteKeyRecoverer{
|
||||||
|
Fetch: fetcherFor(current),
|
||||||
|
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
|
||||||
|
tried++
|
||||||
|
return blobs, 0, nil
|
||||||
|
},
|
||||||
|
MaxRetainedTried: 2,
|
||||||
|
}
|
||||||
|
// A code that opens NEITHER the current package nor any retained one.
|
||||||
|
if _, err := rec.RecoverOffsiteRepoPassword(context.Background(), "a code that opens nothing whatsoever in this test"); err == nil {
|
||||||
|
t.Fatal("expected a refusal")
|
||||||
|
}
|
||||||
|
if tried != 1 {
|
||||||
|
t.Errorf("the retained list was fetched %d times, want exactly 1", tried)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func containsStr(hay, needle string) bool {
|
||||||
|
return len(needle) > 0 && len(hay) >= len(needle) && (func() bool {
|
||||||
|
for i := 0; i+len(needle) <= len(hay); i++ {
|
||||||
|
if hay[i:i+len(needle)] == needle {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
})()
|
||||||
|
}
|
||||||
@@ -0,0 +1,59 @@
|
|||||||
|
// Package httpx holds the one HTTP-transport default this repo may not lose.
|
||||||
|
//
|
||||||
|
// Every client here pins TLS — PBS and PVE by leaf-cert SHA-256, the hub by an optional CA file —
|
||||||
|
// so none of them can use http.DefaultTransport and each hand-rolls its own. Hand-rolling silently
|
||||||
|
// discards DefaultTransport's settings, and one of them is load-bearing:
|
||||||
|
//
|
||||||
|
// Transport: &http.Transport{TLSClientConfig: tlsCfg} // IdleConnTimeout == 0 == NO timeout
|
||||||
|
//
|
||||||
|
// A zero IdleConnTimeout means idle keep-alive connections are retained FOREVER, not "use a sane
|
||||||
|
// default". Combined with a client that is rebuilt on a schedule and dropped (pbsTargetsFromPVE
|
||||||
|
// builds a fresh pbs.Client per cycle), every cycle strands one connection that nothing will ever
|
||||||
|
// close: the abandoned Transport becomes unreachable but its persistConn read-loop goroutine keeps
|
||||||
|
// the socket alive, and an unreachable Transport does not close its connections.
|
||||||
|
//
|
||||||
|
// Measured cost, live: 388 established connections accumulated on ep0's PBS proxy between
|
||||||
|
// 2026-08-18 09:51:22Z and 2026-08-20 08:02:13Z — 194 from each of the two boxes, held open on BOTH
|
||||||
|
// sides, one per agent poll cycle, on a proxy whose descriptor ceiling is 65536. See R-344 and
|
||||||
|
// felhom.eu/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md.
|
||||||
|
package httpx
|
||||||
|
|
||||||
|
import (
|
||||||
|
"crypto/tls"
|
||||||
|
"net/http"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// DefaultIdleConnTimeout is how long an idle keep-alive connection is retained before it is closed.
|
||||||
|
//
|
||||||
|
// It is 90s because that is http.DefaultTransport's own value: the fix for R-344 restores a
|
||||||
|
// standard-library default rather than inventing a number, so there is nothing here to tune and
|
||||||
|
// nothing to justify. It is comfortably shorter than every cadence that drives these clients (the
|
||||||
|
// 15-minute live-snapshot collect and the 6-hour verify loop), so a connection abandoned by one
|
||||||
|
// cycle is closed long before the next.
|
||||||
|
const DefaultIdleConnTimeout = 90 * time.Second
|
||||||
|
|
||||||
|
// NewTransport builds a FRESH *http.Transport pinned to tlsCfg, with the idle-connection timeout
|
||||||
|
// applied.
|
||||||
|
//
|
||||||
|
// Fresh, never shared: each caller pins a different endpoint, and a shared transport would pool
|
||||||
|
// connections across differently pinned servers. Reusing http.DefaultTransport for the same reason
|
||||||
|
// is not an option — it would drop the pin entirely.
|
||||||
|
//
|
||||||
|
// idleConnTimeout <= 0 means USE THE DEFAULT. It deliberately does not mean "no timeout": no-timeout
|
||||||
|
// is the bug this package exists to prevent, and an unset field must never be able to reintroduce
|
||||||
|
// it. Callers pass their configured value straight through; only tests pass a short one.
|
||||||
|
//
|
||||||
|
// Only IdleConnTimeout is set. The other DefaultTransport settings this transport also lacks
|
||||||
|
// (MaxIdleConns, TLSHandshakeTimeout, ExpectContinueTimeout) are deliberately left alone: none of
|
||||||
|
// them accumulates anything, every client bounds its whole request with http.Client.Timeout, and
|
||||||
|
// widening the change would have made the R-344 measurement unattributable.
|
||||||
|
func NewTransport(tlsCfg *tls.Config, idleConnTimeout time.Duration) *http.Transport {
|
||||||
|
if idleConnTimeout <= 0 {
|
||||||
|
idleConnTimeout = DefaultIdleConnTimeout
|
||||||
|
}
|
||||||
|
return &http.Transport{
|
||||||
|
TLSClientConfig: tlsCfg,
|
||||||
|
IdleConnTimeout: idleConnTimeout,
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,74 @@
|
|||||||
|
package httpx
|
||||||
|
|
||||||
|
import (
|
||||||
|
"crypto/tls"
|
||||||
|
"net/http"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// TestNewTransport_ZeroMeansDefaultNeverForever is the whole point of this package.
|
||||||
|
//
|
||||||
|
// http.Transport's zero IdleConnTimeout means "retain idle connections FOREVER". Any code path that
|
||||||
|
// can reach that zero reintroduces R-344, so an unset, zero or negative value must all land on the
|
||||||
|
// default. If someone later "simplifies" NewTransport by passing the argument straight through,
|
||||||
|
// this fails.
|
||||||
|
func TestNewTransport_ZeroMeansDefaultNeverForever(t *testing.T) {
|
||||||
|
for _, tc := range []struct {
|
||||||
|
name string
|
||||||
|
in time.Duration
|
||||||
|
want time.Duration
|
||||||
|
}{
|
||||||
|
{"zero", 0, DefaultIdleConnTimeout},
|
||||||
|
{"negative", -time.Hour, DefaultIdleConnTimeout},
|
||||||
|
{"explicit short value (tests)", 50 * time.Millisecond, 50 * time.Millisecond},
|
||||||
|
{"explicit long value", time.Hour, time.Hour},
|
||||||
|
} {
|
||||||
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
|
got := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, tc.in).IdleConnTimeout
|
||||||
|
if got != tc.want {
|
||||||
|
t.Fatalf("IdleConnTimeout = %v, want %v", got, tc.want)
|
||||||
|
}
|
||||||
|
if got == 0 {
|
||||||
|
t.Fatal("IdleConnTimeout is 0 — that is 'never expire', which is the R-344 defect itself")
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestDefaultIdleConnTimeout_MatchesTheStandardLibrary pins the number to its justification.
|
||||||
|
//
|
||||||
|
// 90s is not a tuned value; it is what http.DefaultTransport uses. Reading it off the standard
|
||||||
|
// library rather than hardcoding 90 means the constant cannot drift away from the reason given for
|
||||||
|
// it in the package doc.
|
||||||
|
func TestDefaultIdleConnTimeout_MatchesTheStandardLibrary(t *testing.T) {
|
||||||
|
std, ok := http.DefaultTransport.(*http.Transport)
|
||||||
|
if !ok {
|
||||||
|
t.Skip("http.DefaultTransport is not an *http.Transport in this Go build")
|
||||||
|
}
|
||||||
|
if DefaultIdleConnTimeout != std.IdleConnTimeout {
|
||||||
|
t.Fatalf("DefaultIdleConnTimeout = %v but http.DefaultTransport uses %v — the doc comment's justification no longer holds",
|
||||||
|
DefaultIdleConnTimeout, std.IdleConnTimeout)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestNewTransport_IsFreshEveryCall guards the pooling property the pinning relies on.
|
||||||
|
//
|
||||||
|
// Each caller pins a DIFFERENT endpoint. A shared transport would pool connections across
|
||||||
|
// differently pinned servers, so returning a package-level singleton would be a security change
|
||||||
|
// dressed as a tidy-up.
|
||||||
|
func TestNewTransport_IsFreshEveryCall(t *testing.T) {
|
||||||
|
a := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
|
||||||
|
b := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
|
||||||
|
if a == b {
|
||||||
|
t.Fatal("NewTransport returned the SAME transport twice — connections would be pooled across differently pinned endpoints")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestNewTransport_KeepsTheTLSConfig — the transport gains a field; it must lose nothing.
|
||||||
|
func TestNewTransport_KeepsTheTLSConfig(t *testing.T) {
|
||||||
|
cfg := &tls.Config{MinVersion: tls.VersionTLS12, InsecureSkipVerify: true} //nolint:gosec // test only
|
||||||
|
if got := NewTransport(cfg, 0).TLSClientConfig; got != cfg {
|
||||||
|
t.Fatalf("TLSClientConfig = %p, want the config passed in (%p) — the pin would be dropped", got, cfg)
|
||||||
|
}
|
||||||
|
}
|
||||||
+73
-2
@@ -15,6 +15,8 @@ import (
|
|||||||
"time"
|
"time"
|
||||||
|
|
||||||
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
|
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||||
)
|
)
|
||||||
|
|
||||||
const reportPath = "/api/v1/host-report"
|
const reportPath = "/api/v1/host-report"
|
||||||
@@ -49,8 +51,11 @@ func NewClient(cfg config.HubConfig, logger *slog.Logger) (*Client, error) {
|
|||||||
tlsCfg.RootCAs = pool
|
tlsCfg.RootCAs = pool
|
||||||
}
|
}
|
||||||
hc := &http.Client{
|
hc := &http.Client{
|
||||||
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
|
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
|
||||||
Transport: &http.Transport{TLSClientConfig: tlsCfg},
|
// R-344, consistency only: this client is built ONCE per process, so it never accumulated
|
||||||
|
// and contributed nothing to the ep0 leak. It carried the same missing default, which over
|
||||||
|
// a tunnel is how one idle connection survives long enough to fail on next use.
|
||||||
|
Transport: httpx.NewTransport(tlsCfg, 0),
|
||||||
}
|
}
|
||||||
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
|
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
|
||||||
}
|
}
|
||||||
@@ -354,3 +359,69 @@ func (c *Client) FetchIdentityEscrow(ctx context.Context) (*IdentityEscrowRespon
|
|||||||
}
|
}
|
||||||
return &out, nil
|
return &out, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// RetainedEscrowPackage is one RETAINED (superseded) sealed identity package. The blob is ciphertext
|
||||||
|
// and is useless without R. `SupersededAt` is the only thing here a human ever sees — it is what lets
|
||||||
|
// the recovery screen name WHICH earlier package a code belongs to.
|
||||||
|
type RetainedEscrowPackage struct {
|
||||||
|
Index int `json:"index"`
|
||||||
|
SupersededAt string `json:"superseded_at"`
|
||||||
|
KeyFingerprint string `json:"key_fingerprint"`
|
||||||
|
IdentityEscrowB64 string `json:"identity_escrow_b64"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// RetainedEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow/retained (hub >= v0.103.0, R-311).
|
||||||
|
//
|
||||||
|
// UnopenableCount is NOT noise. It counts retained packages the hub holds whose key material is absent
|
||||||
|
// (every pre-v0.93.0 row): on a box with those and nothing else, a perfectly correct old recovery code
|
||||||
|
// opens nothing, and the reason is a defect of ours. A caller that ignores this number will tell such a
|
||||||
|
// customer their code is wrong — the exact failure this whole chain exists to stop.
|
||||||
|
type RetainedEscrowResponse struct {
|
||||||
|
HostID string `json:"host_id"`
|
||||||
|
Count int `json:"count"`
|
||||||
|
UnopenableCount int `json:"unopenable_count"`
|
||||||
|
TruncatedCount int `json:"truncated_count"`
|
||||||
|
Packages []RetainedEscrowPackage `json:"packages"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// FetchRetainedIdentityEscrow reads back THIS host's RETAINED sealed identity packages (R-311 —
|
||||||
|
// the retained siblings of FetchIdentityEscrow, self-scoped server-side by the same per-host key).
|
||||||
|
//
|
||||||
|
// SEPARATE FROM FetchIdentityEscrow ON PURPOSE. The ordinary recovery must not pay for this call, and
|
||||||
|
// must not fail because of it: the current package is tried first and alone, and this is reached only
|
||||||
|
// after that has refused. A hub too old to know this route answers 404, which is a CLEAN "none" here
|
||||||
|
// and must never be reported as a failed recovery.
|
||||||
|
func (c *Client) FetchRetainedIdentityEscrow(ctx context.Context) (*RetainedEscrowResponse, error) {
|
||||||
|
if c.hostID == "" {
|
||||||
|
return nil, fmt.Errorf("hub: FetchRetainedIdentityEscrow requires a configured host_id")
|
||||||
|
}
|
||||||
|
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow/retained"
|
||||||
|
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("hub: building retained-escrow request: %w", err)
|
||||||
|
}
|
||||||
|
req.Header.Set("Authorization", "Bearer "+c.apiKey)
|
||||||
|
req.Header.Set("Accept", "application/json")
|
||||||
|
|
||||||
|
resp, err := c.hc.Do(req)
|
||||||
|
if err != nil {
|
||||||
|
return nil, &TransportError{Err: err}
|
||||||
|
}
|
||||||
|
defer resp.Body.Close()
|
||||||
|
|
||||||
|
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
|
||||||
|
if resp.StatusCode == http.StatusNotFound {
|
||||||
|
// A hub older than v0.103.0 has no such route. That is "no retained packages", not a fault —
|
||||||
|
// returning an error here would turn an old hub into a failed recovery on a box whose current
|
||||||
|
// package simply did not open.
|
||||||
|
return &RetainedEscrowResponse{HostID: c.hostID}, nil
|
||||||
|
}
|
||||||
|
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
|
||||||
|
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
|
||||||
|
}
|
||||||
|
var out RetainedEscrowResponse
|
||||||
|
if err := json.Unmarshal(raw, &out); err != nil {
|
||||||
|
return nil, fmt.Errorf("hub: decoding retained escrow fetch: %w", err)
|
||||||
|
}
|
||||||
|
return &out, nil
|
||||||
|
}
|
||||||
|
|||||||
@@ -103,6 +103,7 @@ type Collector struct {
|
|||||||
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
|
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
|
||||||
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
|
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
|
||||||
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
|
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
|
||||||
|
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
|
||||||
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
|
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
|
||||||
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
|
||||||
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
|
||||||
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
|
|||||||
return c
|
return c
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
|
||||||
|
type ControllerSupervisorReporter interface {
|
||||||
|
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
|
||||||
|
}
|
||||||
|
|
||||||
|
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
|
||||||
|
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
|
||||||
|
c.ctrlSup = r
|
||||||
|
return c
|
||||||
|
}
|
||||||
|
|
||||||
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
|
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
|
||||||
// omitted). Returns the collector for chaining.
|
// omitted). Returns the collector for chaining.
|
||||||
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
|
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
|
||||||
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
|
|||||||
if c.pbsdr != nil {
|
if c.pbsdr != nil {
|
||||||
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
|
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
|
||||||
}
|
}
|
||||||
|
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
|
||||||
|
if c.ctrlSup != nil {
|
||||||
|
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
|
||||||
|
}
|
||||||
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
|
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
|
||||||
if c.guestNet != nil {
|
if c.guestNet != nil {
|
||||||
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
|
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
|
||||||
|
|||||||
+41
-2
@@ -68,8 +68,14 @@ type HostReport struct {
|
|||||||
// report is stored opaquely hub-side, so these additive fields need no hub-schema change.
|
// report is stored opaquely hub-side, so these additive fields need no hub-schema change.
|
||||||
// Both are `omitempty` (the Wireguard precedent): in the steady state (no update in flight)
|
// Both are `omitempty` (the Wireguard precedent): in the steady state (no update in flight)
|
||||||
// they are absent — which keeps the cross-repo host-report golden contract byte-stable without
|
// they are absent — which keeps the cross-repo host-report golden contract byte-stable without
|
||||||
// a hub change. They appear only while an update is pending. The hub reads an absent field as
|
// a hub change. They appear only while an update is pending.
|
||||||
// pending=false, the correct default.
|
//
|
||||||
|
// ⚠ CORRECTED 2026-08-08 (R-260). This comment used to end "The hub reads an absent field as
|
||||||
|
// pending=false, the correct default." THE HUB HAS NO FIELD FOR EITHER OF THESE, so it reads
|
||||||
|
// nothing — present or absent — and encoding/json discards them on arrival. The sentence
|
||||||
|
// described an intent, not the code, and it read as settled for long enough that a sweep had to
|
||||||
|
// find it. The emission is correct and stays; the missing consumer is tracked as R-264, and
|
||||||
|
// `felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape.
|
||||||
SelfUpdatePending bool `json:"selfupdate_pending,omitempty"`
|
SelfUpdatePending bool `json:"selfupdate_pending,omitempty"`
|
||||||
SelfUpdatePendingVersion string `json:"selfupdate_pending_version,omitempty"`
|
SelfUpdatePendingVersion string `json:"selfupdate_pending_version,omitempty"`
|
||||||
|
|
||||||
@@ -118,6 +124,39 @@ type HostReport struct {
|
|||||||
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
|
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
|
||||||
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
|
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
|
||||||
OOB *OOBStatus `json:"oob,omitempty"`
|
OOB *OOBStatus `json:"oob,omitempty"`
|
||||||
|
|
||||||
|
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
|
||||||
|
// how many times the agent restarted a dead controller, when last and why, whether it gave up
|
||||||
|
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
|
||||||
|
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
|
||||||
|
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
|
||||||
|
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
|
||||||
|
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
|
||||||
|
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
|
||||||
|
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
|
||||||
|
type ControllerSupervisorStatus struct {
|
||||||
|
Guests []ControllerSupervisorGuest `json:"guests"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// ControllerSupervisorGuest is one supervised guest.
|
||||||
|
type ControllerSupervisorGuest struct {
|
||||||
|
VMID int `json:"vmid"`
|
||||||
|
RestartsTotal int `json:"restarts_total"`
|
||||||
|
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
|
||||||
|
LastReason string `json:"last_reason,omitempty"`
|
||||||
|
Crashloop bool `json:"crashloop"`
|
||||||
|
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
|
||||||
|
Parked bool `json:"parked"`
|
||||||
|
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
|
||||||
|
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
|
||||||
|
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
|
||||||
|
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
|
||||||
|
Restarts24h int `json:"restarts_24h"`
|
||||||
|
SlowCrashloop bool `json:"slow_crashloop"`
|
||||||
|
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
|
||||||
}
|
}
|
||||||
|
|
||||||
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
|
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
|
||||||
|
|||||||
@@ -0,0 +1,133 @@
|
|||||||
|
package localapi
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"io"
|
||||||
|
"log/slog"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
|
||||||
|
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
|
||||||
|
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
|
||||||
|
|
||||||
|
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
|
||||||
|
t.Helper()
|
||||||
|
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
|
||||||
|
var resp struct {
|
||||||
|
Data BackupStatusResponse `json:"data"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||||
|
t.Fatalf("decode: %v (%s)", err, w.Body.String())
|
||||||
|
}
|
||||||
|
return resp.Data.Tiers
|
||||||
|
}
|
||||||
|
|
||||||
|
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
|
||||||
|
t.Helper()
|
||||||
|
srv, err := NewServer(Options{
|
||||||
|
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
|
||||||
|
Storage: fakeStorage{targets: targets},
|
||||||
|
Tokens: staticTokens{"A": 8200},
|
||||||
|
BackupTiers: []BackupTier{
|
||||||
|
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
|
||||||
|
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
|
||||||
|
},
|
||||||
|
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
srv.baseCtx = context.Background()
|
||||||
|
srv.now = func() time.Time { return testNow }
|
||||||
|
return srv
|
||||||
|
}
|
||||||
|
|
||||||
|
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
|
||||||
|
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
|
||||||
|
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
|
||||||
|
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
|
||||||
|
st := &fakeStore{backups: []hub.Backup{
|
||||||
|
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
|
||||||
|
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
|
||||||
|
}}
|
||||||
|
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
|
||||||
|
tiers := tierStatesOf(t, srv)
|
||||||
|
if len(tiers) != 2 {
|
||||||
|
t.Fatalf("want 2 tiers, got %+v", tiers)
|
||||||
|
}
|
||||||
|
local, pbs := tiers[0], tiers[1]
|
||||||
|
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
|
||||||
|
t.Fatalf("local tier lost its successful backup: %+v", local)
|
||||||
|
}
|
||||||
|
if pbs.LastSuccess != nil {
|
||||||
|
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
|
||||||
|
}
|
||||||
|
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
|
||||||
|
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
|
||||||
|
}
|
||||||
|
if pbs.Storage != StoragePresenceAbsent {
|
||||||
|
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
|
||||||
|
}
|
||||||
|
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
|
||||||
|
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
|
||||||
|
var resp struct {
|
||||||
|
Data BackupStatusResponse `json:"data"`
|
||||||
|
}
|
||||||
|
_ = json.Unmarshal(w.Body.Bytes(), &resp)
|
||||||
|
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
|
||||||
|
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// /backup/tiers advertises storage presence, tri-state.
|
||||||
|
func TestBackupTiers_StoragePresence(t *testing.T) {
|
||||||
|
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
|
||||||
|
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
|
||||||
|
var resp struct {
|
||||||
|
Data BackupTiersResponse `json:"data"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
got := map[string]string{}
|
||||||
|
for _, ti := range resp.Data.Tiers {
|
||||||
|
got[ti.Target] = ti.Storage
|
||||||
|
}
|
||||||
|
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
|
||||||
|
t.Fatalf("storage presence wrong: %v", got)
|
||||||
|
}
|
||||||
|
// An unreadable storage view is "unknown", never "absent".
|
||||||
|
srv.storage = tierErrStorage{}
|
||||||
|
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
|
||||||
|
t.Fatalf("unreadable storage view must be unknown, got %q", p)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
type tierErrStorage struct{}
|
||||||
|
|
||||||
|
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
|
||||||
|
return nil, context.DeadlineExceeded
|
||||||
|
}
|
||||||
|
|
||||||
|
// A targeted request keeps the pre-R-517 bytes (no tiers array).
|
||||||
|
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
|
||||||
|
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
|
||||||
|
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
|
||||||
|
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
|
||||||
|
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func containsKey(b []byte, key string) bool {
|
||||||
|
var m struct {
|
||||||
|
Data map[string]json.RawMessage `json:"data"`
|
||||||
|
}
|
||||||
|
_ = json.Unmarshal(b, &m)
|
||||||
|
_, ok := m.Data[key]
|
||||||
|
return ok
|
||||||
|
}
|
||||||
@@ -0,0 +1,453 @@
|
|||||||
|
package localapi
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"sort"
|
||||||
|
"strconv"
|
||||||
|
"strings"
|
||||||
|
"sync"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-523 — the in-guest controller supervisor.
|
||||||
|
//
|
||||||
|
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
|
||||||
|
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
|
||||||
|
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
|
||||||
|
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
|
||||||
|
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
|
||||||
|
// dashboard answered 502 for 33 minutes until the box was power-cycled.
|
||||||
|
//
|
||||||
|
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
|
||||||
|
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
|
||||||
|
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
|
||||||
|
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
|
||||||
|
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
|
||||||
|
// No new privilege.
|
||||||
|
//
|
||||||
|
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
|
||||||
|
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
|
||||||
|
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
|
||||||
|
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
|
||||||
|
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
|
||||||
|
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
|
||||||
|
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
|
||||||
|
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
|
||||||
|
// again after 30 minutes.
|
||||||
|
//
|
||||||
|
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
|
||||||
|
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
|
||||||
|
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
|
||||||
|
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
|
||||||
|
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
|
||||||
|
// record) can never read as a new restart.
|
||||||
|
|
||||||
|
const (
|
||||||
|
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
|
||||||
|
// so a killed controller is restarted 30–60 s after it died.
|
||||||
|
controllerSupervisorInterval = 30 * time.Second
|
||||||
|
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
|
||||||
|
controllerSupervisorConfirm = 2
|
||||||
|
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
|
||||||
|
// controllerCrashloopPause.
|
||||||
|
controllerCrashloopMax = 3
|
||||||
|
controllerCrashloopWindow = 15 * time.Minute
|
||||||
|
controllerCrashloopPause = 30 * time.Minute
|
||||||
|
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
|
||||||
|
// watchdog is indistinguishable from a dead one (standing rule 3).
|
||||||
|
controllerSupervisorHeartbeatEvery = 20
|
||||||
|
|
||||||
|
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
|
||||||
|
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
|
||||||
|
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
|
||||||
|
// the controller container, and a customer inside the guest cannot park the supervisor.
|
||||||
|
ControllerParkedMarker = "controller-parked"
|
||||||
|
|
||||||
|
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
|
||||||
|
|
||||||
|
// R-539 (operator ruling 3 of 2026-09-16) — the SLOW crash loop. The 3-in-15-minutes brake above
|
||||||
|
// cannot see a controller that dies every 20 minutes: no two restarts share its window, so it is
|
||||||
|
// restarted for ever and the only trace is an info event that mails nobody (measured 2026-09-16,
|
||||||
|
// R-531). A second counter over 24 hours raises a WARNING at the fifth restart. It does NOT stop
|
||||||
|
// restarting — the fast brake stays the only brake, unchanged. Every restart the supervisor
|
||||||
|
// performs counts, including one that follows a deliberate operator `docker kill` (measured
|
||||||
|
// 2026-09-15: the supervisor cannot tell a kill from a crash, and a controller that is killed five
|
||||||
|
// times a day is worth a line to the operator either way).
|
||||||
|
controllerSlowCrashloopWindow = 24 * time.Hour
|
||||||
|
controllerSlowCrashloopMax = 5
|
||||||
|
// controllerSlowCounterFile holds the 24-hour restart times and the last raise, per guest, beside
|
||||||
|
// the parked marker.
|
||||||
|
controllerSlowCounterFile = "controller-restarts-24h.json"
|
||||||
|
)
|
||||||
|
|
||||||
|
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
|
||||||
|
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
|
||||||
|
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
|
||||||
|
type controllerSupState struct {
|
||||||
|
notRunningSeen int
|
||||||
|
restarts []time.Time // restart times inside the crash-loop window (pruned)
|
||||||
|
restartsTotal int
|
||||||
|
lastRestartAt time.Time
|
||||||
|
lastReason string
|
||||||
|
crashloopSince time.Time // zero = not in a crash-loop pause
|
||||||
|
parked bool
|
||||||
|
|
||||||
|
// R-539 — the slow counter. PERSISTED, unlike everything above, and the precedent's reason does not
|
||||||
|
// apply to it: persisting the fast record could carry a stale "give up" across the restart that
|
||||||
|
// fixed it, but this record never gives anything up — it only warns. Losing it on an agent restart,
|
||||||
|
// on the other hand, would hide exactly the box it exists for (one whose agent restarts too).
|
||||||
|
restarts24h []time.Time
|
||||||
|
slowCrashloopSince time.Time // the last raise; kept after it ages out, the hub keys on it MOVING
|
||||||
|
}
|
||||||
|
|
||||||
|
type controllerSupervisor struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
guests map[int]*controllerSupState
|
||||||
|
sweeps int
|
||||||
|
}
|
||||||
|
|
||||||
|
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
|
||||||
|
// (staleLock) or the guest executor is not wired.
|
||||||
|
func (s *Server) WatchControllers(ctx context.Context) {
|
||||||
|
if s.staleLock == nil || s.guestExec == nil {
|
||||||
|
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
|
||||||
|
return
|
||||||
|
}
|
||||||
|
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
|
||||||
|
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
|
||||||
|
"crashloop_window", controllerCrashloopWindow.String(),
|
||||||
|
"slow_crashloop_max", controllerSlowCrashloopMax, "slow_crashloop_window", controllerSlowCrashloopWindow.String(),
|
||||||
|
"guests_dir", s.guestsStateDir())
|
||||||
|
t := time.NewTicker(controllerSupervisorInterval)
|
||||||
|
defer t.Stop()
|
||||||
|
for {
|
||||||
|
select {
|
||||||
|
case <-ctx.Done():
|
||||||
|
return
|
||||||
|
case <-t.C:
|
||||||
|
s.ControllerSupervisorTick(ctx)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *Server) guestsStateDir() string {
|
||||||
|
if s.guestsDir != "" {
|
||||||
|
return s.guestsDir
|
||||||
|
}
|
||||||
|
return defaultGuestsStateDir
|
||||||
|
}
|
||||||
|
|
||||||
|
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
|
||||||
|
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
|
||||||
|
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
|
||||||
|
// not stat the file within.
|
||||||
|
func (s *Server) provisionedGuest(vmid int) bool {
|
||||||
|
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
|
||||||
|
return err == nil && fi.IsDir()
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *Server) controllerParked(vmid int) bool {
|
||||||
|
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
|
||||||
|
return err == nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *Server) supState(vmid int) *controllerSupState {
|
||||||
|
if s.ctrlSup.guests == nil {
|
||||||
|
s.ctrlSup.guests = map[int]*controllerSupState{}
|
||||||
|
}
|
||||||
|
st := s.ctrlSup.guests[vmid]
|
||||||
|
if st == nil {
|
||||||
|
st = &controllerSupState{}
|
||||||
|
s.loadSlowCounter(vmid, st)
|
||||||
|
s.ctrlSup.guests[vmid] = st
|
||||||
|
}
|
||||||
|
return st
|
||||||
|
}
|
||||||
|
|
||||||
|
// slowCounterRecord is the on-disk shape of the R-539 counter.
|
||||||
|
type slowCounterRecord struct {
|
||||||
|
Restarts []time.Time `json:"restarts"`
|
||||||
|
SlowCrashloopSince time.Time `json:"slow_crashloop_since,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *Server) slowCounterPath(vmid int) string {
|
||||||
|
return filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), controllerSlowCounterFile)
|
||||||
|
}
|
||||||
|
|
||||||
|
// loadSlowCounter restores the persisted counter into a fresh state. Absent = a clean start; unreadable
|
||||||
|
// or corrupt = a clean start with a WARN (a warning counter must never block supervision).
|
||||||
|
func (s *Server) loadSlowCounter(vmid int, st *controllerSupState) {
|
||||||
|
b, err := os.ReadFile(s.slowCounterPath(vmid))
|
||||||
|
if err != nil {
|
||||||
|
if !os.IsNotExist(err) {
|
||||||
|
s.logger.Warn("controller-supervisor: slow counter unreadable — starting it from zero", "vmid", vmid, "err", err)
|
||||||
|
}
|
||||||
|
return
|
||||||
|
}
|
||||||
|
var rec slowCounterRecord
|
||||||
|
if err := json.Unmarshal(b, &rec); err != nil {
|
||||||
|
s.logger.Warn("controller-supervisor: slow counter corrupt — starting it from zero", "vmid", vmid, "err", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
st.restarts24h = pruneBefore(rec.Restarts, s.clock().Add(-controllerSlowCrashloopWindow))
|
||||||
|
st.slowCrashloopSince = rec.SlowCrashloopSince
|
||||||
|
if len(st.restarts24h) > 0 || !st.slowCrashloopSince.IsZero() {
|
||||||
|
s.logger.Info("controller-supervisor: slow counter restored from disk", "vmid", vmid,
|
||||||
|
"restarts_24h", len(st.restarts24h), "slow_crashloop_since", st.slowCrashloopSince.Format(time.RFC3339))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// saveSlowCounter writes the counter atomically (tmp + rename, 0600). A failure is logged and the
|
||||||
|
// in-memory counter carries on — the next restart retries the write.
|
||||||
|
func (s *Server) saveSlowCounter(vmid int, rec slowCounterRecord) {
|
||||||
|
path := s.slowCounterPath(vmid)
|
||||||
|
b, err := json.Marshal(rec)
|
||||||
|
if err == nil {
|
||||||
|
tmp := path + ".tmp"
|
||||||
|
if err = os.WriteFile(tmp, b, 0o600); err == nil {
|
||||||
|
err = os.Rename(tmp, path)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if err != nil {
|
||||||
|
s.logger.Warn("controller-supervisor: could not persist the slow counter (kept in memory)", "vmid", vmid, "path", path, "err", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
|
||||||
|
// cycle without waiting on the ticker.
|
||||||
|
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
|
||||||
|
if s.staleLock == nil || s.guestExec == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
guests, err := s.staleLock.Guests(ctx)
|
||||||
|
if err != nil {
|
||||||
|
// Ownership unproven ⇒ touch nothing (the guest-power rule).
|
||||||
|
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
var evaluated, down int
|
||||||
|
for _, g := range guests {
|
||||||
|
if ctx.Err() != nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if !s.provisionedGuest(g.VMID) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
evaluated++
|
||||||
|
if !s.superviseOneController(ctx, g.VMID, g.Status) {
|
||||||
|
down++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
s.ctrlSup.mu.Lock()
|
||||||
|
s.ctrlSup.sweeps++
|
||||||
|
sweeps := s.ctrlSup.sweeps
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
|
||||||
|
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
|
||||||
|
"guests_evaluated", evaluated, "controllers_not_running", down)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
|
||||||
|
// known=false means the question could not be answered (pct exec failed for a reason other than a
|
||||||
|
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
|
||||||
|
// running": `docker rm` of the controller is the same outage as a kill.
|
||||||
|
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
|
||||||
|
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
|
||||||
|
if err != nil {
|
||||||
|
msg := strings.ToLower(err.Error() + " " + out)
|
||||||
|
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
|
||||||
|
return false, true, "absent"
|
||||||
|
}
|
||||||
|
return false, false, ""
|
||||||
|
}
|
||||||
|
status = strings.TrimSpace(out)
|
||||||
|
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
|
||||||
|
return status == "running" || status == "restarting", true, status
|
||||||
|
}
|
||||||
|
|
||||||
|
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
|
||||||
|
// allows. Returns false when the controller was observed not running.
|
||||||
|
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
|
||||||
|
now := s.clock()
|
||||||
|
if guestStatus != "running" {
|
||||||
|
s.resetNotRunning(vmid)
|
||||||
|
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
|
||||||
|
}
|
||||||
|
|
||||||
|
running, known, status := s.controllerRunning(ctx, vmid)
|
||||||
|
if !known {
|
||||||
|
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
|
||||||
|
s.resetNotRunning(vmid)
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
parked := s.controllerParked(vmid)
|
||||||
|
s.ctrlSup.mu.Lock()
|
||||||
|
st := s.supState(vmid)
|
||||||
|
st.parked = parked
|
||||||
|
if running {
|
||||||
|
st.notRunningSeen = 0
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
st.notRunningSeen++
|
||||||
|
seen := st.notRunningSeen
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
|
||||||
|
if parked {
|
||||||
|
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
|
||||||
|
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
s.swapMu.Lock()
|
||||||
|
swapping := s.swapInFlight[vmid]
|
||||||
|
s.swapMu.Unlock()
|
||||||
|
if swapping {
|
||||||
|
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
|
||||||
|
"vmid", vmid, "status", status)
|
||||||
|
s.resetNotRunning(vmid)
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
if seen < controllerSupervisorConfirm {
|
||||||
|
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
|
||||||
|
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
lock, _, err := s.staleLock.Lock(ctx, vmid)
|
||||||
|
if err != nil {
|
||||||
|
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
if lock != "" {
|
||||||
|
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
|
||||||
|
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
|
||||||
|
"vmid", vmid, "backup_running", busy, "err", berr)
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// Backoff.
|
||||||
|
s.ctrlSup.mu.Lock()
|
||||||
|
st = s.supState(vmid)
|
||||||
|
if !st.crashloopSince.IsZero() {
|
||||||
|
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
|
||||||
|
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
|
||||||
|
// crash-loop (the hub keys on it moving, not on it clearing).
|
||||||
|
st.restarts = nil
|
||||||
|
st.crashloopSince = time.Time{}
|
||||||
|
}
|
||||||
|
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
|
||||||
|
if len(st.restarts) >= controllerCrashloopMax {
|
||||||
|
st.crashloopSince = now
|
||||||
|
n := len(st.restarts)
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
|
||||||
|
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
|
||||||
|
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
|
||||||
|
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
|
||||||
|
"vmid", vmid, "status", status, "unit", bootstrapUnit)
|
||||||
|
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
|
||||||
|
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
|
||||||
|
reason += "; restart FAILED: " + err.Error()
|
||||||
|
}
|
||||||
|
s.ctrlSup.mu.Lock()
|
||||||
|
st = s.supState(vmid)
|
||||||
|
st.restarts = append(st.restarts, now)
|
||||||
|
st.restartsTotal++
|
||||||
|
st.lastRestartAt = now
|
||||||
|
st.lastReason = reason
|
||||||
|
st.notRunningSeen = 0
|
||||||
|
// R-539: the slow counter. Raise at most once per 24 hours — the hub mails on the raise MOVING.
|
||||||
|
st.restarts24h = append(pruneBefore(st.restarts24h, now.Add(-controllerSlowCrashloopWindow)), now)
|
||||||
|
n24 := len(st.restarts24h)
|
||||||
|
raised := false
|
||||||
|
if n24 >= controllerSlowCrashloopMax && (st.slowCrashloopSince.IsZero() || now.Sub(st.slowCrashloopSince) >= controllerSlowCrashloopWindow) {
|
||||||
|
st.slowCrashloopSince = now
|
||||||
|
raised = true
|
||||||
|
}
|
||||||
|
rec := slowCounterRecord{Restarts: append([]time.Time(nil), st.restarts24h...), SlowCrashloopSince: st.slowCrashloopSince}
|
||||||
|
s.ctrlSup.mu.Unlock()
|
||||||
|
s.saveSlowCounter(vmid, rec)
|
||||||
|
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason, "restarts_24h", n24)
|
||||||
|
if raised {
|
||||||
|
s.logger.Warn("controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop",
|
||||||
|
"vmid", vmid, "restarts_24h", n24, "window", controllerSlowCrashloopWindow.String(), "threshold", controllerSlowCrashloopMax)
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *Server) resetNotRunning(vmid int) {
|
||||||
|
s.ctrlSup.mu.Lock()
|
||||||
|
defer s.ctrlSup.mu.Unlock()
|
||||||
|
if st := s.ctrlSup.guests[vmid]; st != nil {
|
||||||
|
st.notRunningSeen = 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *Server) clock() time.Time {
|
||||||
|
if s.now != nil {
|
||||||
|
return s.now()
|
||||||
|
}
|
||||||
|
return time.Now().UTC()
|
||||||
|
}
|
||||||
|
|
||||||
|
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
|
||||||
|
out := ts[:0]
|
||||||
|
for _, t := range ts {
|
||||||
|
if !t.Before(cutoff) {
|
||||||
|
out = append(out, t)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
|
||||||
|
// Nil when the supervisor is not wired, so the stanza is omitted.
|
||||||
|
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
|
||||||
|
if s.staleLock == nil || s.guestExec == nil {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
s.ctrlSup.mu.Lock()
|
||||||
|
defer s.ctrlSup.mu.Unlock()
|
||||||
|
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
|
||||||
|
for vmid, st := range s.ctrlSup.guests {
|
||||||
|
g := hub.ControllerSupervisorGuest{
|
||||||
|
VMID: vmid,
|
||||||
|
RestartsTotal: st.restartsTotal,
|
||||||
|
LastReason: st.lastReason,
|
||||||
|
Parked: st.parked,
|
||||||
|
Crashloop: !st.crashloopSince.IsZero(),
|
||||||
|
}
|
||||||
|
if !st.lastRestartAt.IsZero() {
|
||||||
|
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
|
||||||
|
}
|
||||||
|
if !st.crashloopSince.IsZero() {
|
||||||
|
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
|
||||||
|
}
|
||||||
|
now := s.clock()
|
||||||
|
g.Restarts24h = len(pruneBefore(append([]time.Time(nil), st.restarts24h...), now.Add(-controllerSlowCrashloopWindow)))
|
||||||
|
if !st.slowCrashloopSince.IsZero() {
|
||||||
|
g.SlowCrashloopSince = st.slowCrashloopSince.UTC().Format(time.RFC3339)
|
||||||
|
g.SlowCrashloop = now.Sub(st.slowCrashloopSince) < controllerSlowCrashloopWindow
|
||||||
|
}
|
||||||
|
out.Guests = append(out.Guests, g)
|
||||||
|
}
|
||||||
|
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
|
||||||
|
return out
|
||||||
|
}
|
||||||
@@ -0,0 +1,350 @@
|
|||||||
|
package localapi
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"errors"
|
||||||
|
"io"
|
||||||
|
"log/slog"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strconv"
|
||||||
|
"sync"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
|
||||||
|
// again", and each guard is pinned by the case where acting would be wrong.
|
||||||
|
|
||||||
|
type supExec struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
status map[int]string // docker .State.Status per vmid; "" = container absent
|
||||||
|
inspectErr error // non-nil = pct exec itself failed (unknown)
|
||||||
|
restarts map[int]int
|
||||||
|
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
|
||||||
|
// stays "exited").
|
||||||
|
onRestart string
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
|
||||||
|
f.mu.Lock()
|
||||||
|
defer f.mu.Unlock()
|
||||||
|
switch {
|
||||||
|
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
|
||||||
|
if f.inspectErr != nil {
|
||||||
|
return "", f.inspectErr
|
||||||
|
}
|
||||||
|
st, ok := f.status[vmid]
|
||||||
|
if !ok || st == "" {
|
||||||
|
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
|
||||||
|
}
|
||||||
|
return st + "\n", nil
|
||||||
|
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
|
||||||
|
if f.restarts == nil {
|
||||||
|
f.restarts = map[int]int{}
|
||||||
|
}
|
||||||
|
f.restarts[vmid]++
|
||||||
|
if f.onRestart != "" {
|
||||||
|
f.status[vmid] = f.onRestart
|
||||||
|
}
|
||||||
|
return "", nil
|
||||||
|
}
|
||||||
|
return "", errors.New("supExec: unexpected args")
|
||||||
|
}
|
||||||
|
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
|
||||||
|
return "", errors.New("supExec: no stdin exec expected")
|
||||||
|
}
|
||||||
|
func (f *supExec) count(vmid int) int {
|
||||||
|
f.mu.Lock()
|
||||||
|
defer f.mu.Unlock()
|
||||||
|
return f.restarts[vmid]
|
||||||
|
}
|
||||||
|
|
||||||
|
type supClock struct{ t time.Time }
|
||||||
|
|
||||||
|
func (c *supClock) now() time.Time { return c.t }
|
||||||
|
|
||||||
|
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
|
||||||
|
t.Helper()
|
||||||
|
dir := t.TempDir()
|
||||||
|
for _, v := range provisioned {
|
||||||
|
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
|
||||||
|
s := &Server{
|
||||||
|
staleLock: ctl,
|
||||||
|
guestExec: ex,
|
||||||
|
guestsDir: dir,
|
||||||
|
swapInFlight: map[int]bool{},
|
||||||
|
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
|
||||||
|
now: clk.now,
|
||||||
|
}
|
||||||
|
return s, clk, dir
|
||||||
|
}
|
||||||
|
|
||||||
|
func runningGuest(vmid int) *fakeGuestPowerCtl {
|
||||||
|
return &fakeGuestPowerCtl{
|
||||||
|
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
|
||||||
|
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
|
||||||
|
// the first (the bootstrap's own rm-f/run window must never be raced).
|
||||||
|
//
|
||||||
|
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
|
||||||
|
// stays 0 → "the killed controller was NOT restarted — this is R-523".
|
||||||
|
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||||
|
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
if n := ex.count(9201); n != 0 {
|
||||||
|
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
|
||||||
|
}
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
if n := ex.count(9201); n != 1 {
|
||||||
|
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
|
||||||
|
}
|
||||||
|
st := s.ControllerSupervisorStatus(context.Background())
|
||||||
|
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
|
||||||
|
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
|
||||||
|
}
|
||||||
|
// Healthy again → no further restart.
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
if n := ex.count(9201); n != 1 {
|
||||||
|
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{}, onRestart: "running"}
|
||||||
|
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
if n := ex.count(9201); n != 1 {
|
||||||
|
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestControllerSupervisor_Guards(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
name string
|
||||||
|
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
|
||||||
|
}{
|
||||||
|
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
|
||||||
|
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
|
||||||
|
panic(err)
|
||||||
|
}
|
||||||
|
}},
|
||||||
|
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
|
||||||
|
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
|
||||||
|
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||||
|
ctl.backupRun = map[int]bool{9201: true}
|
||||||
|
}},
|
||||||
|
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||||
|
ctl.backupErr = errors.New("tasks unreadable")
|
||||||
|
}},
|
||||||
|
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||||
|
ctl.guests[0].Status = "stopped"
|
||||||
|
}},
|
||||||
|
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
|
||||||
|
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
|
||||||
|
}},
|
||||||
|
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
|
||||||
|
ctl.guestsErr = errors.New("api down")
|
||||||
|
}},
|
||||||
|
}
|
||||||
|
for _, tc := range cases {
|
||||||
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||||
|
ctl := runningGuest(9201)
|
||||||
|
s, _, dir := supServer(t, ex, ctl, 9201)
|
||||||
|
tc.setup(s, ex, ctl, dir)
|
||||||
|
for i := 0; i < 4; i++ {
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
}
|
||||||
|
if n := ex.count(9201); n != 0 {
|
||||||
|
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
|
||||||
|
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9202: "exited"}}
|
||||||
|
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
|
||||||
|
for i := 0; i < 3; i++ {
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
}
|
||||||
|
if n := ex.count(9202); n != 0 {
|
||||||
|
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
|
||||||
|
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
|
||||||
|
// again.
|
||||||
|
//
|
||||||
|
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
|
||||||
|
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
|
||||||
|
// failed at "crash-looping controller restarted 10 times".
|
||||||
|
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
|
||||||
|
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
|
||||||
|
s.ControllerSupervisorTick(ctx)
|
||||||
|
clk.t = clk.t.Add(controllerSupervisorInterval)
|
||||||
|
}
|
||||||
|
if n := ex.count(9201); n != controllerCrashloopMax {
|
||||||
|
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
|
||||||
|
}
|
||||||
|
st := s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||||
|
if !st.Crashloop || st.CrashloopSince == "" {
|
||||||
|
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
|
||||||
|
}
|
||||||
|
// Still paused 25 minutes later.
|
||||||
|
clk.t = clk.t.Add(15 * time.Minute)
|
||||||
|
s.ControllerSupervisorTick(ctx)
|
||||||
|
s.ControllerSupervisorTick(ctx)
|
||||||
|
if n := ex.count(9201); n != controllerCrashloopMax {
|
||||||
|
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
|
||||||
|
}
|
||||||
|
// After the pause: tries again.
|
||||||
|
clk.t = clk.t.Add(controllerCrashloopPause)
|
||||||
|
s.ControllerSupervisorTick(ctx)
|
||||||
|
s.ControllerSupervisorTick(ctx)
|
||||||
|
if n := ex.count(9201); n != controllerCrashloopMax+1 {
|
||||||
|
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
|
||||||
|
}
|
||||||
|
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
|
||||||
|
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
|
||||||
|
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
|
||||||
|
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
|
||||||
|
var m map[string][]map[string]any
|
||||||
|
if err := json.Unmarshal(b, &m); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
g := m["guests"][0]
|
||||||
|
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked", "restarts_24h", "slow_crashloop"} {
|
||||||
|
if _, ok := g[k]; !ok {
|
||||||
|
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
// ---- R-539 (operator ruling 3 of 2026-09-16): the SLOW crash loop ----------------------------------
|
||||||
|
|
||||||
|
// supKillOnce kills the controller and lets the supervisor restart it (two confirming sweeps), then
|
||||||
|
// moves the clock on by gap. The container comes back "running", so each restart is a separate act.
|
||||||
|
func supKillOnce(t *testing.T, s *Server, ex *supExec, clk *supClock, gap time.Duration) {
|
||||||
|
t.Helper()
|
||||||
|
before := ex.count(9201)
|
||||||
|
ex.mu.Lock()
|
||||||
|
ex.status[9201] = "exited"
|
||||||
|
ex.mu.Unlock()
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
clk.t = clk.t.Add(controllerSupervisorInterval)
|
||||||
|
s.ControllerSupervisorTick(context.Background())
|
||||||
|
if ex.count(9201) != before+1 {
|
||||||
|
t.Fatalf("kill was not followed by exactly one restart (restarts %d → %d)", before, ex.count(9201))
|
||||||
|
}
|
||||||
|
clk.t = clk.t.Add(gap)
|
||||||
|
}
|
||||||
|
|
||||||
|
// The consequence: a controller that dies every 20 minutes — never three times inside the 15-minute
|
||||||
|
// brake — raises slow_crashloop on the FIFTH restart in 24 hours, and the raise does not move again on
|
||||||
|
// the sixth (the hub mails on movement; once per 24 hours is the ruling).
|
||||||
|
//
|
||||||
|
// RED-PROOF: without the slow counter the stanza never sets slow_crashloop → "five restarts 20 minutes
|
||||||
|
// apart did not raise slow_crashloop — this is R-539".
|
||||||
|
func TestControllerSupervisor_SlowCrashloop(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
|
||||||
|
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
for i := 1; i <= 4; i++ {
|
||||||
|
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||||
|
}
|
||||||
|
g := s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||||
|
if g.Crashloop {
|
||||||
|
t.Fatalf("the 15-minute brake fired on restarts 20 minutes apart — the fixture is wrong: %+v", g)
|
||||||
|
}
|
||||||
|
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
|
||||||
|
t.Fatalf("slow_crashloop raised after only 4 restarts: %+v", g)
|
||||||
|
}
|
||||||
|
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||||
|
g = s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||||
|
if !g.SlowCrashloop || g.SlowCrashloopSince == "" || g.Restarts24h != 5 {
|
||||||
|
t.Fatalf("five restarts 20 minutes apart did not raise slow_crashloop — this is R-539: %+v", g)
|
||||||
|
}
|
||||||
|
first := g.SlowCrashloopSince
|
||||||
|
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||||
|
g = s.ControllerSupervisorStatus(ctx).Guests[0]
|
||||||
|
if g.SlowCrashloopSince != first {
|
||||||
|
t.Fatalf("the raise moved again on the 6th restart (%q → %q) — the operator would be mailed per restart", first, g.SlowCrashloopSince)
|
||||||
|
}
|
||||||
|
if !g.SlowCrashloop {
|
||||||
|
t.Fatalf("slow_crashloop cleared while the loop continues: %+v", g)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The negative control: restarts that never reach five inside any 24 hours never raise it.
|
||||||
|
func TestControllerSupervisor_SpreadRestartsNeverSlowCrashloop(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
|
||||||
|
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
for i := 0; i < 8; i++ { // eight restarts, 7 hours apart: at most 4 inside any 24 hours
|
||||||
|
supKillOnce(t, s, ex, clk, 7*time.Hour)
|
||||||
|
}
|
||||||
|
g := s.ControllerSupervisorStatus(context.Background()).Guests[0]
|
||||||
|
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
|
||||||
|
t.Fatalf("restarts 7 hours apart raised slow_crashloop: %+v", g)
|
||||||
|
}
|
||||||
|
if g.Restarts24h > 4 {
|
||||||
|
t.Fatalf("restarts_24h=%d — the 24-hour window is not pruning", g.Restarts24h)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// An agent restart must not reset the slow counter (the ruling; a box whose AGENT also restarts would
|
||||||
|
// otherwise never reach five). The same state directory, a fresh Server.
|
||||||
|
//
|
||||||
|
// RED-PROOF: keep the counter in memory only → the second Server starts at 0 → "the agent restart
|
||||||
|
// reset the slow counter".
|
||||||
|
func TestControllerSupervisor_SlowCounterSurvivesAgentRestart(t *testing.T) {
|
||||||
|
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
|
||||||
|
s, clk, dir := supServer(t, ex, runningGuest(9201), 9201)
|
||||||
|
for i := 0; i < 4; i++ {
|
||||||
|
supKillOnce(t, s, ex, clk, 20*time.Minute)
|
||||||
|
}
|
||||||
|
s2 := &Server{
|
||||||
|
staleLock: s.staleLock,
|
||||||
|
guestExec: ex,
|
||||||
|
guestsDir: dir,
|
||||||
|
swapInFlight: map[int]bool{},
|
||||||
|
logger: s.logger,
|
||||||
|
now: clk.now,
|
||||||
|
}
|
||||||
|
supKillOnce(t, s2, ex, clk, 20*time.Minute)
|
||||||
|
g := s2.ControllerSupervisorStatus(context.Background()).Guests[0]
|
||||||
|
if g.Restarts24h != 5 || !g.SlowCrashloop {
|
||||||
|
t.Fatalf("the agent restart reset the slow counter: %+v", g)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -98,6 +98,35 @@ func (s *Server) handleRecoverOffsitePassword(w http.ResponseWriter, r *http.Req
|
|||||||
case errors.Is(err, escrow.ErrNoEscrowBlob):
|
case errors.Is(err, escrow.ErrNoEscrowBlob):
|
||||||
s.logger.Warn("local-api: offsite key recovery: the hub holds no sealed bundle for this host", "vmid", vmid)
|
s.logger.Warn("local-api: offsite key recovery: the hub holds no sealed bundle for this host", "vmid", vmid)
|
||||||
writeErr(w, http.StatusNotFound, "the hub holds no sealed recovery bundle for this host — no escrow ceremony has run")
|
writeErr(w, http.StatusNotFound, "the hub holds no sealed recovery bundle for this host — no escrow ceremony has run")
|
||||||
|
// ── R-311 (2026-08-12) — THE CODE IS RIGHT, JUST NOT FOR THE CURRENT PACKAGE. ─────────
|
||||||
|
//
|
||||||
|
// Placed ABOVE the default for the same reason ErrBundleFetch is: the default blames the
|
||||||
|
// customer, and this case is the one where the customer is provably not at fault. The code was
|
||||||
|
// used, it worked, and it opened a package the hub is deliberately keeping.
|
||||||
|
//
|
||||||
|
// 422 rather than 400: the request was well-formed AND the credential was valid — what could
|
||||||
|
// not be processed is the pairing of a correct code with the CURRENT package. A 400 would put
|
||||||
|
// it in the same bucket as a mistype, which is the whole defect. The status is the
|
||||||
|
// machine-readable half; the controller classifies on it and must never parse this sentence.
|
||||||
|
//
|
||||||
|
// The date travels in the body because it is the one fact that lets a customer recognise which
|
||||||
|
// code they are holding. No material, no code, no password — only when that package stopped
|
||||||
|
// being current, and whether it can yield a repository password at all.
|
||||||
|
case errors.Is(err, escrow.ErrCodeOpensRetained):
|
||||||
|
var ro *escrow.RetainedOpenedError
|
||||||
|
match := escrow.RetainedMatch{}
|
||||||
|
if errors.As(err, &ro) {
|
||||||
|
match = ro.Match
|
||||||
|
}
|
||||||
|
s.logger.Info("local-api: offsite key recovery: the code did NOT open the current package but DID open a RETAINED one — the customer is not at fault",
|
||||||
|
"vmid", vmid, "superseded_at", match.SupersededAt, "retained_has_restic_pw", match.HasResticPassword)
|
||||||
|
writeStatus(w, http.StatusUnprocessableEntity, false,
|
||||||
|
map[string]any{
|
||||||
|
"opens_retained": true,
|
||||||
|
"superseded_at": match.SupersededAt,
|
||||||
|
"retained_has_restic_pw": match.HasResticPassword,
|
||||||
|
},
|
||||||
|
"the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held")
|
||||||
case errors.Is(err, escrow.ErrNoResticPassword):
|
case errors.Is(err, escrow.ErrNoResticPassword):
|
||||||
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
|
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
|
||||||
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
|
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
|
||||||
|
|||||||
@@ -185,6 +185,9 @@ type Options struct {
|
|||||||
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
|
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
|
||||||
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
|
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
|
||||||
StaleLock StaleLockController
|
StaleLock StaleLockController
|
||||||
|
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
|
||||||
|
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
|
||||||
|
GuestsStateDir string
|
||||||
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
|
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
|
||||||
ControllerSwapStateDir string
|
ControllerSwapStateDir string
|
||||||
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
|
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
|
||||||
@@ -362,6 +365,12 @@ type Server struct {
|
|||||||
swapMu sync.Mutex
|
swapMu sync.Mutex
|
||||||
swapInFlight map[int]bool
|
swapInFlight map[int]bool
|
||||||
|
|
||||||
|
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
|
||||||
|
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
|
||||||
|
guestExec GuestExecutor
|
||||||
|
guestsDir string
|
||||||
|
ctrlSup controllerSupervisor
|
||||||
|
|
||||||
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
|
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
|
||||||
// detached pipeline runs through (tests inject; production defaults set in NewServer).
|
// detached pipeline runs through (tests inject; production defaults set in NewServer).
|
||||||
netVerifyMu sync.Mutex
|
netVerifyMu sync.Mutex
|
||||||
@@ -484,7 +493,9 @@ func NewServer(o Options) (*Server, error) {
|
|||||||
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
|
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
|
||||||
if o.ControllerSwap != nil {
|
if o.ControllerSwap != nil {
|
||||||
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
|
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
|
||||||
|
s.guestExec = o.ControllerSwap
|
||||||
}
|
}
|
||||||
|
s.guestsDir = o.GuestsStateDir
|
||||||
return s, nil
|
return s, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1084,6 +1095,10 @@ type BackupTierInfo struct {
|
|||||||
Target string `json:"target"`
|
Target string `json:"target"`
|
||||||
CadenceSeconds int64 `json:"cadence_seconds"`
|
CadenceSeconds int64 `json:"cadence_seconds"`
|
||||||
Primary bool `json:"primary"`
|
Primary bool `json:"primary"`
|
||||||
|
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
|
||||||
|
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
|
||||||
|
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
|
||||||
|
Storage string `json:"storage,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
|
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||||
@@ -1093,11 +1108,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
|
|||||||
Target: t.TargetID,
|
Target: t.TargetID,
|
||||||
CadenceSeconds: int64(t.Cadence.Seconds()),
|
CadenceSeconds: int64(t.Cadence.Seconds()),
|
||||||
Primary: t.Primary,
|
Primary: t.Primary,
|
||||||
|
Storage: s.storagePresence(r.Context(), t.TargetID),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
writeOK(w, resp)
|
writeOK(w, resp)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
|
||||||
|
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
|
||||||
|
// successful read that does not list the storage is "absent".
|
||||||
|
func (s *Server) storagePresence(ctx context.Context, target string) string {
|
||||||
|
if s.storage == nil || target == "" {
|
||||||
|
return StoragePresenceUnknown
|
||||||
|
}
|
||||||
|
targets, err := s.storage.Observe(ctx)
|
||||||
|
if err != nil {
|
||||||
|
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
|
||||||
|
return StoragePresenceUnknown
|
||||||
|
}
|
||||||
|
for _, t := range targets {
|
||||||
|
if t.Name == target {
|
||||||
|
return StoragePresencePresent
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return StoragePresenceAbsent
|
||||||
|
}
|
||||||
|
|
||||||
|
const (
|
||||||
|
StoragePresencePresent = "present"
|
||||||
|
StoragePresenceAbsent = "absent"
|
||||||
|
StoragePresenceUnknown = "unknown"
|
||||||
|
)
|
||||||
|
|
||||||
|
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
|
||||||
|
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
|
||||||
|
// result ("presence is not success").
|
||||||
|
type TierBackupState struct {
|
||||||
|
Target string `json:"target"`
|
||||||
|
Primary bool `json:"primary"`
|
||||||
|
Storage string `json:"storage"` // present | absent | unknown
|
||||||
|
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
|
||||||
|
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
|
||||||
|
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
|
||||||
|
// LastSuccessSource is "storage".
|
||||||
|
LastSuccess *hub.Backup `json:"last_success,omitempty"`
|
||||||
|
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
|
||||||
|
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// TierAttempt is the newest recorded attempt on a tier, successful or not.
|
||||||
|
type TierAttempt struct {
|
||||||
|
StartedAt string `json:"started_at"`
|
||||||
|
Success bool `json:"success"`
|
||||||
|
Error string `json:"error,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// tierBackupStates builds the per-tier view for one guest.
|
||||||
|
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
|
||||||
|
out := make([]TierBackupState, 0, len(s.tiers))
|
||||||
|
for _, t := range s.tiers {
|
||||||
|
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
|
||||||
|
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
|
||||||
|
st.LastSuccess, st.LastSuccessSource = b, "record"
|
||||||
|
} else if st.Storage != StoragePresenceAbsent {
|
||||||
|
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
|
||||||
|
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
|
||||||
|
StartedAt: when.UTC().Format(time.RFC3339)}
|
||||||
|
st.LastSuccessSource = "storage"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
|
||||||
|
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
|
||||||
|
}
|
||||||
|
out = append(out, st)
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
// tierFromRequest resolves the `?target=` query parameter to a tier.
|
// tierFromRequest resolves the `?target=` query parameter to a tier.
|
||||||
//
|
//
|
||||||
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
|
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
|
||||||
@@ -1144,6 +1231,10 @@ type BackupStatusResponse struct {
|
|||||||
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
|
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
|
||||||
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
|
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
|
||||||
Target string `json:"target,omitempty"`
|
Target string `json:"target,omitempty"`
|
||||||
|
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
|
||||||
|
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
|
||||||
|
// response exactly as before.
|
||||||
|
Tiers []TierBackupState `json:"tiers,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
|
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||||
@@ -1155,6 +1246,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
|
|||||||
// across ANY target (echo == "" → pickLatestBackup's match-any path).
|
// across ANY target (echo == "" → pickLatestBackup's match-any path).
|
||||||
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
|
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
|
||||||
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
|
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
|
||||||
|
if echo == "" {
|
||||||
|
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
|
||||||
|
}
|
||||||
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
|
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
|
||||||
resp.Phase = job.Phase
|
resp.Phase = job.Phase
|
||||||
resp.JobID = job.JobID
|
resp.JobID = job.JobID
|
||||||
|
|||||||
+14
-2
@@ -10,6 +10,8 @@ import (
|
|||||||
"net/url"
|
"net/url"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||||
)
|
)
|
||||||
|
|
||||||
// Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no
|
// Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no
|
||||||
@@ -31,6 +33,11 @@ type Config struct {
|
|||||||
Secret string // token secret (from <id>.pw)
|
Secret string // token secret (from <id>.pw)
|
||||||
Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy.
|
Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy.
|
||||||
Timeout time.Duration
|
Timeout time.Duration
|
||||||
|
|
||||||
|
// IdleConnTimeout bounds how long this client's idle keep-alive connections are retained.
|
||||||
|
// Zero means httpx.DefaultIdleConnTimeout (90s) — it does NOT mean "no timeout", which is the
|
||||||
|
// R-344 defect. Production leaves it unset; only tests set it, to avoid a 90-second wait.
|
||||||
|
IdleConnTimeout time.Duration
|
||||||
}
|
}
|
||||||
|
|
||||||
// NewClient builds a fingerprint-pinned, token-authed PBS client.
|
// NewClient builds a fingerprint-pinned, token-authed PBS client.
|
||||||
@@ -55,8 +62,13 @@ func NewClient(cfg Config) (*Client, error) {
|
|||||||
authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret,
|
authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret,
|
||||||
namespace: cfg.Namespace,
|
namespace: cfg.Namespace,
|
||||||
http: &http.Client{
|
http: &http.Client{
|
||||||
Timeout: timeout,
|
Timeout: timeout,
|
||||||
Transport: &http.Transport{TLSClientConfig: tlsCfg},
|
// R-344: this transport MUST come from httpx. pbsTargetsFromPVE (cmd/felhom-agent/
|
||||||
|
// main.go) builds a fresh Client every cycle and drops the previous one, so a
|
||||||
|
// transport with no idle timeout strands one connection per cycle, forever, on both
|
||||||
|
// sides. That leaked 388 sockets onto ep0 in 46 hours. Pinned by
|
||||||
|
// TestAbandonedClientsReleaseTheirConnections — do not inline an http.Transport here.
|
||||||
|
Transport: httpx.NewTransport(tlsCfg, cfg.IdleConnTimeout),
|
||||||
},
|
},
|
||||||
}, nil
|
}, nil
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,173 @@
|
|||||||
|
package pbs
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"net"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"strings"
|
||||||
|
"sync"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||||
|
)
|
||||||
|
|
||||||
|
// connCounter is the SERVER-side observer. It counts what the server actually holds, which is the
|
||||||
|
// only thing that answers the question this file exists for: a client that believes it closed a
|
||||||
|
// connection, and a server still holding the socket, is precisely the R-344 shape. Asserting on
|
||||||
|
// anything client-side would be asserting the mechanism instead of the consequence.
|
||||||
|
type connCounter struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
open int
|
||||||
|
total int // every connection ever accepted — how many times the client DIALLED
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *connCounter) hook(_ net.Conn, s http.ConnState) {
|
||||||
|
c.mu.Lock()
|
||||||
|
defer c.mu.Unlock()
|
||||||
|
switch s {
|
||||||
|
case http.StateNew:
|
||||||
|
c.open++
|
||||||
|
c.total++
|
||||||
|
case http.StateClosed, http.StateHijacked:
|
||||||
|
c.open--
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *connCounter) counts() (open, total int) {
|
||||||
|
c.mu.Lock()
|
||||||
|
defer c.mu.Unlock()
|
||||||
|
return c.open, c.total
|
||||||
|
}
|
||||||
|
|
||||||
|
// waitForOpen polls until the server holds want connections, or fails naming what it still holds.
|
||||||
|
func (c *connCounter) waitForOpen(t *testing.T, want int, within time.Duration, what string) {
|
||||||
|
t.Helper()
|
||||||
|
deadline := time.Now().Add(within)
|
||||||
|
for {
|
||||||
|
open, total := c.counts()
|
||||||
|
if open == want {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if time.Now().After(deadline) {
|
||||||
|
t.Fatalf("%s: after %s the server still holds %d open connection(s), want %d (%d dialled in total)",
|
||||||
|
what, within, open, want, total)
|
||||||
|
}
|
||||||
|
time.Sleep(5 * time.Millisecond)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// newCountingPBSServer is newPBSTestServer plus a ConnState hook. Kept separate rather than
|
||||||
|
// changing the shared helper, so the existing tests are untouched by this file.
|
||||||
|
func newCountingPBSServer(t *testing.T, fn http.HandlerFunc) (*httptest.Server, string, *connCounter) {
|
||||||
|
t.Helper()
|
||||||
|
cc := &connCounter{}
|
||||||
|
ts := httptest.NewUnstartedServer(fn)
|
||||||
|
ts.Config.ConnState = cc.hook
|
||||||
|
ts.StartTLS()
|
||||||
|
t.Cleanup(ts.Close)
|
||||||
|
return ts, fingerprintOf(ts), cc
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestAbandonedClientsReleaseTheirConnections is Scenario A, and it is the load-bearing test for
|
||||||
|
// R-344.
|
||||||
|
//
|
||||||
|
// It models what pbsTargetsFromPVE actually does — build a client, use it once, drop it on the
|
||||||
|
// floor without closing anything — and asserts the CONSEQUENCE on the server: the connections go
|
||||||
|
// away. Before the fix every one of these stayed established forever on both sides; 388 of them
|
||||||
|
// accumulated on ep0 in 46 hours.
|
||||||
|
//
|
||||||
|
// Deliberately NOT asserted: that err == nil, or that IdleConnTimeout holds some value. Both were
|
||||||
|
// true of the leaking code.
|
||||||
|
func TestAbandonedClientsReleaseTheirConnections(t *testing.T) {
|
||||||
|
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
w.Write([]byte(`{"data":[]}`))
|
||||||
|
})
|
||||||
|
host, port := hostPort(t, ts.URL)
|
||||||
|
|
||||||
|
const cycles = 5
|
||||||
|
for i := 0; i < cycles; i++ {
|
||||||
|
// One fresh client per "cycle", exactly as pbsTargetsFromPVE builds one per collect.
|
||||||
|
c, err := NewClient(Config{
|
||||||
|
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
|
||||||
|
IdleConnTimeout: 50 * time.Millisecond, // production uses the 90s default
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
|
||||||
|
t.Fatalf("cycle %d: %v", i, err)
|
||||||
|
}
|
||||||
|
_ = c // dropped here — nothing closes it, nothing can
|
||||||
|
}
|
||||||
|
|
||||||
|
if _, total := cc.counts(); total != cycles {
|
||||||
|
t.Fatalf("setup is not modelling the leak: want %d separate dials (one per abandoned client), got %d", cycles, total)
|
||||||
|
}
|
||||||
|
cc.waitForOpen(t, 0, 5*time.Second, "abandoned pbs.Clients")
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestAbandonedClientsReleaseTheirConnections_ProductionDefaultIsUsable pins the value that ships.
|
||||||
|
//
|
||||||
|
// The field being settable is exactly how it could silently become zero again — and zero used to
|
||||||
|
// mean "never expire". This asserts the production path (Config leaving it unset) lands on the
|
||||||
|
// standard-library default, so the leak cannot return through an unset field.
|
||||||
|
func TestPBSClient_UnsetIdleTimeoutUsesTheDefault(t *testing.T) {
|
||||||
|
for _, tc := range []struct {
|
||||||
|
name string
|
||||||
|
cfg time.Duration
|
||||||
|
want time.Duration
|
||||||
|
}{
|
||||||
|
{"unset — the production path", 0, httpx.DefaultIdleConnTimeout},
|
||||||
|
{"explicit zero is NOT no-timeout", 0, httpx.DefaultIdleConnTimeout},
|
||||||
|
{"negative is NOT no-timeout", -time.Second, httpx.DefaultIdleConnTimeout},
|
||||||
|
{"an explicit value is honoured", 3 * time.Second, 3 * time.Second},
|
||||||
|
} {
|
||||||
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
|
c, err := NewClient(Config{
|
||||||
|
Server: "pbs.example", Fingerprint: strings.Repeat("ab", 32),
|
||||||
|
TokenID: "u@pbs!t", Secret: "s", IdleConnTimeout: tc.cfg,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
tr, ok := c.http.Transport.(*http.Transport)
|
||||||
|
if !ok {
|
||||||
|
t.Fatalf("transport is %T, not *http.Transport — the httpx wiring was replaced", c.http.Transport)
|
||||||
|
}
|
||||||
|
if tr.IdleConnTimeout != tc.want {
|
||||||
|
t.Fatalf("IdleConnTimeout = %v, want %v (zero would mean connections are retained FOREVER — that is R-344)", tr.IdleConnTimeout, tc.want)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestPBSClient_KeepAliveStillReuses is Scenario C, and it is the guard against a "fix" that is
|
||||||
|
// worse than the bug.
|
||||||
|
//
|
||||||
|
// Disabling keep-alive entirely would also make the leak go away — by dialling a fresh connection
|
||||||
|
// for every single request, which on a box polling ~40,000 times a day is strictly worse than what
|
||||||
|
// we started with. The fix must retire IDLE connections without stopping reuse.
|
||||||
|
func TestPBSClient_KeepAliveStillReuses(t *testing.T) {
|
||||||
|
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
w.Write([]byte(`{"data":[]}`))
|
||||||
|
})
|
||||||
|
host, port := hostPort(t, ts.URL)
|
||||||
|
|
||||||
|
c, err := NewClient(Config{
|
||||||
|
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
|
||||||
|
IdleConnTimeout: 30 * time.Second, // long enough that reuse is what is being measured
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
for i := 0; i < 3; i++ {
|
||||||
|
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
|
||||||
|
t.Fatalf("request %d: %v", i, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if _, total := cc.counts(); total != 1 {
|
||||||
|
t.Fatalf("one client made 3 sequential requests over %d connection(s), want 1 — keep-alive reuse is broken, which would make the poll load WORSE than the leak", total)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -283,7 +283,36 @@ func (m *Manager) Apply(ctx context.Context, fetched bool, block *hub.WirePBSDR)
|
|||||||
h := descriptorHash(block)
|
h := descriptorHash(block)
|
||||||
cf := m.loadConsumedFailed()
|
cf := m.loadConsumedFailed()
|
||||||
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
|
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
|
||||||
m.setStatus(&hub.PBSDRStatus{State: mk.State, StorageID: block.StorageID, Namespace: block.Namespace, AppliedAt: mk.AppliedAt})
|
// R-221: RE-ASSERT THE SEED, DO NOT REMEMBER IT. The marker records that this descriptor
|
||||||
|
// converged; it says nothing about whether the file the seed writes still exists.
|
||||||
|
//
|
||||||
|
// The two live in different places and die at different times. The marker is host-side
|
||||||
|
// (`<agent-state>/pbsdr/`, markerPath above); the seed's target is `agent.json`, and the
|
||||||
|
// installer's `step_agent_config` renders that file from `base = {}` unless an explicit
|
||||||
|
// `--preserve-from` is passed — it NEVER writes an `escrow` section — then replaces it with
|
||||||
|
// O_TRUNC (felhom-host-install.sh:2396, :2449, :2579; the flag is :1246, defaulting empty at
|
||||||
|
// :256). So a rebuild leaves the marker and takes the seed, the hash still matches, this
|
||||||
|
// branch returns, and `escrow.pbs_storage_id` is never written again. The customer then
|
||||||
|
// cannot run the escrow ceremony AT ALL: handleEscrowPreflight fails the `pbs_storage_id`
|
||||||
|
// row and the wizard refuses, with no way forward from inside the product.
|
||||||
|
//
|
||||||
|
// A rebuild is only the case that was measured. The same hole opens for a hand-edited or
|
||||||
|
// restored config, which is the honest reason this is a seam fix rather than an installer
|
||||||
|
// fix — the seed must be a thing the loop asserts, not a thing it did once.
|
||||||
|
//
|
||||||
|
// COST: this runs on the converged path, i.e. every tick (60 s) forever. It is one small
|
||||||
|
// file read plus a JSON parse — no exec, no network, no Proxmox call — and seedEscrowStorageID
|
||||||
|
// returns early once the value matches. That is the whole reason it is affordable here.
|
||||||
|
//
|
||||||
|
// IT MUST NEVER UN-CONVERGE THE BOX: a failure is a Warn plus a message on the published
|
||||||
|
// status, exactly as finishConverged does it. No marker write, no state change, no retry
|
||||||
|
// storm — the early return below still happens either way.
|
||||||
|
msg := ""
|
||||||
|
if err := m.seedEscrowStorageID(block.StorageID); err != nil {
|
||||||
|
msg = "escrow.pbs_storage_id seed failed: " + err.Error() + " (set it manually before the ceremony)"
|
||||||
|
m.logger.Warn("pbsdr: " + msg)
|
||||||
|
}
|
||||||
|
m.setStatus(&hub.PBSDRStatus{State: mk.State, StorageID: block.StorageID, Namespace: block.Namespace, AppliedAt: mk.AppliedAt, Message: msg})
|
||||||
return // idempotent: this exact descriptor already converged
|
return // idempotent: this exact descriptor already converged
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,232 @@
|
|||||||
|
package pbsdr
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"go/ast"
|
||||||
|
"go/parser"
|
||||||
|
"go/printer"
|
||||||
|
"go/token"
|
||||||
|
"io"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
|
||||||
|
)
|
||||||
|
|
||||||
|
// R-221 — the escrow seed must be ASSERTED on every converged tick, not remembered.
|
||||||
|
//
|
||||||
|
// These tests drive the real Apply() with a real temp-dir agent.json and a call-recording runner.
|
||||||
|
// Calling seedEscrowStorageID directly would prove nothing: the defect IS the early return in
|
||||||
|
// Apply, and a test that steps around it cannot see it.
|
||||||
|
|
||||||
|
// convergedMarker writes a marker whose hash matches the block, i.e. puts the manager on exactly
|
||||||
|
// the idempotent path where the seed used to be skipped.
|
||||||
|
func convergedMarker(t *testing.T, m *Manager, block *hub.WirePBSDR, state string) {
|
||||||
|
t.Helper()
|
||||||
|
if err := m.writeState(m.markerPath(), marker{
|
||||||
|
Hash: descriptorHash(block), State: state, AppliedAt: "2026-08-08T00:00:00Z",
|
||||||
|
}); err != nil {
|
||||||
|
t.Fatalf("write marker: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func escrowStorageID(t *testing.T, cfgPath string) string {
|
||||||
|
t.Helper()
|
||||||
|
raw, err := os.ReadFile(cfgPath)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("read config: %v", err)
|
||||||
|
}
|
||||||
|
var doc struct {
|
||||||
|
Escrow struct {
|
||||||
|
PBSStorageID string `json:"pbs_storage_id"`
|
||||||
|
} `json:"escrow"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(raw, &doc); err != nil {
|
||||||
|
t.Fatalf("parse config: %v", err)
|
||||||
|
}
|
||||||
|
return doc.Escrow.PBSStorageID
|
||||||
|
}
|
||||||
|
|
||||||
|
// drBlock mirrors the existing valid fixture (manager_test.go descriptor()) so that validate()
|
||||||
|
// passes and Apply actually reaches the marker check — the branch these tests are about.
|
||||||
|
func drBlock(storageID string) *hub.WirePBSDR {
|
||||||
|
return &hub.WirePBSDR{
|
||||||
|
Enabled: true, StorageID: storageID, PBSTunnelIP: "10.77.0.1",
|
||||||
|
Datastore: "felhom-offsite", Namespace: "peti", TokenID: "felhom@pbs!peti",
|
||||||
|
Fingerprint: testFP,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SCENARIO A — the seed is re-asserted on a converged box, and nothing else happens.
|
||||||
|
//
|
||||||
|
// THIS IS THE TEST THAT MATTERS. It must fail against the pre-R-221 tree; if it passes there, it is
|
||||||
|
// not testing the defect and that is the finding.
|
||||||
|
func TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls(t *testing.T) {
|
||||||
|
r := &fakeRunner{}
|
||||||
|
st := &fakeStorage{found: true, active: []bool{true}}
|
||||||
|
c := &fakeConsumer{}
|
||||||
|
m, cfgPath := newTestManager(t, r, st, c)
|
||||||
|
|
||||||
|
block := drBlock("felhom-pbs-dr")
|
||||||
|
convergedMarker(t, m, block, "applied")
|
||||||
|
|
||||||
|
// the shape a rebuild leaves behind: escrow section present, pbs_storage_id GONE
|
||||||
|
if got := escrowStorageID(t, cfgPath); got != "" {
|
||||||
|
t.Fatalf("precondition: config already carries a storage id %q", got)
|
||||||
|
}
|
||||||
|
|
||||||
|
m.Apply(context.Background(), true, block)
|
||||||
|
|
||||||
|
if got := escrowStorageID(t, cfgPath); got != "felhom-pbs-dr" {
|
||||||
|
t.Errorf("the converged tick did not re-assert the seed: escrow.pbs_storage_id = %q, want %q.\n"+
|
||||||
|
"This is R-221: the marker survives a rebuild, the descriptor hash still matches, the early "+
|
||||||
|
"return fires and the seed never runs into the config that no longer has it — so the customer "+
|
||||||
|
"cannot run the escrow ceremony at all.", got, "felhom-pbs-dr")
|
||||||
|
}
|
||||||
|
|
||||||
|
// ...and the idempotent path is STILL idempotent. This assertion is not decorative: without it
|
||||||
|
// a "fix" that simply deletes the early return would pass the line above.
|
||||||
|
if calls := r.recorded(); len(calls) != 0 {
|
||||||
|
t.Errorf("a converged tick must execute ZERO Proxmox commands; got %d: %+v", len(calls), calls)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SCENARIO B — an operator's own different value survives, and the warning names both.
|
||||||
|
func TestSeedReasserted_NeverClobbersAnOperatorValue(t *testing.T) {
|
||||||
|
r := &fakeRunner{}
|
||||||
|
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
|
||||||
|
|
||||||
|
if err := os.WriteFile(cfgPath, []byte(
|
||||||
|
`{"log_level":"info","escrow":{"posture":"zero_knowledge","pbs_storage_id":"operator-chosen"},`+
|
||||||
|
`"custom_unknown":{"keep":1}}`), 0o600); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
block := drBlock("hub-chosen")
|
||||||
|
convergedMarker(t, m, block, "applied")
|
||||||
|
|
||||||
|
m.Apply(context.Background(), true, block)
|
||||||
|
|
||||||
|
if got := escrowStorageID(t, cfgPath); got != "operator-chosen" {
|
||||||
|
t.Errorf("a value a person put there was overwritten by the descriptor: got %q, want %q", got, "operator-chosen")
|
||||||
|
}
|
||||||
|
// unknown keys must still round-trip
|
||||||
|
raw, _ := os.ReadFile(cfgPath)
|
||||||
|
if !strings.Contains(string(raw), "custom_unknown") {
|
||||||
|
t.Error("an unknown config key was dropped by the re-assert")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SCENARIO C — the ceremony preflight's live read sees the re-asserted value with NO restart.
|
||||||
|
//
|
||||||
|
// The preflight itself lives in internal/localapi and reads the file through config.Load; what this
|
||||||
|
// asserts is the half that belongs to this package: after a converged tick, THE FILE ON DISK carries
|
||||||
|
// the id, so any live re-read is green. The daemon is never restarted in this test because it is
|
||||||
|
// never started — which is the point.
|
||||||
|
func TestSeedReasserted_IsVisibleOnDiskImmediately(t *testing.T) {
|
||||||
|
r := &fakeRunner{}
|
||||||
|
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
|
||||||
|
block := drBlock("felhom-pbs-dr")
|
||||||
|
convergedMarker(t, m, block, "applied")
|
||||||
|
|
||||||
|
// preflight's predicate BEFORE: storageID == "" → the row is NOT OK
|
||||||
|
if escrowStorageID(t, cfgPath) != "" {
|
||||||
|
t.Fatal("precondition")
|
||||||
|
}
|
||||||
|
m.Apply(context.Background(), true, block)
|
||||||
|
// preflight's predicate AFTER, from the same file the ceremony subprocess loads
|
||||||
|
if id := escrowStorageID(t, cfgPath); id == "" {
|
||||||
|
t.Error("the preflight row would still be NOT OK after a converged tick")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A seed failure must NEVER un-converge the box: no marker rewrite, no state change, and the status
|
||||||
|
// still reports the marker's converged state — with the failure surfaced as a message.
|
||||||
|
func TestSeedReassertFailure_DoesNotUnconverge(t *testing.T) {
|
||||||
|
r := &fakeRunner{}
|
||||||
|
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
|
||||||
|
block := drBlock("felhom-pbs-dr")
|
||||||
|
convergedMarker(t, m, block, "applied")
|
||||||
|
markerBefore, err := os.ReadFile(m.markerPath())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
// make the seed fail in a way it cannot recover from: unparseable config
|
||||||
|
if err := os.WriteFile(cfgPath, []byte(`{ this is not json`), 0o600); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
m.Apply(context.Background(), true, block)
|
||||||
|
|
||||||
|
after, err := os.ReadFile(m.markerPath())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("the marker was removed by a seed failure: %v", err)
|
||||||
|
}
|
||||||
|
if string(after) != string(markerBefore) {
|
||||||
|
t.Error("a seed failure rewrote the convergence marker — it must not touch state")
|
||||||
|
}
|
||||||
|
if calls := r.recorded(); len(calls) != 0 {
|
||||||
|
t.Errorf("a seed failure must not trigger Proxmox work; got %+v", calls)
|
||||||
|
}
|
||||||
|
st := m.Status()
|
||||||
|
if st == nil || st.State != "applied" {
|
||||||
|
t.Errorf("a seed failure must leave the box converged; status = %+v", st)
|
||||||
|
}
|
||||||
|
if st != nil && !strings.Contains(st.Message, "seed failed") {
|
||||||
|
t.Errorf("a seed failure must be surfaced on the status, got message %q", st.Message)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// SEAM WIRING — production must construct the manager with the live config path, or the whole seed
|
||||||
|
// leg is inert. Three shipped defects in this project were fully green while their seam was never
|
||||||
|
// wired, so this walks main.go's AST for the actual call rather than grepping: a commented-out call
|
||||||
|
// satisfies strings.Contains, and an AST walk cannot see a comment.
|
||||||
|
func TestProductionWiring_NewManagerGetsTheLiveConfigPath(t *testing.T) {
|
||||||
|
path := filepath.Join("..", "..", "cmd", "felhom-agent", "main.go")
|
||||||
|
fset := token.NewFileSet()
|
||||||
|
f, err := parser.ParseFile(fset, path, nil, 0) // comments not even collected
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parse main.go: %v", err)
|
||||||
|
}
|
||||||
|
found := false
|
||||||
|
ast.Inspect(f, func(n ast.Node) bool {
|
||||||
|
call, ok := n.(*ast.CallExpr)
|
||||||
|
if !ok {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
sel, ok := call.Fun.(*ast.SelectorExpr)
|
||||||
|
if !ok || sel.Sel.Name != "NewManager" {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
if pkg, ok := sel.X.(*ast.Ident); !ok || pkg.Name != "pbsdr" {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
// signature: (runner, px, hubc, stateDir, secretDir, configPath, logger)
|
||||||
|
if len(call.Args) < 6 {
|
||||||
|
t.Errorf("pbsdr.NewManager called with %d args, expected >= 6", len(call.Args))
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
var buf strings.Builder
|
||||||
|
if err := printNode(&buf, fset, call.Args[5]); err != nil {
|
||||||
|
t.Fatalf("print arg: %v", err)
|
||||||
|
}
|
||||||
|
got := buf.String()
|
||||||
|
if !strings.Contains(got, "SourcePath") {
|
||||||
|
t.Errorf("pbsdr.NewManager's configPath argument is %q, which is not the live config path.\n"+
|
||||||
|
"With an empty or wrong path seedEscrowStorageID returns nil immediately and the entire "+
|
||||||
|
"R-221 fix is inert while every test above still passes.", got)
|
||||||
|
}
|
||||||
|
found = true
|
||||||
|
return false
|
||||||
|
})
|
||||||
|
if !found {
|
||||||
|
t.Error("no pbsdr.NewManager call found in main.go — the manager is not constructed in production")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func printNode(w io.Writer, fset *token.FileSet, n ast.Node) error {
|
||||||
|
return printer.Fprint(w, fset, n)
|
||||||
|
}
|
||||||
@@ -10,6 +10,8 @@ import (
|
|||||||
"net/url"
|
"net/url"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
|
||||||
)
|
)
|
||||||
|
|
||||||
// doer is the minimal HTTP surface the client needs; *http.Client satisfies it.
|
// doer is the minimal HTTP surface the client needs; *http.Client satisfies it.
|
||||||
@@ -65,8 +67,10 @@ func NewClient(cfg Config) (*Client, error) {
|
|||||||
timeout = 30 * time.Second
|
timeout = 30 * time.Second
|
||||||
}
|
}
|
||||||
hc := &http.Client{
|
hc := &http.Client{
|
||||||
Timeout: timeout,
|
Timeout: timeout,
|
||||||
Transport: &http.Transport{TLSClientConfig: tlsCfg},
|
// R-344, consistency only: built ONCE per process, so it never accumulated and contributed
|
||||||
|
// nothing to the ep0 leak. Same missing default, corrected for the same reason.
|
||||||
|
Transport: httpx.NewTransport(tlsCfg, 0),
|
||||||
}
|
}
|
||||||
return &Client{
|
return &Client{
|
||||||
base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json",
|
base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json",
|
||||||
|
|||||||
@@ -43,12 +43,21 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
|||||||
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
|
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
|
||||||
SHARED_INSTRUCTIONS = os.path.join(
|
SHARED_INSTRUCTIONS = os.path.join(
|
||||||
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
|
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
|
||||||
|
# R-389 — shared, like the two above: it lives in felhom.eu/scripts/ and is never copied.
|
||||||
|
SHARED_OBSERVATIONS = os.path.join(
|
||||||
|
os.path.dirname(ROOT), "felhom.eu", "scripts", "observations_gate.py")
|
||||||
|
|
||||||
# (label, absolute script path, args, fast)
|
# (label, absolute script path, args, fast)
|
||||||
GATES = [
|
GATES = [
|
||||||
("reuse-refs", SHARED_REUSE, [ROOT], True),
|
("reuse-refs", SHARED_REUSE, [ROOT], True),
|
||||||
("instructions", SHARED_INSTRUCTIONS, [ROOT], True),
|
("instructions", SHARED_INSTRUCTIONS, [ROOT], True),
|
||||||
("published", os.path.join(ROOT, "scripts", "check-published-versions.py"), [], False),
|
("published", os.path.join(ROOT, "scripts", "check-published-versions.py"), [], False),
|
||||||
|
# R-273: the tag half of a release. Legs 1-2 need no network, so it runs in --fast too — the
|
||||||
|
# missing TAG is what actually broke every install, and the pre-push hook is the earliest place
|
||||||
|
# that can catch it.
|
||||||
|
("release-complete", os.path.join(ROOT, "scripts", "check-release-complete.py"), [], True),
|
||||||
|
# R-389 — a REPORT.md observation with no register row behind it. Fast: stdlib file reads.
|
||||||
|
("observations", SHARED_OBSERVATIONS, [ROOT], True),
|
||||||
]
|
]
|
||||||
|
|
||||||
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
|
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
|
||||||
|
|||||||
@@ -95,6 +95,25 @@ PROBE_CONFIG = "configs/felhom-agent.service"
|
|||||||
|
|
||||||
TAG_RE = re.compile(r"^v(\d+\.\d+\.\d+)$")
|
TAG_RE = re.compile(r"^v(\d+\.\d+\.\d+)$")
|
||||||
|
|
||||||
|
# THE retention number, read from the one file that owns it. A check and the policy it enforces
|
||||||
|
# must read the same number from the same place, or they drift and the drift looks like a defect
|
||||||
|
# in something else — which is exactly what happened on 2026-08-08/09 (R-287).
|
||||||
|
RETENTION_FILE = os.path.join(os.path.dirname(os.path.abspath(__file__)), "retention-policy.json")
|
||||||
|
|
||||||
|
|
||||||
|
def retention_kept():
|
||||||
|
"""How many of the newest generic versions the registry is expected to still serve.
|
||||||
|
|
||||||
|
Fails CLOSED and LOUD: a missing or unreadable policy file makes the check INCONCLUSIVE
|
||||||
|
rather than silently unbounded. An unbounded check would re-create the red this fixed; a
|
||||||
|
silently-bounded one would be worse.
|
||||||
|
"""
|
||||||
|
with open(RETENTION_FILE, encoding="utf-8") as fh:
|
||||||
|
n = json.load(fh)["generic_versions_kept"]
|
||||||
|
if not isinstance(n, int) or n < 1:
|
||||||
|
raise ValueError("generic_versions_kept must be a positive int, got %r" % (n,))
|
||||||
|
return n
|
||||||
|
|
||||||
tried = []
|
tried = []
|
||||||
|
|
||||||
|
|
||||||
@@ -189,7 +208,28 @@ def main():
|
|||||||
print(" no v<semver> tags in this repo yet — nothing to check, and nothing proven")
|
print(" no v<semver> tags in this repo yet — nothing to check, and nothing proven")
|
||||||
print("\ncheck-published-versions: NOTHING TO CHECK")
|
print("\ncheck-published-versions: NOTHING TO CHECK")
|
||||||
return 0
|
return 0
|
||||||
print(" %d released version(s) to verify: %s" % (len(versions), ", ".join(versions)))
|
all_versions = versions
|
||||||
|
try:
|
||||||
|
keep = retention_kept()
|
||||||
|
except Exception as e:
|
||||||
|
inconclusive("cannot read the retention policy (%s): %s" % (RETENTION_FILE, e))
|
||||||
|
|
||||||
|
# Bound the assertion to what the registry is expected to still hold. Sorted by SEMVER, not
|
||||||
|
# lexically: "0.9.0" > "0.10.0" as strings, and that would silently drop the wrong end.
|
||||||
|
def _key(v):
|
||||||
|
return tuple(int(x) for x in v.split("."))
|
||||||
|
versions = sorted(all_versions, key=_key)[-keep:]
|
||||||
|
dropped = [v for v in all_versions if v not in versions]
|
||||||
|
|
||||||
|
print(" %d released version(s); retention policy keeps the newest %d" % (len(all_versions), keep))
|
||||||
|
print(" verifying: %s" % ", ".join(versions))
|
||||||
|
if dropped:
|
||||||
|
# NEVER silent. A bounded check that does not say what it stopped covering is how a
|
||||||
|
# narrowing becomes permanent by accident.
|
||||||
|
print(" NOT ASSERTED (older than the retention window, and therefore not expected to be")
|
||||||
|
print(" downloadable): %s" % ", ".join(dropped))
|
||||||
|
print(" ^ these versions still have git TAGS and are still installable in the sense that")
|
||||||
|
print(" their configs resolve; what is no longer asserted is the BINARY's presence.")
|
||||||
|
|
||||||
bad = []
|
bad = []
|
||||||
for v in versions:
|
for v in versions:
|
||||||
|
|||||||
@@ -0,0 +1,124 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""check-release-complete.py — the version at the head of CHANGELOG.md is a COMPLETE release.
|
||||||
|
|
||||||
|
THE DEFECT THIS IS A MACHINE FOR (2026-08-08/09, R-273). Agent v0.128.0 was built, tested,
|
||||||
|
CHANGELOG'd and published to the package registry — and its git tag was never pushed. The hub then
|
||||||
|
vouched it, and because felhom-host-install.sh fetches an agent's config files from
|
||||||
|
`raw/tag/v<version>/configs/`, EVERY fresh install and every reinstall died at step 5 of 8, as root,
|
||||||
|
on a virgin machine, for the better part of a day.
|
||||||
|
|
||||||
|
`scripts/release-agent.sh` already warns about exactly this, in as many words:
|
||||||
|
"a released version without a git tag 404s a box mid-install, as root"
|
||||||
|
The warning was there, it was correct, and the step was still missed. **So the fix is a machine and
|
||||||
|
not a reminder** — that is the whole point of this file.
|
||||||
|
|
||||||
|
WHAT IT ASSERTS, for the newest `## vX.Y.Z` in CHANGELOG.md:
|
||||||
|
1. a git tag `vX.Y.Z` EXISTS, and
|
||||||
|
2. it points at a commit that is an ANCESTOR OF (or equal to) the tip it was released from — a tag
|
||||||
|
parked on an unrelated commit is not a release, and
|
||||||
|
3. the generic package for X.Y.Z is DOWNLOADABLE.
|
||||||
|
|
||||||
|
(3) needs the network. (1) and (2) do not, and they are the half that actually failed — so this gate
|
||||||
|
is useful offline and says so rather than going quiet.
|
||||||
|
|
||||||
|
EXIT CODES, matching this repo's other gates: 0 clean, 1 convicted, 2 inconclusive. An unreachable
|
||||||
|
registry is INCONCLUSIVE for leg 3 only; legs 1 and 2 still run and can still convict.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import urllib.error
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||||
|
GITEA_BASE = os.environ.get("GITEA_BASE", "https://gitea.dooplex.hu").rstrip("/")
|
||||||
|
OWNER, PKG = "admin", "felhom-agent"
|
||||||
|
HEAD_RE = re.compile(r"^##\s+v?(\d+\.\d+\.\d+)\b", re.M)
|
||||||
|
|
||||||
|
|
||||||
|
def git(*args):
|
||||||
|
return subprocess.run(("git",) + args, cwd=ROOT, capture_output=True, text=True)
|
||||||
|
|
||||||
|
|
||||||
|
def head_version():
|
||||||
|
ch = os.path.join(ROOT, "CHANGELOG.md")
|
||||||
|
if not os.path.exists(ch):
|
||||||
|
return None
|
||||||
|
m = HEAD_RE.search(open(ch, encoding="utf-8").read())
|
||||||
|
return m.group(1) if m else None
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
print("check-release-complete — the newest CHANGELOG version is a complete release")
|
||||||
|
v = head_version()
|
||||||
|
if not v:
|
||||||
|
print(" no '## vX.Y.Z' heading in CHANGELOG.md — nothing to check, and nothing proven")
|
||||||
|
return 0
|
||||||
|
tag = "v" + v
|
||||||
|
print(" newest CHANGELOG version: %s" % tag)
|
||||||
|
|
||||||
|
problems, inconclusive = [], []
|
||||||
|
|
||||||
|
# ---- leg 1 + 2: the tag, and where it points. Offline-capable. -----------------------------
|
||||||
|
r = git("rev-parse", "-q", "--verify", "refs/tags/%s^{commit}" % tag)
|
||||||
|
if r.returncode != 0:
|
||||||
|
# A shallow CI clone has no tags of its own; ask the remote before convicting, so this
|
||||||
|
# gate does not fire on a clone shape rather than on a real defect.
|
||||||
|
ls = git("ls-remote", "--tags", "origin", "refs/tags/%s" % tag)
|
||||||
|
if ls.returncode != 0:
|
||||||
|
inconclusive.append("cannot reach origin to look for tag %s: %s"
|
||||||
|
% (tag, ls.stderr.strip()[:120]))
|
||||||
|
elif not ls.stdout.strip():
|
||||||
|
problems.append(
|
||||||
|
"TAG %s DOES NOT EXIST. The installer fetches this version's configs from\n"
|
||||||
|
" %s/%s/felhom-agent/raw/tag/%s/configs/ — without the tag every install\n"
|
||||||
|
" 404s mid-run, as root. Fix: git tag -a %s <released-commit> && git push origin %s"
|
||||||
|
% (tag, GITEA_BASE, OWNER, tag, tag, tag))
|
||||||
|
else:
|
||||||
|
print(" ok tag %s exists on origin (not in this shallow clone)" % tag)
|
||||||
|
else:
|
||||||
|
sha = r.stdout.strip()
|
||||||
|
anc = git("merge-base", "--is-ancestor", sha, "HEAD")
|
||||||
|
if anc.returncode == 0:
|
||||||
|
print(" ok tag %s -> %s, an ancestor of HEAD" % (tag, sha[:10]))
|
||||||
|
else:
|
||||||
|
problems.append("tag %s points at %s, which is NOT an ancestor of HEAD — a tag parked "
|
||||||
|
"on an unrelated commit is not a release" % (tag, sha[:10]))
|
||||||
|
|
||||||
|
# ---- leg 3: the package. Needs the network. ------------------------------------------------
|
||||||
|
url = "%s/api/packages/%s/generic/%s/%s/%s" % (GITEA_BASE, OWNER, PKG, v, PKG)
|
||||||
|
req = urllib.request.Request(url, method="HEAD")
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(req, timeout=25) as resp:
|
||||||
|
if resp.status == 200:
|
||||||
|
print(" ok package %s is downloadable" % v)
|
||||||
|
else:
|
||||||
|
problems.append("package %s returned HTTP %s at %s" % (v, resp.status, url))
|
||||||
|
except urllib.error.HTTPError as e:
|
||||||
|
if e.code == 404:
|
||||||
|
problems.append("PACKAGE %s IS NOT PUBLISHED (HTTP 404 at %s).\n"
|
||||||
|
" Fix: bash scripts/release-agent.sh %s" % (v, url, v))
|
||||||
|
else:
|
||||||
|
inconclusive.append("registry returned HTTP %s for %s" % (e.code, v))
|
||||||
|
except Exception as e:
|
||||||
|
inconclusive.append("registry unreachable (%s) — leg 3 not checked; legs 1-2 still ran" % e)
|
||||||
|
|
||||||
|
if problems:
|
||||||
|
print("\ncheck-release-complete: INCOMPLETE RELEASE")
|
||||||
|
for p in problems:
|
||||||
|
print(" - " + p)
|
||||||
|
return 1
|
||||||
|
if inconclusive:
|
||||||
|
print("\ncheck-release-complete: INCONCLUSIVE — an undetermined result is never a pass")
|
||||||
|
for i in inconclusive:
|
||||||
|
print(" - " + i)
|
||||||
|
return 2
|
||||||
|
print("\ncheck-release-complete: %s is tagged, placed and published." % tag)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
{
|
||||||
|
"_comment": [
|
||||||
|
"THE retention number for published agent artifacts. One file, read by everything that",
|
||||||
|
"depends on it, because a check and the policy it enforces must read the same number from the",
|
||||||
|
"same place or they drift — and the drift looks like a defect in something else.",
|
||||||
|
"",
|
||||||
|
"WHAT WENT WRONG WITHOUT IT (2026-08-08/09). The registry stopped serving felhom-agent",
|
||||||
|
"0.120.0 and older, while scripts/check-published-versions.py demanded that EVERY git tag",
|
||||||
|
"still be downloadable. Both rules are individually sensible; together they are impossible.",
|
||||||
|
"CI went red at a commit whose own run had been green the day before, on a true finding that",
|
||||||
|
"no one could act on. The red will return at the next publish unless the two read one number.",
|
||||||
|
"",
|
||||||
|
"HOW THE NUMBER WAS ARRIVED AT — stated honestly, because it is weaker than it looks.",
|
||||||
|
"generic_versions_kept is 10 because that is what the registry demonstrably holds today",
|
||||||
|
"(felhom-agent 0.121.0..0.128.0 = 10 versions, queried 2026-08-09). It is an OBSERVED state,",
|
||||||
|
"NOT a ruling anyone has been able to locate: no register row records a package prune, R-210",
|
||||||
|
"is WAITING-ON-OPERATOR and says 'Nothing was deleted; this is a list, not an action', and it",
|
||||||
|
"concerns local Docker images rather than this registry. Container packages currently hold 19",
|
||||||
|
"each, so there is no uniform ten-per-package cap visible either. See R-287.",
|
||||||
|
"",
|
||||||
|
"SO THIS FILE IS A FLOOR, NOT A LICENCE. It says: CI may assume nothing older than the newest",
|
||||||
|
"N generic versions is still downloadable. It does NOT authorise deleting anything, and the",
|
||||||
|
"operator should confirm or replace the number — at which point this file changes and both",
|
||||||
|
"readers follow it in the same commit.",
|
||||||
|
"",
|
||||||
|
"THE DEEPER BOUND, recorded so a future session does not have to re-derive it: the principled",
|
||||||
|
"limit is the hub's vouched min_agent floor (0.127.0 on 2026-08-09). Nothing can install an",
|
||||||
|
"agent below it — the hub refuses to vouch one and boxes update to the floor — so a released",
|
||||||
|
"version below the floor being un-downloadable costs nothing real. Bounding on the floor would",
|
||||||
|
"be better than bounding on a count, and it needs the gate to read the hub, which is network",
|
||||||
|
"the gate does not have today. Filed as the follow-up in R-287.",
|
||||||
|
"",
|
||||||
|
"NEVER retire a git TAG to satisfy this. felhom-host-install.sh fetches an agent's config",
|
||||||
|
"files from raw/tag/v<version>/configs/, so deleting a tag retires the ability to install that",
|
||||||
|
"version at all — a strictly worse act than an un-downloadable binary."
|
||||||
|
],
|
||||||
|
"generic_versions_kept": 10,
|
||||||
|
"readers": [
|
||||||
|
"scripts/check-published-versions.py — bounds its assertion to the newest N versions",
|
||||||
|
"documentation/runbooks/registry-retention.md (felhom.eu) — the prune procedure"
|
||||||
|
],
|
||||||
|
"recorded": "2026-08-09",
|
||||||
|
"recorded_by": "CC, from the registry's observed state; NOT from a located operator ruling"
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user